CountyCaseField Ledger contour markCountyCase

Executive Summary

Two federal systems produce county-year road-death counts, but they count different populations. The Fatality Analysis Reporting System counts fatalities from qualifying fatal crashes in the county where the crash happened. Federal mortality data count certified motor-vehicle traffic deaths by the county where the person lived. Neither is the corrected or authoritative version of the other.

Across 755 unsuppressed Ohio and Pennsylvania county-years from 2015 to 2024 the two counts differ by a median of 16.0% of their own average, with the resident-death count the larger in the typical county-year (median signed divergence -9.2%). The state totals provide a separate measure of alignment: in every one of the twenty state-years, the published resident-death total runs above the crash-location total, at a ratio between 1.019 and 1.113, median 1.066. Measuring each county-year against the count that ratio would imply produces a second, separately reported benchmark — a median absolute residual of 14.4%, with a median signed residual of -3.1%.

The next question is whether particular counties are reliably far apart, so a caveat could attach to a place. They are not: grouping the residual by county across the 73 counties with at least five eligible years gives an intraclass correlation of 0.179 against a prespecified floor of 0.300. The evidence does not support treating residual divergence as a stable property of identifiable counties, so this paper publishes no county ranking and no named county figure.

The most important limitation is who is missing. The publisher suppresses small county cells, so 795 of the 1,550 county-years are withheld and never estimated here. Everything below describes the unsuppressed, larger-county subset, and up to 24.0% of a state-year's published resident road deaths sits inside counties outside it.

The permitted implication is a source-selection one: anyone quoting a county road-death figure can state how far the two counts typically run apart and caveat accordingly, without treating either as an undercount of the other.

Business Question

For Ohio and Pennsylvania county-years from 2015 to 2024 where the mortality publisher releases a county cell: how far apart do the crash-location fatality count and the resident motor-vehicle death count run, how does that compare with the definitional alignment already visible between the two state totals, and does the separately measured residual concentrate stably in identifiable counties?

The reader is a professional research reader who needs governed public-data comparisons. This is a reference benchmark; it names no actor who must act, and proposes no action.

The question matters because county road-death figures are often quoted without naming which system produced them. Each publisher exposes its own side through a query tool, and the crash-census publisher's leading-causes reporting places crash fatalities against certified causes of death — at national grain. Published work comparing the two systems nationally found their totals close in aggregate: 458,071 motor-vehicle traffic deaths in multiple-cause-of-death data against 452,318 in the crash census over 1999–2009. What has been missing is the local magnitude, and a reporter writing about one county has had no governed figure for it.

The supported branch is the one observed: local divergence is material and measurable. The null branch would have been equally useful: close agreement would have shown that the systems were interchangeable for this comparison.

Data & Scope

The crash-location measure comes from the Fatality Analysis Reporting System, published by the National Highway Traffic Safety Administration and distributed through its National Center for Statistics and Analysis: a census of qualifying fatal traffic crashes, summing the fatalities on those crashes by the county of the crash. Its scope note is explicit — qualifying fatal crashes only, not all police-reported crashes, and not all motor-vehicle resident deaths.

The resident measure comes from CDC WONDER's underlying-cause-of-death data, published by the Centers for Disease Control and Prevention: the published count of motor-vehicle traffic deaths of county residents, selected by a single underlying cause. Its scope note is equally explicit — suppressed cells remain undisclosed, and this is not a count of crash-location deaths.

Four mismatches between them are irreducible and disclosed rather than resolved.

The universe is county-years in Ohio and Pennsylvania, 2015 through 2024, for which the publisher releases an unsuppressed county cell. There is no denominator anywhere, so no exposure, rate, or per-capita reading is computed or permitted. Each source is aggregated to county-year first and only the aggregates compared; no record-level join is performed and no individual or incident linkage is asserted.

Selection is dominated by the publisher's suppression rule, reported rather than worked around.

LedgerCounty-years
County-year keys in the mortality table, Ohio and Pennsylvania 2015–20241,550
Cells the publisher releases755
Cells the publisher withholds795
Published cells excluded for falling below the minimum count0
Eligible county-years analysed755
Distinct counties contributing at least one eligible county-year121

Suppressed cells are excluded, never imputed, bounded, or read as zero. The crash-location side of a suppressed county-year is withheld from any paired display too, because showing it would narrow the suppressed range. On the crash-census side 21,877 qualifying fatal-crash records were retained; neither source contributed a row whose county and state codes disagreed.

Method

The comparison uses a symmetric measure, because neither system may serve as the denominator of truth. For each eligible county-year the divergence is the difference between the two counts as a share of their own average: (crash-location − resident) / ((crash-location + resident) / 2). Positive means the crash-location count is larger, negative that the resident count is. The materiality threshold, fixed before any result was inspected, is 0.100 — where quoting one system rather than the other changes a headline county figure by more than rounding.

A cell enters only if the larger count is at least 10 — a floor aligned with the publisher's own suppression boundary rather than added on top, because below ten deaths the statistic measures integer granularity rather than system disagreement. No published cell fell below it.

The definitional alignment is measured separately, at state level, where it is directly observable: the ratio of the published state resident-death total to the state crash-location total. Each eligible county-year then gets a definition-expected resident count — its crash-location count scaled by that state-year ratio — and a residual, computed with the same symmetric formula against it.

The three quantities — divergence, ratio, residual — are reported side by side as three distinct measurements. None is subtracted from, divided by, or expressed as a share of another anywhere in this paper. The ratio is a descriptive state-level alignment factor and nothing else: not a correction, not an adjustment, not an estimate of true deaths, and not a mechanism operating inside any county.

Whether divergence is a stable property of places is tested with a one-way random-effects intraclass correlation of the residual grouped by county, over counties with at least five eligible years, against a floor of 0.300.

Four sensitivity tests were prespecified: the minimum count moved to 5, 20, and 30 alongside the primary 10; each year and then each state dropped in turn; both one-sided forms recomputed; and the withheld share of each state-year's published resident deaths measured directly. A ladder tier can affect the verdict only if it independently clears the same coverage floors as the primary analysis; tiers that do not are published in full and excluded from grading. Detail is in the method notes below.

Findings

Both the raw divergence and the separately defined residual are material.

Grouped histogram of county-year divergence between the two official road-death counts, Ohio and Pennsylvania 2015 to 2024. Filled bars show raw divergence and outlined bars the residual against the definition-expected count, across the same 755 eligible county-years; both distributions are wide and centred slightly below zero.

Scroll chart horizontally to read all labels.

Raw divergence and the residual measured against the definition-expected count, side by side over the same 755 eligible county-years and never merged. Median absolute raw divergence 16.0%, median absolute residual 14.4%; both exceed the 0.100 materiality threshold.
MeasureScopeCounty-years10th pct25th pctMedian75th pct90th pctMedian absolute
Raw divergenceBoth states755-0.400000-0.235294-0.0923080.0571430.1818180.160000
Raw divergenceOhio376-0.473389-0.285714-0.1202410.0000000.1818180.175827
Raw divergencePennsylvania379-0.363636-0.200000-0.0606060.0740740.2001060.142857
ResidualBoth states755-0.344435-0.175653-0.0307690.1129250.2575440.143750
ResidualOhio376-0.385949-0.207373-0.0455130.0911530.2593860.152588
ResidualPennsylvania379-0.308708-0.146279-0.0137820.1296420.2498590.135821

Two features carry the practical meaning. The distribution is wide — a quarter of eligible county-years sit at or below -0.235294 on the raw measure, a tenth at or below -0.400000 — and it is asymmetric around zero, the same fact the median signed divergence of -9.2% expresses.

The state-level alignment is the second quantity, directly observable.

Line chart of the state-level definitional ratio between the two official road-death counts for Ohio and Pennsylvania, 2015 to 2024. Both lines stay above a reference line at 1.00 in every year, with Ohio generally higher and more variable than Pennsylvania.

Scroll chart horizontally to read all labels.

Published state resident-death total divided by the state crash-location total. A value of 1.00 means the two state totals are equal. The ratio is a descriptive state-level alignment factor, never a correction.

Across the twenty state-years the ratio runs from 1.018886 to 1.113074 with a median of 1.066364, sitting above 1.00 in both states in every year and running higher and more variable in Ohio than in Pennsylvania. The consistent direction of that state-level difference is itself useful context for readers.

The third quantity is the residual, material at 14.4% median absolute against the same 0.100 threshold. By state, its median absolute value is 0.152588 in Ohio and 0.135821 in Pennsylvania.

The stronger hypothesis this research set out to test failed, and informatively. Residual divergence does not concentrate in identifiable counties: the intraclass correlation is 0.178839 against the prespecified 0.300 floor. That is what forbids a county list — a reader cannot look up "this county runs 20% apart" and expect it to hold next year.

Those are the measurements. The interpretation is limited: the benchmark establishes how far apart the two systems typically run at county-year grain in these two states, and that the distance does not attach to particular places. It establishes nothing about why, and attributes no error or fault to either system.

Association Analysis

One association analysis is justified by the question this paper asks: the stability test already reported, the intraclass correlation of the residual grouped by county.

ComponentValueMeaning
Counties with five or more eligible years73the stability group
County-year residual values in those groups657observations entering the test
Effective group size8.996744scaling term for unequal group sizes
Mean square between counties0.132036variation of county means around the grand mean
Mean square within counties0.044616variation within each county
Intraclass correlation0.178839against a prespecified floor of 0.300

The ICC is well short of the prespecified level for a reusable county-level caveat. Counties with fewer than five eligible years are outside the test; the method notes below give how many years each of the 121 counties contributes, unnamed.

No further association test is justified. Correlating divergence against county characteristics, modelling commuter flows, or testing the pooled median for significance would each import an explanatory or causal inference this evidence cannot support. Association is not causation, and here the test returns a null: a place-based caveat is not available, which is not the same as another factor being responsible.

Business Implications

This is a reference benchmark. Its value is in what a reader can now state precisely, not in any action it recommends.

A reader quoting a county road-death figure for Ohio or Pennsylvania can name the system and the typical distance to the other: a median of 16.0% of the two counts' average, usually with the resident count larger. The choice of source can materially change a county-year figure; the state-level ratio supplies separate context about how the two state totals align.

A reader can also decline a comparison the evidence does not support. Because the residual does not concentrate by county, no "this county is a known outlier" caveat is available. The correct caveat is distributional: the two counts typically differ by roughly a sixth of their average, in either direction, broadly rather than in particular places.

Several readings are ruled out wherever this benchmark is used. Neither system is correct, more complete, or authoritative; neither is undercounting the other, and no county has a missing resident or death inferred here. The state-level ratio is not a correction factor and must not be applied to a published count to produce an adjusted figure. No share of the divergence is explained, removed, or accounted for by any quantity here: the raw and residual benchmarks are two independent measurements, not a decomposition to be netted against each other. Nothing supports a rate, exposure, or per-capita reading: no denominator is used.

Limitations

The two systems measure different universes, so divergence is definitional evidence rather than error attribution. The four mismatches — place, universe, attribution, timing — are disclosed above and not resolved by this design.

The result describes the unsuppressed, larger-county subset only: 795 of the 1,550 county-years are withheld by the publisher, and up to 24.0% of a state-year's published resident road deaths sits inside them. Suppressed cells are never estimated, imputed, named, or paired with the crash-location value, so how far the systems run apart inside the smallest counties stays unknown.

Counties below the minimum count or with a suppressed mortality cell are never named with a divergence figure, and because residual divergence is not a stable county property, no county is named with one at all. The state-level ratio is a descriptive alignment factor, never a correction, and no exposure, rate, risk-per-mile, or per-capita interpretation is permitted. The scope is Ohio and Pennsylvania county-years 2015–2024 at a single data snapshot, so nothing generalizes to other states, years, or vintages.

Four questions remain open. Why a particular county-year diverges as much as it does — no county-level mechanism is established. How far apart the systems run inside suppressed counties. Whether the same magnitudes hold elsewhere. And how much of any county-year's divergence comes from each of the four mismatches; only the aggregate alignment factor is quantified, and it is not a county mechanism.

Sources and method notes

Sources

All external sources were verified on 11 August 2026.

Field definitions

TermDefinition
Crash-location countSum of fatalities recorded on qualifying fatal crashes, grouped by the county of the crash.
Resident countPublished count of certified motor-vehicle traffic deaths of residents of the county.
Eligible county-yearA county-year whose mortality cell is published and whose larger count is at least 10.
Divergence(crash-location − resident) ÷ ((crash-location + resident) ÷ 2).
State-level definitional ratioPublished state resident-death total ÷ state crash-location total, for the same state and year.
Definition-expected resident countCrash-location count × the state-level ratio for that state and year.
Residual(definition-expected − resident) ÷ ((definition-expected + resident) ÷ 2).
Intraclass correlationOne-way random-effects ICC(1) of the residual grouped by county, over counties with at least five eligible years.

Method and sensitivity notes

Coverage floors and observed values: at least 200 eligible county-years, observed 755; at least 30 counties with five or more eligible years, observed 73; both states present in every year, satisfied; withheld share of a state-year's published resident deaths below 0.500, observed maximum 0.240476.

Minimum-count sensitivity. A tier is graded only if it independently clears those floors.

Minimum countCellsCounties ≥ 5 yearsDivergenceMaterialResidualMaterialICCStableTier role
5755730.160000yes0.143750yes0.178839nodisclosure only
10755730.160000yes0.143750yes0.178839nograded
20378380.144330yes0.130257yes0.238263nograded
30217220.137931yes0.098361no0.250501nounder-covered disclosure

The two graded tiers agree on all three conclusions. The tier at 5 is disclosure-only because it crosses the publisher's own suppression boundary. The tier at 30 is published in full and excluded from grading for one recorded reason: it retains only 22 counties with five or more eligible years against an inherited floor of 30. Its residual of 0.098361 is disclosed as not material at the unchanged 0.100 threshold. No argument in this paper rests on how close that value is to the threshold.

Leave-one-out. Dropping each year in turn and then each state in turn gives 12 replicates. All 12 reproduce every primary conclusion.

Alternative metrics. Recomputing with one-sided denominators gives a median absolute value of 0.153846 divided by the resident count and 0.165138 divided by the crash-location count. Both are material at the 0.100 threshold, so the conclusion does not depend on the symmetric form.

Suppression coverage. The maximum share of a state-year's published resident deaths sitting inside withheld counties is 0.240476, against a stop bound of 0.500.

County recurrence. Of the 121 counties contributing at least one eligible county-year, 18 contribute one year, 16 two, 8 three, 6 four, and 73 contribute five or more. Only that last group enters the stability test. No county is named.

Chart notes

The distribution figure uses 0.1-wide bands from -0.6 to 0.6 with one explicit tail band at each end; the bands are fixed in code and were not chosen after inspecting a result. Series are separable by colour and by fill, so neither figure depends on colour perception. The chart-data files carry the exact labels, values, series order, units, stored rounding, and display rounding behind each figure.