Executive Summary
Two federal systems produce county-year road-death counts, but they count different populations. The Fatality Analysis Reporting System counts fatalities from qualifying fatal crashes in the county where the crash happened. Federal mortality data count certified motor-vehicle traffic deaths by the county where the person lived. Neither is the corrected or authoritative version of the other.
Across 755 unsuppressed Ohio and Pennsylvania county-years from 2015 to 2024 the two counts differ by a median of 16.0% of their own average, with the resident-death count the larger in the typical county-year (median signed divergence -9.2%). The state totals provide a separate measure of alignment: in every one of the twenty state-years, the published resident-death total runs above the crash-location total, at a ratio between 1.019 and 1.113, median 1.066. Measuring each county-year against the count that ratio would imply produces a second, separately reported benchmark — a median absolute residual of 14.4%, with a median signed residual of -3.1%.
The next question is whether particular counties are reliably far apart, so a caveat could attach to a place. They are not: grouping the residual by county across the 73 counties with at least five eligible years gives an intraclass correlation of 0.179 against a prespecified floor of 0.300. The evidence does not support treating residual divergence as a stable property of identifiable counties, so this paper publishes no county ranking and no named county figure.
The most important limitation is who is missing. The publisher suppresses small county cells, so 795 of the 1,550 county-years are withheld and never estimated here. Everything below describes the unsuppressed, larger-county subset, and up to 24.0% of a state-year's published resident road deaths sits inside counties outside it.
The permitted implication is a source-selection one: anyone quoting a county road-death figure can state how far the two counts typically run apart and caveat accordingly, without treating either as an undercount of the other.
Business Question
For Ohio and Pennsylvania county-years from 2015 to 2024 where the mortality publisher releases a county cell: how far apart do the crash-location fatality count and the resident motor-vehicle death count run, how does that compare with the definitional alignment already visible between the two state totals, and does the separately measured residual concentrate stably in identifiable counties?
The reader is a professional research reader who needs governed public-data comparisons. This is a reference benchmark; it names no actor who must act, and proposes no action.
The question matters because county road-death figures are often quoted without naming which system produced them. Each publisher exposes its own side through a query tool, and the crash-census publisher's leading-causes reporting places crash fatalities against certified causes of death — at national grain. Published work comparing the two systems nationally found their totals close in aggregate: 458,071 motor-vehicle traffic deaths in multiple-cause-of-death data against 452,318 in the crash census over 1999–2009. What has been missing is the local magnitude, and a reporter writing about one county has had no governed figure for it.
The supported branch is the one observed: local divergence is material and measurable. The null branch would have been equally useful: close agreement would have shown that the systems were interchangeable for this comparison.
Data & Scope
The crash-location measure comes from the Fatality Analysis Reporting System, published by the National Highway Traffic Safety Administration and distributed through its National Center for Statistics and Analysis: a census of qualifying fatal traffic crashes, summing the fatalities on those crashes by the county of the crash. Its scope note is explicit — qualifying fatal crashes only, not all police-reported crashes, and not all motor-vehicle resident deaths.
The resident measure comes from CDC WONDER's underlying-cause-of-death data, published by the Centers for Disease Control and Prevention: the published count of motor-vehicle traffic deaths of county residents, selected by a single underlying cause. Its scope note is equally explicit — suppressed cells remain undisclosed, and this is not a count of crash-location deaths.
Four mismatches between them are irreducible and disclosed rather than resolved.
- Place. One counts where the crash happened, the other where the person lived. Commuter and through-traffic counties diverge mechanically — expected, not error.
- Universe. The crash census admits only qualifying crashes on a public traffic way; deaths on private property or non-traffic ways fall outside it while the death certificate records them.
- Attribution. The mortality file selects a single underlying cause of death; the crash census records a crash-involved death inside its qualifying window.
- Timing. One is dated by year of death, the other by year of crash, so a death crossing a year boundary lands in different years.
The universe is county-years in Ohio and Pennsylvania, 2015 through 2024, for which the publisher releases an unsuppressed county cell. There is no denominator anywhere, so no exposure, rate, or per-capita reading is computed or permitted. Each source is aggregated to county-year first and only the aggregates compared; no record-level join is performed and no individual or incident linkage is asserted.
Selection is dominated by the publisher's suppression rule, reported rather than worked around.
| Ledger | County-years |
|---|---|
| County-year keys in the mortality table, Ohio and Pennsylvania 2015–2024 | 1,550 |
| Cells the publisher releases | 755 |
| Cells the publisher withholds | 795 |
| Published cells excluded for falling below the minimum count | 0 |
| Eligible county-years analysed | 755 |
| Distinct counties contributing at least one eligible county-year | 121 |
Suppressed cells are excluded, never imputed, bounded, or read as zero. The crash-location side of a suppressed county-year is withheld from any paired display too, because showing it would narrow the suppressed range. On the crash-census side 21,877 qualifying fatal-crash records were retained; neither source contributed a row whose county and state codes disagreed.
Method
The comparison uses a symmetric measure, because neither system may serve as the denominator of truth. For each eligible county-year the divergence is the difference between the two counts as a share of their own average: (crash-location − resident) / ((crash-location + resident) / 2). Positive means the crash-location count is larger, negative that the resident count is. The materiality threshold, fixed before any result was inspected, is 0.100 — where quoting one system rather than the other changes a headline county figure by more than rounding.
A cell enters only if the larger count is at least 10 — a floor aligned with the publisher's own suppression boundary rather than added on top, because below ten deaths the statistic measures integer granularity rather than system disagreement. No published cell fell below it.
The definitional alignment is measured separately, at state level, where it is directly observable: the ratio of the published state resident-death total to the state crash-location total. Each eligible county-year then gets a definition-expected resident count — its crash-location count scaled by that state-year ratio — and a residual, computed with the same symmetric formula against it.
The three quantities — divergence, ratio, residual — are reported side by side as three distinct measurements. None is subtracted from, divided by, or expressed as a share of another anywhere in this paper. The ratio is a descriptive state-level alignment factor and nothing else: not a correction, not an adjustment, not an estimate of true deaths, and not a mechanism operating inside any county.
Whether divergence is a stable property of places is tested with a one-way random-effects intraclass correlation of the residual grouped by county, over counties with at least five eligible years, against a floor of 0.300.
Four sensitivity tests were prespecified: the minimum count moved to 5, 20, and 30 alongside the primary 10; each year and then each state dropped in turn; both one-sided forms recomputed; and the withheld share of each state-year's published resident deaths measured directly. A ladder tier can affect the verdict only if it independently clears the same coverage floors as the primary analysis; tiers that do not are published in full and excluded from grading. Detail is in the method notes below.
Findings
Both the raw divergence and the separately defined residual are material.
Scroll chart horizontally to read all labels.
| Measure | Scope | County-years | 10th pct | 25th pct | Median | 75th pct | 90th pct | Median absolute |
|---|---|---|---|---|---|---|---|---|
| Raw divergence | Both states | 755 | -0.400000 | -0.235294 | -0.092308 | 0.057143 | 0.181818 | 0.160000 |
| Raw divergence | Ohio | 376 | -0.473389 | -0.285714 | -0.120241 | 0.000000 | 0.181818 | 0.175827 |
| Raw divergence | Pennsylvania | 379 | -0.363636 | -0.200000 | -0.060606 | 0.074074 | 0.200106 | 0.142857 |
| Residual | Both states | 755 | -0.344435 | -0.175653 | -0.030769 | 0.112925 | 0.257544 | 0.143750 |
| Residual | Ohio | 376 | -0.385949 | -0.207373 | -0.045513 | 0.091153 | 0.259386 | 0.152588 |
| Residual | Pennsylvania | 379 | -0.308708 | -0.146279 | -0.013782 | 0.129642 | 0.249859 | 0.135821 |
Two features carry the practical meaning. The distribution is wide — a quarter of eligible county-years sit at or below -0.235294 on the raw measure, a tenth at or below -0.400000 — and it is asymmetric around zero, the same fact the median signed divergence of -9.2% expresses.
The state-level alignment is the second quantity, directly observable.
Scroll chart horizontally to read all labels.
Across the twenty state-years the ratio runs from 1.018886 to 1.113074 with a median of 1.066364, sitting above 1.00 in both states in every year and running higher and more variable in Ohio than in Pennsylvania. The consistent direction of that state-level difference is itself useful context for readers.
The third quantity is the residual, material at 14.4% median absolute against the same 0.100 threshold. By state, its median absolute value is 0.152588 in Ohio and 0.135821 in Pennsylvania.
The stronger hypothesis this research set out to test failed, and informatively. Residual divergence does not concentrate in identifiable counties: the intraclass correlation is 0.178839 against the prespecified 0.300 floor. That is what forbids a county list — a reader cannot look up "this county runs 20% apart" and expect it to hold next year.
Those are the measurements. The interpretation is limited: the benchmark establishes how far apart the two systems typically run at county-year grain in these two states, and that the distance does not attach to particular places. It establishes nothing about why, and attributes no error or fault to either system.
Association Analysis
One association analysis is justified by the question this paper asks: the stability test already reported, the intraclass correlation of the residual grouped by county.
| Component | Value | Meaning |
|---|---|---|
| Counties with five or more eligible years | 73 | the stability group |
| County-year residual values in those groups | 657 | observations entering the test |
| Effective group size | 8.996744 | scaling term for unequal group sizes |
| Mean square between counties | 0.132036 | variation of county means around the grand mean |
| Mean square within counties | 0.044616 | variation within each county |
| Intraclass correlation | 0.178839 | against a prespecified floor of 0.300 |
The ICC is well short of the prespecified level for a reusable county-level caveat. Counties with fewer than five eligible years are outside the test; the method notes below give how many years each of the 121 counties contributes, unnamed.
No further association test is justified. Correlating divergence against county characteristics, modelling commuter flows, or testing the pooled median for significance would each import an explanatory or causal inference this evidence cannot support. Association is not causation, and here the test returns a null: a place-based caveat is not available, which is not the same as another factor being responsible.
Business Implications
This is a reference benchmark. Its value is in what a reader can now state precisely, not in any action it recommends.
A reader quoting a county road-death figure for Ohio or Pennsylvania can name the system and the typical distance to the other: a median of 16.0% of the two counts' average, usually with the resident count larger. The choice of source can materially change a county-year figure; the state-level ratio supplies separate context about how the two state totals align.
A reader can also decline a comparison the evidence does not support. Because the residual does not concentrate by county, no "this county is a known outlier" caveat is available. The correct caveat is distributional: the two counts typically differ by roughly a sixth of their average, in either direction, broadly rather than in particular places.
Several readings are ruled out wherever this benchmark is used. Neither system is correct, more complete, or authoritative; neither is undercounting the other, and no county has a missing resident or death inferred here. The state-level ratio is not a correction factor and must not be applied to a published count to produce an adjusted figure. No share of the divergence is explained, removed, or accounted for by any quantity here: the raw and residual benchmarks are two independent measurements, not a decomposition to be netted against each other. Nothing supports a rate, exposure, or per-capita reading: no denominator is used.
Limitations
The two systems measure different universes, so divergence is definitional evidence rather than error attribution. The four mismatches — place, universe, attribution, timing — are disclosed above and not resolved by this design.
The result describes the unsuppressed, larger-county subset only: 795 of the 1,550 county-years are withheld by the publisher, and up to 24.0% of a state-year's published resident road deaths sits inside them. Suppressed cells are never estimated, imputed, named, or paired with the crash-location value, so how far the systems run apart inside the smallest counties stays unknown.
Counties below the minimum count or with a suppressed mortality cell are never named with a divergence figure, and because residual divergence is not a stable county property, no county is named with one at all. The state-level ratio is a descriptive alignment factor, never a correction, and no exposure, rate, risk-per-mile, or per-capita interpretation is permitted. The scope is Ohio and Pennsylvania county-years 2015–2024 at a single data snapshot, so nothing generalizes to other states, years, or vintages.
Four questions remain open. Why a particular county-year diverges as much as it does — no county-level mechanism is established. How far apart the systems run inside suppressed counties. Whether the same magnitudes hold elsewhere. And how much of any county-year's divergence comes from each of the four mismatches; only the aggregate alignment factor is quantified, and it is not a county mechanism.
Sources and method notes
Sources
- National Center for Statistics and Analysis, NHTSA, tools, publications, and data — publisher and distribution surface for the Fatality Analysis Reporting System.
- NHTSA, Motor Vehicle Traffic Crashes as a Leading Cause of Death and Potential Years of Life Lost — the publisher's own comparison of crash-census fatalities against certified causes of death, 2012–2022, at national grain.
- I-Jen P. Castle, Hsiao-Ye Yi, Ralph W. Hingson and Aaron M. White, State Variation in Underreporting of Alcohol Involvement on Death Certificates: Motor Vehicle Traffic Crash Fatalities as an Example, Journal of Studies on Alcohol and Drugs, 2014 — 458,071 motor-vehicle traffic deaths in multiple-cause-of-death data against 452,318 in the crash census for 1999–2009.
All external sources were verified on 11 August 2026.
Field definitions
| Term | Definition |
|---|---|
| Crash-location count | Sum of fatalities recorded on qualifying fatal crashes, grouped by the county of the crash. |
| Resident count | Published count of certified motor-vehicle traffic deaths of residents of the county. |
| Eligible county-year | A county-year whose mortality cell is published and whose larger count is at least 10. |
| Divergence | (crash-location − resident) ÷ ((crash-location + resident) ÷ 2). |
| State-level definitional ratio | Published state resident-death total ÷ state crash-location total, for the same state and year. |
| Definition-expected resident count | Crash-location count × the state-level ratio for that state and year. |
| Residual | (definition-expected − resident) ÷ ((definition-expected + resident) ÷ 2). |
| Intraclass correlation | One-way random-effects ICC(1) of the residual grouped by county, over counties with at least five eligible years. |
Method and sensitivity notes
Coverage floors and observed values: at least 200 eligible county-years, observed 755; at least 30 counties with five or more eligible years, observed 73; both states present in every year, satisfied; withheld share of a state-year's published resident deaths below 0.500, observed maximum 0.240476.
Minimum-count sensitivity. A tier is graded only if it independently clears those floors.
| Minimum count | Cells | Counties ≥ 5 years | Divergence | Material | Residual | Material | ICC | Stable | Tier role |
|---|---|---|---|---|---|---|---|---|---|
| 5 | 755 | 73 | 0.160000 | yes | 0.143750 | yes | 0.178839 | no | disclosure only |
| 10 | 755 | 73 | 0.160000 | yes | 0.143750 | yes | 0.178839 | no | graded |
| 20 | 378 | 38 | 0.144330 | yes | 0.130257 | yes | 0.238263 | no | graded |
| 30 | 217 | 22 | 0.137931 | yes | 0.098361 | no | 0.250501 | no | under-covered disclosure |
The two graded tiers agree on all three conclusions. The tier at 5 is disclosure-only because it crosses the publisher's own suppression boundary. The tier at 30 is published in full and excluded from grading for one recorded reason: it retains only 22 counties with five or more eligible years against an inherited floor of 30. Its residual of 0.098361 is disclosed as not material at the unchanged 0.100 threshold. No argument in this paper rests on how close that value is to the threshold.
Leave-one-out. Dropping each year in turn and then each state in turn gives 12 replicates. All 12 reproduce every primary conclusion.
Alternative metrics. Recomputing with one-sided denominators gives a median absolute value of 0.153846 divided by the resident count and 0.165138 divided by the crash-location count. Both are material at the 0.100 threshold, so the conclusion does not depend on the symmetric form.
Suppression coverage. The maximum share of a state-year's published resident deaths sitting inside withheld counties is 0.240476, against a stop bound of 0.500.
County recurrence. Of the 121 counties contributing at least one eligible county-year, 18 contribute one year, 16 two, 8 three, 6 four, and 73 contribute five or more. Only that last group enters the stability test. No county is named.
Chart notes
The distribution figure uses 0.1-wide bands from -0.6 to 0.6 with one explicit tail band at each end; the bands are fixed in code and were not chosen after inspecting a result. Series are separable by colour and by fill, so neither figure depends on colour perception. The chart-data files carry the exact labels, values, series order, units, stored rounding, and display rounding behind each figure.