n_units_reporting is category-conditional and cannot be read as a response rate #36

Closed
opened 2026-08-05 10:51:59 -04:00 by jared · 0 comments
Owner

Filed while building the client-facing Southern API guide (cog_explorer/docs/superpowers/specs/2026-08-05-south-api-guide-design.md). Verified live 2026-08-05. Verdict: definitional — the numbers are right; the field name and its absent documentation invite a specific wrong reading.

Symptom

provenance$coverage's n_units_reporting is category-conditional. It counts units with rows for the requested category, which conflates two very different things:

  • a unit that was not collected that year (sampling), and
  • a unit that was collected and genuinely spends nothing in that category.

Nothing in the field name, the schema, or the docs distinguishes them, and the natural reading of n_units_reporting / n_units_expected is "response rate" — which it is not.

Evidence

Georgia cities, category = "Police", via the API's geographic rollup (cog_geographic_rollup()):

Year n_units_reporting n_units_expected is_census_year
FY2022 393 567 true
FY2024 94 567 false

FY2022 is a complete census year. Collection is not partial. Yet the ratio reads 69.3%, which as a response rate would be alarming. It is not a response rate: the 174-city gap is overwhelmingly real zeros — Georgia cities that contract policing to the county sheriff and therefore report no police spending at all.

The FY2024 figure mixes both effects, and there is currently no way to separate them.

Why this matters

The disclosure added for #13 is genuinely good and the guide depends on it. But the field is one an analyst will reach for to build exactly the wrong thing: a data-quality or response-rate dashboard, or a coverage-weighted adjustment, both of which would be badly biased by the real-zero component.

Concretely, dividing an aggregate by n_units_reporting to get a "mean per reporting unit" is wrong for any category that a meaningful share of units legitimately does not fund.

Suggested

  1. Document the semantics in cog_geographic_rollup() / cog_peer_compare() @return, and mirror it to the API's /data-dictionary: n_units_reporting counts units with rows for this category, not units collected in this year. It is not a response rate and must not be used as one.
  2. Consider a second counter — n_units_collected, units present in the corpus that year for any category. With both, the two effects separate cleanly: n_units_collected / n_units_expected is the true collection rate, and n_units_reporting / n_units_collected is category participation among collected units. That is the pair an analyst actually wants, and both are derivable from data already in hand.
  3. The comparison that is currently valid — the same category across a census and a sample year, where the real-zero component is roughly constant and the difference is the sampling — is worth stating explicitly as the supported use, since it is the one the guide relies on.

Item 1 is the minimum and should land regardless. Item 2 is the fix that makes the field safe to use.

Filed while building the **client-facing Southern API guide** (`cog_explorer/docs/superpowers/specs/2026-08-05-south-api-guide-design.md`). Verified live 2026-08-05. Verdict: **definitional** — the numbers are right; the field name and its absent documentation invite a specific wrong reading. ## Symptom `provenance$coverage`'s `n_units_reporting` is **category-conditional**. It counts units with rows *for the requested category*, which conflates two very different things: - a unit that was not collected that year (sampling), and - a unit that was collected and genuinely spends nothing in that category. Nothing in the field name, the schema, or the docs distinguishes them, and the natural reading of `n_units_reporting / n_units_expected` is "response rate" — which it is not. ## Evidence Georgia cities, `category = "Police"`, via the API's geographic rollup (`cog_geographic_rollup()`): | Year | `n_units_reporting` | `n_units_expected` | `is_census_year` | |---|---|---|---| | FY2022 | 393 | 567 | **true** | | FY2024 | 94 | 567 | false | FY2022 is a **complete census year**. Collection is not partial. Yet the ratio reads 69.3%, which as a response rate would be alarming. It is not a response rate: the 174-city gap is overwhelmingly real zeros — Georgia cities that contract policing to the county sheriff and therefore report no police spending at all. The FY2024 figure mixes both effects, and there is currently no way to separate them. ## Why this matters The disclosure added for #13 is genuinely good and the guide depends on it. But the field is one an analyst will reach for to build exactly the wrong thing: a data-quality or response-rate dashboard, or a coverage-weighted adjustment, both of which would be badly biased by the real-zero component. Concretely, dividing an aggregate by `n_units_reporting` to get a "mean per reporting unit" is wrong for any category that a meaningful share of units legitimately does not fund. ## Suggested 1. **Document the semantics** in `cog_geographic_rollup()` / `cog_peer_compare()` `@return`, and mirror it to the API's `/data-dictionary`: `n_units_reporting` counts units with **rows for this category**, not units collected in this year. It is not a response rate and must not be used as one. 2. **Consider a second counter** — `n_units_collected`, units present in the corpus that year for *any* category. With both, the two effects separate cleanly: `n_units_collected / n_units_expected` is the true collection rate, and `n_units_reporting / n_units_collected` is category participation among collected units. That is the pair an analyst actually wants, and both are derivable from data already in hand. 3. The comparison that **is** currently valid — the same category across a census and a sample year, where the real-zero component is roughly constant and the difference is the sampling — is worth stating explicitly as the supported use, since it is the one the guide relies on. Item 1 is the minimum and should land regardless. Item 2 is the fix that makes the field safe to use.
jared added the south-guidekodor/fixkodor/fixseverity/mediumverdict/definitional labels 2026-08-05 10:51:59 -04:00
kodor added the kodor/triagedkodor/triaged labels 2026-08-10 02:04:59 -04:00
jared added the
origin
review
type
debt
ws
corpus
labels 2026-08-23 16:07:40 -04:00
jared closed this issue 2026-09-09 11:51:06 -04:00
jared referenced this issue from a commit 2026-09-09 13:38:27 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#36