n_units_reporting is category-conditional and cannot be read as a response rate #36
Closed
opened 2026-08-05 10:51:59 -04:00 by jared
·
0 comments
No Branch/Tag Specified
main
ci/mirror-canonical-tags
chore/release-47-badges-mirror-pr
docs/readme-perf-remeasure-56
feat/pagination-search-balances-57
feat/duckdb-threads-60
feat/cohort-predicates-58
fix/windows-backslash-paths
ci/mirror-to-github
ci/github-actions-matrix
feat/public-release-0.3.0
chore/fixture-sb203
ci/apt-https
fix/pushdown-pagination
feat/all-categories-37
fix/partial-coverage-signposting-9
fix/schema-v7
fix/cog-categories-balance-subtype
feat/cog-balances-25
feat/revenue-concepts-12
feat/expenditure-concepts-11
feat/coverage-disclosure-13
feat/complete-argument-18
fix/kodor-batch-14-15-16
fix/all-scoped-series-breaks-19
fix/regen-fixture-corpus-18
test/walkthrough-findings
feat/expenditure-concept
fix/3-url-trailing-slash
feat/phase-r3-signposting
fix/fixture-option-b-aggregates
feat/phase-r2-harmonization
feat/phase-r1-forward
feat/cog-gov-search-basket-mode
v0.4.0
Labels
Clear labels
kodor
kodor/feature-proposal
kodor/fix
kodor/needs-review
kodor/triaged
madison-walkthrough
severity/high
severity/low
severity/medium
south-guide
verdict/defect
verdict/definitional
kodor
kodor/feature-proposal
kodor/fix
kodor/needs-review
kodor/triaged
Kodor should process this issue
Kodor has written a feature proposal
Kodor should implement a fix (assigned to Kodor)
Kodor's work or failure needs Jared's review
Kodor has already triaged this issue (skip)
Surfaced while building the client-facing Southern API guide
needs
human
Cannot move without a person -- a decision, a check an agent cannot make, something outside the repo
origin
client
Came from a client ask
origin
obligation
Created by a change elsewhere
origin
review
Came from human review
origin
roborev
Promoted from a roborev finding
type
chore
Maintenance with no behaviour change
type
debt
Owed work -- docs, tests, cleanup a change obligated
type
decision
Needs a decision before work can proceed
type
defect
Something is wrong
type
feature
New capability
ws
api
Query verbs and results
ws
corpus
Corpus, mirror, provenance
ws
docs
Vignettes and guides
Assign a task to kodor
Kodor thinks this needs a feature.
Kodor should fix this
Kodor thinks the user is ready to review this.
Kodor is done with this issue.
Milestone
No items
No Milestone
Projects
Clear projects
No projects
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Civilytics/uscogdata#36
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Filed while building the client-facing Southern API guide (
cog_explorer/docs/superpowers/specs/2026-08-05-south-api-guide-design.md). Verified live 2026-08-05. Verdict: definitional — the numbers are right; the field name and its absent documentation invite a specific wrong reading.Symptom
provenance$coverage'sn_units_reportingis category-conditional. It counts units with rows for the requested category, which conflates two very different things:Nothing in the field name, the schema, or the docs distinguishes them, and the natural reading of
n_units_reporting / n_units_expectedis "response rate" — which it is not.Evidence
Georgia cities,
category = "Police", via the API's geographic rollup (cog_geographic_rollup()):n_units_reportingn_units_expectedis_census_yearFY2022 is a complete census year. Collection is not partial. Yet the ratio reads 69.3%, which as a response rate would be alarming. It is not a response rate: the 174-city gap is overwhelmingly real zeros — Georgia cities that contract policing to the county sheriff and therefore report no police spending at all.
The FY2024 figure mixes both effects, and there is currently no way to separate them.
Why this matters
The disclosure added for #13 is genuinely good and the guide depends on it. But the field is one an analyst will reach for to build exactly the wrong thing: a data-quality or response-rate dashboard, or a coverage-weighted adjustment, both of which would be badly biased by the real-zero component.
Concretely, dividing an aggregate by
n_units_reportingto get a "mean per reporting unit" is wrong for any category that a meaningful share of units legitimately does not fund.Suggested
cog_geographic_rollup()/cog_peer_compare()@return, and mirror it to the API's/data-dictionary:n_units_reportingcounts units with rows for this category, not units collected in this year. It is not a response rate and must not be used as one.n_units_collected, units present in the corpus that year for any category. With both, the two effects separate cleanly:n_units_collected / n_units_expectedis the true collection rate, andn_units_reporting / n_units_collectedis category participation among collected units. That is the pair an analyst actually wants, and both are derivable from data already in hand.Item 1 is the minimum and should land regardless. Item 2 is the fix that makes the field safe to use.
jared referenced this issue2026-08-23 16:14:53 -04:00