feat(coverage): n_units_collected separates sampling from real zeros (#36)
provenance$coverage's n_units_reporting is category-conditional: it counts governments with rows for the SPECIFIC requested category, which conflates two different things -- a government never collected that year (sampling), and one collected but genuinely spending nothing in that category (a real zero). FY2012 Georgia Police is the motivating case from the issue: a complete census year reads as a 69% "response rate" because most of the gap is cities that contract policing to the county sheriff, not non-response. Adds a second counter, n_units_collected: how many of the caller's expected cohort appear in the corpus that year for ANY category. n_units_collected / n_units_expected is the true collection rate; n_units_reporting / n_units_collected is category participation among collected units. cog_geographic_rollup() and cog_peer_compare() both carry it; cog_explain() prints it alongside n_units_reporting. Two real bugs caught and fixed while finishing this (both against the already-written, previously-uncommitted draft): - .coverage_table()'s candidate list for the collection query was derived from the category-filtered result rows, not the caller's full expected cohort. A government with zero rows in the requested category across every requested year never appears in that result, so it was silently excluded from n_units_collected too -- collapsing the new counter back to the old, broken one for exactly the governments it exists to count. Fixed by threading an explicit `expected_ids` (all_govids / peer_govids) through instead. - The collection query hardcoded long_view = "spending_long_harmonized", which does not exist on a corpus with schema_version < 5 (R/basis.R resolves basis = "raw" there; R/views.R only registers the harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare() would hard-error on a corpus vintage the package otherwise explicitly supports. Fixed by deriving long_view from the basis cog_spending() actually resolved (prov$basis) via the existing .select_long_view() helper, matching how every other basis-aware query in the package already does this. Also: cog_explain()'s general "complete census only in years ending in 2 or 7" footnote was gated on the OLD counter's absence, making it permanently unreachable now that both callers always supply the new one -- ungated it, since the explanation is orthogonal to which counter set is present. Dropped a dead conditional branch, fixed two stale roxygen blocks in R/peers.R/R/rollup.R still describing the old two-counter model, fixed the same staleness in README.md, and switched two `uscogdata:::` self-references to the package's own convention of calling internal helpers unqualified. 1101 tests pass (2 skipped live-corpus), including new direct regression tests for both bugs above (one exercising a government collected-but-absent from a category result, one running the full rollup/peer-compare path against a doctored schema_version 4 corpus). Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW -- all addressed above). Closes #36. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -227,7 +227,8 @@ A statewide total resting on a fifth of the universe looks exactly like one
|
||||
resting on all of it, so every multi-government result now says which it is:
|
||||
|
||||
```r
|
||||
attr(rollup, "provenance")$coverage # per-year n_units_reporting, is_census_year
|
||||
attr(rollup, "provenance")$coverage
|
||||
# per-year n_units_expected, n_units_collected, n_units_reporting, is_census_year
|
||||
```
|
||||
|
||||
`cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` take a
|
||||
@@ -235,8 +236,14 @@ attr(rollup, "provenance")$coverage # per-year n_units_reporting, is_census_ye
|
||||
`"consistent"` (only units reporting in every requested year, a balanced
|
||||
panel).
|
||||
|
||||
`n_units_reporting` is **category-conditional**, and it is not a response rate. A government that was surveyed and genuinely spends
|
||||
nothing in the requested category is indistinguishable from one never surveyed.
|
||||
`n_units_reporting` is **category-conditional**: it counts governments with
|
||||
rows for the *specific* category you asked for, so a government that was
|
||||
surveyed and genuinely spends nothing in that category is indistinguishable
|
||||
from one never surveyed — it is not a response rate on its own.
|
||||
`n_units_collected` is the number that separates them: governments present in
|
||||
the corpus that year for *any* category. `n_units_collected / n_units_expected`
|
||||
is the true collection rate; `n_units_reporting / n_units_collected` is
|
||||
category participation among collected units.
|
||||
|
||||
### Absent cells mean two different things
|
||||
|
||||
|
||||
Reference in New Issue
Block a user