feat(coverage): n_units_collected separates sampling from real zeros (#36) #71

Merged
jared merged 1 commits from fix/coverage-collected-36 into main 2026-09-09 11:51:06 -04:00
Owner

provenance$coverage's n_units_reporting is category-conditional: it counts governments with rows for the SPECIFIC requested category, conflating "never collected that year" (sampling) with "collected but genuinely spends nothing in this category" (a real zero). Motivating case: FY2012 Georgia Police reads as a 69% "response rate" in a complete census year, when most of the gap is cities that contract policing to the county sheriff.

Adds a second counter, n_units_collected: how many of the caller's expected cohort appear in the corpus that year for ANY category. n_units_collected / n_units_expected is the true collection rate; n_units_reporting / n_units_collected is category participation among collected units. cog_geographic_rollup() and cog_peer_compare() both carry it; cog_explain() prints it.

Two real bugs caught and fixed while finishing this draft, both verified against the fixture and covered by new regression tests:

  1. The collection query's candidate list was derived from category-filtered result rows instead of the caller's full expected cohort -- silently excluding exactly the "collected but real zero" governments the counter exists to count. Fixed by threading expected_ids through explicitly.
  2. The collection query hardcoded spending_long_harmonized, which doesn't exist on a corpus with schema_version < 5 -- would hard-error on a vintage the package otherwise explicitly supports. Fixed by deriving the view from the basis cog_spending() actually resolved.

Also fixes a now-permanently-unreachable footnote in cog_explain(), a dead conditional branch, two stale roxygen blocks (R/peers.R, R/rollup.R) and one in README.md still describing the old two-counter model.

1101 tests pass (2 skipped live-corpus). Independently reviewed by a code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW -- all addressed).

Closes #36.

🤖 Generated with Claude Code

provenance$coverage's n_units_reporting is category-conditional: it counts governments with rows for the SPECIFIC requested category, conflating "never collected that year" (sampling) with "collected but genuinely spends nothing in this category" (a real zero). Motivating case: FY2012 Georgia Police reads as a 69% "response rate" in a complete census year, when most of the gap is cities that contract policing to the county sheriff. Adds a second counter, `n_units_collected`: how many of the caller's expected cohort appear in the corpus that year for ANY category. `n_units_collected / n_units_expected` is the true collection rate; `n_units_reporting / n_units_collected` is category participation among collected units. `cog_geographic_rollup()` and `cog_peer_compare()` both carry it; `cog_explain()` prints it. Two real bugs caught and fixed while finishing this draft, both verified against the fixture and covered by new regression tests: 1. The collection query's candidate list was derived from category-filtered result rows instead of the caller's full expected cohort -- silently excluding exactly the "collected but real zero" governments the counter exists to count. Fixed by threading `expected_ids` through explicitly. 2. The collection query hardcoded `spending_long_harmonized`, which doesn't exist on a corpus with `schema_version < 5` -- would hard-error on a vintage the package otherwise explicitly supports. Fixed by deriving the view from the basis `cog_spending()` actually resolved. Also fixes a now-permanently-unreachable footnote in `cog_explain()`, a dead conditional branch, two stale roxygen blocks (R/peers.R, R/rollup.R) and one in README.md still describing the old two-counter model. 1101 tests pass (2 skipped live-corpus). Independently reviewed by a code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW -- all addressed). Closes #36. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
jared added 1 commit 2026-09-09 11:46:45 -04:00
feat(coverage): n_units_collected separates sampling from real zeros (#36)
R-CMD-check / check (push) Successful in 3m39s
R-CMD-check / check (pull_request) Successful in 3m42s
b41d5ee2aa
provenance$coverage's n_units_reporting is category-conditional: it
counts governments with rows for the SPECIFIC requested category, which
conflates two different things -- a government never collected that
year (sampling), and one collected but genuinely spending nothing in
that category (a real zero). FY2012 Georgia Police is the motivating
case from the issue: a complete census year reads as a 69% "response
rate" because most of the gap is cities that contract policing to the
county sheriff, not non-response.

Adds a second counter, n_units_collected: how many of the caller's
expected cohort appear in the corpus that year for ANY category.
n_units_collected / n_units_expected is the true collection rate;
n_units_reporting / n_units_collected is category participation among
collected units. cog_geographic_rollup() and cog_peer_compare() both
carry it; cog_explain() prints it alongside n_units_reporting.

Two real bugs caught and fixed while finishing this (both against the
already-written, previously-uncommitted draft):

- .coverage_table()'s candidate list for the collection query was
  derived from the category-filtered result rows, not the caller's
  full expected cohort. A government with zero rows in the requested
  category across every requested year never appears in that result,
  so it was silently excluded from n_units_collected too -- collapsing
  the new counter back to the old, broken one for exactly the
  governments it exists to count. Fixed by threading an explicit
  `expected_ids` (all_govids / peer_govids) through instead.
- The collection query hardcoded long_view = "spending_long_harmonized",
  which does not exist on a corpus with schema_version < 5 (R/basis.R
  resolves basis = "raw" there; R/views.R only registers the
  harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare()
  would hard-error on a corpus vintage the package otherwise explicitly
  supports. Fixed by deriving long_view from the basis cog_spending()
  actually resolved (prov$basis) via the existing .select_long_view()
  helper, matching how every other basis-aware query in the package
  already does this.

Also: cog_explain()'s general "complete census only in years ending in
2 or 7" footnote was gated on the OLD counter's absence, making it
permanently unreachable now that both callers always supply the new
one -- ungated it, since the explanation is orthogonal to which
counter set is present. Dropped a dead conditional branch, fixed two
stale roxygen blocks in R/peers.R/R/rollup.R still describing the old
two-counter model, fixed the same staleness in README.md, and switched
two `uscogdata:::` self-references to the package's own convention of
calling internal helpers unqualified.

1101 tests pass (2 skipped live-corpus), including new direct
regression tests for both bugs above (one exercising a government
collected-but-absent from a category result, one running the full
rollup/peer-compare path against a doctored schema_version 4 corpus).

Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW
-- all addressed above).

Closes #36.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
jared merged commit 9eaa759ccb into main 2026-09-09 11:51:06 -04:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#71