feat(coverage): n_units_collected separates sampling from real zeros (#36)
provenance$coverage's n_units_reporting is category-conditional: it counts governments with rows for the SPECIFIC requested category, which conflates two different things -- a government never collected that year (sampling), and one collected but genuinely spending nothing in that category (a real zero). FY2012 Georgia Police is the motivating case from the issue: a complete census year reads as a 69% "response rate" because most of the gap is cities that contract policing to the county sheriff, not non-response. Adds a second counter, n_units_collected: how many of the caller's expected cohort appear in the corpus that year for ANY category. n_units_collected / n_units_expected is the true collection rate; n_units_reporting / n_units_collected is category participation among collected units. cog_geographic_rollup() and cog_peer_compare() both carry it; cog_explain() prints it alongside n_units_reporting. Two real bugs caught and fixed while finishing this (both against the already-written, previously-uncommitted draft): - .coverage_table()'s candidate list for the collection query was derived from the category-filtered result rows, not the caller's full expected cohort. A government with zero rows in the requested category across every requested year never appears in that result, so it was silently excluded from n_units_collected too -- collapsing the new counter back to the old, broken one for exactly the governments it exists to count. Fixed by threading an explicit `expected_ids` (all_govids / peer_govids) through instead. - The collection query hardcoded long_view = "spending_long_harmonized", which does not exist on a corpus with schema_version < 5 (R/basis.R resolves basis = "raw" there; R/views.R only registers the harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare() would hard-error on a corpus vintage the package otherwise explicitly supports. Fixed by deriving long_view from the basis cog_spending() actually resolved (prov$basis) via the existing .select_long_view() helper, matching how every other basis-aware query in the package already does this. Also: cog_explain()'s general "complete census only in years ending in 2 or 7" footnote was gated on the OLD counter's absence, making it permanently unreachable now that both callers always supply the new one -- ungated it, since the explanation is orthogonal to which counter set is present. Dropped a dead conditional branch, fixed two stale roxygen blocks in R/peers.R/R/rollup.R still describing the old two-counter model, fixed the same staleness in README.md, and switched two `uscogdata:::` self-references to the package's own convention of calling internal helpers unqualified. 1101 tests pass (2 skipped live-corpus), including new direct regression tests for both bugs above (one exercising a government collected-but-absent from a category result, one running the full rollup/peer-compare path against a doctored schema_version 4 corpus). Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW -- all addressed above). Closes #36. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
+76
-6
@@ -84,22 +84,92 @@
|
||||
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
|
||||
#' missing year fails to make.
|
||||
#'
|
||||
#' `n_units_reporting` describes the result the caller actually received, so
|
||||
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
|
||||
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
|
||||
#' `n_units_reporting` is the number that tells the truth.
|
||||
#' Three counters are returned, each answering a different question:
|
||||
#'
|
||||
#' * `n_units_expected` -- the universe the caller named (govids passed in,
|
||||
#' or peers for cog_peer_compare). "How many governments did you ask
|
||||
#' about?"
|
||||
#' * `n_units_collected` -- how many of those appear in the corpus at all
|
||||
#' that year, in ANY category. This is a statement about survey collection,
|
||||
#' independent of what was asked for: "of the governments you named, how
|
||||
#' many did Census actually collect data from this year?" It separates
|
||||
#' sampling (not collected) from real zeros (collected but spends nothing
|
||||
#' in your category).
|
||||
#' * `n_units_reporting` -- how many of those appear with rows for the
|
||||
#' SPECIFIC category you requested. This is always <= n_units_collected:
|
||||
#' a government can be collected but have no rows for "Police" because it
|
||||
#' contracts policing to the county sheriff, not because it wasn't
|
||||
#' surveyed.
|
||||
#'
|
||||
#' `n_units_reporting` therefore conflates two very different things: a unit
|
||||
#' that was not collected (sampling) and a unit that was collected but spends
|
||||
#' nothing in that category. The ratio n_units_collected / n_units_expected is
|
||||
#' the true collection rate; n_units_reporting / n_units_collected measures
|
||||
#' category participation among collected units.
|
||||
#'
|
||||
#' `is_census_year` is a statement about the SURVEY CALENDAR, never a claim of
|
||||
#' completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
|
||||
#' cities report. The counters are what tell the truth.
|
||||
#'
|
||||
#' @param con Active DuckDB connection (used to look up n_units_collected).
|
||||
#' @param long_view The verb's own long view, used for the collection query;
|
||||
#' NULL skips the lookup and leaves n_units_collected as NA_integer_.
|
||||
#' @param expected_ids The full EXPECTED cohort (govids the caller named),
|
||||
#' used as the candidate list for the collection query. Required alongside
|
||||
#' `con`/`long_view` for a correct count -- see the note below on why it
|
||||
#' must not be derived from `result`/`rows`. `NULL`, or non-`NULL` but
|
||||
#' empty after dropping `NA`/`""` entries, skips the lookup and leaves
|
||||
#' n_units_collected as NA_integer_.
|
||||
#' @noRd
|
||||
.coverage_table <- function(result, years, n_expected,
|
||||
id_col = "canonical_govid", rows = NULL) {
|
||||
id_col = "canonical_govid", rows = NULL,
|
||||
con = NULL, long_view = NULL,
|
||||
expected_ids = NULL) {
|
||||
years <- sort(unique(as.integer(years)))
|
||||
src <- if (is.null(rows)) result else rows
|
||||
reporting <- vapply(years, function(y) {
|
||||
ids <- src[[id_col]][as.integer(src$year) == y]
|
||||
length(unique(ids[!is.na(ids)]))
|
||||
}, integer(1))
|
||||
|
||||
# n_units_collected: count EXPECTED cohort members present in the corpus
|
||||
# for ANY category that year, not just the requested one. This separates
|
||||
# sampling (not collected at all) from real zeros (collected but no rows
|
||||
# for this category). Only computed when a connection, long_view, AND
|
||||
# expected_ids are all provided; otherwise NA_integer_.
|
||||
#
|
||||
# The candidate list MUST be expected_ids, not derived from `result`/
|
||||
# `rows`: a government with zero rows in the requested category across
|
||||
# EVERY requested year never appears in `result` at all, so deriving
|
||||
# candidates from it would silently exclude exactly the "collected but
|
||||
# real zero" governments this counter exists to count -- collapsing
|
||||
# n_units_collected back to n_units_reporting for precisely the case #36
|
||||
# was filed over.
|
||||
if (!is.null(con) && !is.null(long_view) && length(expected_ids) > 0L) {
|
||||
cohort_chr <- .sql_lit_chr(unique(expected_ids[!is.na(expected_ids) &
|
||||
nzchar(expected_ids)]))
|
||||
years_lit <- paste(years, collapse = ",")
|
||||
collected_q <- sprintf(
|
||||
"SELECT year, COUNT(DISTINCT canonical_govid) AS n
|
||||
FROM %s
|
||||
WHERE canonical_govid IN (%s)
|
||||
AND year IN (%s)
|
||||
GROUP BY year",
|
||||
long_view, cohort_chr, years_lit
|
||||
)
|
||||
collected_df <- DBI::dbGetQuery(con, collected_q)
|
||||
collected_map <- setNames(collected_df$n, as.integer(collected_df$year))
|
||||
collected <- vapply(years, function(y) {
|
||||
val <- collected_map[as.character(y)]
|
||||
if (is.na(val)) 0L else as.integer(val)
|
||||
}, integer(1))
|
||||
} else {
|
||||
collected <- rep(NA_integer_, length(years))
|
||||
}
|
||||
|
||||
tibble::tibble(
|
||||
year = years,
|
||||
n_units_collected = collected,
|
||||
n_units_reporting = as.integer(reporting),
|
||||
n_units_expected = rep(as.integer(n_expected), length(years)),
|
||||
is_census_year = .is_census_year(years)
|
||||
|
||||
+24
-6
@@ -166,12 +166,30 @@ cog_explain <- function(result, format = c("print", "list")) {
|
||||
cli::cli_h2("Reporting coverage")
|
||||
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
|
||||
cov <- prov$coverage
|
||||
cli::cli_ul(sprintf(
|
||||
"%d: %d of %d units reporting (%.0f%%) -- %s year",
|
||||
cov$year, cov$n_units_reporting, cov$n_units_expected,
|
||||
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
|
||||
ifelse(cov$is_census_year, "census", "sample")
|
||||
))
|
||||
has_collected <- "n_units_collected" %in% names(cov)
|
||||
if (has_collected) {
|
||||
# Three counters: collected separates sampling from real zeros;
|
||||
# reporting is category-conditional and never a response rate.
|
||||
cli::cli_ul(sprintf(
|
||||
"%d: %d of %d units collected, %d reporting in this category -- %s year",
|
||||
cov$year,
|
||||
cov$n_units_collected,
|
||||
cov$n_units_expected,
|
||||
cov$n_units_reporting,
|
||||
ifelse(cov$is_census_year, "census", "sample")
|
||||
))
|
||||
} else {
|
||||
cli::cli_ul(sprintf(
|
||||
"%d: %d of %d units reporting (%.0f%%) -- %s year",
|
||||
cov$year, cov$n_units_reporting, cov$n_units_expected,
|
||||
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
|
||||
ifelse(cov$is_census_year, "census", "sample")
|
||||
))
|
||||
}
|
||||
# Explains what the per-row "-- sample year" tag means, regardless of
|
||||
# which branch above rendered it -- not gated on has_collected, which
|
||||
# would make this permanently unreachable now that both real callers
|
||||
# (cog_geographic_rollup(), cog_peer_compare()) always supply it.
|
||||
if (any(!cov$is_census_year)) {
|
||||
cli::cli_text(
|
||||
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
|
||||
|
||||
@@ -195,11 +195,13 @@ cog_find_peers <- function(target_govid,
|
||||
#' a balanced panel.
|
||||
#'
|
||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
#' statement about the **survey calendar**, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
#' report. `n_units_reporting` is the number that tells the truth.
|
||||
#' `n_units_expected`, `n_units_collected`, `n_units_reporting` and
|
||||
#' `is_census_year`, and `provenance$coverage_mode` records the mode.
|
||||
#' `is_census_year` is a statement about the **survey calendar**, never a
|
||||
#' claim of completeness: FY1967 is a census year in which only 97 of
|
||||
#' Wisconsin's 608 cities report. `n_units_reporting` is
|
||||
#' category-conditional and is not a response rate on its own -- see
|
||||
#' "Reading `coverage`" below for what each counter answers.
|
||||
#'
|
||||
#' The comparison target is exempt from `"consistent"` balancing -- it is the
|
||||
#' subject of the comparison, not a member of the cohort -- and the
|
||||
@@ -241,14 +243,23 @@ cog_find_peers <- function(target_govid,
|
||||
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||
#' ```
|
||||
#' @section Reading `coverage`:
|
||||
#' `provenance$coverage` reports `n_units_reporting` against
|
||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
#' it counts cohort members with rows for the category you asked for, not
|
||||
#' cohort members collected that year.** A government that was surveyed and
|
||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
||||
#' that was never surveyed.
|
||||
#' `provenance$coverage` carries three per-year counters:
|
||||
#'
|
||||
#' * `n_units_expected` -- how many governments you asked about.
|
||||
#' * `n_units_collected` -- how many of those appear in the corpus at all
|
||||
#' that year (in ANY category), separating sampling from real zeros.
|
||||
#' * `n_units_reporting` -- how many have rows for the SPECIFIC category you
|
||||
#' requested. This is always <= n_units_collected: a government can be
|
||||
#' collected but have no rows for "Police" because it contracts policing
|
||||
#' to the county sheriff, not because it wasn't surveyed.
|
||||
#'
|
||||
#' **`n_units_reporting` is category-conditional** and therefore **not a
|
||||
#' response rate**: `n_units_reporting / n_units_expected` conflates sampling
|
||||
#' (never collected) with real zeros (collected but spends nothing in your
|
||||
#' category). Use `n_units_collected / n_units_expected` for the true
|
||||
#' collection rate, and `n_units_reporting / n_units_collected` for category
|
||||
#' participation among collected units.
|
||||
#'
|
||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
#' contract policing to the county sheriff, not non-response.
|
||||
@@ -325,10 +336,20 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
||||
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
|
||||
# reported in FY2019". Including the target would inflate every count by one
|
||||
# and make a cohort that has entirely stopped reporting look non-empty.
|
||||
# n_units_collected is looked up against the spending long view matching
|
||||
# whatever basis cog_spending() actually resolved above (prov$basis) --
|
||||
# NOT hardcoded to spending_long_harmonized, which does not exist on a
|
||||
# corpus with schema_version < 5 (R/basis.R resolves basis = "raw" there,
|
||||
# and only *_long, not *_long_harmonized, is registered; see R/views.R).
|
||||
# The con comes from .ensure_session() already called inside cog_spending().
|
||||
con <- .ensure_session()
|
||||
prov$coverage_mode <- coverage
|
||||
prov$coverage <- .coverage_table(
|
||||
out, years, length(peer_govids),
|
||||
rows = r[r$role == "peer", , drop = FALSE]
|
||||
rows = r[r$role == "peer", , drop = FALSE],
|
||||
con = con,
|
||||
long_view = .select_long_view("spending_annotated", prov$basis),
|
||||
expected_ids = peer_govids
|
||||
)
|
||||
attr(out, "provenance") <- prov
|
||||
out
|
||||
|
||||
+36
-13
@@ -49,11 +49,13 @@
|
||||
#' a balanced panel.
|
||||
#'
|
||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
#' statement about the **survey calendar**, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
#' report. `n_units_reporting` is the number that tells the truth.
|
||||
#' `n_units_expected`, `n_units_collected`, `n_units_reporting` and
|
||||
#' `is_census_year`, and `provenance$coverage_mode` records the mode.
|
||||
#' `is_census_year` is a statement about the **survey calendar**, never a
|
||||
#' claim of completeness: FY1967 is a census year in which only 97 of
|
||||
#' Wisconsin's 608 cities report. `n_units_reporting` is
|
||||
#' category-conditional and is not a response rate on its own -- see
|
||||
#' "Reading `coverage`" below for what each counter answers.
|
||||
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
||||
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
||||
@@ -61,14 +63,23 @@
|
||||
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
|
||||
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
||||
#' @section Reading `coverage`:
|
||||
#' `provenance$coverage` reports `n_units_reporting` against
|
||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
#' it counts governments with rows for the category you asked for, not
|
||||
#' governments collected that year.** A government that was surveyed and
|
||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
||||
#' that was never surveyed.
|
||||
#' `provenance$coverage` carries three per-year counters:
|
||||
#'
|
||||
#' * `n_units_expected` -- how many governments you asked about.
|
||||
#' * `n_units_collected` -- how many of those appear in the corpus at all
|
||||
#' that year (in ANY category), separating sampling from real zeros.
|
||||
#' * `n_units_reporting` -- how many have rows for the SPECIFIC category you
|
||||
#' requested. This is always <= n_units_collected: a government can be
|
||||
#' collected but have no rows for "Police" because it contracts policing
|
||||
#' to the county sheriff, not because it wasn't surveyed.
|
||||
#'
|
||||
#' **`n_units_reporting` is category-conditional** and therefore **not a
|
||||
#' response rate**: `n_units_reporting / n_units_expected` conflates sampling
|
||||
#' (never collected) with real zeros (collected but spends nothing in your
|
||||
#' category). Use `n_units_collected / n_units_expected` for the true
|
||||
#' collection rate, and `n_units_reporting / n_units_collected` for category
|
||||
#' participation among collected units.
|
||||
#'
|
||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
#' contract policing to the county sheriff, not non-response.
|
||||
@@ -137,8 +148,20 @@ cog_geographic_rollup <- function(govids, category, years,
|
||||
# n_units_expected is the universe the CALLER named -- the govids passed in
|
||||
# -- not the national universe. That is what makes the ratio meaningful:
|
||||
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
|
||||
# n_units_collected is looked up against the spending long view matching
|
||||
# whatever basis cog_spending() actually resolved above (prov$basis) --
|
||||
# NOT hardcoded to spending_long_harmonized, which does not exist on a
|
||||
# corpus with schema_version < 5 (R/basis.R resolves basis = "raw" there,
|
||||
# and only *_long, not *_long_harmonized, is registered; see R/views.R).
|
||||
# The con comes from .ensure_session() already called inside cog_spending().
|
||||
con <- .ensure_session()
|
||||
prov$coverage_mode <- coverage
|
||||
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
|
||||
prov$coverage <- .coverage_table(
|
||||
r, years, length(unique(all_govids)),
|
||||
con = con,
|
||||
long_view = .select_long_view("spending_annotated", prov$basis),
|
||||
expected_ids = all_govids
|
||||
)
|
||||
attr(r, "provenance") <- prov
|
||||
|
||||
r
|
||||
|
||||
Reference in New Issue
Block a user