Multi-government aggregates disclose no reporting coverage: add the coverage argument and always-on coverage metadata #13

Closed
opened 2026-07-29 00:05:28 -04:00 by jared · 0 comments
Owner

Filed from the Madison walkthrough audit (cog_explorer/docs/walkthroughs/FINDINGS.md, 2026-07-28). Verdict: definitional — the verbs do exactly what "sum/match what's there" should do; the gap is that the return value never says what "there" was.

Root cause

The Census of Governments is a complete census only in years ending in 2 and 7. In every other year it is a sample, and the sample size varies enormously. Neither cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() has any concept of "the universe": each sums or labels whichever canonical_govids happen to have rows in the requested years and returns that with no column, flag, or provenance entry distinguishing "every government reported" from "5% of governments reported."

Both findings are one problem seen from two verbs, and the owner has settled one design that resolves both.

Findings resolved

Finding Severity Summary
F-020 high cog_geographic_rollup() (and any multi-government aggregate built the same way) silently sums whichever governments reported that year, with no signal that the reporting universe varies 5.1%-99.5% year to year
F-023 high cog_find_peers() fixes peer-cohort membership at a single cohort_year, but nothing in cog_peer_compare()'s return says how many of those peers actually reported in the years requested — for a small city's cohort, that can collapse to 3 of 15

Reproduction (verbatim from FINDINGS.md, verified against the live corpus)

F-020

wi_cities <- cog_gov_search(name = NULL, state = "WI", type = "city")
nrow(wi_cities)   #> 608

wi_roll <- cog_geographic_rollup(
  govids = list(city = wi_cities$canonical_govid), category = NULL,
  years = 1967:2023
)
wi_year <- wi_roll |> dplyr::group_by(year) |>
  dplyr::summarise(wi_total = sum(amt_nominal, na.rm = TRUE),
                   n_govs = dplyr::n_distinct(canonical_govid))
wi_year |> dplyr::filter(year %in% c(1972, 1982, 1983, 2001, 2002, 2003, 2022))
#> 1972 (census): 576 govs, $955,736,000
#> 1982 (census): 586 govs, $2,472,492,000
#> 1983 (sample): 211 govs, $1,806,167,000
#> 2001 (sample):  31 govs, $2,148,955,000
#> 2002 (census): 591 govs, $5,155,306,000
#> 2003 (sample):  31 govs, $2,334,711,000
#> 2022 (census): 605 govs, $9,059,909,000

# Cross-checked against the RAW `long` table directly, NOT through
# cog_geographic_rollup(), which is the function under test:
con <- uscogdata:::.ensure_session()
DBI::dbGetQuery(con, "
  SELECT year, COUNT(DISTINCT canonical_govid) n_govs_raw, SUM(amt)*1000 total_raw
  FROM long WHERE type = 2 AND fips_state = 55
    AND LEFT(item_code,1) IN ('E','F','G') AND NOT is_aggregate
    AND year IN (1982,1983,2001,2002,2003,2022) GROUP BY year ORDER BY year
")
#> 1982: 580 govs, $2,466,880,000   1983: 209 govs, $1,801,544,000
#> 2001:  30 govs, $2,143,334,000   2002: 585 govs, $5,122,876,000
#> 2003:  30 govs, $2,329,492,000   2022: 605 govs, $9,059,909,000 (exact match)

mad_share <- wi_year |> dplyr::left_join(mad_year, by = "year") |>
  dplyr::mutate(share = mad_total / wi_total)
cor(mad_share$share, mad_share$n_govs)          #> -0.699 (p = 2.9e-9)
cor(mad_share$share, log(mad_share$n_govs))     #> -0.796 (p = 3.8e-13)

F-023

# Madison (large city): cohort fixed at FY2023 population, requested 1997-2023
peers_mad <- cog_find_peers("552025209777", year = 2023L, max_peers = 15)
attr(peers_mad, "cohort_year")                          #> 2023
mad_cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers_mad,
                            category = NULL, years = 1997:2023,
                            per_capita = TRUE, adjust_to_year = 2023L)
mad_cmp |> dplyr::filter(role == "peer") |>
  dplyr::group_by(year) |>
  dplyr::summarise(n = dplyr::n_distinct(canonical_govid)) |>
  dplyr::pull(n) |> range()
#> 15 15   -- all 15 peers report in every one of the 27 years, 1997-2023

# A small Wisconsin city (CHILTON CITY, 552015177095, ACS pop 4,017):
peers_chilton <- cog_find_peers("552015177095", year = 2023L, max_peers = 15)
nrow(peers_chilton)                                      #> 15
chilton_cmp <- cog_peer_compare(target_govid = "552015177095",
                                peers = peers_chilton, category = NULL,
                                years = 1997:2023, per_capita = TRUE,
                                adjust_to_year = 2023L)
by_year <- chilton_cmp |> dplyr::filter(role == "peer") |>
  dplyr::group_by(year) |>
  dplyr::summarise(n = dplyr::n_distinct(canonical_govid))
by_year |> dplyr::filter(year %in% c(1997,2001,2002,2003,2007,2012,2017,2022))
#> 1997:15  2001: 3  2002:15  2003: 3  2007:15  2012:15  2017:15  2022:15
range(by_year$n)                                          #> 3 15

# Cross-checked against the RAW `long` table directly, NOT through
# cog_peer_compare() (the function under test):
#> 2001: 3   2002: 15   2003: 3   -- exact match; not a cog_spending()-side
#>                                   filtering artifact

Both reproduce on the bundled fixture corpus (2011/2012/2019/2020), which is what the test asserts against: Wisconsin's 608-city universe rolls up 597 governments in FY2012 (a census year) but only 152 / 112 / 114 in FY2011 / FY2019 / FY2020; and Chilton's 15-peer cohort taken at FY2012 reports 15 of 15 in FY2012 and 3 of 15 in FY2019 and FY2020.

Why it matters

A caller who does not independently know the Census of Governments survey calendar has no way to learn from the return value alone that a given year's "statewide total" rests on a fraction of the actual governments. 43 of 55 years (78%) are sample years. The cleanest illustration: FY2002 (census, 591 govs, Madison's share 5.89%) to FY2003 (sample, 31 govs) — Madison's own total is essentially flat (+0.6%) while its apparent share more than doubles to 13.1% on the denominator collapse alone. The peer side is arguably worse exposure: the reassurance a Madison-scale user gets from a stable-looking peer band is not a property of cog_peer_compare() — it is a property of Madison being large. Governments matched to a small target sit in exactly the population band most exposed to the sample cycle, and building a peer comparison for one's own small city is, if anything, a more common use of this package than a full geographic rollup.

The machinery to say so already half-exists: cog_geographic_rollup()'s provenance records excluded_govids for the population case — a far smaller effect than the 5%-to-99% coverage swing that has no analog at all.

The agreed design (settled 2026-07-28 by the project owner — not open for re-litigation)

A coverage argument on cog_geographic_rollup(), cog_peer_compare()/cog_find_peers(), and their cog-api equivalents:

  • coverage = "all" — every unit that reported that year. Today's behaviour, and the default, kept for backward compatibility so nothing currently calling these verbs breaks.
  • coverage = "census" — restrict to census years only (years ending in 2 or 7).
  • coverage = "consistent" — restrict to units reporting in every requested year, producing a balanced panel.

Independent of which mode is chosen, every result carries always-on coverage metadata — n_units_reporting, n_units_expected, is_census_year — so that even the default "all" mode can no longer mislead silently.

Motivating principle, stated by the owner: using these verbs correctly must not require the user to know that the Census of Governments is a complete census only in years ending in 2 and 7 — that fact about survey design belongs in the tooling, not in the analyst's head.

Definition of done

  1. coverage = c("all", "census", "consistent") implemented on cog_geographic_rollup(), cog_find_peers(), and cog_peer_compare(), defaulting to "all".
  2. Every result from those verbs carries per-year n_units_reporting, n_units_expected, and is_census_year, regardless of mode — reachable programmatically, not only via cog_explain().
  3. Note that CENSUS_YEARS alone is not a fully reliable proxy for complete coverage: FY1967 reports only 97 of Wisconsin's 608 type-2 governments (16.0%), a nationwide pattern for that vintage (type-2 nationwide: 3,765 reporting in 1967 vs. 18,517 in 1972). is_census_year must be a statement about the survey calendar; n_units_reporting is what actually tells the truth.
  4. Test goes green: tests/testthat/test-coverage-disclosure.R → test_that("multi-government aggregates disclose reporting coverage on every result", ...). Against the fixture it asserts a Wisconsin all-cities rollup reports n_units_expected == 608 in every year with n_units_reporting of 597 (FY2012, is_census_year == TRUE) and 112 (FY2019, is_census_year == FALSE); that Chilton's FY2012 15-peer cohort reports n_units_reporting == 3 for FY2019; and that coverage = "consistent" returns only units present in every requested year. Remove the skip() on line 1 of the test body to activate.

Cross-reference

cog-api carries the same gap on its own surface, with an extra defect: provenance.scope.govids_missing — the field that looks built to answer exactly this — reports [] for a cohort member that contributed no rows. Tracked there as finding F-032.

Severity: high. Verdict: definitional.

Filed from the **Madison walkthrough audit** (`cog_explorer/docs/walkthroughs/FINDINGS.md`, 2026-07-28). Verdict: **definitional** — the verbs do exactly what "sum/match what's there" should do; the gap is that the return value never says what "there" was. ## Root cause The Census of Governments is a **complete census only in years ending in 2 and 7**. In every other year it is a sample, and the sample size varies enormously. Neither `cog_geographic_rollup()` nor `cog_peer_compare()`/`cog_find_peers()` has any concept of "the universe": each sums or labels whichever `canonical_govid`s happen to have rows in the requested years and returns that with no column, flag, or `provenance` entry distinguishing "every government reported" from "5% of governments reported." Both findings are one problem seen from two verbs, and the owner has settled one design that resolves both. ## Findings resolved | Finding | Severity | Summary | |---|---|---| | **F-020** | high | `cog_geographic_rollup()` (and any multi-government aggregate built the same way) silently sums whichever governments reported that year, with no signal that the reporting universe varies 5.1%-99.5% year to year | | **F-023** | high | `cog_find_peers()` fixes peer-cohort *membership* at a single `cohort_year`, but nothing in `cog_peer_compare()`'s return says how many of those peers actually reported in the years requested — for a small city's cohort, that can collapse to 3 of 15 | ## Reproduction (verbatim from FINDINGS.md, verified against the live corpus) **F-020** ```r wi_cities <- cog_gov_search(name = NULL, state = "WI", type = "city") nrow(wi_cities) #> 608 wi_roll <- cog_geographic_rollup( govids = list(city = wi_cities$canonical_govid), category = NULL, years = 1967:2023 ) wi_year <- wi_roll |> dplyr::group_by(year) |> dplyr::summarise(wi_total = sum(amt_nominal, na.rm = TRUE), n_govs = dplyr::n_distinct(canonical_govid)) wi_year |> dplyr::filter(year %in% c(1972, 1982, 1983, 2001, 2002, 2003, 2022)) #> 1972 (census): 576 govs, $955,736,000 #> 1982 (census): 586 govs, $2,472,492,000 #> 1983 (sample): 211 govs, $1,806,167,000 #> 2001 (sample): 31 govs, $2,148,955,000 #> 2002 (census): 591 govs, $5,155,306,000 #> 2003 (sample): 31 govs, $2,334,711,000 #> 2022 (census): 605 govs, $9,059,909,000 # Cross-checked against the RAW `long` table directly, NOT through # cog_geographic_rollup(), which is the function under test: con <- uscogdata:::.ensure_session() DBI::dbGetQuery(con, " SELECT year, COUNT(DISTINCT canonical_govid) n_govs_raw, SUM(amt)*1000 total_raw FROM long WHERE type = 2 AND fips_state = 55 AND LEFT(item_code,1) IN ('E','F','G') AND NOT is_aggregate AND year IN (1982,1983,2001,2002,2003,2022) GROUP BY year ORDER BY year ") #> 1982: 580 govs, $2,466,880,000 1983: 209 govs, $1,801,544,000 #> 2001: 30 govs, $2,143,334,000 2002: 585 govs, $5,122,876,000 #> 2003: 30 govs, $2,329,492,000 2022: 605 govs, $9,059,909,000 (exact match) mad_share <- wi_year |> dplyr::left_join(mad_year, by = "year") |> dplyr::mutate(share = mad_total / wi_total) cor(mad_share$share, mad_share$n_govs) #> -0.699 (p = 2.9e-9) cor(mad_share$share, log(mad_share$n_govs)) #> -0.796 (p = 3.8e-13) ``` **F-023** ```r # Madison (large city): cohort fixed at FY2023 population, requested 1997-2023 peers_mad <- cog_find_peers("552025209777", year = 2023L, max_peers = 15) attr(peers_mad, "cohort_year") #> 2023 mad_cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers_mad, category = NULL, years = 1997:2023, per_capita = TRUE, adjust_to_year = 2023L) mad_cmp |> dplyr::filter(role == "peer") |> dplyr::group_by(year) |> dplyr::summarise(n = dplyr::n_distinct(canonical_govid)) |> dplyr::pull(n) |> range() #> 15 15 -- all 15 peers report in every one of the 27 years, 1997-2023 # A small Wisconsin city (CHILTON CITY, 552015177095, ACS pop 4,017): peers_chilton <- cog_find_peers("552015177095", year = 2023L, max_peers = 15) nrow(peers_chilton) #> 15 chilton_cmp <- cog_peer_compare(target_govid = "552015177095", peers = peers_chilton, category = NULL, years = 1997:2023, per_capita = TRUE, adjust_to_year = 2023L) by_year <- chilton_cmp |> dplyr::filter(role == "peer") |> dplyr::group_by(year) |> dplyr::summarise(n = dplyr::n_distinct(canonical_govid)) by_year |> dplyr::filter(year %in% c(1997,2001,2002,2003,2007,2012,2017,2022)) #> 1997:15 2001: 3 2002:15 2003: 3 2007:15 2012:15 2017:15 2022:15 range(by_year$n) #> 3 15 # Cross-checked against the RAW `long` table directly, NOT through # cog_peer_compare() (the function under test): #> 2001: 3 2002: 15 2003: 3 -- exact match; not a cog_spending()-side #> filtering artifact ``` Both reproduce on the **bundled fixture corpus** (2011/2012/2019/2020), which is what the test asserts against: Wisconsin's 608-city universe rolls up 597 governments in FY2012 (a census year) but only 152 / 112 / 114 in FY2011 / FY2019 / FY2020; and Chilton's 15-peer cohort taken at FY2012 reports 15 of 15 in FY2012 and **3 of 15** in FY2019 and FY2020. ## Why it matters A caller who does not independently know the Census of Governments survey calendar has no way to learn from the return value alone that a given year's "statewide total" rests on a fraction of the actual governments. 43 of 55 years (78%) are sample years. The cleanest illustration: FY2002 (census, 591 govs, Madison's share 5.89%) to FY2003 (sample, 31 govs) — Madison's own total is essentially flat (+0.6%) while its apparent share **more than doubles to 13.1%** on the denominator collapse alone. The peer side is arguably worse exposure: the reassurance a Madison-scale user gets from a stable-looking peer band is not a property of `cog_peer_compare()` — it is a property of Madison being large. Governments matched to a small target sit in exactly the population band most exposed to the sample cycle, and building a peer comparison for one's own small city is, if anything, a more common use of this package than a full geographic rollup. The machinery to say so already half-exists: `cog_geographic_rollup()`'s provenance records `excluded_govids` for the population case — a far smaller effect than the 5%-to-99% coverage swing that has no analog at all. ## The agreed design (settled 2026-07-28 by the project owner — not open for re-litigation) A `coverage` argument on `cog_geographic_rollup()`, `cog_peer_compare()`/`cog_find_peers()`, and their `cog-api` equivalents: * **`coverage = "all"`** — every unit that reported that year. Today's behaviour, and the **default**, kept for backward compatibility so nothing currently calling these verbs breaks. * **`coverage = "census"`** — restrict to census years only (years ending in 2 or 7). * **`coverage = "consistent"`** — restrict to units reporting in *every* requested year, producing a balanced panel. Independent of which mode is chosen, every result carries **always-on coverage metadata** — `n_units_reporting`, `n_units_expected`, `is_census_year` — so that even the default `"all"` mode can no longer mislead silently. Motivating principle, stated by the owner: **using these verbs correctly must not require the user to know that the Census of Governments is a complete census only in years ending in 2 and 7** — that fact about survey design belongs in the tooling, not in the analyst's head. ## Definition of done 1. `coverage = c("all", "census", "consistent")` implemented on `cog_geographic_rollup()`, `cog_find_peers()`, and `cog_peer_compare()`, defaulting to `"all"`. 2. Every result from those verbs carries per-year `n_units_reporting`, `n_units_expected`, and `is_census_year`, regardless of mode — reachable programmatically, not only via `cog_explain()`. 3. Note that `CENSUS_YEARS` alone is **not** a fully reliable proxy for complete coverage: FY1967 reports only 97 of Wisconsin's 608 type-2 governments (16.0%), a nationwide pattern for that vintage (type-2 nationwide: 3,765 reporting in 1967 vs. 18,517 in 1972). `is_census_year` must be a statement about the survey calendar; `n_units_reporting` is what actually tells the truth. 4. Test goes green: **`tests/testthat/test-coverage-disclosure.R`** → `test_that("multi-government aggregates disclose reporting coverage on every result", ...)`. Against the fixture it asserts a Wisconsin all-cities rollup reports `n_units_expected == 608` in every year with `n_units_reporting` of **597** (FY2012, `is_census_year == TRUE`) and **112** (FY2019, `is_census_year == FALSE`); that Chilton's FY2012 15-peer cohort reports `n_units_reporting == 3` for FY2019; and that `coverage = "consistent"` returns only units present in every requested year. Remove the `skip()` on line 1 of the test body to activate. ## Cross-reference **`cog-api`** carries the same gap on its own surface, with an extra defect: `provenance.scope.govids_missing` — the field that looks built to answer exactly this — reports `[]` for a cohort member that contributed no rows. Tracked there as finding F-032. **Severity: high. Verdict: definitional.**
jared added the severity/highverdict/definitionalmadison-walkthrough labels 2026-07-29 00:05:28 -04:00
jared closed this issue 2026-07-30 12:06:53 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#13