Files
uscogdata/man/cog_peer_compare.Rd
jaredandClaude Sonnet 5 b41d5ee2aa
R-CMD-check / check (push) Successful in 3m39s
R-CMD-check / check (pull_request) Successful in 3m42s
feat(coverage): n_units_collected separates sampling from real zeros (#36)
provenance$coverage's n_units_reporting is category-conditional: it
counts governments with rows for the SPECIFIC requested category, which
conflates two different things -- a government never collected that
year (sampling), and one collected but genuinely spending nothing in
that category (a real zero). FY2012 Georgia Police is the motivating
case from the issue: a complete census year reads as a 69% "response
rate" because most of the gap is cities that contract policing to the
county sheriff, not non-response.

Adds a second counter, n_units_collected: how many of the caller's
expected cohort appear in the corpus that year for ANY category.
n_units_collected / n_units_expected is the true collection rate;
n_units_reporting / n_units_collected is category participation among
collected units. cog_geographic_rollup() and cog_peer_compare() both
carry it; cog_explain() prints it alongside n_units_reporting.

Two real bugs caught and fixed while finishing this (both against the
already-written, previously-uncommitted draft):

- .coverage_table()'s candidate list for the collection query was
  derived from the category-filtered result rows, not the caller's
  full expected cohort. A government with zero rows in the requested
  category across every requested year never appears in that result,
  so it was silently excluded from n_units_collected too -- collapsing
  the new counter back to the old, broken one for exactly the
  governments it exists to count. Fixed by threading an explicit
  `expected_ids` (all_govids / peer_govids) through instead.
- The collection query hardcoded long_view = "spending_long_harmonized",
  which does not exist on a corpus with schema_version < 5 (R/basis.R
  resolves basis = "raw" there; R/views.R only registers the
  harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare()
  would hard-error on a corpus vintage the package otherwise explicitly
  supports. Fixed by deriving long_view from the basis cog_spending()
  actually resolved (prov$basis) via the existing .select_long_view()
  helper, matching how every other basis-aware query in the package
  already does this.

Also: cog_explain()'s general "complete census only in years ending in
2 or 7" footnote was gated on the OLD counter's absence, making it
permanently unreachable now that both callers always supply the new
one -- ungated it, since the explanation is orthogonal to which
counter set is present. Dropped a dead conditional branch, fixed two
stale roxygen blocks in R/peers.R/R/rollup.R still describing the old
two-counter model, fixed the same staleness in README.md, and switched
two `uscogdata:::` self-references to the package's own convention of
calling internal helpers unqualified.

1101 tests pass (2 skipped live-corpus), including new direct
regression tests for both bugs above (one exercising a government
collected-but-absent from a category result, one running the full
rollup/peer-compare path against a doctored schema_version 4 corpus).

Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW
-- all addressed above).

Closes #36.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:46:17 -04:00

141 lines
6.0 KiB
R
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/peers.R
\name{cog_peer_compare}
\alias{cog_peer_compare}
\title{Compare a target government against a peer set}
\usage{
cog_peer_compare(
target_govid,
peers,
category,
years,
per_capita = TRUE,
adjust_to_year = NULL,
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
)
}
\arguments{
\item{target_govid}{Character scalar.}
\item{peers}{A tibble from [cog_find_peers()] or a character vector of
`canonical_govid`s.}
\item{category}{Character scalar or vector.}
\item{years}{Integer vector.}
\item{per_capita}{Default `TRUE` — peer compare usually normalizes by
population.}
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across peer sets counts
intergovernmental transfers twice; `"primary"` and `"direct"` combine
safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_expected`, `n_units_collected`, `n_units_reporting` and
`is_census_year`, and `provenance$coverage_mode` records the mode.
`is_census_year` is a statement about the **survey calendar**, never a
claim of completeness: FY1967 is a census year in which only 97 of
Wisconsin's 608 cities report. `n_units_reporting` is
category-conditional and is not a response rate on its own -- see
"Reading `coverage`" below for what each counter answers.
The comparison target is exempt from `"consistent"` balancing -- it is the
subject of the comparison, not a member of the cohort -- and the
`summary_*` quantiles are computed AFTER the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against
the cohort size: "3 of your 15 peers reported in FY2019".}
}
\value{
Tibble matching [cog_spending()]'s columns, plus a `role`
column taking values `"target"`, `"peer"`, `"summary_p25"`,
`"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
among target+peers at `max(years)`, NA for other rows), and
`cohort_year` (the year used to build the peer cohort, read from
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
`cohort_year`, and `cohort_govids`.
**The `summary_*` rows are per-category quantiles: they are not additive.**
Each one is computed **within each `(year, spend_subtype,
category)` cell** across the peer set, so a `summary_p50` row is *the
median peer's value in that one category*, not *the value of the median
peer's total*. The median peer for Police and the median peer for Fire
are usually different governments, so summing `summary_*` rows across
categories does not give any peer's total and misstates the band it
appears to describe — measured at −32.7% to +251.0% across 24 years on
one cohort, with a sign flip at FY2012.
Facet by `role` **and** `category` (the documented use, and what the
rows are built for). For a genuine "median peer's total spending" line,
sum each peer's own categories first and take the quantile of those
per-government totals:
```r
library(dplyr)
cmp |>
filter(role %in% c("target", "peer")) |>
group_by(year, role, canonical_govid) |>
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
filter(role == "peer") |>
group_by(year) |>
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
```
}
\description{
Pulls spending for the target plus a peer set (either a
[cog_find_peers()] result or a character vector of `canonical_govid`) and
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
`summary_p75`) so the result can be faceted by `role` in a single ggplot
call. Those summary rows are quantiles **within each category**, not
quantiles of each peer's total — see the `@return` section before summing
them.
}
\section{Reading `coverage`}{
`provenance$coverage` carries three per-year counters:
* `n_units_expected` -- how many governments you asked about.
* `n_units_collected` -- how many of those appear in the corpus at all
that year (in ANY category), separating sampling from real zeros.
* `n_units_reporting` -- how many have rows for the SPECIFIC category you
requested. This is always <= n_units_collected: a government can be
collected but have no rows for "Police" because it contracts policing
to the county sheriff, not because it wasn't surveyed.
**`n_units_reporting` is category-conditional** and therefore **not a
response rate**: `n_units_reporting / n_units_expected` conflates sampling
(never collected) with real zeros (collected but spends nothing in your
category). Use `n_units_collected / n_units_expected` for the true
collection rate, and `n_units_reporting / n_units_collected` for category
participation among collected units.
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
`category = "Police"`; the 174-city gap is overwhelmingly cities that
contract policing to the county sheriff, not non-response.
The comparison that *is* valid is the same category across a census year
(ending in 2 or 7) and a sample year, where the real-zero component is
roughly constant and the difference reflects the survey cycle. `is_census_year`
marks which is which.
}