Closes uscogdata#36. It counts governments with rows for the requested category, so a surveyed government that genuinely spends nothing there is indistinguishable from one never surveyed. In FY2022, a complete census year, Georgia reports 393 of 567 cities for Police -- the gap is cities that contract to the sheriff. Documents the comparison that IS valid: same category, census year vs sample year.
130 lines
5.5 KiB
R
130 lines
5.5 KiB
R
% Generated by roxygen2: do not edit by hand
|
||
% Please edit documentation in R/peers.R
|
||
\name{cog_peer_compare}
|
||
\alias{cog_peer_compare}
|
||
\title{Compare a target government against a peer set}
|
||
\usage{
|
||
cog_peer_compare(
|
||
target_govid,
|
||
peers,
|
||
category,
|
||
years,
|
||
per_capita = TRUE,
|
||
adjust_to_year = NULL,
|
||
expenditure_concept = c("primary", "direct", "total"),
|
||
coverage = c("all", "census", "consistent")
|
||
)
|
||
}
|
||
\arguments{
|
||
\item{target_govid}{Character scalar.}
|
||
|
||
\item{peers}{A tibble from [cog_find_peers()] or a character vector of
|
||
`canonical_govid`s.}
|
||
|
||
\item{category}{Character scalar or vector.}
|
||
|
||
\item{years}{Integer vector.}
|
||
|
||
\item{per_capita}{Default `TRUE` — peer compare usually normalizes by
|
||
population.}
|
||
|
||
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
|
||
|
||
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
|
||
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
||
refused here because combining Total across peer sets counts
|
||
intergovernmental transfers twice; `"primary"` and `"direct"` combine
|
||
safely.}
|
||
|
||
\item{coverage}{How to handle the Census of Governments survey cycle,
|
||
which is a **complete census only in years ending in 2 and 7** -- every
|
||
other year is a sample, and the sample varies enormously (on the bundled
|
||
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||
and 112 in FY2019).
|
||
|
||
* `"all"` (default) -- every unit that reported that year. Unchanged
|
||
behaviour, so existing code keeps working.
|
||
* `"census"` -- census years only. Aborts if the requested range holds
|
||
none, rather than silently returning nothing.
|
||
* `"consistent"` -- only units reporting in *every* requested year, giving
|
||
a balanced panel.
|
||
|
||
Regardless of mode, `provenance$coverage` always carries per-year
|
||
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
||
statement about the **survey calendar**, never a claim of completeness:
|
||
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||
report. `n_units_reporting` is the number that tells the truth.
|
||
|
||
The comparison target is exempt from `"consistent"` balancing -- it is the
|
||
subject of the comparison, not a member of the cohort -- and the
|
||
`summary_*` quantiles are computed AFTER the filter, so they describe the
|
||
cohort actually returned. `n_units_reporting` counts peers only, against
|
||
the cohort size: "3 of your 15 peers reported in FY2019".}
|
||
}
|
||
\value{
|
||
Tibble matching [cog_spending()]'s columns, plus a `role`
|
||
column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
||
`"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
||
among target+peers at `max(years)`, NA for other rows), and
|
||
`cohort_year` (the year used to build the peer cohort, read from
|
||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||
`cohort_year`, and `cohort_govids`.
|
||
|
||
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
||
Each one is computed **within each `(year, spend_subtype,
|
||
category)` cell** across the peer set, so a `summary_p50` row is *the
|
||
median peer's value in that one category*, not *the value of the median
|
||
peer's total*. The median peer for Police and the median peer for Fire
|
||
are usually different governments, so summing `summary_*` rows across
|
||
categories does not give any peer's total and misstates the band it
|
||
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||
one cohort, with a sign flip at FY2012.
|
||
|
||
Facet by `role` **and** `category` (the documented use, and what the
|
||
rows are built for). For a genuine "median peer's total spending" line,
|
||
sum each peer's own categories first and take the quantile of those
|
||
per-government totals:
|
||
|
||
```r
|
||
library(dplyr)
|
||
cmp |>
|
||
filter(role %in% c("target", "peer")) |>
|
||
group_by(year, role, canonical_govid) |>
|
||
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||
filter(role == "peer") |>
|
||
group_by(year) |>
|
||
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||
```
|
||
}
|
||
\description{
|
||
Pulls spending for the target plus a peer set (either a
|
||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||
call. Those summary rows are quantiles **within each category**, not
|
||
quantiles of each peer's total — see the `@return` section before summing
|
||
them.
|
||
}
|
||
\section{Reading `coverage`}{
|
||
|
||
`provenance$coverage` reports `n_units_reporting` against
|
||
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||
it counts cohort members with rows for the category you asked for, not
|
||
cohort members collected that year.** A government that was surveyed and
|
||
genuinely spends nothing in that category is indistinguishable here from one
|
||
that was never surveyed.
|
||
|
||
The ratio is therefore **not a response rate** and must not be used as one.
|
||
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||
`category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||
contract policing to the county sheriff, not non-response.
|
||
|
||
The comparison that *is* valid is the same category across a census year
|
||
(ending in 2 or 7) and a sample year, where the real-zero component is
|
||
roughly constant and the difference reflects the survey cycle. `is_census_year`
|
||
marks which is which.
|
||
}
|
||
|