The three kodor/fix issues, taken over after a day with no branch, PR or comment on any of them. Batched because each is single-file with a committed acceptance test, and two share documentation surfaces. #16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight into regexp_matches() unescaped, while basket mode in the same file already routed it through .escape_regex() with the comment "so `name` is treated as a literal substring". Two failure modes, both HTTP 200 through the API: a government could not be found by its own complete name when that name contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a bare "." matched all 608 Wisconsin cities. Malformed pattern text reached the engine as an error, which cog-api surfaced as a 500 -- reachable by typing a real name one character at a time ("Athens-Clarke County (bal"). Utility mode now calls the escaper that already existed. Roxygen updated: utility mode is documented as a literal case-insensitive substring match, and the basket-mode "substring fallback" step no longer describes itself as a regex either. BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop working, because there is no regex left to anchor. Two existing tests used "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now search for those characters literally. Updated to the bare names, which still resolve to exactly one row each once scoped by state/type (verified, not assumed). There is no exact-match option in utility mode any more -- noted on the issue, since that is a real if small capability loss. #15 (F-004) -- the raw Census files report thousands of dollars; this package multiplies by 1000 and returns full US dollars. Correct, and already stated in ?cog_spending / ?cog_revenue @return, in provenance, and in cog-api's data-dictionary. Absent from every surface a reader meets FIRST. Added to README.md as its own section and to both vignettes' openings. The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule ("All raw `amt` values are in $1,000s") without scoping it to the raw column -- a reader applying that to amt_nominal overstates by 1000x and gets a plausible-looking number rather than an obvious error. Fixed there too; that directory has no git remote, so it rides in no PR and is left uncommitted for the owner. #14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately inside each (year, spend_subtype, category) cell, so a summary_p50 row is "the median peer's value in that one category", never "the value of the median peer's total" -- the median peer for Police and for Fire are usually different governments. Summing them across categories misstated a total-spending band by -32.7% to +251.0% across 24 years, with a sign flip at FY2012. The verb is right and its documented use (facet by role AND category) is unaffected, so the fix is @return prose plus a worked snippet showing the correct computation: sum each peer's own categories first, then take the quantile of those per-government totals. This is the R-side counterpart of cog-api#9, fixed on the API surface earlier today; the wording is deliberately consistent across the two. Note the phrase "not additive" has to stay on one roxygen source line -- the test greps the generated Rd, where a line wrap turns it into "not additive" and stops matching. Cost one red run to find. man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with no content change. The two Rd files kept carry only the edits above. Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips are #11, #12 and #13.
82 lines
3.1 KiB
R
82 lines
3.1 KiB
R
% Generated by roxygen2: do not edit by hand
|
||
% Please edit documentation in R/peers.R
|
||
\name{cog_peer_compare}
|
||
\alias{cog_peer_compare}
|
||
\title{Compare a target government against a peer set}
|
||
\usage{
|
||
cog_peer_compare(
|
||
target_govid,
|
||
peers,
|
||
category,
|
||
years,
|
||
per_capita = TRUE,
|
||
adjust_to_year = NULL,
|
||
expenditure_concept = c("direct", "total")
|
||
)
|
||
}
|
||
\arguments{
|
||
\item{target_govid}{Character scalar.}
|
||
|
||
\item{peers}{A tibble from [cog_find_peers()] or a character vector of
|
||
`canonical_govid`s.}
|
||
|
||
\item{category}{Character scalar or vector.}
|
||
|
||
\item{years}{Integer vector.}
|
||
|
||
\item{per_capita}{Default `TRUE` — peer compare usually normalizes by
|
||
population.}
|
||
|
||
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
|
||
|
||
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
|
||
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||
single-government queries but cannot be used here because combining Total
|
||
across peer sets counts intergovernmental transfers twice.}
|
||
}
|
||
\value{
|
||
Tibble matching [cog_spending()]'s columns, plus a `role`
|
||
column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
||
`"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
||
among target+peers at `max(years)`, NA for other rows), and
|
||
`cohort_year` (the year used to build the peer cohort, read from
|
||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||
`cohort_year`, and `cohort_govids`.
|
||
|
||
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
||
Each one is computed **within each `(year, spend_subtype,
|
||
category)` cell** across the peer set, so a `summary_p50` row is *the
|
||
median peer's value in that one category*, not *the value of the median
|
||
peer's total*. The median peer for Police and the median peer for Fire
|
||
are usually different governments, so summing `summary_*` rows across
|
||
categories does not give any peer's total and misstates the band it
|
||
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||
one cohort, with a sign flip at FY2012.
|
||
|
||
Facet by `role` **and** `category` (the documented use, and what the
|
||
rows are built for). For a genuine "median peer's total spending" line,
|
||
sum each peer's own categories first and take the quantile of those
|
||
per-government totals:
|
||
|
||
```r
|
||
library(dplyr)
|
||
cmp |>
|
||
filter(role %in% c("target", "peer")) |>
|
||
group_by(year, role, canonical_govid) |>
|
||
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||
filter(role == "peer") |>
|
||
group_by(year) |>
|
||
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||
```
|
||
}
|
||
\description{
|
||
Pulls spending for the target plus a peer set (either a
|
||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||
call. Those summary rows are quantiles **within each category**, not
|
||
quantiles of each peer's total — see the `@return` section before summing
|
||
them.
|
||
}
|