Files
uscogdata/tests/testthat/test-peer-summary-scope.R
T
jared d006dea6e4
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00

45 lines
2.4 KiB
R

# Madison walkthrough audit -- finding F-021. Tracked as uscogdata#14.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# .peer_summary_rows() computes stats::quantile() separately INSIDE each
# (year, spend_subtype, category) cell. A summary_p50 row is therefore "the
# median peer's value in that one category", not "the value of the median
# peer's total". Summing those rows across categories -- the obvious move for a
# caller who wants one peer-median total line and reads only the column names --
# misstated a total-spending band by -32.7% to +251.0% across the 24 years the
# audit tested, with a sign flip at FY2012.
#
# The verb is not wrong and its documented use (faceting by role AND category)
# is unaffected, so the fix is documentation: one sentence in @return.
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
rd <- paste(readLines(testthat::test_path("..", "..", "man", "cog_peer_compare.Rd"),
warn = FALSE), collapse = " ")
# The @return section must say the quantile is computed within each cell...
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
# ...and must warn that the rows are not additive across category.
expect_match(rd, "not additive|do(es)? not sum|cannot be summed", ignore.case = TRUE)
# ...naming the grouping explicitly.
expect_match(rd, "spend_subtype", fixed = TRUE)
# Pin the mechanism numerically so a future refactor that quietly changes the
# quantile grouping fails here rather than silently invalidating the sentence
# above. Fixture: Madison, 10 peers found at FY2020, category = NULL.
peers <- cog_find_peers("552025209777", year = 2020L, max_peers = 10L)
cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers,
category = NULL, years = 2020L, per_capita = TRUE)
naive <- sum(cmp$amt_per_capita_nominal[cmp$role == "summary_p50"], na.rm = TRUE)
peer_rows <- cmp[cmp$role == "peer", ]
per_gov <- tapply(peer_rows$amt_per_capita_nominal, peer_rows$canonical_govid,
sum, na.rm = TRUE)
correct <- unname(stats::quantile(per_gov, 0.5, na.rm = TRUE))
expect_equal(round(naive), 6180) # summing the built-in summary rows
expect_equal(round(correct), 2043) # quantile of each peer's OWN total
expect_gt(naive / correct, 2) # a +200% misstatement on this cohort
})