The three kodor/fix issues, taken over after a day with no branch, PR or comment on any of them. Batched because each is single-file with a committed acceptance test, and two share documentation surfaces. #16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight into regexp_matches() unescaped, while basket mode in the same file already routed it through .escape_regex() with the comment "so `name` is treated as a literal substring". Two failure modes, both HTTP 200 through the API: a government could not be found by its own complete name when that name contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a bare "." matched all 608 Wisconsin cities. Malformed pattern text reached the engine as an error, which cog-api surfaced as a 500 -- reachable by typing a real name one character at a time ("Athens-Clarke County (bal"). Utility mode now calls the escaper that already existed. Roxygen updated: utility mode is documented as a literal case-insensitive substring match, and the basket-mode "substring fallback" step no longer describes itself as a regex either. BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop working, because there is no regex left to anchor. Two existing tests used "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now search for those characters literally. Updated to the bare names, which still resolve to exactly one row each once scoped by state/type (verified, not assumed). There is no exact-match option in utility mode any more -- noted on the issue, since that is a real if small capability loss. #15 (F-004) -- the raw Census files report thousands of dollars; this package multiplies by 1000 and returns full US dollars. Correct, and already stated in ?cog_spending / ?cog_revenue @return, in provenance, and in cog-api's data-dictionary. Absent from every surface a reader meets FIRST. Added to README.md as its own section and to both vignettes' openings. The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule ("All raw `amt` values are in $1,000s") without scoping it to the raw column -- a reader applying that to amt_nominal overstates by 1000x and gets a plausible-looking number rather than an obvious error. Fixed there too; that directory has no git remote, so it rides in no PR and is left uncommitted for the owner. #14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately inside each (year, spend_subtype, category) cell, so a summary_p50 row is "the median peer's value in that one category", never "the value of the median peer's total" -- the median peer for Police and for Fire are usually different governments. Summing them across categories misstated a total-spending band by -32.7% to +251.0% across 24 years, with a sign flip at FY2012. The verb is right and its documented use (facet by role AND category) is unaffected, so the fix is @return prose plus a worked snippet showing the correct computation: sum each peer's own categories first, then take the quantile of those per-government totals. This is the R-side counterpart of cog-api#9, fixed on the API surface earlier today; the wording is deliberately consistent across the two. Note the phrase "not additive" has to stay on one roxygen source line -- the test greps the generated Rd, where a line wrap turns it into "not additive" and stops matching. Cost one red run to find. man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with no content change. The two Rd files kept carry only the edits above. Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips are #11, #12 and #13.
58 lines
3.2 KiB
R
58 lines
3.2 KiB
R
# Madison walkthrough audit -- finding F-025. Tracked as uscogdata#16.
|
|
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
#
|
|
# cog_gov_search()'s UTILITY mode interpolates `name` into
|
|
# regexp_matches(gov_name, <name>, 'i')
|
|
# unescaped (R/search.R:102), while BASKET mode in the same file already routes
|
|
# it through .escape_regex() (R/search.R:307) with the comment "so `name` is
|
|
# treated as a literal substring". Two failure modes result:
|
|
# correctness -- a real government cannot be found by its own exact name, and
|
|
# a single "." matches everything (HTTP 200 both ways via the API);
|
|
# robustness -- malformed regex reaches the engine and errors, which cog-api
|
|
# surfaces as a 500, reachable by typing a real name one
|
|
# character at a time.
|
|
#
|
|
# NOT asserted here: the finding's `q=St. Louis` example. Under correct literal
|
|
# matching that search still returns 0 rows, because the stored name is
|
|
# "ST LOUIS CITY" with no period -- it demonstrates today's over-matching
|
|
# semantics, not a row the fix makes findable.
|
|
|
|
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
|
|
|
# -- correctness (1): a government must be findable by its own exact name ---
|
|
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
|
# grouping, so its own complete name matches nothing.
|
|
fredonia <- cog_gov_search(name = "FREDONIA (BRISCOE) CITY")
|
|
expect_equal(nrow(fredonia), 1L)
|
|
expect_equal(fredonia$canonical_govid, "052117184386")
|
|
expect_equal(cog_gov_search(name = "FREDONIA (BRISCOE)")$canonical_govid,
|
|
"052117184386")
|
|
|
|
# -- correctness (2): a metacharacter must not become a wildcard ------------
|
|
# No Wisconsin city or village name contains a literal period -- established
|
|
# against the raw registry below, NOT through the verb under test. A literal
|
|
# search for "." must therefore return nothing; today it returns all 608.
|
|
con <- DBI::dbConnect(duckdb::duckdb())
|
|
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
xwalk <- paste0(sub("/$", "", Sys.getenv("USCOGDATA_URL")),
|
|
"/data/canonical_fips_xwalk.parquet")
|
|
with_dot <- DBI::dbGetQuery(con, paste0(
|
|
"SELECT COUNT(*) n FROM read_parquet('", xwalk, "') ",
|
|
"WHERE fips_state = '55' AND govs_type = 2 AND gov_name LIKE '%.%'"))
|
|
expect_equal(as.integer(with_dot$n[[1]]), 0L)
|
|
|
|
expect_equal(nrow(cog_gov_search(name = ".", state = "WI", type = "city")), 0L)
|
|
expect_equal(nrow(cog_gov_search(name = "M.dison", state = "WI", type = "city")), 0L)
|
|
expect_equal(nrow(cog_gov_search(name = "Mad(i|o)son", state = "WI", type = "city")), 0L)
|
|
|
|
# A metacharacter-free name still resolves exactly as before.
|
|
expect_equal(nrow(cog_gov_search(name = "Madison", state = "WI", type = "city")), 1L)
|
|
|
|
# -- robustness: malformed pattern text returns no rows, and does not error --
|
|
# "[" alone, and "Athens-Clarke County (bal" -- an in-progress substring of
|
|
# ATHENS-CLARKE COUNTY (BALANCE), a real government -- both currently raise
|
|
# (DuckDB: "Invalid Input Error: missing ]").
|
|
expect_equal(nrow(cog_gov_search(name = "[")), 0L)
|
|
expect_equal(nrow(cog_gov_search(name = "Athens-Clarke County (bal")), 0L)
|
|
})
|