Six skipped tests, one per issue opened from the Madison walkthrough audit (docs/walkthroughs/FINDINGS.md in cog_explorer). Each asserts the desired behaviour, so it fails today and goes green when the fix lands; each is guarded by a single skip() naming its issue and finding IDs, so the suite stays green and activating a test is a one-line deletion. test-expenditure-concepts.R #11 F-012, F-017, F-018 test-revenue-concept-insurance-trust.R #12 F-014 test-coverage-disclosure.R #13 F-020, F-023 test-peer-summary-scope.R #14 F-021 test-amount-units-documented.R #15 F-004 test-gov-search-literal-match.R #16 F-025 helper-walkthrough-raw.R reads the corpus's long parquet partitions directly, bypassing uscogdata's SQL views. Every expected amount comes from there rather than from the verb under test - verifying an absence through the filter that creates it proves nothing, which was the most common defect in the audit itself. Verified: with the skips removed all six fail (or error) against the bundled fixture; with them in place the full suite is 576 pass / 0 fail / 6 skip.
47 lines
2.5 KiB
R
47 lines
2.5 KiB
R
# Madison walkthrough audit -- finding F-004. Tracked as uscogdata#15.
|
|
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
#
|
|
# The raw Census files report thousands of dollars; this package multiplies by
|
|
# 1000 and returns full US dollars. That is the friendlier choice and is not
|
|
# wrong -- but cog_explorer's CLAUDE.md states "All raw `amt` values are in
|
|
# $1,000s", so a reader who applies that rule to amt_nominal overstates every
|
|
# figure by 1000x, and gets a plausible-looking number rather than an obvious
|
|
# error. The audit rates this the highest-consequence definitional gap it found.
|
|
#
|
|
# Deliberately NOT asserted here: man/cog_spending.Rd and man/cog_revenue.Rd,
|
|
# which ALREADY carry the statement in their @return sections (verified
|
|
# 2026-07-29), as does cog-api's data-dictionary.md (since 2b71b41). The gap is
|
|
# in the surfaces a reader meets first and in cog_explorer's own conventions
|
|
# doc -- see uscogdata#15 for the full surface-by-surface table and for the two
|
|
# secondary tasks (cog_explorer/CLAUDE.md, which has no git remote, and
|
|
# cog-api's llms.txt, which is silent on units).
|
|
|
|
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
|
testthat::skip("Blocked on uscogdata#15 (finding F-004)")
|
|
|
|
says_units <- function(path) {
|
|
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
|
grepl("full US dollars|full U\\.S\\. dollars", txt, ignore.case = TRUE) &&
|
|
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
|
}
|
|
|
|
expect_true(says_units(testthat::test_path("..", "..", "README.md")))
|
|
expect_true(says_units(testthat::test_path("..", "..", "vignettes", "total-spending.Rmd")))
|
|
expect_true(says_units(testthat::test_path("..", "..", "vignettes",
|
|
"population-denominators.Rmd")))
|
|
|
|
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
|
# The expected raw amount is read straight from the corpus's parquet
|
|
# partitions -- never through cog_spending(), which is the thing being
|
|
# described. Madison FY2020: E/F/G = 623,347 ($1,000s) -> $623,347,000.
|
|
raw_thousands <- wt_raw_amt("552025209777", 2020L, prefixes = c("E", "F", "G"))
|
|
expect_equal(raw_thousands, 623347)
|
|
|
|
returned <- cog_spending(govid = "552025209777", years = 2020L)
|
|
expect_equal(sum(returned$amt_nominal), raw_thousands * 1000)
|
|
|
|
units <- attr(returned, "provenance")$transformations$units_conversion
|
|
expect_true(units$applied)
|
|
expect_equal(units$multiplier, 1000)
|
|
})
|