fix: gate direct-suppressed flag/note on an actually-covering recipe
R-CMD-check / check (push) Successful in 2m53s
R-CMD-check / check (pull_request) Successful in 2m53s

.detect_direct_suppressed() equated "no Direct sibling row" with "Direct
was suppressed", but the dominant real cause is a government with
genuinely no direct spending in that category (e.g. a state funding K-12
entirely through school districts) -- correct, ordinary data, not
suppression. Measured: 32 of 50 states false-flagged on a clean FY2019
category = NULL total query, and all 141 flagged rows across 50 states x
{2011, 2019} fell back to "no covering recipe found" instead of naming one
-- including AL Corrections, which names corrections_combined correctly
when category is supplied explicitly.

Both the flag and its row note are now gated on a harmonization recipe
actually covering that exact (year, canonical_govid, category) triple, via
a new .covering_recipes() helper that runs the same generic recipe join
per-row regardless of whether the caller supplied a category filter.
.notes_column() takes the precomputed note vector directly instead of
searching a category-gated suggestions list; .direct_suppressed_note() is
removed (its "no recipe found" fallback no longer applies -- if no recipe
covers a triple, it isn't suppression).

Also recomputes two total-spending.Rmd figures the prior wave never
actually reconciled with its own "measured against the fixture" caption:
State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M
(flat 188.3%, now 144%-189% varying by year).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-07-27 13:05:58 -04:00
co-authored by Claude Opus 5
parent e7d3a7a310
commit d258cef8c5
3 changed files with 213 additions and 48 deletions
+123 -43
View File
@@ -251,18 +251,22 @@ cog_spending <- function(govid, years, category = NULL,
# C1(b): when expenditure_concept = "total", flag any row where the IG # C1(b): when expenditure_concept = "total", flag any row where the IG
# leg has dollars but the Direct leg has none for that same (year, # leg has dollars but the Direct leg has none for that same (year,
# canonical_govid, category) -- the UNION'd figure there is # canonical_govid, category) AND a harmonization recipe actually recovers
# intergovernmental money ALONE, not Direct + IG, and both the row-level # the missing Direct dollars for that exact triple -- see
# notes and the provenance must say so rather than pass silently as a # .detect_direct_suppressed() for why bare Direct-row absence alone is NOT
# plausible Total. # sufficient (the dominant real cause is a government that simply has no
direct_suppressed <- if (identical(expenditure_concept, "total")) { # direct spending in that category, which is correct, ordinary data). When
.detect_direct_suppressed(result, subtype_col) # a covering recipe is found, both the row-level notes and the provenance
# say so rather than pass silently as a plausible Total.
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
.detect_direct_suppressed(con, result, subtype_col)
} else { } else {
rep(FALSE, nrow(result)) list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
} }
direct_suppressed <- direct_suppressed_info$flag
direct_suppressed_flag <- isTRUE(any(direct_suppressed)) direct_suppressed_flag <- isTRUE(any(direct_suppressed))
result$notes <- .notes_column(result, direct_suppressed, suggestions) result$notes <- .notes_column(result, direct_suppressed_info$notes)
# Determine expenditure_concept_note: only non-empty for "total", explains # Determine expenditure_concept_note: only non-empty for "total", explains
# how the IG leg was assembled from legacy-era aggregates. When the Direct # how the IG leg was assembled from legacy-era aggregates. When the Direct
@@ -502,47 +506,126 @@ cog_spending <- function(govid, years, category = NULL,
#' Detect rows where expenditure_concept = "total" is reporting the #' Detect rows where expenditure_concept = "total" is reporting the
#' intergovernmental leg with NO Direct counterpart in the same (year, #' intergovernmental leg with NO Direct counterpart in the same (year,
#' canonical_govid, category) group -- i.e. the Direct leg is suppressed #' canonical_govid, category) group AND a harmonization recipe actually
#' (typically a legacy aggregate-only family, see C1(a) above) rather than #' recovers the missing Direct dollars for that exact (year, canonical_govid,
#' genuinely zero. `TRUE` only for the `spend_subtype == "intergovernmental"` #' category) triple.
#' row(s) in each such group. #'
#' Bare Direct-row absence is deliberately NOT sufficient on its own: the
#' dominant real cause of "no Direct sibling row" is a government that simply
#' has no direct spending in that category (e.g. a state that funds K-12
#' entirely through school districts), which is correct, ordinary data, not
#' suppression. Genuine suppression -- a legacy aggregate-only family whose
#' Direct-leg basis query excludes it by construction (spending_long/
#' spending_long_harmonized both filter NOT is_aggregate) -- always has a
#' covering harmonization recipe, because that is exactly what the recipe
#' catalog exists to recover (see R/suggestions.R and `cog_recipes()`). So
#' checking "does a recipe actually cover this triple" cleanly separates the
#' two cases instead of conflating them.
#'
#' Returns `list(flag, notes)`, both the same length as `result`: `flag` is
#' `TRUE` only for the `spend_subtype == "intergovernmental"` row(s) in a
#' suppressed group, and `notes` names the recovering recipe(s) for those
#' rows (`NA` everywhere else).
#' @noRd #' @noRd
.detect_direct_suppressed <- function(result, subtype_col) { .detect_direct_suppressed <- function(con, result, subtype_col) {
n <- nrow(result) n <- nrow(result)
if (n == 0L) return(logical(0)) empty_notes <- rep(NA_character_, n)
if (n == 0L) return(list(flag = logical(0), notes = character(0)))
is_ig <- result[[subtype_col]] %in% "intergovernmental" is_ig <- result[[subtype_col]] %in% "intergovernmental"
if (!any(is_ig)) return(rep(FALSE, n)) if (!any(is_ig)) return(list(flag = rep(FALSE, n), notes = empty_notes))
key <- paste(result$year, result$canonical_govid, result$category, sep = "\r") key <- paste(result$year, result$canonical_govid, result$category, sep = "\r")
has_direct <- key %in% unique(key[!is_ig]) has_direct <- key %in% unique(key[!is_ig])
is_ig & !has_direct candidate <- is_ig & !has_direct
}
#' Build the notes text for a direct-suppressed row: names the recipe that flag <- rep(FALSE, n)
#' recovers the missing Direct component when one of the (already notes <- empty_notes
#' Direct-leg-scoped, see C1(a)) suggestions covers this row's year, or a if (!any(candidate)) return(list(flag = flag, notes = notes))
#' generic fallback when no such recipe was found.
#' @noRd idx <- which(candidate)
.direct_suppressed_note <- function(year, suggestions) { rows <- unique(result[idx, c("year", "canonical_govid", "category")])
matching <- Filter(function(s) { covering <- .covering_recipes(con, rows)
ay <- s$available_years cov_key <- paste(covering$year, covering$canonical_govid, covering$category,
!is.null(ay) && length(ay) == 2L && year >= ay[1] && year <= ay[2] sep = "\r")
}, suggestions)
if (length(matching) == 0L) { for (i in idx) {
return(paste( k <- paste(result$year[i], result$canonical_govid[i], result$category[i],
"Direct component is unavailable through this basis for this year", sep = "\r")
"(legacy aggregate-only family); no covering recipe found in this", m <- match(k, cov_key)
"corpus -- see cog_recipes()." if (is.na(m)) next
)) ids <- covering$recipe_ids[[m]]
} if (length(ids) == 0L) next
ids <- sort(unique(vapply(matching, function(s) s$recipe_id, character(1)))) flag[i] <- TRUE
sprintf( notes[i] <- sprintf(
"Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).", "Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).",
paste(ids, collapse = "', '") paste(sort(unique(ids)), collapse = "', '")
) )
} }
list(flag = flag, notes = notes)
}
#' For each (year, canonical_govid, category) triple potentially affected by
#' a suppressed Direct leg, find the harmonization recipe(s) that (a) cover
#' this `category` (share a component item_code via `summary_categories`,
#' excluding any recipe that is itself entirely intergovernmental M/L -- the
#' same exclusion `.build_suggestions()` applies, see I2) and (b) actually
#' produce a `long` row for this exact (canonical_govid, year) via the same
#' generic join `.run_recipe()` uses (component year_min/year_max +
#' gov_type_scope, no is_aggregate filter -- a recipe's whole point is to
#' recover data that's aggregate-only). Adds a list-column `recipe_ids`
#' (possibly length-0) to `rows`.
#' @noRd
.covering_recipes <- function(con, rows) {
rows$recipe_ids <- vector("list", nrow(rows))
cats <- unique(rows$category[!is.na(rows$category)])
if (length(cats) == 0L) return(rows)
cand <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT sc.category, r.recipe_id
FROM harmonization_recipes r
JOIN summary_categories sc ON sc.item_code = r.component_code
WHERE sc.category IN (%s)
AND r.recipe_id NOT IN (
SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE LEFT(component_code, 1) IN ('M', 'L')
)",
.sql_lit_chr(cats)
))
if (nrow(cand) == 0L) return(rows)
recipe_ids_all <- unique(cand$recipe_id)
govids <- unique(rows$canonical_govid)
years <- unique(rows$year)
covered <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.canonical_govid, l.year
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%s)
AND l.canonical_govid IN (%s)
AND l.year IN (%s)",
.sql_lit_chr(recipe_ids_all), .sql_lit_chr(govids), paste(years, collapse = ",")
))
for (i in seq_len(nrow(rows))) {
cat_i <- rows$category[i]
if (is.na(cat_i)) next
cat_recipe_ids <- cand$recipe_id[cand$category == cat_i]
if (length(cat_recipe_ids) == 0L) next
sub <- covered[covered$canonical_govid == rows$canonical_govid[i] &
covered$year == rows$year[i] &
covered$recipe_id %in% cat_recipe_ids, ]
rows$recipe_ids[[i]] <- sort(unique(sub$recipe_id))
}
rows
}
#' @noRd #' @noRd
.notes_column <- function(result, direct_suppressed = NULL, suggestions = list()) { .notes_column <- function(result, direct_suppressed_notes = NULL) {
n <- nrow(result) n <- nrow(result)
if (n == 0L) return(character(0)) if (n == 0L) return(character(0))
parts <- vector("list", 3L) parts <- vector("list", 3L)
@@ -562,11 +645,8 @@ cog_spending <- function(govid, years, category = NULL,
} else { } else {
rep(NA_character_, n) rep(NA_character_, n)
} }
parts[[3]] <- if (!is.null(direct_suppressed) && any(direct_suppressed)) { parts[[3]] <- if (!is.null(direct_suppressed_notes)) {
vapply(seq_len(n), function(i) { direct_suppressed_notes
if (!isTRUE(direct_suppressed[i])) return(NA_character_)
.direct_suppressed_note(result$year[i], suggestions)
}, character(1))
} else { } else {
rep(NA_character_, n) rep(NA_character_, n)
} }
+79
View File
@@ -367,6 +367,85 @@ test_that("C1(b): expenditure_concept_direct_suppressed is FALSE when the Direct
grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"]))) grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"])))
}) })
# M/I fix: .detect_direct_suppressed() was equating "no Direct sibling row"
# with "Direct was suppressed", but the dominant real cause is a government
# that simply has no direct spending in that category -- correct, ordinary
# data. The fix gates the flag (and its row note) on a harmonization recipe
# ACTUALLY covering that exact (year, canonical_govid, category) triple.
test_that("M/I: true positive, category supplied explicitly (unchanged behavior)", {
al <- "010000226085"
t_cat <- suppressMessages(cog_spending(
al, years = 2011, category = "Corrections", expenditure_concept = "total"
))
expect_true(attr(t_cat, "provenance")$expenditure_concept_direct_suppressed)
expect_match(t_cat$notes, "corrections_combined", fixed = TRUE)
expect_match(t_cat$notes, "unavailable", fixed = TRUE)
})
test_that("M/I: true positive, category = NULL now also names the recipe (was the fallback bug)", {
# Root bug: .build_suggestions() short-circuits to list() when category is
# NULL, so the note previously always hit its "no covering recipe found"
# fallback here even though corrections_combined genuinely covers this row.
al <- "010000226085"
t_null <- suppressMessages(cog_spending(
al, years = 2011, category = NULL, expenditure_concept = "total"
))
corr_row <- t_null[t_null$category %in% "Corrections", ]
expect_equal(nrow(corr_row), 1L)
expect_true(attr(t_null, "provenance")$expenditure_concept_direct_suppressed)
expect_match(corr_row$notes, "corrections_combined", fixed = TRUE)
expect_match(corr_row$notes, "unavailable", fixed = TRUE)
expect_false(grepl("no covering recipe found", corr_row$notes, fixed = TRUE))
})
test_that("M/I: false positive -- Virginia Education K-12 FY2019 total is NOT flagged", {
# States fund K-12 through school districts, so the Direct leg (E12/F12/
# G12) is genuinely, correctly zero -- not suppressed. Must not be flagged
# and must carry no suppression note.
va <- "510000227542"
t_va <- suppressMessages(cog_spending(
va, years = 2019, category = "Education K-12", expenditure_concept = "total"
))
expect_equal(nrow(t_va), 1L)
expect_equal(t_va$spend_subtype, "intergovernmental")
expect_equal(t_va$amt_nominal, 8028179000)
expect_false(isTRUE(attr(t_va, "provenance")$expenditure_concept_direct_suppressed))
expect_false(nzchar(t_va$notes) && grepl("unavailable", t_va$notes))
})
test_that("M/I: false positive by construction -- 'Other Education' has no E/F/G code, never flagged", {
# "Other Education" maps only to M21/L21 in summary_categories -- there is
# no E/F/G code for it in this corpus at all, so no Direct-recovering
# recipe can exist and it must never be flagged, in any fixture year.
con <- uscogdata:::.ensure_session()
years_all <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
states <- DBI::dbGetQuery(con,
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
oe <- suppressMessages(cog_spending(
states, years = years_all, category = "Other Education",
expenditure_concept = "total"
))
expect_false(isTRUE(attr(oe, "provenance")$expenditure_concept_direct_suppressed))
expect_false(any(nzchar(oe$notes) & grepl("unavailable", oe$notes)))
})
test_that("M/I: a clean FY2019 category = NULL total query flags far fewer than the pre-fix 32/50 states", {
con <- uscogdata:::.ensure_session()
states <- DBI::dbGetQuery(con,
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
r <- suppressMessages(cog_spending(
states, years = 2019, category = NULL, expenditure_concept = "total"
))
ig <- r[r$spend_subtype == "intergovernmental", ]
flagged <- ig[nzchar(ig$notes) & grepl("unavailable", ig$notes), ]
expect_lt(length(unique(flagged$canonical_govid)), 32L)
# Every remaining flagged row must actually name a covering recipe --
# never the old no-recipe-found fallback.
expect_true(all(grepl("recipe = '", flagged$notes, fixed = TRUE)))
expect_false(any(grepl("no covering recipe found", flagged$notes, fixed = TRUE)))
})
test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", { test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", {
with_corpus_missing_ig_categories({ with_corpus_missing_ig_categories({
con <- uscogdata:::.ensure_session() con <- uscogdata:::.ensure_session()
+10 -4
View File
@@ -161,13 +161,17 @@ share of a government's own Direct spending is:
| Government type | Intergovernmental / Direct | | Government type | Intergovernmental / Direct |
|---|---| |---|---|
| State | 17.2% | | State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) |
| County | 3.4%-5.1% (varies by year) | | County | 3.4%-5.1% (varies by year) |
| City | 2.6%-3.1% (varies by year) | | City | 2.6%-3.1% (varies by year) |
So the Direct/Total choice matters overwhelmingly for **state** governments So the Direct/Total choice matters overwhelmingly for **state** governments
-- a state's Total genuinely differs from its Direct by a meaningful margin, -- a state's Total genuinely differs from its Direct by a meaningful margin,
while for a county or city the two are close. That's also why the mistake while for a county or city the two are close. The state range is also far
wider than a single flat figure would suggest: legacy wide-era years (2011:
48.4%) carry proportionally more intergovernmental spending than the modern
era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly
3x larger a decade earlier than it is today. That's also why the mistake
this vignette warns about is easy to make unnoticed at the county/city level this vignette warns about is easy to make unnoticed at the county/city level
and costly at the state level: rolling up every government in a state using and costly at the state level: rolling up every government in a state using
`total` instead of `direct` overstates the true figure -- measured at 7.6% `total` instead of `direct` overstates the true figure -- measured at 7.6%
@@ -186,8 +190,10 @@ its own spending, just routed to a different kind of recipient. On the
bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for
state governments (a state has no "payments to the state government" leg of state governments (a state has no "payments to the state government" leg of
its own) but is 43%-51% the size of `M` for counties (varies by year) and its own) but is 43%-51% the size of `M` for counties (varies by year) and
188.3% the size of `M` for cities -- so a `total` that omitted `L` would 144%-189% the size of `M` for cities (varies by year; 166% pooled across
silently undercount Total specifically for local governments. all four) -- so a `total` that omitted `L` would silently undercount Total
specifically for local governments, and for cities `L` is often the
*larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes both legs (excluding `cog_spending(expenditure_concept = "total")` includes both legs (excluding
the `L--` family-total rollup row, which would double-count its own the `L--` family-total rollup row, which would double-count its own
components). components).