From d258cef8c52f6f2790dff18767d210af0e32e10a Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 27 Jul 2026 13:05:58 -0400 Subject: [PATCH] fix: gate direct-suppressed flag/note on an actually-covering recipe .detect_direct_suppressed() equated "no Direct sibling row" with "Direct was suppressed", but the dominant real cause is a government with genuinely no direct spending in that category (e.g. a state funding K-12 entirely through school districts) -- correct, ordinary data, not suppression. Measured: 32 of 50 states false-flagged on a clean FY2019 category = NULL total query, and all 141 flagged rows across 50 states x {2011, 2019} fell back to "no covering recipe found" instead of naming one -- including AL Corrections, which names corrections_combined correctly when category is supplied explicitly. Both the flag and its row note are now gated on a harmonization recipe actually covering that exact (year, canonical_govid, category) triple, via a new .covering_recipes() helper that runs the same generic recipe join per-row regardless of whether the caller supplied a category filter. .notes_column() takes the precomputed note vector directly instead of searching a category-gated suggestions list; .direct_suppressed_note() is removed (its "no recipe found" fallback no longer applies -- if no recipe covers a triple, it isn't suppression). Also recomputes two total-spending.Rmd figures the prior wave never actually reconciled with its own "measured against the fixture" caption: State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M (flat 188.3%, now 144%-189% varying by year). Co-Authored-By: Claude Opus 5 (1M context) --- R/spending.R | 168 ++++++++++++++++------ tests/testthat/test-expenditure-concept.R | 79 ++++++++++ vignettes/total-spending.Rmd | 14 +- 3 files changed, 213 insertions(+), 48 deletions(-) diff --git a/R/spending.R b/R/spending.R index 7edd710..80a9064 100644 --- a/R/spending.R +++ b/R/spending.R @@ -251,18 +251,22 @@ cog_spending <- function(govid, years, category = NULL, # C1(b): when expenditure_concept = "total", flag any row where the IG # leg has dollars but the Direct leg has none for that same (year, - # canonical_govid, category) -- the UNION'd figure there is - # intergovernmental money ALONE, not Direct + IG, and both the row-level - # notes and the provenance must say so rather than pass silently as a - # plausible Total. - direct_suppressed <- if (identical(expenditure_concept, "total")) { - .detect_direct_suppressed(result, subtype_col) + # canonical_govid, category) AND a harmonization recipe actually recovers + # the missing Direct dollars for that exact triple -- see + # .detect_direct_suppressed() for why bare Direct-row absence alone is NOT + # sufficient (the dominant real cause is a government that simply has no + # direct spending in that category, which is correct, ordinary data). When + # a covering recipe is found, both the row-level notes and the provenance + # say so rather than pass silently as a plausible Total. + direct_suppressed_info <- if (identical(expenditure_concept, "total")) { + .detect_direct_suppressed(con, result, subtype_col) } else { - rep(FALSE, nrow(result)) + list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result))) } + direct_suppressed <- direct_suppressed_info$flag direct_suppressed_flag <- isTRUE(any(direct_suppressed)) - result$notes <- .notes_column(result, direct_suppressed, suggestions) + result$notes <- .notes_column(result, direct_suppressed_info$notes) # Determine expenditure_concept_note: only non-empty for "total", explains # how the IG leg was assembled from legacy-era aggregates. When the Direct @@ -502,47 +506,126 @@ cog_spending <- function(govid, years, category = NULL, #' Detect rows where expenditure_concept = "total" is reporting the #' intergovernmental leg with NO Direct counterpart in the same (year, -#' canonical_govid, category) group -- i.e. the Direct leg is suppressed -#' (typically a legacy aggregate-only family, see C1(a) above) rather than -#' genuinely zero. `TRUE` only for the `spend_subtype == "intergovernmental"` -#' row(s) in each such group. +#' canonical_govid, category) group AND a harmonization recipe actually +#' recovers the missing Direct dollars for that exact (year, canonical_govid, +#' category) triple. +#' +#' Bare Direct-row absence is deliberately NOT sufficient on its own: the +#' dominant real cause of "no Direct sibling row" is a government that simply +#' has no direct spending in that category (e.g. a state that funds K-12 +#' entirely through school districts), which is correct, ordinary data, not +#' suppression. Genuine suppression -- a legacy aggregate-only family whose +#' Direct-leg basis query excludes it by construction (spending_long/ +#' spending_long_harmonized both filter NOT is_aggregate) -- always has a +#' covering harmonization recipe, because that is exactly what the recipe +#' catalog exists to recover (see R/suggestions.R and `cog_recipes()`). So +#' checking "does a recipe actually cover this triple" cleanly separates the +#' two cases instead of conflating them. +#' +#' Returns `list(flag, notes)`, both the same length as `result`: `flag` is +#' `TRUE` only for the `spend_subtype == "intergovernmental"` row(s) in a +#' suppressed group, and `notes` names the recovering recipe(s) for those +#' rows (`NA` everywhere else). #' @noRd -.detect_direct_suppressed <- function(result, subtype_col) { +.detect_direct_suppressed <- function(con, result, subtype_col) { n <- nrow(result) - if (n == 0L) return(logical(0)) + empty_notes <- rep(NA_character_, n) + if (n == 0L) return(list(flag = logical(0), notes = character(0))) is_ig <- result[[subtype_col]] %in% "intergovernmental" - if (!any(is_ig)) return(rep(FALSE, n)) + if (!any(is_ig)) return(list(flag = rep(FALSE, n), notes = empty_notes)) + key <- paste(result$year, result$canonical_govid, result$category, sep = "\r") has_direct <- key %in% unique(key[!is_ig]) - is_ig & !has_direct -} + candidate <- is_ig & !has_direct -#' Build the notes text for a direct-suppressed row: names the recipe that -#' recovers the missing Direct component when one of the (already -#' Direct-leg-scoped, see C1(a)) suggestions covers this row's year, or a -#' generic fallback when no such recipe was found. -#' @noRd -.direct_suppressed_note <- function(year, suggestions) { - matching <- Filter(function(s) { - ay <- s$available_years - !is.null(ay) && length(ay) == 2L && year >= ay[1] && year <= ay[2] - }, suggestions) - if (length(matching) == 0L) { - return(paste( - "Direct component is unavailable through this basis for this year", - "(legacy aggregate-only family); no covering recipe found in this", - "corpus -- see cog_recipes()." - )) + flag <- rep(FALSE, n) + notes <- empty_notes + if (!any(candidate)) return(list(flag = flag, notes = notes)) + + idx <- which(candidate) + rows <- unique(result[idx, c("year", "canonical_govid", "category")]) + covering <- .covering_recipes(con, rows) + cov_key <- paste(covering$year, covering$canonical_govid, covering$category, + sep = "\r") + + for (i in idx) { + k <- paste(result$year[i], result$canonical_govid[i], result$category[i], + sep = "\r") + m <- match(k, cov_key) + if (is.na(m)) next + ids <- covering$recipe_ids[[m]] + if (length(ids) == 0L) next + flag[i] <- TRUE + notes[i] <- sprintf( + "Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).", + paste(sort(unique(ids)), collapse = "', '") + ) } - ids <- sort(unique(vapply(matching, function(s) s$recipe_id, character(1)))) - sprintf( - "Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).", - paste(ids, collapse = "', '") - ) + list(flag = flag, notes = notes) +} + +#' For each (year, canonical_govid, category) triple potentially affected by +#' a suppressed Direct leg, find the harmonization recipe(s) that (a) cover +#' this `category` (share a component item_code via `summary_categories`, +#' excluding any recipe that is itself entirely intergovernmental M/L -- the +#' same exclusion `.build_suggestions()` applies, see I2) and (b) actually +#' produce a `long` row for this exact (canonical_govid, year) via the same +#' generic join `.run_recipe()` uses (component year_min/year_max + +#' gov_type_scope, no is_aggregate filter -- a recipe's whole point is to +#' recover data that's aggregate-only). Adds a list-column `recipe_ids` +#' (possibly length-0) to `rows`. +#' @noRd +.covering_recipes <- function(con, rows) { + rows$recipe_ids <- vector("list", nrow(rows)) + cats <- unique(rows$category[!is.na(rows$category)]) + if (length(cats) == 0L) return(rows) + + cand <- DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT sc.category, r.recipe_id + FROM harmonization_recipes r + JOIN summary_categories sc ON sc.item_code = r.component_code + WHERE sc.category IN (%s) + AND r.recipe_id NOT IN ( + SELECT DISTINCT recipe_id FROM harmonization_recipes + WHERE LEFT(component_code, 1) IN ('M', 'L') + )", + .sql_lit_chr(cats) + )) + if (nrow(cand) == 0L) return(rows) + + recipe_ids_all <- unique(cand$recipe_id) + govids <- unique(rows$canonical_govid) + years <- unique(rows$year) + covered <- DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT r.recipe_id, l.canonical_govid, l.year + FROM long l + JOIN harmonization_recipes r + ON l.item_code = r.component_code + AND l.year BETWEEN r.year_min AND r.year_max + AND (r.gov_type_scope = 'all' + OR (r.gov_type_scope = 'state' AND l.type = 0) + OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) + WHERE r.recipe_id IN (%s) + AND l.canonical_govid IN (%s) + AND l.year IN (%s)", + .sql_lit_chr(recipe_ids_all), .sql_lit_chr(govids), paste(years, collapse = ",") + )) + + for (i in seq_len(nrow(rows))) { + cat_i <- rows$category[i] + if (is.na(cat_i)) next + cat_recipe_ids <- cand$recipe_id[cand$category == cat_i] + if (length(cat_recipe_ids) == 0L) next + sub <- covered[covered$canonical_govid == rows$canonical_govid[i] & + covered$year == rows$year[i] & + covered$recipe_id %in% cat_recipe_ids, ] + rows$recipe_ids[[i]] <- sort(unique(sub$recipe_id)) + } + rows } #' @noRd -.notes_column <- function(result, direct_suppressed = NULL, suggestions = list()) { +.notes_column <- function(result, direct_suppressed_notes = NULL) { n <- nrow(result) if (n == 0L) return(character(0)) parts <- vector("list", 3L) @@ -562,11 +645,8 @@ cog_spending <- function(govid, years, category = NULL, } else { rep(NA_character_, n) } - parts[[3]] <- if (!is.null(direct_suppressed) && any(direct_suppressed)) { - vapply(seq_len(n), function(i) { - if (!isTRUE(direct_suppressed[i])) return(NA_character_) - .direct_suppressed_note(result$year[i], suggestions) - }, character(1)) + parts[[3]] <- if (!is.null(direct_suppressed_notes)) { + direct_suppressed_notes } else { rep(NA_character_, n) } diff --git a/tests/testthat/test-expenditure-concept.R b/tests/testthat/test-expenditure-concept.R index b19f374..95954ff 100644 --- a/tests/testthat/test-expenditure-concept.R +++ b/tests/testthat/test-expenditure-concept.R @@ -367,6 +367,85 @@ test_that("C1(b): expenditure_concept_direct_suppressed is FALSE when the Direct grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"]))) }) +# M/I fix: .detect_direct_suppressed() was equating "no Direct sibling row" +# with "Direct was suppressed", but the dominant real cause is a government +# that simply has no direct spending in that category -- correct, ordinary +# data. The fix gates the flag (and its row note) on a harmonization recipe +# ACTUALLY covering that exact (year, canonical_govid, category) triple. + +test_that("M/I: true positive, category supplied explicitly (unchanged behavior)", { + al <- "010000226085" + t_cat <- suppressMessages(cog_spending( + al, years = 2011, category = "Corrections", expenditure_concept = "total" + )) + expect_true(attr(t_cat, "provenance")$expenditure_concept_direct_suppressed) + expect_match(t_cat$notes, "corrections_combined", fixed = TRUE) + expect_match(t_cat$notes, "unavailable", fixed = TRUE) +}) + +test_that("M/I: true positive, category = NULL now also names the recipe (was the fallback bug)", { + # Root bug: .build_suggestions() short-circuits to list() when category is + # NULL, so the note previously always hit its "no covering recipe found" + # fallback here even though corrections_combined genuinely covers this row. + al <- "010000226085" + t_null <- suppressMessages(cog_spending( + al, years = 2011, category = NULL, expenditure_concept = "total" + )) + corr_row <- t_null[t_null$category %in% "Corrections", ] + expect_equal(nrow(corr_row), 1L) + expect_true(attr(t_null, "provenance")$expenditure_concept_direct_suppressed) + expect_match(corr_row$notes, "corrections_combined", fixed = TRUE) + expect_match(corr_row$notes, "unavailable", fixed = TRUE) + expect_false(grepl("no covering recipe found", corr_row$notes, fixed = TRUE)) +}) + +test_that("M/I: false positive -- Virginia Education K-12 FY2019 total is NOT flagged", { + # States fund K-12 through school districts, so the Direct leg (E12/F12/ + # G12) is genuinely, correctly zero -- not suppressed. Must not be flagged + # and must carry no suppression note. + va <- "510000227542" + t_va <- suppressMessages(cog_spending( + va, years = 2019, category = "Education K-12", expenditure_concept = "total" + )) + expect_equal(nrow(t_va), 1L) + expect_equal(t_va$spend_subtype, "intergovernmental") + expect_equal(t_va$amt_nominal, 8028179000) + expect_false(isTRUE(attr(t_va, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(nzchar(t_va$notes) && grepl("unavailable", t_va$notes)) +}) + +test_that("M/I: false positive by construction -- 'Other Education' has no E/F/G code, never flagged", { + # "Other Education" maps only to M21/L21 in summary_categories -- there is + # no E/F/G code for it in this corpus at all, so no Direct-recovering + # recipe can exist and it must never be flagged, in any fixture year. + con <- uscogdata:::.ensure_session() + years_all <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year + states <- DBI::dbGetQuery(con, + "SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid + oe <- suppressMessages(cog_spending( + states, years = years_all, category = "Other Education", + expenditure_concept = "total" + )) + expect_false(isTRUE(attr(oe, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(any(nzchar(oe$notes) & grepl("unavailable", oe$notes))) +}) + +test_that("M/I: a clean FY2019 category = NULL total query flags far fewer than the pre-fix 32/50 states", { + con <- uscogdata:::.ensure_session() + states <- DBI::dbGetQuery(con, + "SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid + r <- suppressMessages(cog_spending( + states, years = 2019, category = NULL, expenditure_concept = "total" + )) + ig <- r[r$spend_subtype == "intergovernmental", ] + flagged <- ig[nzchar(ig$notes) & grepl("unavailable", ig$notes), ] + expect_lt(length(unique(flagged$canonical_govid)), 32L) + # Every remaining flagged row must actually name a covering recipe -- + # never the old no-recipe-found fallback. + expect_true(all(grepl("recipe = '", flagged$notes, fixed = TRUE))) + expect_false(any(grepl("no covering recipe found", flagged$notes, fixed = TRUE))) +}) + test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", { with_corpus_missing_ig_categories({ con <- uscogdata:::.ensure_session() diff --git a/vignettes/total-spending.Rmd b/vignettes/total-spending.Rmd index 83707f3..0fb8b33 100644 --- a/vignettes/total-spending.Rmd +++ b/vignettes/total-spending.Rmd @@ -161,13 +161,17 @@ share of a government's own Direct spending is: | Government type | Intergovernmental / Direct | |---|---| -| State | 17.2% | +| State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) | | County | 3.4%-5.1% (varies by year) | | City | 2.6%-3.1% (varies by year) | So the Direct/Total choice matters overwhelmingly for **state** governments -- a state's Total genuinely differs from its Direct by a meaningful margin, -while for a county or city the two are close. That's also why the mistake +while for a county or city the two are close. The state range is also far +wider than a single flat figure would suggest: legacy wide-era years (2011: +48.4%) carry proportionally more intergovernmental spending than the modern +era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly +3x larger a decade earlier than it is today. That's also why the mistake this vignette warns about is easy to make unnoticed at the county/city level and costly at the state level: rolling up every government in a state using `total` instead of `direct` overstates the true figure -- measured at 7.6% @@ -186,8 +190,10 @@ its own spending, just routed to a different kind of recipient. On the bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has no "payments to the state government" leg of its own) but is 43%-51% the size of `M` for counties (varies by year) and -188.3% the size of `M` for cities -- so a `total` that omitted `L` would -silently undercount Total specifically for local governments. +144%-189% the size of `M` for cities (varies by year; 166% pooled across +all four) -- so a `total` that omitted `L` would silently undercount Total +specifically for local governments, and for cities `L` is often the +*larger* of the two legs. `cog_spending(expenditure_concept = "total")` includes both legs (excluding the `L--` family-total rollup row, which would double-count its own components).