47 KiB
Partial-Coverage Signposting Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Fire a harmonization-recipe suggestion when a requested category still returns rows but structurally excludes component dollars the corpus does hold — closing uscogdata#9, where cog_spending(category = "Public Welfare") silently drops aggregate-published E67/E68 in legacy years.
Architecture: .build_suggestions() keeps its recipe-candidate query, M/L exclusion, and IG-counterpart attachment untouched. Its sole year test — "the result has zero rows in this year" — gains a parallel qualifying path: "a component code carries nonzero dollars for this government in this year that the verb's own long view structurally excludes." Reachability is decided by anti-joining the real view (spending_long_harmonized / revenue_long_harmonized) rather than restating its WHERE clause, so the trigger tracks the view if it changes. Both verbs inherit the fix from the shared .verb_spendrev() call site.
Tech Stack: R (pkg uscogdata), DuckDB via DBI, testthat 3e, cli. Downstream: cog-api (plumber).
Global Constraints
- No new package dependencies.
withrstays Suggests-only, test-only. - All tests run offline against the bundled fixture at
inst/extdata/fixture_corpus/(years 2011, 2012, 2019, 2020). No live corpus, no credentials. Guard every corpus-touching test withskip_if_no_corpus(). - Every dollar figure in this plan is measured, not estimated — reproduced against the bundled fixture on 2026-08-04. Assert them exactly.
- Raw
amtis in $1,000s; multiply by1000.0in SQL. Values reaching provenance are full US dollars. - Suggestions only run under
basis = "harmonized"and a non-NULLcategory— the early return atR/suggestions.R:56is unchanged and must stay. - Functions under 50 lines, files under 400 lines.
R/suggestions.Ris currently 251 lines and must stay under 400. - Use
.sql_lit_chr()for every string literal interpolated into SQL. Never interpolate a user-supplied value as a SQL identifier. - Commit after each task. Conventional-commit prefixes (
fix:,test:,docs:,chore:).
Reference Data (measured 2026-08-04, bundled fixture)
Government ids are canonical_govid.
| Case | gov | year | category | today | after |
|---|---|---|---|---|---|
| Target defect | 061037123085 (LA County) |
2011 | Public Welfare | suggestions = list(), reports $3,185,943,000 |
2 suggestions; welfare_cash_e67_wide $1,803,872,000 (E67), welfare_cash_e68_wide $271,589,000 (E68) |
| Revenue twin | 020000227749 (Alaska state) |
2011 | Miscellaneous Revenue | suggestions = list(), reports $943,842,000 from U11,U20,U30 |
1 suggestion; rents_royalties_u4_wide $1,899,995,000 (U4-) |
| Regression | 061037123085 |
2011 | Corrections | 3 suggestions, empty_year |
same 3, now carrying $1,371,460,000 (E05), $17,373,000 (F05), $884,000 (G05) |
| Negative | 061037123085 |
2019 | Public Welfare | list() |
list() |
| Negative | higher_ed_e18_wide, general_gov_e89_wide |
any | any | never fire | never fire (E18/E89 are leaf-and-classified in 2011) |
Structural invariant the anti-join rests on: no recipe component is ever renamed by harmonization. Measured: 0 rows in long where item_code is a recipe component and harmonized_code IS NOT NULL AND harmonized_code <> item_code. All 19 components that differ have harmonized_code IS NULL, and every one is is_aggregate = TRUE.
Pre-verified safe — these existing assertions do not regress (0 new fires each):
test-recipes.R:196Broward 2019:2020 Correctionstest-expenditure-concept.R:293AL state 2019 Policetest-recipes.R:82Browardrecipe = "corrections_combined"— safe by construction; the recipe branch setssuggestions <- list()atR/spending.R:376and never calls.build_suggestions()
File Structure
- Modify
R/suggestions.R— add.suppressed_components(); widen.build_suggestions(); extend.inform_suggestions(). Owns the whole trigger. - Modify
R/spending.R— add.select_long_view()beside.select_view()(~line 511); pass the long-view name at the.build_suggestions()call site (~line 396). - Modify
R/explain.R— render the new detail in the Suggestions section (~line 127). - Modify
inst/schemas/provenance-v1.json— document the four new per-suggestion fields. - Modify
tests/testthat/test-recipes.R— all new reader tests (this file already owns signposting tests). - Modify
NEWS.md— user-facing entry. - Modify (cog-api)
api/tests/testthat/test-handlers-governments.R— assert the fields survive the envelope.
Task 0: Pre-flight — clean branch off main
Files: none (git only)
Interfaces:
- Consumes: nothing
- Produces: a clean
fix/partial-coverage-signposting-9branch based onorigin/main; a recorded baseline test count later tasks compare against
The working tree is currently on
ci/apt-httpswith two uncommitted fixture modifications (inst/extdata/fixture_corpus/data/series_breaks.parquet,inst/extdata/fixture_corpus/manifest.json). Do not carry them into this branch and do not discard them without checking with the user — they are unrelated to this work.
- Step 1: Inspect the uncommitted fixture changes
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata
git status --short
git diff --stat
Expected: exactly the two fixture files listed above, modified.
- Step 2: Stash them, and say so in the stash message
git stash push -m "WIP fixture series_breaks/manifest, unrelated to #9" \
inst/extdata/fixture_corpus/data/series_breaks.parquet \
inst/extdata/fixture_corpus/manifest.json
git status --short
Expected: clean tree. Report the stash ref to the user so it is not lost.
- Step 3: Branch off current
main
git fetch origin
git checkout -b fix/partial-coverage-signposting-9 origin/main
git log --oneline -1
- Step 4: Record the baseline test count
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -20
Expected: a green suite. Write the exact PASS/FAIL/SKIP/WARN numbers into the task notes — every later task compares against them. Do not proceed if the baseline is already red; report to the user instead.
Task 1: .suppressed_components() + the structural guard
Files:
- Modify:
R/suggestions.R(append after.build_suggestions(), before.attach_ig_counterparts()) - Modify:
R/spending.R:511-513(add.select_long_view()beside.select_view()) - Test:
tests/testthat/test-recipes.R(append at end)
Interfaces:
-
Consumes:
.sql_lit_chr(x)(existing,R/config.R);.select_view(view_base, basis)(existing,R/spending.R:511) -
Produces:
.select_long_view(view_base, basis)→ character(1)."spending_annotated"+"harmonized"→"spending_long_harmonized"..suppressed_components(con, candidates, govid, years, long_view)→ tibble with columnsrecipe_id(chr),year(dbl),suppressed_amount(dbl, full US dollars),suppressed_codes(chr, comma-joined sorted item codes). Zero rows when nothing is suppressed. Aborts on along_viewoutside the allowlist.
-
Step 1: Write the failing tests
Append to tests/testthat/test-recipes.R:
# --- uscogdata#9: partial-coverage signposting ------------------------------
test_that("no recipe component is ever renamed by harmonization", {
# The suppression trigger anti-joins the verb's long view on item_code.
# That is only sound because harmonization never rewrites a recipe
# component's code -- every component whose harmonized_code differs has
# harmonized_code IS NULL (and is aggregate-flagged). If this ever fails,
# .suppressed_components() would report reachable dollars as suppressed.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS renamed FROM long
WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes)
AND harmonized_code IS NOT NULL
AND harmonized_code <> item_code")$renamed
expect_equal(as.integer(n), 0L)
})
test_that(".select_long_view maps annotated view bases to their long views", {
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "harmonized"),
"spending_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("revenue_annotated", "harmonized"),
"revenue_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "raw"),
"spending_long")
})
test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2011L,
long_view = "spending_long_harmonized")
expect_s3_class(s, "tbl_df")
expect_equal(nrow(s), 2L)
s <- s[order(s$recipe_id), ]
expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
expect_equal(s$suppressed_amount, c(1803872000, 271589000))
expect_equal(s$suppressed_codes, c("E67", "E68"))
})
test_that(".suppressed_components finds nothing in a modern year", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2019L,
long_view = "spending_long_harmonized")
expect_equal(nrow(s), 0L)
})
test_that(".suppressed_components rejects a long_view outside the allowlist", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
expect_error(
uscogdata:::.suppressed_components(
con, candidates = "welfare_cash_e67_wide", govid = "061037123085",
years = 2011L, long_view = "long; DROP TABLE x"),
class = "uscogdata_internal_error")
})
- Step 2: Run to verify they fail
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30
Expected: FAIL — .select_long_view and .suppressed_components are not found. The "no recipe component is ever renamed" test should already PASS (it asserts existing corpus structure).
- Step 3: Add
.select_long_view()toR/spending.R
Insert immediately after .select_view() (currently ends at line 513):
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
#' anti-joins the LONG view rather than the annotated one: they have identical
#' row membership (the annotated views are the long views plus LEFT JOINs, see
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
#' the suppression test is asking about.
#' @noRd
.select_long_view <- function(view_base, basis) {
.select_view(sub("_annotated$", "_long", view_base), basis)
}
- Step 4: Add
.suppressed_components()toR/suggestions.R
Insert after .build_suggestions() closes (currently line 132), before the .attach_ig_counterparts() roxygen block:
#' Measure, per (recipe, year), the component dollars this government holds
#' that the calling verb's own long view structurally excludes.
#'
#' This is the second qualifying path for a suggestion (uscogdata#9). The
#' first -- row absence -- only fires when a category returns NOTHING in a
#' requested year, which is how Corrections behaves in the wide era. Public
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
#' 2011, and absent from `summary_categories` entirely) are dropped. The
#' caller gets a plausible number a third too low, silently.
#'
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
#' view rather than restating its WHERE clause, so this stays correct if
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
#' anti-join is keyed on `item_code`, which is sound only because
#' harmonization never renames a recipe component -- asserted by the "no
#' recipe component is ever renamed by harmonization" test in
#' tests/testthat/test-recipes.R.
#'
#' Note what this deliberately does NOT count as suppressed: a component
#' excluded from the RESULT for scoping reasons -- because it belongs to a
#' different `category`, or because `expenditure_concept` narrowed the
#' subtypes -- is still present in the view, so it never fires. Suggesting a
#' recipe is a coverage fix, not a category redefinition. Measured on the
#' bundled fixture, this keeps `higher_ed_e18_wide` and `general_gov_e89_wide`
#' silent (E18/E89 are leaf-and-classified even in the wide era) and confines
#' every fire to 2011.
#'
#' @param con Active DuckDB connection.
#' @param candidates Character vector of recipe ids to measure.
#' @param govid Character vector of canonical_govid values.
#' @param years Integer vector of requested years.
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
#' nothing is suppressed.
#' @noRd
.suppressed_components <- function(con, candidates, govid, years, long_view) {
empty <- tibble::tibble(
recipe_id = character(0), year = numeric(0),
suppressed_amount = numeric(0), suppressed_codes = character(0)
)
if (length(candidates) == 0L) return(empty)
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
# never be quoted safely. It is always internally derived from a fixed
# view_base, so an off-allowlist value is a programming error, not input.
if (!long_view %in% c("spending_long", "spending_long_harmonized",
"revenue_long", "revenue_long_harmonized")) {
cli::cli_abort(
"Internal error: unexpected `long_view` {.val {long_view}}.",
class = "uscogdata_internal_error"
)
}
sql <- sprintf(
"SELECT r.recipe_id,
l.year,
SUM(l.amt) * 1000.0 AS suppressed_amount,
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
AS suppressed_codes
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%1$s)
AND l.canonical_govid IN (%2$s)
AND l.year IN (%3$s)
AND l.amt <> 0
AND NOT EXISTS (
SELECT 1 FROM %4$s v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year
AND v.item_code = l.item_code
)
GROUP BY 1, 2
ORDER BY 1, 2",
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","), long_view
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
- Step 5: Run to verify they pass
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30
Expected: PASS, 0 failures.
- Step 6: Commit
git add R/suggestions.R R/spending.R tests/testthat/test-recipes.R
git commit -m "feat: measure structurally-suppressed recipe component dollars (#9)"
Task 2: Widen .build_suggestions() and add the four fields
Files:
- Modify:
R/suggestions.R:54-132(.build_suggestions()) and its header comment block at lines 1-31 - Modify:
R/spending.R:396-398(call site) - Test:
tests/testthat/test-recipes.R
Interfaces:
-
Consumes:
.suppressed_components(con, candidates, govid, years, long_view)and.select_long_view(view_base, basis)from Task 1 -
Produces:
.build_suggestions(con, govid, years, category, result, basis, flow_prefixes, long_view)— one new trailing argument. Each returned suggestion is a list with the five existing keys (recipe_id,label,available_years,hint,ig_recipe_id) plus:trigger— chr(1),"empty_year"or"suppressed_component"."empty_year"wins when both apply.suppressed_amount— dbl(1), full US dollars summed across requested years,0when none.suppressed_years— integer vector, sorted,integer(0)when none.suppressed_codes— chr vector, sorted unique,character(0)when none.
-
Step 1: Write the failing tests
Append to tests/testthat/test-recipes.R:
test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", {
# The bug: E74/E79 return rows for FY2011, so there is no row-absence gap,
# so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were
# dropped for being aggregate-published. LA County reports $3,185,943,000
# and omits $2,075,461,000, a 39% understatement, silently.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 2L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]]
expect_equal(e67$trigger, "suppressed_component")
expect_equal(e67$suppressed_amount, 1803872000)
expect_equal(e67$suppressed_years, 2011L)
expect_equal(e67$suppressed_codes, "E67")
expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'")
e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]]
expect_equal(e68$trigger, "suppressed_component")
expect_equal(e68$suppressed_amount, 271589000)
expect_equal(e68$suppressed_codes, "E68")
})
test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", {
# Corrections is the case that already worked: zero rows in FY2011, so the
# row-absence path fires. It must keep firing, keep trigger = "empty_year",
# keep its IG counterpart -- and now also report what was suppressed.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Corrections"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 3L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("corrections_combined", "corrections_capital_combined",
"corrections_other_capital_combined"))
expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year"))
cc <- sugg[[which(ids == "corrections_combined")]]
expect_equal(cc$suppressed_amount, 1371460000)
expect_equal(cc$suppressed_codes, "E05")
expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined")
})
test_that("uscogdata#9: the revenue verb inherits the same trigger", {
# Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from
# U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-`
# rents and royalties -- the omission is LARGER than the reported figure.
skip_if_no_corpus()
r <- suppressMessages(
cog_revenue("020000227749", years = 2011L,
category = "Miscellaneous Revenue"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 1L)
expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide")
expect_equal(sugg[[1]]$trigger, "suppressed_component")
expect_equal(sugg[[1]]$suppressed_amount, 1899995000)
expect_equal(sugg[[1]]$suppressed_codes, "U4-")
# A revenue recipe must never be handed an M/L expenditure counterpart.
expect_null(sugg[[1]]$ig_recipe_id)
})
test_that("uscogdata#9: no partial-coverage fire in a modern year", {
skip_if_no_corpus()
r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare")
expect_length(attr(r, "provenance")$suggestions, 0L)
})
test_that("uscogdata#9: leaf-and-classified wide-era families never fire", {
# higher_ed_e18_wide and general_gov_e89_wide are the control group: their
# components (E16/E18, E85/E89) are ordinary classified leaves even in the
# wide era, so widening the trigger must leave them silent. This is the
# measurement that refutes "it would fire on every category in every legacy
# year" -- corpus-wide on the fixture, these two produce zero suppressed rows.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide')
AND l.amt <> 0
AND NOT EXISTS (
SELECT 1 FROM spending_long_harmonized v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year AND v.item_code = l.item_code)")$n
expect_equal(as.integer(n), 0L)
})
- Step 2: Run to verify they fail
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -40
Expected: FAIL — the Public Welfare and revenue tests get length(sugg) == 0; the Corrections test fails on the missing suppressed_amount. The two negative-control tests should already PASS.
- Step 3: Replace the body of
.build_suggestions()
In R/suggestions.R, change the signature and everything from result_years <- through the closing .attach_ig_counterparts(...) call. The candidates query (lines 70-81) is unchanged — do not touch it.
.build_suggestions <- function(con, govid, years, category, result, basis,
flow_prefixes, long_view) {
if (!identical(basis, "harmonized") || is.null(category)) return(list())
# ... candidates query unchanged ...
if (length(candidates) == 0L) return(list())
result_years <- if (is.null(result) || nrow(result) == 0L) {
integer(0)
} else {
unique(as.integer(result$year))
}
gap_years <- setdiff(as.integer(years), result_years)
# Path 2 (uscogdata#9): component dollars this government holds that the
# verb's own view structurally excludes. Measured across ALL requested
# years, not just gap years -- the whole point is that a year with rows can
# still be missing dollars.
supp <- .suppressed_components(con, candidates, govid, years, long_view)
if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
"SELECT recipe_id, any_value(label) AS label,
MIN(year_min) AS year_min, MAX(year_max) AS year_max
FROM harmonization_recipes
WHERE recipe_id IN (%s)
GROUP BY recipe_id",
.sql_lit_chr(candidates)
)))
# Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
# covers for this government, restricted to the gap years.
covered <- if (length(gap_years) == 0L) {
data.frame(recipe_id = character(0), year = integer(0))
} else {
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.year
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%s)
AND l.canonical_govid IN (%s)
AND l.year IN (%s)",
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(gap_years, collapse = ",")
))
}
suggestions <- list()
for (rid in candidates) {
empty_hit <- rid %in% covered$recipe_id
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
supp_hit <- nrow(s_rows) > 0L
if (!empty_hit && !supp_hit) next
m <- meta[meta$recipe_id == rid, ]
suggestions[[length(suggestions) + 1L]] <- list(
recipe_id = rid,
label = m$label[[1]],
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
hint = sprintf("re-run with recipe = '%s'", rid),
# An empty year is the stronger claim -- the category returned nothing
# at all -- so it wins when both paths qualify. The suppressed_* fields
# are still populated, so an empty_year fire also reports its dollars.
trigger = if (empty_hit) "empty_year" else "suppressed_component",
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
suppressed_years = if (supp_hit) {
sort(unique(as.integer(s_rows$year)))
} else {
integer(0)
},
suppressed_codes = if (supp_hit) {
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
} else {
character(0)
}
)
}
.attach_ig_counterparts(con, suggestions, flow_prefixes)
}
- Step 4: Update the file header comment
In R/suggestions.R, replace the first paragraph (lines 1-5) so the file's stated contract matches its behavior:
# R/suggestions.R
# Recipe-component-driven signposting. When a basis = "harmonized" query for
# a category comes back incomplete in some requested year -- and a
# harmonization recipe would actually fill it for this government -- surface
# that recipe as a suggestion. "Incomplete" has two forms, and a recipe
# qualifies on either:
# 1. empty_year -- the result has no rows at all in that year.
# 2. suppressed_component -- the result HAS rows, but a component code
# carries dollars the verb's own long view structurally excludes
# (aggregate-published, or absent from summary_categories). This is
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
# aggregate-only E67/E68, so form 1 never fired and the caller got a
# number a third too low with no signpost at all.
- Step 5: Update the call site in
R/spending.R
At lines 396-398, pass the long view:
suggestions <- .build_suggestions(con, govid, years, category,
direct_leg_result,
resolved$basis, flow_prefixes,
.select_long_view(view_base, resolved$basis))
- Step 6: Run the recipes tests
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30
Expected: PASS, 0 failures.
- Step 7: Run the full suite — this is the regression gate
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -30
Expected: 0 FAIL, and PASS count = Task 0 baseline plus the new tests. The three pre-verified expect_length(suggestions, 0L) assertions (test-recipes.R:196, test-recipes.R:82, test-expenditure-concept.R:293) must still pass. If any fails, stop and report — it means the trigger is firing somewhere the measurement said it would not.
- Step 8: Commit
git add R/suggestions.R R/spending.R tests/testthat/test-recipes.R
git commit -m "fix: signpost aggregate-suppressed components in a category that still has rows (#9)"
Task 3: Surface the dollars in the cli message and cog_explain()
Files:
- Modify:
R/suggestions.R:237-251(.inform_suggestions()) - Modify:
R/explain.R:125-132 - Test:
tests/testthat/test-recipes.R
Interfaces:
-
Consumes: the suggestion fields from Task 2
-
Produces: no new functions;
.inform_suggestions()andcog_explain()rendersuppressed_amount/suppressed_years/suppressed_codeswhensuppressed_amount > 0 -
Step 1: Write the failing tests
Append to tests/testthat/test-recipes.R:
test_that("uscogdata#9: the cli message reports the suppressed dollars", {
skip_if_no_corpus()
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"1,803,872,000", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"FY2011", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"E67", fixed = TRUE)
})
test_that("uscogdata#9: cog_explain() reports the suppressed dollars", {
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
out <- paste(capture.output(cog_explain(r), type = "message"),
capture.output(cog_explain(r)), collapse = "\n")
expect_match(out, "271,589,000", fixed = TRUE)
})
- Step 2: Run to verify they fail
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30
Expected: FAIL — no dollar figures in either output.
- Step 3: Extend
.inform_suggestions()
Replace the body in R/suggestions.R:
.inform_suggestions <- function(suggestions) {
bullets <- vapply(suggestions, function(s) {
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
s$available_years[1], s$available_years[2], s$hint)
# Only present when dollars were actually measured as excluded. An
# empty_year fire can carry them too -- the year had no rows AND the
# component was suppressed -- which is strictly more informative.
if (isTRUE(s$suppressed_amount > 0)) {
bullet <- paste0(bullet, sprintf(
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
if (!is.null(s$ig_recipe_id)) {
bullet <- paste0(bullet, sprintf(
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
}
bullet
}, character(1))
cli::cli_inform(c(
i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
stats::setNames(bullets, rep("*", length(bullets)))
))
}
The header changed from "Coverage gap detected for the requested years" because a partial-coverage fire is not a gap — the year has rows, they are just short. Verified 2026-08-04: no test in
tests/asserts the old string, so this is safe.
- Step 4: Extend the
cog_explain()renderer
Replace lines 125-132 of R/explain.R:
if (length(prov$suggestions) > 0L) {
cli::cli_h2("Suggestions")
sugg_lines <- vapply(prov$suggestions, function(s) {
line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
s$available_years[1], s$available_years[2], s$hint)
if (isTRUE(s$suppressed_amount > 0)) {
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
line
}, character(1))
cli::cli_ul(sugg_lines)
}
- Step 5: Run to verify they pass
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes|explain")' 2>&1 | tail -30
Expected: PASS, 0 failures.
- Step 6: Commit
git add R/suggestions.R R/explain.R tests/testthat/test-recipes.R
git commit -m "feat: report suppressed component dollars in the signpost message (#9)"
Task 4: Document the contract — schema, NEWS, roxygen
Files:
- Modify:
inst/schemas/provenance-v1.json:36 - Modify:
NEWS.md(new section at top, under the# uscogdata 0.1.0 (development)heading) - Test:
tests/testthat/test-recipes.R
Interfaces:
-
Consumes: the suggestion fields from Task 2
-
Produces: no code interfaces; the schema is the published contract cog-api reads against
-
Step 1: Write the failing test
Append to tests/testthat/test-recipes.R:
test_that("the provenance schema documents the suggestion trigger fields", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE)
props <- sch$properties$suggestions$items$properties
expect_true(all(c("trigger", "suppressed_amount", "suppressed_years",
"suppressed_codes") %in% names(props)))
expect_setequal(unlist(props$trigger$enum),
c("empty_year", "suppressed_component"))
})
- Step 2: Run to verify it fails
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -20
Expected: FAIL — sch$properties$suggestions$items is NULL.
- Step 3: Replace the
suggestionsproperty ininst/schemas/provenance-v1.json
Replace line 36 ("suggestions": { "type": "array" },) with:
"suggestions": {
"type": "array",
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
"items": {
"type": "object",
"required": ["recipe_id", "label", "available_years", "hint", "trigger",
"suppressed_amount", "suppressed_years", "suppressed_codes"],
"properties": {
"recipe_id": { "type": "string" },
"label": { "type": "string" },
"available_years": {
"type": "array",
"items": { "type": "integer" },
"description": "[year_min, year_max] of the recipe's component coverage."
},
"hint": { "type": "string" },
"ig_recipe_id": {
"type": ["string", "null"],
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
},
"trigger": {
"type": "string",
"enum": ["empty_year", "suppressed_component"],
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars the verb's long view structurally excludes -- aggregate-published, or absent from summary_categories. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way."
},
"suppressed_amount": {
"type": "number",
"description": "Full US dollars this government holds in the recipe's component codes that the result excludes, summed across the requested years. 0 when nothing is suppressed."
},
"suppressed_years": {
"type": "array",
"items": { "type": "integer" },
"description": "The requested years contributing to suppressed_amount."
},
"suppressed_codes": {
"type": "array",
"items": { "type": "string" },
"description": "The excluded component item codes, sorted."
}
}
}
},
- Step 4: Add the NEWS.md entry
Insert immediately after the # uscogdata 0.1.0 (development) line:
## Signposting now catches partially-suppressed categories
* A coverage suggestion used to fire only when a category returned **no rows
at all** in a requested year. That missed the more dangerous case: a
category that still returns rows while silently dropping component codes
the wide era publishes only as aggregates (#9). `cog_spending(category =
"Public Welfare")` for FY2011 returned a plausible figure that omitted
`E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true
$5,261,404,000, a 39% understatement, with `provenance$suggestions` empty.
* Suggestions now also fire on **partial** coverage, and every suggestion
carries `trigger` (`"empty_year"` or `"suppressed_component"`),
`suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller
can see how much is missing and decide whether to re-run with the recipe.
* `cog_revenue()` gets the same fix through the shared verb path. Alaska's
FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping
$1,899,995,000 of aggregate-published `U4-` rents and royalties.
* The trigger stays recipe-driven, so it only fires where a harmonization
recipe actually exists to name the fix. Measured on the bundled fixture,
every fire lands in the wide era; `higher_ed_e18_wide` and
`general_gov_e89_wide` stay silent, because their components are ordinary
classified leaves even pre-2012.
- Step 5: Regenerate docs and run the full suite
/usr/bin/Rscript -e 'devtools::document()'
/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -30
Expected: 0 FAIL. devtools::document() should produce no man/ changes (all edits are @noRd or non-roxygen); if it does, include them in the commit.
- Step 6: Commit
git add inst/schemas/provenance-v1.json NEWS.md tests/testthat/test-recipes.R man/
git commit -m "docs: document the suggestion trigger and suppressed-dollar fields (#9)"
Task 5: Ship the reader — check, push, PR
Files: none (verification + git)
Interfaces:
-
Consumes: everything from Tasks 1-4
-
Produces: a green CI run and an open PR on
gitea.civilytics.org/Civilytics/uscogdata -
Step 1: Run R CMD check
/usr/bin/Rscript -e 'rcmdcheck::rcmdcheck(args = c("--no-manual", "--as-cran"), error_on = "warning")' 2>&1 | tail -40
Expected: 0 errors, 0 warnings. Notes about fixture size are pre-existing.
- Step 2: Confirm the file-length constraint still holds
wc -l R/suggestions.R R/spending.R R/explain.R
Expected: R/suggestions.R under 400 lines. If it has crossed, split .suppressed_components() into R/suppression.R and re-run the suite before continuing.
- Step 3: Push and open the PR
git push -u origin fix/partial-coverage-signposting-9
tea pr create --title "fix: signpost partially-suppressed categories (#9)" \
--description "Closes #9.
\`.build_suggestions()\` fired only when a category returned zero rows in a requested year. Public Welfare is the failure mode that missed: E74/E79 still return rows for legacy years, so no row-absence gap existed, while aggregate-published E67/E68 were dropped by \`spending_long\`'s \`NOT is_aggregate\` filter. LA County FY2011 reported \$3,185,943,000 and omitted \$2,075,461,000 -- a 39% understatement -- with \`provenance\$suggestions\` empty.
A recipe now also qualifies when a component code carries dollars the verb's own long view structurally excludes, measured per government by anti-joining the real view. Each suggestion carries \`trigger\`, \`suppressed_amount\`, \`suppressed_years\` and \`suppressed_codes\`.
\`cog_revenue()\` inherits it: Alaska FY2011 Miscellaneous Revenue dropped \$1,899,995,000 of \`U4-\`.
Blast radius measured on the bundled fixture: every fire lands in the wide era, none in 2012/2019/2020, and the leaf-and-classified control families (\`higher_ed_e18_wide\`, \`general_gov_e89_wide\`) stay silent."
- Step 4: Wait for CI green, then merge
tea pr list
Report the PR number and CI status to the user. Do not merge without confirming CI is green — merging is what cog-api's next build picks up.
Task 6: cog-api — assert the fields survive the envelope, then deploy
Files:
- Modify:
~/Nextcloud/Civilytics/Code/Civilytics/cog-api/api/tests/testthat/test-handlers-governments.R(append near the existing suggestions test at line 162)
Interfaces:
-
Consumes: the merged reader from Task 5. cog-api needs no handler code change —
prov$suggestionsis lifted verbatim into the envelope atapi/R/handlers_query.R:100andapi/R/handlers_governments.R:131,195,317. -
Produces: a deployed API whose
/spendingresponses carry the new fields -
Step 1: Confirm the reader ref is not pinned to a stale SHA
cog-api's CI and build clone uscogdata at USCOGDATA_REF (a Gitea repo variable, default main — see .gitea/workflows/ci.yml:99-121).
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api
tea api -X GET /repos/Civilytics/cog-api/actions/variables
Expected: USCOGDATA_REF is absent or main. If it is pinned to a commit SHA, stop and report to the user — the deploy will silently ship the old reader, and every assertion below will fail for the wrong reason.
- Step 2: Write the failing test
Append to api/tests/testthat/test-handlers-governments.R:
test_that("suppressed-component suggestion fields survive the envelope (uscogdata#9)", {
# LA County FY2011 Public Welfare returns rows but drops aggregate-published
# E67/E68. The reader signposts it; the envelope must not flatten the detail
# away, because a Tableau consumer only ever sees the envelope.
res <- handle_gov_spending("061037123085", years = "2011",
category = "Public Welfare",
adjust = NULL, base_year = NULL, per_capita = NULL,
subtype = NULL, limit = "1000", page = "0")
ids <- vapply(res$suggestions, function(s) s$recipe_id, character(1))
expect_true("welfare_cash_e67_wide" %in% ids)
hit <- res$suggestions[[which(ids == "welfare_cash_e67_wide")]]
expect_equal(hit$trigger, "suppressed_component")
expect_equal(hit$suppressed_amount, 1803872000)
expect_equal(hit$suppressed_codes, "E67")
})
- Step 3: Install the merged reader locally and run the test
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api
/usr/bin/Rscript -e 'remotes::install_local("../cog_explorer/uscogdata", upgrade = "never")'
cd api/tests/testthat && /usr/bin/Rscript -e 'testthat::test_local(filter = "handlers-governments")' 2>&1 | tail -30
Expected: PASS. If it fails with length(ids) == 0, the installed reader is stale — re-run the install step.
- Step 4: Run the full cog-api suite
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api/api/tests/testthat
/usr/bin/Rscript -e 'testthat::test_local(reporter = "summary")' 2>&1 | tail -30
Expected: 0 FAIL. Record the count.
- Step 5: Commit, push, PR
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api
git checkout -b test/suppressed-component-fields
git add api/tests/testthat/test-handlers-governments.R
git commit -m "test: assert suppressed-component suggestion fields survive the envelope (uscogdata#9)"
git push -u origin test/suppressed-component-fields
tea pr create --title "test: assert uscogdata#9 suggestion fields survive the envelope" \
--description "uscogdata#9 adds \`trigger\`, \`suppressed_amount\`, \`suppressed_years\` and \`suppressed_codes\` to each \`prov\$suggestions\` entry. cog-api lifts \`prov\$suggestions\` verbatim, so no handler change is needed -- this pins that pass-through so a future envelope refactor cannot silently drop the detail.
Merging this is the prod deploy that picks up the new reader."
- Step 6: Merge after CI green, then verify the live endpoint
Merging a cog-api PR is the production deploy, and
deploy.ymldoes not wait for CI. Confirm CI is green before merging, not after.
tea pr list
# after merge + deploy settles:
curl -s "https://<cog-api-host>/v1/governments/061037123085/spending?years=2011&category=Public%20Welfare" \
| python3 -c "import json,sys; d=json.load(sys.stdin); print(json.dumps(d['suggestions'], indent=2))"
Expected: two suggestions, welfare_cash_e67_wide carrying "suppressed_amount": 1803872000. Report the live output to the user.
Task 7: Close out the issues
Files: none (issue tracker)
Interfaces:
-
Consumes: the merged and deployed work from Tasks 5-6
-
Produces: uscogdata#9 closed with a corrected record
-
Step 1: Comment on uscogdata#9 with the resolution
cd ~/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata
tea comment 9 "**Fixed.** The trigger was the problem, not the crosswalk.
\`.build_suggestions()\` already identified \`welfare_cash_e67_wide\` and \`welfare_cash_e68_wide\` as candidates -- \`J67\`/\`J68\` are Public Welfare members, so the candidate query matched. It then threw them away at the \`gap_years\` early return, because E74/E79 returned rows and there was no empty year to detect.
A recipe now also qualifies when a component code carries dollars the verb's own long view structurally excludes -- aggregate-published, or absent from \`summary_categories\`. Reachability is decided by anti-joining the real view rather than restating its WHERE clause, which is sound because no recipe component is ever renamed by harmonization (now asserted).
Each suggestion carries \`trigger\`, \`suppressed_amount\`, \`suppressed_years\`, \`suppressed_codes\`. \`cog_revenue()\` inherits the fix.
**On the 'would fire on every category in every legacy year' worry** -- measured, it does not. Candidates stay gated on the 24-recipe catalog, so it only fires where an actionable recipe exists. On the bundled fixture every fire lands in the wide era, none in 2012/2019/2020, and \`higher_ed_e18_wide\`/\`general_gov_e89_wide\` never fire at all because E18/E89 are ordinary classified leaves pre-2012.
**The 'minor' sub-item is withdrawn, not implemented.** \`aggregate_fallback\` is no longer vestigial: \`.build_verb_sql()\` uses \`bool_or(is_aggregate)\` (R/spending.R:611), and \`ig_long\` deliberately keeps aggregate rows, so the flag is live for the intergovernmental leg -- \`test-expenditure-concept.R:99\` asserts aggregate-sourced IG dollars report TRUE. Removing it would break that disclosure.
**Reproduction, now signposted** (bundled fixture, LA County):
\`\`\`r
r <- cog_spending(\"061037123085\", years = 2011, category = \"Public Welfare\")
attr(r, \"provenance\")\$suggestions[[1]]\$suppressed_amount # 1803872000
\`\`\`"
- Step 2: Close the issue
tea issue close 9
tea issue 9 | head -5
Expected: state closed.
- Step 3: Restore the stashed fixture changes
git checkout ci/apt-https
git stash list
git stash pop
git status --short
Expected: the two fixture files modified again, as they were before Task 0. Report to the user that they are restored and still uncommitted.
Self-Review
Spec coverage — every element of the approved design maps to a task:
| Design element | Task |
|---|---|
Second trigger in .build_suggestions(), candidate query untouched |
2 |
| Anti-join the real view, not a restated predicate | 1 |
| Structural invariant (no renamed components) asserted | 1 |
| Scoping-excluded ≠ suppressed; F-007 stays out | 1 (roxygen), 2 (control test) |
Four new suggestion fields, empty_year precedence |
2 |
| Both verbs via the shared call site | 2 (revenue test) |
cli message + cog_explain() |
3 |
| Provenance schema documents the fields (no v1 bump) | 4 |
| LA County / Alaska / Corrections / 2019 / control tests | 2, 3 |
| cog-api pass-through + redeploy | 6 |
Close #9, withdraw the aggregate_fallback sub-item |
7 |
Placeholder scan — no TBD/TODO; every code step carries real code; every assertion carries a measured value.
Type consistency — .suppressed_components() returns recipe_id/year/suppressed_amount/suppressed_codes in Task 1 and is consumed under exactly those names in Task 2. .select_long_view(view_base, basis) is defined in Task 1 and called with that signature in Task 2. suppressed_codes is comma-joined inside the tibble (Task 1) and split into a character vector for the suggestion (Task 2) — deliberate, and the tests assert both shapes correctly.