Closes#6. Repo 2 of 3 for the "total spending" semantics work; follows pipeline PR #59 (merged). Design + owner rulings: cog_pipeline/docs/superpowers/specs/2026-07-25-total-spending-semantics-design.md. Plan: .superpowers/plans/2026-07-27-expenditure-concept-reader.md.
What this does
Adds expenditure_concept = c("direct", "total") to cog_spending().
"direct" is the default and is numerically identical to prior behavior — verified by comparing 578 result rows cell-for-cell against the pre-change commit: 0 differences.
"total" adds the government's intergovernmental payments — M (to local governments) plus L (to state governments) — as a second leg, surfacing as a third spend_subtype, "intergovernmental".
The cross-government verbs refuse "total" (owner ruling R1). cog_geographic_rollup() and cog_peer_compare() abort with class uscogdata_concept_not_aggregatable and a message that names the fix and explains the mechanism.
Also: provenance records the concept; the recipe-suggestion machinery names the intergovernmental counterpart; the inert "K" flow prefix is dropped (it matches zero rows corpus-wide).
The two facts the IG leg rests on
Both measured against the published corpus, both now pinned by tests.
1. Aggregate IG rows carry no harmonized_code — 2,918,024 M rows and 2,188,518 L rows. Harmonized space is leaf-only by construction. So the IG view joins on COALESCE(harmonized_code, item_code); a plain IS NOT NULL filter would drop every legacy IG aggregate.
2. Aggregate IG codes are year-disjoint from their components, so the IG view can safely omit NOT is_aggregate. M47 appears through 2011 with M94 absent; from 2012 the reverse. Same shape for M89/M91-93, L89/L91-93, M05/M04. Under the naive filter only 29.9% of legacy IG dollars survive. L-- — the IG-to-state family total — is excluded, and a reviewer confirmed it equals the sum of every L-NN code exactly, per government, across all 6,422 governments in 2011.
A data-invariant test now fails loudly if a future corpus ever back-fills components into an aggregate's year.
Guards
Four, composing without overlap, each with a typed condition class:
verb-level refusal in the two cross-government verbs;
view_base guard inside the shared .verb_spendrev(), so cog_revenue() can never receive the expenditure IG leg;
recipe + "total" conflict abort;
a corpus precondition check that aborts if summary_categories predates pipeline #59, rather than silently returning $0.
What the reviews caught
Two Criticals that the per-task reviews missed, both found by the whole-branch review and both fixed:
The two legs had asymmetric era coverage. The IG leg skips NOT is_aggregate; the Direct leg keeps it. In the wide era that returned an IG-only figure that looked like a total — AL Corrections FY2011 reported $31,358,000 against a true Direct of $521,651,000 — and because gap detection ran on the UNION'd result, the surviving IG rows cancelled the coverage-gap suggestion that would have caught it. Gap detection now runs on the Direct leg only, and affected rows carry a note naming the recovery recipe plus a testable provenance flag.
No precondition check on the corpus. Against a pre-#59 corpus, "total" returned $0 for a $5.15B figure, or collapsed $6.8B into an NA-subtype row invisible to the filter every test and doc uses. Now a typed abort.
A follow-up round then fixed the suppression flag firing on genuinely-zero Direct legs: it was TRUE for 32 of 50 states on a clean FY2019 query, describing correct modern data as a legacy artifact. Both the flag and the note are now gated on a recipe actually covering the exact (year, govid, category) triple. Clean FY2019 now flags 0 of 50; across {2011, 2019} the 50 true positives all name a recipe, where previously 141 of 141 said none was found.
Verification
Suite: PASS 576 | FAIL 0 | SKIP 0 | WARN 0 (baseline 467 on the pre-branch fixture).
Fixture regenerated from the merged pipeline publish tree — 194 → 260 category rows.
Both vignettes knit offline; every chunk executes, including a live error = TRUE chunk showing the real refusal.
DESCRIPTION and NAMESPACE unchanged.
Mutation testing confirmed each invariant's test genuinely discriminates.
Notes for review
cog_categories() now surfaces the intergovernmental subtype. That was an adjudicated call, not an accident: it is a discovery verb, and after this change users see that subtype in results, so hiding it would misreport the data model. The category values themselves are unchanged.
main was stale locally when this branch was cut, so origin/main (PR #8) is merged in as 54ece11 rather than rebased — rebasing would have rewritten the commit SHAs the task reviews verified against. The merge is conflict-free and touches nothing this branch modifies.
Known follow-ups, none blocking: the README's vignette() pointer will not resolve in an installed package (^vignettes$ is in .Rbuildignore); the vignette hard-codes the fixture path that a documented release step would .Rbuildignore; cog_spending() accepts a govid vector with "total" and warns only in roxygen, not at runtime.
Please review; not merging without your sign-off.
Closes #6. Repo 2 of 3 for the "total spending" semantics work; follows pipeline PR #59 (merged). Design + owner rulings: `cog_pipeline/docs/superpowers/specs/2026-07-25-total-spending-semantics-design.md`. Plan: `.superpowers/plans/2026-07-27-expenditure-concept-reader.md`.
## What this does
Adds **`expenditure_concept = c("direct", "total")`** to `cog_spending()`.
- **`"direct"`** is the default and is numerically identical to prior behavior — verified by comparing 578 result rows cell-for-cell against the pre-change commit: **0 differences**.
- **`"total"`** adds the government's intergovernmental payments — `M` (to local governments) plus `L` (to state governments) — as a second leg, surfacing as a third `spend_subtype`, `"intergovernmental"`.
**The cross-government verbs refuse `"total"`** (owner ruling R1). `cog_geographic_rollup()` and `cog_peer_compare()` abort with class `uscogdata_concept_not_aggregatable` and a message that names the fix and explains the mechanism.
Also: provenance records the concept; the recipe-suggestion machinery names the intergovernmental counterpart; the inert `"K"` flow prefix is dropped (it matches zero rows corpus-wide).
## The two facts the IG leg rests on
Both measured against the published corpus, both now pinned by tests.
**1. Aggregate IG rows carry no `harmonized_code`** — 2,918,024 `M` rows and 2,188,518 `L` rows. Harmonized space is leaf-only by construction. So the IG view joins on `COALESCE(harmonized_code, item_code)`; a plain `IS NOT NULL` filter would drop every legacy IG aggregate.
**2. Aggregate IG codes are year-disjoint from their components**, so the IG view can safely omit `NOT is_aggregate`. `M47` appears through 2011 with `M94` absent; from 2012 the reverse. Same shape for `M89`/`M91-93`, `L89`/`L91-93`, `M05`/`M04`. Under the naive filter only **29.9%** of legacy IG dollars survive. `L--` — the IG-to-state family total — is excluded, and a reviewer confirmed it equals the sum of every `L-NN` code **exactly, per government, across all 6,422 governments** in 2011.
A data-invariant test now fails loudly if a future corpus ever back-fills components into an aggregate's year.
## Guards
Four, composing without overlap, each with a typed condition class:
- verb-level refusal in the two cross-government verbs;
- `view_base` guard inside the shared `.verb_spendrev()`, so `cog_revenue()` can never receive the expenditure IG leg;
- `recipe` + `"total"` conflict abort;
- a corpus precondition check that aborts if `summary_categories` predates pipeline #59, rather than silently returning $0.
## What the reviews caught
Two Criticals that the per-task reviews missed, both found by the whole-branch review and both fixed:
- **The two legs had asymmetric era coverage.** The IG leg skips `NOT is_aggregate`; the Direct leg keeps it. In the wide era that returned an IG-only figure that looked like a total — AL Corrections FY2011 reported $31,358,000 against a true Direct of $521,651,000 — and because gap detection ran on the UNION'd result, the surviving IG rows cancelled the coverage-gap suggestion that would have caught it. Gap detection now runs on the Direct leg only, and affected rows carry a note naming the recovery recipe plus a testable provenance flag.
- **No precondition check on the corpus.** Against a pre-#59 corpus, `"total"` returned $0 for a $5.15B figure, or collapsed $6.8B into an `NA`-subtype row invisible to the filter every test and doc uses. Now a typed abort.
A follow-up round then fixed the suppression flag firing on genuinely-zero Direct legs: it was TRUE for 32 of 50 states on a clean FY2019 query, describing correct modern data as a legacy artifact. Both the flag and the note are now gated on a recipe actually covering the exact `(year, govid, category)` triple. Clean FY2019 now flags **0 of 50**; across `{2011, 2019}` the 50 true positives all name a recipe, where previously 141 of 141 said none was found.
## Verification
- **Suite: `PASS 576 | FAIL 0 | SKIP 0 | WARN 0`** (baseline 467 on the pre-branch fixture).
- Fixture regenerated from the merged pipeline publish tree — 194 → 260 category rows.
- Both vignettes knit offline; every chunk executes, including a live `error = TRUE` chunk showing the real refusal.
- `DESCRIPTION` and `NAMESPACE` unchanged.
- Mutation testing confirmed each invariant's test genuinely discriminates.
## Notes for review
`cog_categories()` now surfaces the `intergovernmental` subtype. That was an adjudicated call, not an accident: it is a discovery verb, and after this change users see that subtype in results, so hiding it would misreport the data model. The `category` values themselves are unchanged.
`main` was stale locally when this branch was cut, so `origin/main` (PR #8) is merged in as `54ece11` rather than rebased — rebasing would have rewritten the commit SHAs the task reviews verified against. The merge is conflict-free and touches nothing this branch modifies.
Known follow-ups, none blocking: the README's `vignette()` pointer will not resolve in an installed package (`^vignettes$` is in `.Rbuildignore`); the vignette hard-codes the fixture path that a documented release step would `.Rbuildignore`; `cog_spending()` accepts a govid vector with `"total"` and warns only in roxygen, not at runtime.
Please review; not merging without your sign-off.
Seven TDD tasks. Records the two measured facts the design rests on: aggregate
IG rows carry no harmonized_code (so the IG leg must COALESCE item_code), and
aggregate IG codes are year-disjoint from their modern leaf components (so
skipping NOT is_aggregate cannot double-count). Baseline measured at PASS 467.
Plan lives under .superpowers/ because docs/ is the gitignored pkgdown output
dir; .Rbuildignore'd so it never ships in the package tarball.
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).
cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.
Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
Nine review items on the expenditure_concept = direct|total feature:
- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
wide-era data puts a $0 leaf row in the same group as the real aggregate
row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
scoped to the aggregate flag rather than bare code presence (M89/L89
continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
retained, NULL harmonized_code coalesced, L-- excluded), rather than
relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
necessary (not just incidental) against a corpus whose `long` genuinely
lacks a harmonized_code column, and rewrote the misleading "v5-only
parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
test, guarded .verb_spendrev() against total on a non-spending
view_base, added a roxygen caveat against summing total across levels
of government, and replaced an uncheckable corpus-wide SQL comment
figure with a fixture-verifiable one.
Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
Owner ruling R1. Combining Census Total across governments counts
intergovernmental transfers twice, and these results land in Tableau where a
warning would be invisible -- so this is a hard error whose message names the
fix and the reason.
- 'both cross-government verbs still accept the direct default' now tests both verbs
- 'the refusal message names the fix and the reason' now asserts both functions name
themselves correctly in their error messages (cog_geographic_rollup vs cog_peer_compare)
Addresses coordinator feedback to prevent test coverage gaps and ensure the helper's
verb name argument is pinned correctly.
Review found the suffix-set match alone is unsafe: revenue-side recipes
(ig_federal_b47_wide, ig_state_c47_wide, ig_local_d47_wide, and their *_89
siblings) coincidentally share exact suffix sets with M/L expenditure
recipes despite representing a different flow direction. Reachable today via
a mis-scoped cog_spending(category = "IG Federal") call, not just
cog_revenue(). Thread flow_prefixes (same parameter .build_harmonization_block
already uses) through .build_suggestions()/.attach_ig_counterparts() and
require a firing recipe's own prefixes to be both in the calling verb's flow
family and within {E,F,G} before searching the M/L catalog.
Task 7 (final) of the expenditure_concept plan. The vignette leads with the
two archetype questions -- a single government's own trend (either concept
works, held fixed across years) vs a cross-government rollup (direct only,
with the refusal error from cog_geographic_rollup() shown and explained) --
walked through with code that runs against the bundled fixture corpus
(years 2011/2012/2019/2020, substituting for "2017 vs today"). Explains the
double-counting mechanism (a state's M44 payment to a county is the same
dollar as the county's own E44/F44), why Total = Direct + M + L rather than
Direct + M, and the composition rules (expenditure_concept is orthogonal to
basis, mutually exclusive with recipe). README gets a short pointer section
with the one-line rule.
Local main was stale at fa40266 when this branch was created, so it was
missing 748ca4a. Merging rather than rebasing to preserve the reviewed
commit SHAs recorded in the SDD ledger.
C1: spending_long/spending_long_harmonized filter NOT is_aggregate but
ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a
legacy aggregate-only family (e.g. Corrections pre-2012) can survive on
the IG leg while Direct is suppressed. expenditure_concept = "total"
then UNIONs an IG-only figure that reads as a plausible Total, and the
coverage-gap suggestion machinery -- fed the UNION'd result -- saw the
surviving IG row as coverage and stayed silent.
(a) .build_suggestions() is now fed a Direct-leg-only view of the
result (IG rows filtered out before the gap-years computation),
so the recipe hints fire for "total" exactly as they do for
"direct".
(b) Any row where IG has dollars but Direct has none for the same
(year, canonical_govid, category) is now flagged: the row's
`notes` name the recovering recipe (drawn from the Direct-leg
suggestions), and provenance gains an explicit
`expenditure_concept_direct_suppressed` boolean plus an appended
warning on `expenditure_concept_note` -- both cheap for a
downstream consumer (cog-api passes provenance through verbatim)
to test, rather than silently asserting Direct + IG when that
arithmetic didn't happen.
Measured before/after on AL state government, Corrections, 2011:
"total" already correctly returns the corpus's actual IG-only figure
($31,358,000, vs. true Direct of $521,651,000 via recipe =
"corrections_combined"), but before this fix it did so with 0
suggestions and an unqualified "Total = Direct + IG" note; after, it
fires 3 recipe hints and both the row notes and provenance say plainly
that Direct is unavailable through this basis.
C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59
with no schema_version bump, so schema_version can't gate "total" --
a pre-#59 corpus can report any supported schema_version and still
have zero M/L category rows, in which case ig_annotated's LEFT JOIN
silently produces NA category/spend_subtype (0 rows for a specific
category, or one invisible NA-subtype group for category = NULL). New
.require_ig_categories() checks summary_categories directly and aborts
with class uscogdata_ig_categories_unsupported, naming PR #59 and
directing the user to a newer corpus.
Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose
synthetic summary_categories carries only one E36 row) by asserting
the new guard fires against that same connection, rather than leaving
the two silently contradictory.
.build_suggestions()'s candidate query picks recipes by component_code
matching the requested category's summary_categories rows, with no
flow-prefix filter. Task 1's M04/M05 category rows share the
"Corrections" category with the Direct-flavored E04/E05, so
corrections_ig_local_combined (entirely M-prefixed) became a raw
top-level candidate for a plain (Direct) cog_spending() call.
Following that hint would silently return intergovernmental dollars
under provenance$expenditure_concept = "direct".
Task 6's flow-family gate in .attach_ig_counterparts() already protects
the *counterpart* lookup (deciding whether a firing suggestion gets an
ig_recipe_id attached) but never touched the candidate list itself.
Exclude any recipe with an M/L-prefixed component from candidates
unconditionally -- an M/L recipe should never be a coverage-gap filler
for either verb, which is a stronger guarantee than the counterpart
gate's flow_prefixes check.
Confirmed via the full suite: before this fix, a Direct cog_spending()
call for category = "Corrections" printed "corrections_ig_local_combined
... re-run with recipe = 'corrections_ig_local_combined'" as its own
suggestion; after, it appears only as the "intergovernmental
counterpart" annotation on corrections_combined and its capital-outlay
siblings. The pre-existing "IG Federal" mis-scoped test (revenue-side
B-prefixed recipes) is unaffected -- those aren't M/L, so they remain
valid candidates with ig_recipe_id still gated to NULL.
Adds regression coverage for the three preceding fixes:
- "total" on a legacy aggregate-only family (AL Corrections 2011) now
fires recipe suggestions, flags provenance$expenditure_concept_direct_
suppressed, and names a recovering recipe in the affected row's notes,
plus a contrast test confirming the flag stays FALSE when the Direct
leg is present.
- expenditure_concept = "total" aborts with class
uscogdata_ig_categories_unsupported against a corpus whose
summary_categories carries no M/L rows (new
with_corpus_missing_ig_categories() fixture helper), is unaffected for
"direct" on the same corpus, and still works on a corpus that does
carry M/L rows.
- an M/L recipe (corrections_ig_local_combined) no longer appears as a
raw suggestion for a Direct-flavored cog_spending() call.
.print_provenance() printed "Basis:" but nothing about Direct vs Total
-- the most consequential switch this branch adds to cog_spending() was
invisible in the package's designated "what am I looking at" verb. Add
a "Concept: direct|total (<note>)" line next to Basis, and surface a
cli warning when provenance$expenditure_concept_direct_suppressed is
TRUE (see the C1 fix), so the suppressed-Direct case is visible in the
human-readable explain output, not just in the structured provenance.
M5: README.md described the bundled fixture as a "3.6 MB two-year slice
(2019 + 2020)"; it's now a 15 MB four-year slice (2011, 2012, 2019,
2020), matching the regenerated fixture and the vignette's own
description.
M2: total-spending.Rmd cited County 1.8% / City 0.8% intergovernmental-
to-Direct and County 91.6% L/M, all roughly 2x off against the bundled
fixture. Measured directly against the fixture (all 50 states, each of
its four years): County IG/Direct 3.4%-5.1%, City IG/Direct 2.6%-3.1%,
County L/M 43%-51% (all varying by year). State 17.2%, AL 7.6%,
national 11.6%, and City L/M 188.3% were re-checked and left as-is.
M1: the "Why Total = Direct + M + L" paragraph described money a local
government *receives* and the state "redistributing as M" -- backwards.
M and L are both the *queried* government's own payments *out*: M to
other local governments, L up to its state. Rewrote the explanation;
the conclusion and non-M2-flagged figures are unchanged.
.detect_direct_suppressed() equated "no Direct sibling row" with "Direct
was suppressed", but the dominant real cause is a government with
genuinely no direct spending in that category (e.g. a state funding K-12
entirely through school districts) -- correct, ordinary data, not
suppression. Measured: 32 of 50 states false-flagged on a clean FY2019
category = NULL total query, and all 141 flagged rows across 50 states x
{2011, 2019} fell back to "no covering recipe found" instead of naming one
-- including AL Corrections, which names corrections_combined correctly
when category is supplied explicitly.
Both the flag and its row note are now gated on a harmonization recipe
actually covering that exact (year, canonical_govid, category) triple, via
a new .covering_recipes() helper that runs the same generic recipe join
per-row regardless of whether the caller supplied a category filter.
.notes_column() takes the precomputed note vector directly instead of
searching a category-gated suggestions list; .direct_suppressed_note() is
removed (its "no recipe found" fallback no longer applies -- if no recipe
covers a triple, it isn't suppression).
Also recomputes two total-spending.Rmd figures the prior wave never
actually reconciled with its own "measured against the fixture" caption:
State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M
(flat 188.3%, now 144%-189% varying by year).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jared
merged commit 1f257812b6 into main2026-07-27 13:13:02 -04:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes #6. Repo 2 of 3 for the "total spending" semantics work; follows pipeline PR #59 (merged). Design + owner rulings:
cog_pipeline/docs/superpowers/specs/2026-07-25-total-spending-semantics-design.md. Plan:.superpowers/plans/2026-07-27-expenditure-concept-reader.md.What this does
Adds
expenditure_concept = c("direct", "total")tocog_spending()."direct"is the default and is numerically identical to prior behavior — verified by comparing 578 result rows cell-for-cell against the pre-change commit: 0 differences."total"adds the government's intergovernmental payments —M(to local governments) plusL(to state governments) — as a second leg, surfacing as a thirdspend_subtype,"intergovernmental".The cross-government verbs refuse
"total"(owner ruling R1).cog_geographic_rollup()andcog_peer_compare()abort with classuscogdata_concept_not_aggregatableand a message that names the fix and explains the mechanism.Also: provenance records the concept; the recipe-suggestion machinery names the intergovernmental counterpart; the inert
"K"flow prefix is dropped (it matches zero rows corpus-wide).The two facts the IG leg rests on
Both measured against the published corpus, both now pinned by tests.
1. Aggregate IG rows carry no
harmonized_code— 2,918,024Mrows and 2,188,518Lrows. Harmonized space is leaf-only by construction. So the IG view joins onCOALESCE(harmonized_code, item_code); a plainIS NOT NULLfilter would drop every legacy IG aggregate.2. Aggregate IG codes are year-disjoint from their components, so the IG view can safely omit
NOT is_aggregate.M47appears through 2011 withM94absent; from 2012 the reverse. Same shape forM89/M91-93,L89/L91-93,M05/M04. Under the naive filter only 29.9% of legacy IG dollars survive.L--— the IG-to-state family total — is excluded, and a reviewer confirmed it equals the sum of everyL-NNcode exactly, per government, across all 6,422 governments in 2011.A data-invariant test now fails loudly if a future corpus ever back-fills components into an aggregate's year.
Guards
Four, composing without overlap, each with a typed condition class:
view_baseguard inside the shared.verb_spendrev(), socog_revenue()can never receive the expenditure IG leg;recipe+"total"conflict abort;summary_categoriespredates pipeline #59, rather than silently returning $0.What the reviews caught
Two Criticals that the per-task reviews missed, both found by the whole-branch review and both fixed:
NOT is_aggregate; the Direct leg keeps it. In the wide era that returned an IG-only figure that looked like a total — AL Corrections FY2011 reported $31,358,000 against a true Direct of $521,651,000 — and because gap detection ran on the UNION'd result, the surviving IG rows cancelled the coverage-gap suggestion that would have caught it. Gap detection now runs on the Direct leg only, and affected rows carry a note naming the recovery recipe plus a testable provenance flag."total"returned $0 for a $5.15B figure, or collapsed $6.8B into anNA-subtype row invisible to the filter every test and doc uses. Now a typed abort.A follow-up round then fixed the suppression flag firing on genuinely-zero Direct legs: it was TRUE for 32 of 50 states on a clean FY2019 query, describing correct modern data as a legacy artifact. Both the flag and the note are now gated on a recipe actually covering the exact
(year, govid, category)triple. Clean FY2019 now flags 0 of 50; across{2011, 2019}the 50 true positives all name a recipe, where previously 141 of 141 said none was found.Verification
PASS 576 | FAIL 0 | SKIP 0 | WARN 0(baseline 467 on the pre-branch fixture).error = TRUEchunk showing the real refusal.DESCRIPTIONandNAMESPACEunchanged.Notes for review
cog_categories()now surfaces theintergovernmentalsubtype. That was an adjudicated call, not an accident: it is a discovery verb, and after this change users see that subtype in results, so hiding it would misreport the data model. Thecategoryvalues themselves are unchanged.mainwas stale locally when this branch was cut, soorigin/main(PR #8) is merged in as54ece11rather than rebased — rebasing would have rewritten the commit SHAs the task reviews verified against. The merge is conflict-free and touches nothing this branch modifies.Known follow-ups, none blocking: the README's
vignette()pointer will not resolve in an installed package (^vignettes$is in.Rbuildignore); the vignette hard-codes the fixture path that a documented release step would.Rbuildignore;cog_spending()accepts a govid vector with"total"and warns only in roxygen, not at runtime.Please review; not merging without your sign-off.
Review found the suffix-set match alone is unsafe: revenue-side recipes (ig_federal_b47_wide, ig_state_c47_wide, ig_local_d47_wide, and their *_89 siblings) coincidentally share exact suffix sets with M/L expenditure recipes despite representing a different flow direction. Reachable today via a mis-scoped cog_spending(category = "IG Federal") call, not just cog_revenue(). Thread flow_prefixes (same parameter .build_harmonization_block already uses) through .build_suggestions()/.attach_ig_counterparts() and require a firing recipe's own prefixes to be both in the calling verb's flow family and within {E,F,G} before searching the M/L catalog.C1: spending_long/spending_long_harmonized filter NOT is_aggregate but ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a legacy aggregate-only family (e.g. Corrections pre-2012) can survive on the IG leg while Direct is suppressed. expenditure_concept = "total" then UNIONs an IG-only figure that reads as a plausible Total, and the coverage-gap suggestion machinery -- fed the UNION'd result -- saw the surviving IG row as coverage and stayed silent. (a) .build_suggestions() is now fed a Direct-leg-only view of the result (IG rows filtered out before the gap-years computation), so the recipe hints fire for "total" exactly as they do for "direct". (b) Any row where IG has dollars but Direct has none for the same (year, canonical_govid, category) is now flagged: the row's `notes` name the recovering recipe (drawn from the Direct-leg suggestions), and provenance gains an explicit `expenditure_concept_direct_suppressed` boolean plus an appended warning on `expenditure_concept_note` -- both cheap for a downstream consumer (cog-api passes provenance through verbatim) to test, rather than silently asserting Direct + IG when that arithmetic didn't happen. Measured before/after on AL state government, Corrections, 2011: "total" already correctly returns the corpus's actual IG-only figure ($31,358,000, vs. true Direct of $521,651,000 via recipe = "corrections_combined"), but before this fix it did so with 0 suggestions and an unqualified "Total = Direct + IG" note; after, it fires 3 recipe hints and both the row notes and provenance say plainly that Direct is unavailable through this basis. C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59 with no schema_version bump, so schema_version can't gate "total" -- a pre-#59 corpus can report any supported schema_version and still have zero M/L category rows, in which case ig_annotated's LEFT JOIN silently produces NA category/spend_subtype (0 rows for a specific category, or one invisible NA-subtype group for category = NULL). New .require_ig_categories() checks summary_categories directly and aborts with class uscogdata_ig_categories_unsupported, naming PR #59 and directing the user to a newer corpus. Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose synthetic summary_categories carries only one E36 row) by asserting the new guard fires against that same connection, rather than leaving the two silently contradictory..detect_direct_suppressed() equated "no Direct sibling row" with "Direct was suppressed", but the dominant real cause is a government with genuinely no direct spending in that category (e.g. a state funding K-12 entirely through school districts) -- correct, ordinary data, not suppression. Measured: 32 of 50 states false-flagged on a clean FY2019 category = NULL total query, and all 141 flagged rows across 50 states x {2011, 2019} fell back to "no covering recipe found" instead of naming one -- including AL Corrections, which names corrections_combined correctly when category is supplied explicitly. Both the flag and its row note are now gated on a harmonization recipe actually covering that exact (year, canonical_govid, category) triple, via a new .covering_recipes() helper that runs the same generic recipe join per-row regardless of whether the caller supplied a category filter. .notes_column() takes the precomputed note vector directly instead of searching a category-gated suggestions list; .direct_suppressed_note() is removed (its "no recipe found" fallback no longer applies -- if no recipe covers a triple, it isn't suppression). Also recomputes two total-spending.Rmd figures the prior wave never actually reconciled with its own "measured against the fixture" caption: State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M (flat 188.3%, now 144%-189% varying by year). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>