.detect_direct_suppressed() equated "no Direct sibling row" with "Direct
was suppressed", but the dominant real cause is a government with
genuinely no direct spending in that category (e.g. a state funding K-12
entirely through school districts) -- correct, ordinary data, not
suppression. Measured: 32 of 50 states false-flagged on a clean FY2019
category = NULL total query, and all 141 flagged rows across 50 states x
{2011, 2019} fell back to "no covering recipe found" instead of naming one
-- including AL Corrections, which names corrections_combined correctly
when category is supplied explicitly.
Both the flag and its row note are now gated on a harmonization recipe
actually covering that exact (year, canonical_govid, category) triple, via
a new .covering_recipes() helper that runs the same generic recipe join
per-row regardless of whether the caller supplied a category filter.
.notes_column() takes the precomputed note vector directly instead of
searching a category-gated suggestions list; .direct_suppressed_note() is
removed (its "no recipe found" fallback no longer applies -- if no recipe
covers a triple, it isn't suppression).
Also recomputes two total-spending.Rmd figures the prior wave never
actually reconciled with its own "measured against the fixture" caption:
State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M
(flat 188.3%, now 144%-189% varying by year).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds regression coverage for the three preceding fixes:
- "total" on a legacy aggregate-only family (AL Corrections 2011) now
fires recipe suggestions, flags provenance$expenditure_concept_direct_
suppressed, and names a recovering recipe in the affected row's notes,
plus a contrast test confirming the flag stays FALSE when the Direct
leg is present.
- expenditure_concept = "total" aborts with class
uscogdata_ig_categories_unsupported against a corpus whose
summary_categories carries no M/L rows (new
with_corpus_missing_ig_categories() fixture helper), is unaffected for
"direct" on the same corpus, and still works on a corpus that does
carry M/L rows.
- an M/L recipe (corrections_ig_local_combined) no longer appears as a
raw suggestion for a Direct-flavored cog_spending() call.
Review found the suffix-set match alone is unsafe: revenue-side recipes
(ig_federal_b47_wide, ig_state_c47_wide, ig_local_d47_wide, and their *_89
siblings) coincidentally share exact suffix sets with M/L expenditure
recipes despite representing a different flow direction. Reachable today via
a mis-scoped cog_spending(category = "IG Federal") call, not just
cog_revenue(). Thread flow_prefixes (same parameter .build_harmonization_block
already uses) through .build_suggestions()/.attach_ig_counterparts() and
require a firing recipe's own prefixes to be both in the calling verb's flow
family and within {E,F,G} before searching the M/L catalog.
- 'both cross-government verbs still accept the direct default' now tests both verbs
- 'the refusal message names the fix and the reason' now asserts both functions name
themselves correctly in their error messages (cog_geographic_rollup vs cog_peer_compare)
Addresses coordinator feedback to prevent test coverage gaps and ensure the helper's
verb name argument is pinned correctly.
Owner ruling R1. Combining Census Total across governments counts
intergovernmental transfers twice, and these results land in Tableau where a
warning would be invisible -- so this is a hard error whose message names the
fix and the reason.
Nine review items on the expenditure_concept = direct|total feature:
- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
wide-era data puts a $0 leaf row in the same group as the real aggregate
row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
scoped to the aggregate flag rather than bare code presence (M89/L89
continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
retained, NULL harmonized_code coalesced, L-- excluded), rather than
relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
necessary (not just incidental) against a corpus whose `long` genuinely
lacks a harmonized_code column, and rewrote the misleading "v5-only
parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
test, guarded .verb_spendrev() against total on a non-spending
view_base, added a roxygen caveat against summing total across levels
of government, and replaced an uncheckable corpus-wide SQL comment
figure with a fixture-verifiable one.
Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.