Commit Graph
4 Commits
Author SHA1 Message Date
jared 1b2294e3a0 fix: require a DIFFERENT recipe component to cover a per-code gap
.recipe_coverage()'s covered_years were computed once per recipe as a
union across ALL of its components (aggregate rows included), without
excluding the component currently being tested for a gap. So a code
whose only representation in a year was its own wide-era aggregate row
satisfied its own "covered" check -- self-coverage, not the "other
components" review-doc 0.3's criterion actually specifies ("...has no
rows ... but other components do").

.recipe_coverage() now returns (recipe_id, component_code, year)
triples instead of collapsing across components, and
.recipe_component_gapped() excludes the component under test before
checking coverage, so a gap only fires when a genuinely different
sibling component has data in that year.

Adds the boundary test this gap in coverage let slip through untested:
Broward FY2011 alone, where E05/F05/G05 each report solely as their own
wide-era aggregate row and E04/F04/G04 don't exist as codes before 2012
corpus-wide, so none of the three Corrections recipes have any OTHER
component to cover them -- must produce zero suggestions. The existing
2011-2012 combined test still passes, now firing because of the 2012
E05-gapped/E04-covers pair rather than 2011's self-coverage. Updates the
header comment to state the other-component requirement explicitly.
2026-07-23 12:22:23 -04:00
jared 267bc24fee feat: narrow harmonization signposting to per-code gap detection
.build_suggestions() previously flagged a recipe only when the WHOLE
category result had zero rows in a requested year, so a multi-code
category where one recipe component was genuinely gapped never fired
if any sibling code (same recipe or not) had data that year. Each
recipe's own in-category component is now checked individually -- a
component fires when it has no rows in a requested (in-scope) year the
recipe's own generic join otherwise covers, even when the overall
category result looks complete.

Decomposes .build_suggestions() into .category_recipe_components/
.recipe_meta/.component_presence/.recipe_coverage/.recipe_component_gapped
helpers, drops the now-unused `result` param, and rewrites the header
comment to describe the new, deliberately wider scope plus the
per-government `covered` guard that still filters recipes with no data
at all (ordinary reporting variance vs. a real format-boundary gap).

Tests pin the multi-code case the coarse check missed (Cleburne County
FY2012: G05 gapped, G04 covers, masked because E04/E05 have data) next
to the still-guarded no-recipe-coverage case (F04/F05 both absent), and
update the Broward 2019-2020 case to its new, correct expectation (fires
for corrections_combined/corrections_other_capital_combined, still
silent for corrections_capital_combined) plus a fresh true-full-coverage
negative case (Maricopa County).
2026-07-23 12:22:22 -04:00
jared 77f48047b1 fix: exercise real harmonized-view SQL in tests; unambiguous recipe provenance
R-CMD-check / check (pull_request) Successful in 3m5s
R-CMD-check / check (push) Successful in 3m1s
test-views.R's harmonized-view test previously ran a hand-rolled REPLACE
query with no WHERE clause, so a regression in any of
inst/sql/22-spending_long_harmonized.sql / 23-revenue_long_harmonized.sql's
three predicates (NOT is_aggregate, harmonized_code IS NOT NULL, the
E/F/G/K or T/A/U/B/C/D prefix filter) would go uncaught. Replaced it with a
test that reads the real SQL files off disk, substitutes {url} exactly as
.register_views() does, and executes them (plus their 10-long.sql
dependency) against a synthetic hive-partitioned parquet tree written via
DuckDB's own COPY ... TO (FORMAT PARQUET) (no arrow dependency, matching
this package's existing convention). Ten rows are crafted so each predicate
is independently falsifiable by a specific row; manually broke each
predicate in turn to confirm the test fails exactly as expected, then
restored the SQL files (see the task report for the RED-phase transcript).

Also fixes a provenance ambiguity: a recipe= query bypasses
spending_annotated(_harmonized)/revenue_annotated(_harmonized) entirely
(.run_recipe() joins `long` directly), so basis= has no effect on it, but
provenance was still reporting basis = "harmonized"/"raw" (whatever the
argument resolved to) with harmonization$applied = FALSE alongside it --
misleading, since it looks like harmonization was evaluated and found
nothing to exclude rather than "not applicable here." Recipe results now
report basis = "recipe" with an inert harmonization block carrying an
explicit note, regardless of what basis= was passed.
2026-07-18 23:55:48 -04:00
jared 4de915b557 feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.

Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).

Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
2026-07-18 23:33:02 -04:00