fix: address Task 3 code review (bool_or, invariant tests, guards, docs)

Nine review items on the expenditure_concept = direct|total feature:

- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
  bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
  dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
  wide-era data puts a $0 leaf row in the same group as the real aggregate
  row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
  across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
  aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
  scoped to the aggregate flag rather than bare code presence (M89/L89
  continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
  pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
  retained, NULL harmonized_code coalesced, L-- excluded), rather than
  relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
  necessary (not just incidental) against a corpus whose `long` genuinely
  lacks a harmonized_code column, and rewrote the misleading "v5-only
  parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
  test, guarded .verb_spendrev() against total on a non-spending
  view_base, added a roxygen caveat against summing total across levels
  of government, and replaced an uncheckable corpus-wide SQL comment
  figure with a fixture-verifiable one.

Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
This commit is contained in:
2026-07-27 10:00:07 -04:00
parent fefd4fe969
commit e2088458e1
6 changed files with 322 additions and 19 deletions
+20 -10
View File
@@ -1,15 +1,25 @@
# R/views.R
# SQL files whose view definitions read schema-v5-only parquet tables
# (harmonization_map.parquet, harmonization_recipes.parquet,
# series_breaks.parquet) or select from views built on top of them. DuckDB's
# read_parquet() resolves the file at CREATE VIEW time (even for a view, it
# still needs the source schema) and errors immediately -- "IO Error: No
# files found" -- if the path doesn't exist, so these cannot be registered
# unconditionally against a v4 corpus the way the rest of inst/sql/ is.
# Registration is therefore gated on manifest$schema_version >= 5; verb-level
# *usage* of the resulting views is separately gated by .resolve_basis() /
# .require_schema_v5().
# SQL files that cannot be registered unconditionally against a v4 corpus,
# for one of two distinct reasons -- both fail at CREATE VIEW time (DuckDB
# resolves a view's source schema eagerly, even though it defers execution),
# so a v4 corpus can't tolerate either unconditionally:
#
# (a) Missing FILE. 33-/34-/35- read_parquet() a v5-only parquet table
# (harmonization_map.parquet, harmonization_recipes.parquet,
# series_breaks.parquet) that doesn't exist at all on a v4 corpus --
# "IO Error: No files found".
#
# (b) Missing COLUMN. 22-/23-/25- reference `long.harmonized_code`, a
# column that does not exist on a v4 corpus's `long` table (harmonized
# space was introduced in schema v5) -- "Binder Error: Referenced
# column harmonized_code not found". 42-/43-/45- are on this list only
# because they SELECT s.* FROM the (a)/(b) views above, so they'd fail
# to resolve their own source view if it weren't already skipped.
#
# Registration is therefore gated on manifest$schema_version >= 5 for all of
# them; verb-level *usage* of the resulting views is separately gated by
# .resolve_basis() / .require_schema_v5().
.harmonization_view_files <- c(
"22-spending_long_harmonized.sql",
"23-revenue_long_harmonized.sql",