Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:
<= FY2011 dense_source absent => Census published $0
>= FY2012 sparse_source absent => not reported, unknown
A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.
complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.
Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.
Scoping decisions, each of which would invent rows if taken loosely:
- The grid is per government TYPE (code_set.type). Filling against the
union of all types would give a county cells like "state IG transfer to
school districts", indistinguishable from real census zeros.
- NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
grid offers cells those views never return, so each would fill as a
phantom $0.
- Filling happens BEFORE per_capita and inflation, so a census_zero stays
0 through both and a not_reported stays NA rather than becoming 0.
Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.
Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.
Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.
Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.
Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
Nine review items on the expenditure_concept = direct|total feature:
- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
wide-era data puts a $0 leaf row in the same group as the real aggregate
row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
scoped to the aggregate flag rather than bare code presence (M89/L89
continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
retained, NULL harmonized_code coalesced, L-- excluded), rather than
relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
necessary (not just incidental) against a corpus whose `long` genuinely
lacks a harmonized_code column, and rewrote the misleading "v5-only
parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
test, guarded .verb_spendrev() against total on a non-spending
view_base, added a roxygen caveat against summing total across levels
of government, and replaced an uncheckable corpus-wide SQL comment
figure with a fixture-verifiable one.
Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.