Metadata tables refreshed from the Phase Q4 corpus (schema 4, pipeline_commit
a082b26): canonical_fips_xwalk.parquet now carries the extended population
bridge (pop_confidence exact 15.7% -> 97.3%); canonical_alias.parquet reflects
the Q3 rename continuations. 2019/2020 long partitions unchanged (continuations
remap only pre-2017 predecessor rows). Suite 336 PASS / 0 FAIL / 0 WARN.
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:
- Salt Lake County now resolves to 20 distinct cities/townships instead
of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
a separate gap, documented in cog_pipeline).
manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.
All 283 uscogdata tests pass against the new fixture.
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.
setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).
helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.
test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).
Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
Seven DuckDB views register on session open: long (raw), spending_long
and revenue_long (prefix-filtered, NOT is_aggregate per reader-spec §4),
canonical_fips_xwalk and summary_categories (identity), spending_annotated
and revenue_annotated (LEFT JOIN xwalk + categories for verb composition).
File prefix `NN-` enforces creation order so *_annotated views resolve
their *_long dependencies.
Deviation from plan: series_breaks/cpi_annual/legacy_aggregate_map and
*_with_transforms views are deferred — their backing parquets are not
in the v0.1 corpus (manifest.files.metadata only lists canonical_fips_xwalk
and summary_categories). The verbs will compute CPI adjustment verb-side
against a bundled cpi table in a later task.