Source: cog_pipeline publish_cache built 2026-07-23T16:06:45Z at 4f992a0
(pipeline PR #43, issue #28 Option B ruling: legacy aggregate families now
publish only the H2-designated Direct flavor; Census Total = code + M-code).
Fixture delta, verified against the prior partition: year=2011 loses
173,394 non-designated aggregate rows (3,037,606 -> 2,864,212; -05 max
rows/gov 2 -> 1), leaf rows byte-identical; 2012/2019/2020 partitions,
metadata parquets, and docs unchanged. manifest.json resyncs sha256 /
row_count / size_bytes for the changed partition.
Full suite vs the regenerated fixture: 467 PASS / 0 FAIL / 0 WARN / 0 SKIP
— zero pin adjudications needed (reader verbs filter is_aggregate rows and
no main test pins legacy aggregate counts).
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.
Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
* test-manifest.R: v6 accepted; boundary rejection moves to v7.
* test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
Education NEC local, SB184-186). Broward's 2011 partition zero-pads
exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
Metadata tables refreshed from the Phase Q4 corpus (schema 4, pipeline_commit
a082b26): canonical_fips_xwalk.parquet now carries the extended population
bridge (pop_confidence exact 15.7% -> 97.3%); canonical_alias.parquet reflects
the Q3 rename continuations. 2019/2020 long partitions unchanged (continuations
remap only pre-2017 predecessor rows). Suite 336 PASS / 0 FAIL / 0 WARN.
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:
- Salt Lake County now resolves to 20 distinct cities/townships instead
of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
a separate gap, documented in cog_pipeline).
manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.
All 283 uscogdata tests pass against the new fixture.
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.
setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).
helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.
test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).
Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.