Adds a third comparison arm to measure_signposting_rate(): Task 19c's
first per-code pass (git ref da72bf3, self-coverage allowed) alongside
the existing coarse (b0df1ec) and live corrected per-code arms, pulled
verbatim via the same git-show mechanism (renamed
.measure_load_coarse_impl -> .measure_load_git_impl since it now loads
more than the coarse arm).
Reports both the original delta (self-coverage-allowed rate minus
coarse) and the corrected delta (live per-code rate minus coarse), plus
the self-coverage share of the original delta (queries that fired ONLY
because a component's own aggregate row satisfied its own coverage
check). Verifies percode-fired is always a subset of selfcov-fired
(stopifnot) -- the corrected arm is a strict narrowing of the buggy one,
so the decomposition is exact rather than approximate.
data-raw/measure_signposting_rate.R runs every summary_categories
category x a (seeded, deterministic) sample of up to 20 governments x
the widest pre/post-2012 year span the active corpus actually supports,
through both the R2 coarse .build_suggestions() (pulled verbatim from
git ref b0df1ec, evaluated in an isolated env parented on the uscogdata
namespace) and the current per-code version, and reports the suggestion
rate and delta under each, overall and by category.
Parameterized by USCOGDATA_URL (defaults to the bundled fixture when
unset) so it can be re-run against the staged/full corpus later. Detects
and reports when the active corpus can't fill a full 3-year pre/3-year
post-2012 design instead of padding or fabricating years. This script
measures the coarse-vs-per-code tradeoff; it does not rule on what
suggestion-rate increase is an acceptable amount of added noise -- that
is Jared's call at Checkpoint R3.
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
Adds R/sysdata.rda with the annual-average CPIAUCSL index (1947-2026,
80 years, sourced from FRED) and an internal .inflate() helper that
converts nominal amounts between two years via the ratio of CPI values.
Bundling CPI in the package (rather than publishing a cpi_annual.parquet
in the corpus) matches the reader-specification intent: real-dollar
conversion is a verb-level option, not a corpus-level artifact, so the
target_year stays flexible at query time.
data-raw/cpi_annual.R carries the one-shot FRED refresh used to build
sysdata.rda. Re-run when the CPI series needs to roll forward.
Verified: CPI(2000)/CPI(2021) ≈ 0.635, matching the ~0.63 sanity
anchor in cog_explorer's existing inflation logic.
Tests: tests/testthat/test-adjust.R covers the known 2000→2021 ≈ 1.574x
anchor, vectorized from_year, error cases for out-of-range years, and
NA-amount preservation. 14 pass / 0 fail.