The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.
Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.
Three test repairs, each a real consequence of sparsification rather than
a number to bump:
test-categories.R "assistance" joined the spending subtype
vocabulary with the J-prefix codes.
test-spending.R The harmonization block counts rows that
exist. Broward's E21/F21/G21 were zero-pads
and are gone, so the anchor moves to FL state,
whose three NA-mapped rows carry $2.83B --
the amount accounting was previously asserted
only against 0 and could not have caught a
bug. Broward keeps a test of its own, now
asserting the zero-pads are absent.
test-expenditure-concept.R Coverage-gap suggestions are presence-based.
AL state's only FY2011 B47 cell was an
explicit zero, so ig_federal_b47_wide stopped
being a candidate there; FL state carries a
real amount, so the counterpart guard is
exercised against a suggestion that fires.
test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.
Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.
Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
* test-manifest.R: v6 accepted; boundary rejection moves to v7.
* test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
Education NEC local, SB184-186). Broward's 2011 partition zero-pads
exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.
Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).
Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
Swaps every hardcoded 9-char canonical_govid literal (Broward County,
Fort Lauderdale City, Florida/Alabama state govts, Bexar/Tarrant/Wayne
counties, San Diego/Oakland/Miami/Austin cities) for its 12-char Phase P
equivalent, resolved by name+type+state against the regenerated fixture
xwalk. Also updates two gov_name search patterns that no longer match
under Phase P canonical naming ("FLORIDA STATE GOVT" -> "FLORIDA"; the
"Miami" substring test now pins type = "city" since MIAMI-DADE COUNTY's
canonical name now also contains "Miami", which would otherwise make the
match ambiguous across govs_types instead of resolving via largest-pop).
Underlying per-year population figures for Broward County and Alabama
are unchanged, so no expected data-value literals needed recomputation.
Suite: 126 test blocks / 336 expectations, 0 FAIL / 0 WARN / 0 SKIP.
Updates transformations\$per_capita with the new denominator_source string,
popyear_range, and pop_source_counts. .attach_per_capita stashes
popyear_range on the result; .verb_spendrev strips the helper attr after
provenance is built.
.notes_column now joins multiple per-row notes with '; '. Adds the
'No population denominator available for this gov type' note when
pop_source is 'unavailable'. Gracefully handles absent pop_source
(per_capita = FALSE). Two new tests: one corpus-level (census_f33
branch) and one synthetic unit test covering multi-note concatenation.
Adds a RED test asserting that cog_spending(per_capita = TRUE) divides
by the per-year F-33 population (Broward 2019: 1,935,878; 2020: 1,952,778)
rather than the static ACS value (1,940,907). Uses absolute-tolerance
expect_true(abs(...) < 1) instead of expect_equal(tolerance=1) because
testthat 3 treats the tolerance argument as relative.
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.
setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).
helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.
test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).
Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
Two UX fixes surfaced by first real-user use:
1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
a character vector OR a data.frame with a canonical_govid column
(e.g. output of cog_gov_search() or cog_find_peers()). Shared
.coerce_govid_input() helper in session.R. This lets the natural
pipe work:
cog_gov_search('MIAMI', state='FL', type='city') |>
cog_spending(years=2022, category='Police')
cog_peer_compare already accepted a data.frame for the peer arg;
behavior there is unchanged.
2. .check_govids_in_scope() message reworded. The old text led with
'v0.1 covers gov_types 0-3' which falsely implied the missing govids
were scope-excluded types when the more common real cause is a typo
or a guessed value. New message leads with typo + pre-2017 PID,
mentions scope exclusion as one possibility, and points at
cog_gov_search() as the recovery path.
Tests: 165 pass / 0 fail. check 0E/0W/0N.
Three pieces:
1. cog_gov_search: name/state/type search over canonical_fips_xwalk
for resolving human-readable place names into canonical_govids.
Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
0-3 or name ('state','county','city','township') for type. Types
4/5 emit an explanatory cli message and return an empty tibble
(v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
50 states + DC + territories; FIPS 66 = GU (not GA).
2. cog_mirror: downloads manifest-listed files to a local directory
with SHA-256 idempotency (files with matching hash return status
'cached'). Supports HTTP and local-path fixture URLs. Round-trip
test: mirror + re-open against the mirror + query Broward 2020
returns identical results.
3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
queries canonical_fips_xwalk for the requested govids, emits a
cli_inform listing any missing ones, and records the found/missing
sets under provenance$scope. Wired into cog_spending (and
transitively into cog_revenue, cog_geographic_rollup,
cog_peer_compare via their cog_spending calls).
Also: dropped dbplyr from Imports (unused).
Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
Three core query verbs over the spending_annotated / revenue_annotated
DuckDB views. Each verb accepts vector govid, vector years, optional
category filter, per_capita flag, and adjust_to_year for CPI-U
real-dollar conversion (bundled index).
Amounts are returned in full USD (SUM(amt) * 1000) so callers can
freely rescale to millions/billions. The $1,000s -> $USD conversion
is recorded in provenance$transformations$units_conversion.
Every result carries an attr(., "provenance") list matching
inst/schemas/provenance-v1.json. cog_explain() prints the structured
form via cli or returns the raw list for MCP/JSON consumers.
Also: .fetch_or_cache_manifest() now handles local fixture paths so
tests can point USCOGDATA_FIXTURE_URL at the pipeline publish_cache/
without a working HTTP server.
Tests: 80 pass / 0 fail. devtools::check() 0E/0W/2N (both notes
pre-existing / environmental).