.build_series_break_refs() matches `fin_code IN (<codes in the result>)`. No row's item_code is ever the literal "ALL", so the four corpus-wide entries could never match and reached no user: SB085 1977 dollar precision across the 1976/1977 boundary SB087 2002 imputation exclusion FY2002-2006 SB194 2012 dense -> sparse representation change SB086 2017 government id scheme change SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv carries an ALL @ 2012 entry describing the representation change, SO cog_explain() surfaces it". The entry shipped; the reader dropped it. A query spanning FY2011 -> FY2012 crosses the boundary where an absent cell stops meaning "Census published $0" and starts meaning "not reported", and nothing said so. Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs() on the break_year window alone -- which codes a result happens to contain is irrelevant to a caveat about the corpus. A separate field rather than more entries in series_break_refs, because an ALL caveat qualifies the whole result and folding the two together invites reading it as a caveat about one series; .build_series_break_refs() now excludes 'ALL' explicitly so the two stay disjoint by construction. cog_explain() prints them under their own "Corpus-wide caveats" heading, and cog-api passes provenance through verbatim, so the field reaches the API with no change there. On the year rule: all four entries are BOUNDARY caveats -- their own join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not being comparable across FY2012, of pre- vs post-2017 ids -- so the same `break_year BETWEEN min(years) AND max(years)` rule the code-specific path uses is the right one, and matches the issue's DoD 1. The issue's DoD 3 also asks that a FY2011 query surface SB085; that cannot hold under DoD 1 and does not hold under any reading of SB085's text, whose boundary is 1976/1977. Tested with a range that actually spans it, and flagged on the issue. Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's bundled fixture, which predates the break being catalogued. Suite: 606 pass / 0 fail / 6 skip (was 594/0/6). cog-api 357 / 0 / 8, unchanged.
8.2 KiB
8.2 KiB
uscogdata 0.1.0 (development)
Corpus-wide series breaks now reach users (corpus_break_refs)
- Four catalogued series breaks carry
fin_code = "ALL"— caveats about the corpus as a whole rather than about one item code.series_break_refsis built by matchingfin_codeagainst the item codes in the result, and no row'sitem_codeis ever the literal"ALL", so none of them could ever be surfaced:SB085(dollar precision across the 1976/1977 boundary),SB087(imputation exclusion from FY2002),SB194(the dense → sparse representation change at FY2012) andSB086(the government id scheme change at FY2017). - Provenance gains
corpus_break_refs, selected on the break-year window alone and disjoint fromseries_break_refsby construction, so a consumer can tell a whole-result caveat from a break in one series.cog_explain()prints them under their own "Corpus-wide caveats" heading. cog-api passes provenance through verbatim, so the field appears there without an API change. SB194is the one that made this urgent: a query spanning FY2011 → FY2012 crosses the boundary where an absent cell stops meaning "Census published$0" and starts meaning "not reported", and until now nothing said so.
Bundled fixture regenerated against the sparsified corpus
inst/extdata/fixture_corpus/now tracks the corpus published on 2026-07-29 (pipeline_commit 83f9715, schema v6). The wide era no longer stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of which none are$0. Absence now means two different things — in adense_sourceyear (≤ FY2011) an absent cell means Census published$0; in asparse_sourceyear (≥ FY2012) it means not reported. The corpus carries that rule in two new tables the fixture now ships,representation.parquetandcode_set.parquet, alongsidecensus_collection_coverage.parquetandlineage_events.parquet(all ten publish-tree metadata tables, up from six). Catalogued upstream as series breakSB194.cog_categories()gains anassistancespending subtype: the J-prefix aid/benefit codes (J19,J67,J68,J85) are categorised now that the upstream crosswalk covers every flow code carrying dollars.- Two consequences worth knowing about, both visible in provenance rather
than in returned dollars. The harmonization block's
na_rows_excludedcounts only rows that exist, so wide-era codes that were zero-padded no longer appear there. Coverage-gapsuggestionsare presence-based for the same reason, so a recipe whose component codes were all$0for a given government-year is no longer suggested for it. tests/testthat/test-fixture-vintage.Rpins these structural facts, so a fixture left behind by a future publish fails loudly instead of letting the suite pass against a corpus that no longer exists.
Breaking: corpus schema_version 4 (Phase P canonical ids)
- The package now requires corpus
schema_version = 4(MinCorpusSchema/MaxCorpusSchemainDESCRIPTIONare both4); older corpora built against schema 3 are rejected bycog_open()with a clear version-mismatch error.canonical_govidis now uniformly 12 characters across every vintage the corpus covers (previously a mix of 9-char legacy ids and 12-char FIPS ids depending on source year) — every hardcodedcanonical_govidliteral from a pre-Phase-P corpus is now invalid and must be re-resolved viacog_gov_search()or the newcanonical_aliaslookup table.canonical_fips_xwalkgains four columns (legacy_govs_id,census_geoid,id_source;confidenceis renamed topop_confidence) and a companioncanonical_aliastable ships in the corpus for mapping legacy/alternate ids onto the current canonical namespace. The bundled fixture corpus (inst/extdata/fixture_corpus/) has been regenerated against the Phase P publish tree, now ships the fullcanonical_fips_xwalkandcanonical_aliasmaster tables alongside the 2019-2020 long partitions, and is reproducible viadata-raw/regenerate_fixture_corpus.R.
Clearer errors when USCOGDATA_URL is unconfigured or returns non-JSON
cog_open()now aborts with theuscogdata_url_not_configurederror class when the resolved corpus URL still contains the placeholderREPLACE_WITH_SHARE_TOKENsentinel (or is empty). The message lists both remediation paths (Sys.setenv(USCOGDATA_URL = ...)andoptions(uscogdata.url = ...)) and points at the bundled fixture for offline testing. Previously the package proceeded to fetch the placeholder URL, cached the resulting HTML welcome page, and failed downstream with a crypticjsonlitelexical-error..fetch_or_cache_manifest()now parses the HTTP response body before persisting it. Non-JSON responses (login pages, 404 HTML) raiseuscogdata_invalid_manifestwith the URL, Content-Type, and underlying parse error — and never write to the on-disk cache.- Manifest cache writes are now atomic (write to
manifest.json.tmp.<pid>incache_dir, thenfile.renameover the target), so an interrupted fetch cannot replace a previously-good cache. - Existing caches with non-JSON content (poisoned by the prior code path) are silently refetched instead of returning a parse error to the caller.
- Local
USCOGDATA_URLpaths whosemanifest.jsonis not valid JSON now surface the sameuscogdata_invalid_manifestclass with file context.
Per-capita denominators now use per-year Census F-33 population
cog_spending()andcog_revenue()previously divided all years' amounts by a single ACS 2018-2022 estimate (canonical_fips_xwalk.population_acs), producing biased per-capita values for time-series analysis. They now divide by the F-33populationrecorded on each gov-year via the newgov_population_yearlyview. Result tibbles gain apop_sourcecolumn with values"census_f33"or"unavailable".notesis updated to concatenate multiple notes with"; ".
Peer cohorts can be set to a chosen year
cog_find_peers()adds ayearargument (default: most recent year for which the target has an observed population ingov_population_yearly). The returned column previously namedpopulation_acsis nowpopulationand reflects the cohort year's vintage. The cohort year is attached to the returned tibble asattr(x, "cohort_year").cog_peer_compare()now stamps acohort_yearcolumn on its result (read from the peers tibble's attribute) and recordscohort_yearpluscohort_govidsin provenance. When the caller supplies a bare character vector instead of acog_find_peers()result,cohort_yearisNA.
Rollups exclude govs missing population
cog_geographic_rollup(per_capita = TRUE)drops rows whose government haspop_source == "unavailable"and records the dropped govids inprovenance$rollup$excluded_govids. This excludes special districts (type 4) and school districts (type 5) from per-capita rollups by design.
New: vignette and provenance metadata
- New vignette
population-denominatorscovers the four population sources, the type-4/5 coverage gap, the popyear quirk, and how to build moving-window peer cohorts manually. - Provenance gains
transformations$per_capita$popyear_rangeandpop_source_counts.cog_explain()renders both.
New features
cog_gov_search()gains a basket mode: passing vectorname/state/typearguments resolves multiple place names in one call and returns a tibble of canonical rows in input order, ready to pipe intocog_spending()/cog_revenue(). Per-row resolution follows an exact-then-substring matching algorithm with deterministic disambiguation; ambiguous and missing entries are surfaced via a sidecar audit tibble plus a single console summary message.- New exports
cog_basket_resolution()andcog_basket_unresolved()expose the basket sidecar for iterative query refinement.
Breaking changes
- The first formal of
cog_gov_search()was renamed frompatterntoname. All existing call sites incog_explorer/and the package itself use positional first-arg, so this rename is non-breaking in practice. Callers that passpattern = ...by name must update toname = ....