Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.
Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
* test-manifest.R: v6 accepted; boundary rejection moves to v7.
* test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
Education NEC local, SB184-186). Broward's 2011 partition zero-pads
exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:
1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
when the resolved URL is empty or still contains the
`REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
`Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
existing caches with non-JSON content are silently refetched instead
of returning a parse error to the caller.
Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().
Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:
type = NULL | "spending" | "revenue"
pattern = regex matched case-insensitively on category
The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.
Also: fix @noRd placement in session.R so devtools::document() stops
warning.
Tests: +16 new / 181 total pass. check 0E/0W/0N.
Two UX fixes surfaced by first real-user use:
1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
a character vector OR a data.frame with a canonical_govid column
(e.g. output of cog_gov_search() or cog_find_peers()). Shared
.coerce_govid_input() helper in session.R. This lets the natural
pipe work:
cog_gov_search('MIAMI', state='FL', type='city') |>
cog_spending(years=2022, category='Police')
cog_peer_compare already accepted a data.frame for the peer arg;
behavior there is unchanged.
2. .check_govids_in_scope() message reworded. The old text led with
'v0.1 covers gov_types 0-3' which falsely implied the missing govids
were scope-excluded types when the more common real cause is a typo
or a guessed value. New message leads with typo + pre-2017 PID,
mentions scope exclusion as one possibility, and points at
cog_gov_search() as the recovery path.
Tests: 165 pass / 0 fail. check 0E/0W/0N.
Three pieces:
1. cog_gov_search: name/state/type search over canonical_fips_xwalk
for resolving human-readable place names into canonical_govids.
Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
0-3 or name ('state','county','city','township') for type. Types
4/5 emit an explanatory cli message and return an empty tibble
(v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
50 states + DC + territories; FIPS 66 = GU (not GA).
2. cog_mirror: downloads manifest-listed files to a local directory
with SHA-256 idempotency (files with matching hash return status
'cached'). Supports HTTP and local-path fixture URLs. Round-trip
test: mirror + re-open against the mirror + query Broward 2020
returns identical results.
3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
queries canonical_fips_xwalk for the requested govids, emits a
cli_inform listing any missing ones, and records the found/missing
sets under provenance$scope. Wired into cog_spending (and
transitively into cog_revenue, cog_geographic_rollup,
cog_peer_compare via their cog_spending calls).
Also: dropped dbplyr from Imports (unused).
Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
Tracks the cog_pipeline corpus bump from 2 → 3 (year=2012 partition is
now modern-only; v2 shipped a broken mixed-source partition). session
open aborts with a clear message if a reader hits a legacy v2 corpus.