Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.
Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).
Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
cog_geographic_rollup(per_capita = TRUE) now drops rows whose government
has no observed population for that year (pop_source == 'unavailable'),
matching the spec's exclusion rule. Records included/excluded govids in
provenance$rollup.
Adds optional 'year' argument (defaults to most recent observed year for
the target). Filters and ranks candidates by gov_population_yearly.population
at that year. Returned column renamed population_acs -> population.
Cohort year attached as attr(x, 'cohort_year').
Adds .resolve_cohort_year() helper. Updates test assertions to use
'population' column name. Regenerates man/cog_find_peers.Rd.
Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.
Sidecar accessors for basket-mode results. cog_basket_resolution()
returns the full resolution tibble (drops candidates list-col by
default for readable printing). cog_basket_unresolved() filters to
ambiguous/no_match rows for iterative refinement.
Pre-rename in preparation for basket mode. All existing callers in this
package and cog_explorer/ pass the first argument positionally, so this
rename is non-breaking. No deprecation alias added per design spec
(no external consumers; package is pre-release v0.1.0).
@param roxygen also updated to match the new formal; man/cog_gov_search.Rd
regenerated via devtools::document().
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().
Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:
type = NULL | "spending" | "revenue"
pattern = regex matched case-insensitively on category
The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.
Also: fix @noRd placement in session.R so devtools::document() stops
warning.
Tests: +16 new / 181 total pass. check 0E/0W/0N.
Three pieces:
1. cog_gov_search: name/state/type search over canonical_fips_xwalk
for resolving human-readable place names into canonical_govids.
Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
0-3 or name ('state','county','city','township') for type. Types
4/5 emit an explanatory cli message and return an empty tibble
(v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
50 states + DC + territories; FIPS 66 = GU (not GA).
2. cog_mirror: downloads manifest-listed files to a local directory
with SHA-256 idempotency (files with matching hash return status
'cached'). Supports HTTP and local-path fixture URLs. Round-trip
test: mirror + re-open against the mirror + query Broward 2020
returns identical results.
3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
queries canonical_fips_xwalk for the requested govids, emits a
cli_inform listing any missing ones, and records the found/missing
sets under provenance$scope. Wired into cog_spending (and
transitively into cog_revenue, cog_geographic_rollup,
cog_peer_compare via their cog_spending calls).
Also: dropped dbplyr from Imports (unused).
Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
cog_find_peers selects peers from canonical_fips_xwalk by same-type,
same-state, and population-range (ratio or absolute) criteria,
ordered by |log(pop_ratio)| ascending.
cog_peer_compare accepts the find_peers result (or a plain character
vector of govids), pulls spending for target + peers via cog_spending,
and appends summary rows (summary_p25/p50/p75) so the whole result
can be faceted by `role` in a single ggplot call. Summary rows honor
per_capita + adjust_to_year by picking the right value column.
target_rank reports the target's rank among target+peers at max(years).
Provenance is rewritten with verb = cog_peer_compare and peer_count.
Also: globalVariables('.data') in zzz.R to silence R CMD check on
tidy-eval pronouns.
Tests: 21 new / 120 total pass. devtools::check() 0E/0W/2N.
Wraps cog_spending across a named list of state/county/city layers,
tagging each row with its `layer` and attaching a scope_note that
documents geographic-scope caveats (state totals are statewide, county
totals include areas outside a listed city, city proper excludes
special districts). Per-capita uses each layer's own population from
the canonical_fips_xwalk.
Provenance is inherited from cog_spending but rewritten to reflect
the outer verb (verb, call, layers).
Tests: 19 new / 99 total pass. devtools::check() 0E/0W/2N.
Three core query verbs over the spending_annotated / revenue_annotated
DuckDB views. Each verb accepts vector govid, vector years, optional
category filter, per_capita flag, and adjust_to_year for CPI-U
real-dollar conversion (bundled index).
Amounts are returned in full USD (SUM(amt) * 1000) so callers can
freely rescale to millions/billions. The $1,000s -> $USD conversion
is recorded in provenance$transformations$units_conversion.
Every result carries an attr(., "provenance") list matching
inst/schemas/provenance-v1.json. cog_explain() prints the structured
form via cli or returns the raw list for MCP/JSON consumers.
Also: .fetch_or_cache_manifest() now handles local fixture paths so
tests can point USCOGDATA_FIXTURE_URL at the pipeline publish_cache/
without a working HTTP server.
Tests: 80 pass / 0 fail. devtools::check() 0E/0W/2N (both notes
pre-existing / environmental).