8 Commits
Author SHA1 Message Date
jared 0fbae00e27 feat: name a cohort by state/type predicate instead of a 40k-id IN list
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00
jared a67735f121 chore: release metadata -- author of record, URLs, schema ceiling, 0.3.0
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.

MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.

0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
2026-08-08 17:46:39 -04:00
jared 61b9c95731 chore: release 0.2.0
R-CMD-check / check (push) Successful in 3m34s
R-CMD-check / check (pull_request) Successful in 6m30s
Bumps the minor version because 'All Categories' adds public surface
without breaking any existing call.

The bump is load-bearing, not cosmetic: cog-api installs this package
with install_local(), which no-ops when the version already matches.
Without it, Phase 1 would silently test against the 0.1.0 reader and
pass while proving nothing.
2026-08-05 12:04:24 -04:00
jared 7818cd2b1a feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.

Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.

Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
2026-07-18 23:19:17 -04:00
jared e635a1fc9e feat!: require corpus schema_version 4 (Phase P canonical ids)
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
2026-07-11 09:31:33 -04:00
jared 377eed1240 feat: cog_gov_search + cog_mirror + scope-aware verb behavior
Three pieces:

1. cog_gov_search: name/state/type search over canonical_fips_xwalk
   for resolving human-readable place names into canonical_govids.
   Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
   0-3 or name ('state','county','city','township') for type. Types
   4/5 emit an explanatory cli message and return an empty tibble
   (v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
   50 states + DC + territories; FIPS 66 = GU (not GA).

2. cog_mirror: downloads manifest-listed files to a local directory
   with SHA-256 idempotency (files with matching hash return status
   'cached'). Supports HTTP and local-path fixture URLs. Round-trip
   test: mirror + re-open against the mirror + query Broward 2020
   returns identical results.

3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
   queries canonical_fips_xwalk for the requested govids, emits a
   cli_inform listing any missing ones, and records the found/missing
   sets under provenance$scope. Wired into cog_spending (and
   transitively into cog_revenue, cog_geographic_rollup,
   cog_peer_compare via their cog_spending calls).

Also: dropped dbplyr from Imports (unused).

Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
2026-04-24 11:38:32 -04:00
jared cea7a241c5 chore: require corpus schema_version = 3
Tracks the cog_pipeline corpus bump from 2 → 3 (year=2012 partition is
now modern-only; v2 shipped a broken mixed-source partition). session
open aborts with a clear message if a reader hits a legacy v2 corpus.
2026-04-23 14:35:55 -04:00
jared f7035049e7 feat: package skeleton — DESCRIPTION, NAMESPACE, session/manifest/cache/views
Minimal skeleton for uscogdata v0.1. Internal session layer with lazy
cog_open(), manifest fetch+validate+cache, view registration placeholder.
Depends on DuckDB >=1.0, httr2, jsonlite. inst/schemas/provenance-v1.json
ships the structured provenance JSON Schema.
2026-04-23 09:05:54 -04:00