cog_open() connected with a bare dbConnect() and set no resource pragmas, so
DuckDB claimed every visible core. Right for one interactive session on a
dedicated machine; wrong for a server, where cog-api runs two replicas on an
8-core host budgeted 4 and each replica independently claims all 8.
USCOGDATA_DUCKDB_THREADS and USCOGDATA_DUCKDB_MEMORY_LIMIT now resolve through
.cfg() -- inheriting the env var > option > default precedence USCOGDATA_URL
already had -- and are applied as pragmas when the connection is created.
Unset issues NO pragma, so an unconfigured session is byte-identical to before.
That negative property is asserted directly against a connection opened the
pre-change way rather than against a hardcoded core count.
.cfg() returns an env var as character, so both resolvers coerce and validate
rather than trusting the type: sprintf("SET threads TO %d", "4") would
otherwise abort inside the connection path with an error naming the pragma
instead of the setting the operator got wrong.
Replaces cog-api's getFromNamespace(".ensure_session", "uscogdata") workaround,
which depended on a private name and on the session already being open.
Final-review findings F-1, F-2, F-6, F-8 (plus the F-9 @return reword,
which shares R/balances.R).
F-2: .validate_balance_inputs() checked 2 of cog_balances()' 7 arguments.
years = integer(0) leaked a raw DuckDB 'Parser Error ... AND year IN ()'
with the generated SQL echoed back; govid = character(0) and a non-character
category returned 0 rows with no error at all; recipe = c("a","b") threw
'the condition has length > 1' from inside .validate_recipe_id(). Replaced
with a call to the money verbs' own .validate_verb_inputs() (R/spending.R),
which validates the exact superset needed. Deleted the local copy rather
than extending it -- two validators is how they drift. Placed AFTER
.coerce_govid_input(), because .validate_verb_inputs() asserts
is.character(govid) and a data-frame govid is not unwrapped before that.
This is helper reuse of the same kind as .build_verb_sql()/.attach_per_capita();
the verb still does NOT route through .verb_spendrev().
F-1: falls out of F-2 for free -- the recipe/category mutual-exclusivity
guard lives inside .validate_verb_inputs(). Previously recipe silently
discarded category AND overwrote provenance$category with the recipe label,
so a caller asking for Fund Balances got X40/Z77 insurance-trust holdings
with no trace of the dropped filter.
F-6: cog_explain() rendered every provenance caveat block except
balance_caveats. Since .emit_balance_caveats() fires at most once per
session -- and is routinely consumed by a suppressMessages() call or an
unread knitr chunk -- cog_explain() is the only surface left for a caller
who deliberately audits the result. Added a 'Holdings caveats' section
guarded on !is.null(prov$balance_caveats). Also relabels the cosmetic
'Concept: NA' line on balance results as 'not applicable (holdings are a
stock, not a flow)'.
F-8: the coverage-window query has no govid and no year predicate -- its
answer depends only on the mounted corpus -- yet it scanned all of
balance_long on every call (35% of verb runtime on the fixture, and a
per-request throughput ceiling for cog-api#26). Memoised in
.uscogdata_env$balance_coverage_windows, invalidated by cog_close(), the
same pattern as .uscogdata_env$manifest.
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.
Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
* test-manifest.R: v6 accepted; boundary rejection moves to v7.
* test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
Education NEC local, SB184-186). Broward's 2011 partition zero-pads
exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:
1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
when the resolved URL is empty or still contains the
`REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
`Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
existing caches with non-JSON content are silently refetched instead
of returning a parse error to the caller.
Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().
Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:
type = NULL | "spending" | "revenue"
pattern = regex matched case-insensitively on category
The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.
Also: fix @noRd placement in session.R so devtools::document() stops
warning.
Tests: +16 new / 181 total pass. check 0E/0W/0N.
Two UX fixes surfaced by first real-user use:
1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
a character vector OR a data.frame with a canonical_govid column
(e.g. output of cog_gov_search() or cog_find_peers()). Shared
.coerce_govid_input() helper in session.R. This lets the natural
pipe work:
cog_gov_search('MIAMI', state='FL', type='city') |>
cog_spending(years=2022, category='Police')
cog_peer_compare already accepted a data.frame for the peer arg;
behavior there is unchanged.
2. .check_govids_in_scope() message reworded. The old text led with
'v0.1 covers gov_types 0-3' which falsely implied the missing govids
were scope-excluded types when the more common real cause is a typo
or a guessed value. New message leads with typo + pre-2017 PID,
mentions scope exclusion as one possibility, and points at
cog_gov_search() as the recovery path.
Tests: 165 pass / 0 fail. check 0E/0W/0N.
Three pieces:
1. cog_gov_search: name/state/type search over canonical_fips_xwalk
for resolving human-readable place names into canonical_govids.
Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
0-3 or name ('state','county','city','township') for type. Types
4/5 emit an explanatory cli message and return an empty tibble
(v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
50 states + DC + territories; FIPS 66 = GU (not GA).
2. cog_mirror: downloads manifest-listed files to a local directory
with SHA-256 idempotency (files with matching hash return status
'cached'). Supports HTTP and local-path fixture URLs. Round-trip
test: mirror + re-open against the mirror + query Broward 2020
returns identical results.
3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
queries canonical_fips_xwalk for the requested govids, emits a
cli_inform listing any missing ones, and records the found/missing
sets under provenance$scope. Wired into cog_spending (and
transitively into cog_revenue, cog_geographic_rollup,
cog_peer_compare via their cog_spending calls).
Also: dropped dbplyr from Imports (unused).
Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
Tracks the cog_pipeline corpus bump from 2 → 3 (year=2012 partition is
now modern-only; v2 shipped a broken mixed-source partition). session
open aborts with a clear message if a reader hits a legacy v2 corpus.