Commit Graph
78 Commits
Author SHA1 Message Date
jared 887acf7e81 test: add missing cog_peer_compare coverage to expenditure_concept tests
- 'both cross-government verbs still accept the direct default' now tests both verbs
- 'the refusal message names the fix and the reason' now asserts both functions name
  themselves correctly in their error messages (cog_geographic_rollup vs cog_peer_compare)

Addresses coordinator feedback to prevent test coverage gaps and ensure the helper's
verb name argument is pinned correctly.
2026-07-27 10:19:55 -04:00
jared 81fd1a5279 feat: refuse expenditure_concept = total in the cross-government verbs
Owner ruling R1. Combining Census Total across governments counts
intergovernmental transfers twice, and these results land in Tableau where a
warning would be invisible -- so this is a hard error whose message names the
fix and the reason.
2026-07-27 10:14:16 -04:00
jared e2088458e1 fix: address Task 3 code review (bool_or, invariant tests, guards, docs)
Nine review items on the expenditure_concept = direct|total feature:

- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
  bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
  dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
  wide-era data puts a $0 leaf row in the same group as the real aggregate
  row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
  across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
  aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
  scoped to the aggregate flag rather than bare code presence (M89/L89
  continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
  pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
  retained, NULL harmonized_code coalesced, L-- excluded), rather than
  relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
  necessary (not just incidental) against a corpus whose `long` genuinely
  lacks a harmonized_code column, and rewrote the misleading "v5-only
  parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
  test, guarded .verb_spendrev() against total on a non-spending
  view_base, added a roxygen caveat against summing total across levels
  of government, and replaced an uncheckable corpus-wide SQL comment
  figure with a fixture-verifiable one.

Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
2026-07-27 10:00:07 -04:00
jared fefd4fe969 feat: expenditure_concept = direct|total in cog_spending()
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
2026-07-27 09:28:17 -04:00
jared 7ed1da9b79 fix: drop the inert K prefix from the spending flow prefixes
K matches zero rows corpus-wide (audited pipeline-side). Numerically inert;
removed so the code stops implying a prefix the data never had.
2026-07-27 09:16:41 -04:00
jared 9240a18ea3 chore: regenerate fixture corpus with the intergovernmental category rows
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).

cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.

Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
2026-07-27 09:11:10 -04:00
jared 46fed3a241 docs: Task 1 amendment — cog_categories is a third consumer of summary_categories 2026-07-27 09:09:42 -04:00
jared c9d1a05d4f chore: gitignore .superpowers/sdd working artifacts, keep plans/ tracked 2026-07-27 09:02:54 -04:00
jared fcecd62a03 docs: implementation plan for uscogdata expenditure_concept (repo 2 of 3)
Seven TDD tasks. Records the two measured facts the design rests on: aggregate
IG rows carry no harmonized_code (so the IG leg must COALESCE item_code), and
aggregate IG codes are year-disjoint from their modern leaf components (so
skipping NOT is_aggregate cannot double-count). Baseline measured at PASS 467.

Plan lives under .superpowers/ because docs/ is the gitignored pkgdown output
dir; .Rbuildignore'd so it never ships in the package tarball.
2026-07-27 09:02:42 -04:00
jared fa40266d07 Merge pull request 'Regenerate fixture corpus from the Option B (single-flavor aggregate) publish tree' (#7) from fix/fixture-option-b-aggregates into main
R-CMD-check / check (push) Successful in 2m34s
Reviewed-on: #7
2026-07-23 12:21:22 -04:00
jared 3583c05852 chore: regenerate fixture corpus from the Option B (single-flavor aggregate) publish tree
R-CMD-check / check (push) Successful in 2m49s
R-CMD-check / check (pull_request) Successful in 2m38s
Source: cog_pipeline publish_cache built 2026-07-23T16:06:45Z at 4f992a0
(pipeline PR #43, issue #28 Option B ruling: legacy aggregate families now
publish only the H2-designated Direct flavor; Census Total = code + M-code).

Fixture delta, verified against the prior partition: year=2011 loses
173,394 non-designated aggregate rows (3,037,606 -> 2,864,212; -05 max
rows/gov 2 -> 1), leaf rows byte-identical; 2012/2019/2020 partitions,
metadata parquets, and docs unchanged. manifest.json resyncs sha256 /
row_count / size_bytes for the changed partition.

Full suite vs the regenerated fixture: 467 PASS / 0 FAIL / 0 WARN / 0 SKIP
— zero pin adjudications needed (reader verbs filter is_aggregate rows and
no main test pins legacy aggregate counts).
2026-07-23 12:16:05 -04:00
jared bd53230ae7 Merge branch 'feat/schema-v6-support'
R-CMD-check / check (push) Successful in 19m38s
2026-07-22 13:16:11 -04:00
jaredandClaude Opus 4.8 e813ffd3aa feat: accept corpus schema v6 (FIPS geography harmonization); v6 fixture
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.

Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
  * test-manifest.R: v6 accepted; boundary rejection moves to v7.
  * test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
    extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
    Education NEC local, SB184-186). Broward's 2011 partition zero-pads
    exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
    excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 13:16:11 -04:00
jared b0df1ec668 Merge pull request 'Phase R2: basis= harmonized/raw, recipes, signposting (schema 4/5 dual-accept)' (#5) from feat/phase-r2-harmonization into main
R-CMD-check / check (push) Successful in 2m47s
Reviewed-on: #5
2026-07-19 11:05:05 -04:00
jared 77f48047b1 fix: exercise real harmonized-view SQL in tests; unambiguous recipe provenance
R-CMD-check / check (pull_request) Successful in 3m5s
R-CMD-check / check (push) Successful in 3m1s
test-views.R's harmonized-view test previously ran a hand-rolled REPLACE
query with no WHERE clause, so a regression in any of
inst/sql/22-spending_long_harmonized.sql / 23-revenue_long_harmonized.sql's
three predicates (NOT is_aggregate, harmonized_code IS NOT NULL, the
E/F/G/K or T/A/U/B/C/D prefix filter) would go uncaught. Replaced it with a
test that reads the real SQL files off disk, substitutes {url} exactly as
.register_views() does, and executes them (plus their 10-long.sql
dependency) against a synthetic hive-partitioned parquet tree written via
DuckDB's own COPY ... TO (FORMAT PARQUET) (no arrow dependency, matching
this package's existing convention). Ten rows are crafted so each predicate
is independently falsifiable by a specific row; manually broke each
predicate in turn to confirm the test fails exactly as expected, then
restored the SQL files (see the task report for the RED-phase transcript).

Also fixes a provenance ambiguity: a recipe= query bypasses
spending_annotated(_harmonized)/revenue_annotated(_harmonized) entirely
(.run_recipe() joins `long` directly), so basis= has no effect on it, but
provenance was still reporting basis = "harmonized"/"raw" (whatever the
argument resolved to) with harmonization$applied = FALSE alongside it --
misleading, since it looks like harmonization was evaluated and found
nothing to exclude rather than "not applicable here." Recipe results now
report basis = "recipe" with an inert harmonization block carrying an
explicit note, regardless of what basis= was passed.
2026-07-18 23:55:48 -04:00
jared 4de915b557 feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.

Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).

Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
2026-07-18 23:33:02 -04:00
jared 7818cd2b1a feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.

Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.

Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
2026-07-18 23:19:17 -04:00
jared 3b725770d2 Merge pull request 'Phase R1: cog_manifest() accessor + CPI coverage pins' (#4) from feat/phase-r1-forward into main
R-CMD-check / check (push) Successful in 2m40s
Reviewed-on: #4
2026-07-18 18:07:13 -04:00
jared 4d61692f05 feat: export cog_manifest() accessor + pin CPI coverage 1967-present
R-CMD-check / check (push) Successful in 2m47s
R-CMD-check / check (pull_request) Successful in 2m57s
2026-07-18 13:40:54 -04:00
jared 70cf553828 chore: regenerate fixture corpus from cog_pipeline Phase Q4 publish (a082b26)
R-CMD-check / check (push) Successful in 2m18s
Metadata tables refreshed from the Phase Q4 corpus (schema 4, pipeline_commit
a082b26): canonical_fips_xwalk.parquet now carries the extended population
bridge (pop_confidence exact 15.7% -> 97.3%); canonical_alias.parquet reflects
the Q3 rename continuations. 2019/2020 long partitions unchanged (continuations
remap only pre-2017 predecessor rows). Suite 336 PASS / 0 FAIL / 0 WARN.
2026-07-13 19:36:10 -04:00
jared 3c55447308 Merge feat/phase-p-schema-4: corpus schema_version 4 (Phase P canonical ids)
R-CMD-check / check (push) Successful in 2m14s
2026-07-11 14:11:53 -04:00
jared 92c9a7382e test: re-baseline canonical_govid literals to 12-char namespace
Swaps every hardcoded 9-char canonical_govid literal (Broward County,
Fort Lauderdale City, Florida/Alabama state govts, Bexar/Tarrant/Wayne
counties, San Diego/Oakland/Miami/Austin cities) for its 12-char Phase P
equivalent, resolved by name+type+state against the regenerated fixture
xwalk. Also updates two gov_name search patterns that no longer match
under Phase P canonical naming ("FLORIDA STATE GOVT" -> "FLORIDA"; the
"Miami" substring test now pins type = "city" since MIAMI-DADE COUNTY's
canonical name now also contains "Miami", which would otherwise make the
match ambiguous across govs_types instead of resolving via largest-pop).
Underlying per-year population figures for Broward County and Alabama
are unchanged, so no expected data-value literals needed recomputation.
Suite: 126 test blocks / 336 expectations, 0 FAIL / 0 WARN / 0 SKIP.
2026-07-11 09:32:02 -04:00
jared 570a9408a2 feat: regenerate fixture corpus from Phase P publish tree + committed regen script
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
2026-07-11 09:31:49 -04:00
jared e635a1fc9e feat!: require corpus schema_version 4 (Phase P canonical ids)
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
2026-07-11 09:31:33 -04:00
jared 874347242b fix(manifest): actionable errors when USCOGDATA_URL is unset or returns non-JSON
R-CMD-check / check (push) Successful in 2m5s
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:

1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
   when the resolved URL is empty or still contains the
   `REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
   `Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
   remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
   it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
   Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
   existing caches with non-JSON content are silently refetched instead
   of returning a parse error to the caller.

Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
2026-05-27 11:57:31 -04:00
jared 0dd3f15ada fix(ci): use knitr::rmarkdown vignette engine + ignore gitea/CLAUDE
R-CMD-check / check (push) Successful in 1m44s
R CMD build (run by rcmdcheck before R CMD check) rebuilds vignettes
from source regardless of --no-vignettes. Our vignette declared
%\VignetteEngine{knitr::knitr} which requires the 'markdown' package
that isn't in CI's dependency tree. Switch to %\VignetteEngine{knitr::rmarkdown},
which uses the already-Suggests-listed 'rmarkdown' package and matches
the output: rmarkdown::html_vignette directive in the YAML header.

Also add ^\.gitea$ and ^CLAUDE\.md$ to .Rbuildignore so R CMD check
stops emitting the "hidden file" / "non-standard top-level file" notes.

Local rcmdcheck (mirroring CI's exact args) now reports
0 errors / 0 warnings / 0 notes.

The pdflatex notice in the build log is unrelated — R CMD build prints
"Not building PDF manual" and continues; with --no-manual it's silenced
entirely.
2026-04-29 19:25:16 -04:00
jared 919548685b polish(per-year-pop): expand popyear in cog_explain + propagate pop_range
R-CMD-check / check (push) Failing after 1m30s
Final-review followups:

1. cog_explain rendered the popyear range as raw 2-digit values
   ("popyear range: 19-20"), which a user could read as the years 19-20.
   Added .expand_popyear() helper to format as 4-digit calendar years
   (2019-2020). Pivot at 70 to handle pre-2000 vintages if the corpus
   ever extends backward.

2. cog_peer_compare provenance was missing pop_range and is_ratio,
   omitted from the spec-required reproducibility metadata.
   cog_find_peers now stamps both as tibble attributes; cog_peer_compare
   reads them through to provenance$pop_range and provenance$is_ratio.

Test coverage extended: explain test asserts the 4-digit format and
rejects the old 2-digit form; peer-compare test asserts pop_range +
is_ratio propagate end-to-end.

326 PASS / 0 FAIL.
2026-04-29 19:16:48 -04:00
jared 716cfe25e5 build: ignore vignette build artifacts (doc/, Meta/)
Auto-added by devtools::document() during the per-year-population
documentation pass.
2026-04-29 19:10:15 -04:00
jared 33c0274727 docs(news): per-year population denominators (unreleased)
Summarizes the per-capita and peer-cohort behavior changes for users
upgrading from earlier 0.1 snapshots.
2026-04-29 19:09:00 -04:00
jared c46354f049 docs(vignette): population denominators rationale + usage
Explains the four population sources, why F-33 is the default, type-4/5
coverage gap, the popyear quirk, and how to build moving-window peer
cohorts manually.
2026-04-29 19:05:38 -04:00
jared 916212c327 feat(explain): render new per-capita provenance fields
cog_explain() now prints denominator_source, popyear_range, and
pop_source_counts under the Transformations section.
2026-04-29 19:02:37 -04:00
jared a2ced368f5 feat(provenance): record per-year denominator metadata
Updates transformations\$per_capita with the new denominator_source string,
popyear_range, and pop_source_counts. .attach_per_capita stashes
popyear_range on the result; .verb_spendrev strips the helper attr after
provenance is built.
2026-04-29 19:00:16 -04:00
jared b7ebb4cd88 feat(rollup): drop unavailable-pop rows + provenance audit
cog_geographic_rollup(per_capita = TRUE) now drops rows whose government
has no observed population for that year (pop_source == 'unavailable'),
matching the spec's exclusion rule. Records included/excluded govids in
provenance$rollup.
2026-04-29 17:50:59 -04:00
jared c334be7706 test(rollup): per-year denominator + provenance expectations (failing) 2026-04-29 17:50:07 -04:00
jared cadce8d528 docs(peers): document cohort_year on cog_peer_compare return 2026-04-29 17:45:05 -04:00
jared a92450ff76 feat(peers): stamp cohort_year on cog_peer_compare results
Reads attr(peers, 'cohort_year') when the caller passed a cog_find_peers()
tibble; NA when the caller passed a bare character vector. Stamped as a
constant column on the result and recorded in provenance alongside the
cohort govids.
2026-04-29 17:44:25 -04:00
jared 54dd40a61d feat(peers): cog_find_peers uses per-year population
Adds optional 'year' argument (defaults to most recent observed year for
the target). Filters and ranks candidates by gov_population_yearly.population
at that year. Returned column renamed population_acs -> population.
Cohort year attached as attr(x, 'cohort_year').

Adds .resolve_cohort_year() helper. Updates test assertions to use
'population' column name. Regenerates man/cog_find_peers.Rd.
2026-04-29 17:26:32 -04:00
jared 807ed35cb7 test(peers): per-year cohort expectations (failing)
Three RED tests that drive Task 6's cog_find_peers() rewrite:
- defaults year to most recent observed (expects cohort_year attr + population column)
- honors explicit year= argument (expects cohort_year attr)
- errors with "no observed population" for unobserved year
2026-04-29 17:17:52 -04:00
jared 9ae46746c0 refactor(notes): simplify aggregate-fallback predicate in .notes_column
The plan-supplied predicate combined three redundant checks
(is.null + any + %in% TRUE). Element-wise behavior was correct via
scalar recycling, but the form was confusing — a code-quality reviewer
misread it as a multi-row false-positive bug. Simplify to mirror the
parts[[2]] structure: gate on column presence, then element-wise
%in% TRUE check. Equivalent semantics, fewer ways to misread.
2026-04-29 17:08:56 -04:00
jared 9238b04b69 feat(notes): concatenate notes; flag unavailable population
.notes_column now joins multiple per-row notes with '; '. Adds the
'No population denominator available for this gov type' note when
pop_source is 'unavailable'. Gracefully handles absent pop_source
(per_capita = FALSE). Two new tests: one corpus-level (census_f33
branch) and one synthetic unit test covering multi-note concatenation.
2026-04-29 17:03:05 -04:00
jared cbc867bed1 docs(per-capita): refresh roxygen for per-year denominator + pop_source 2026-04-29 16:58:27 -04:00
jared 4ea0583d3a feat(per-capita): use per-year F-33 population in spending verbs
cog_spending(per_capita = TRUE) and cog_revenue(per_capita = TRUE) now
divide each year's amount by that gov-year's population from
gov_population_yearly (drawn from long.population) instead of a single
static ACS 2018-2022 value. Adds pop_source column with values
'census_f33' or 'unavailable'.
2026-04-29 16:36:22 -04:00
jared e4a105013e test(spending): document fixture-pop origin in per-year-denominator test 2026-04-29 16:29:24 -04:00
jared 21b3d66c0e test(spending): per-year denominator expectation (failing)
Adds a RED test asserting that cog_spending(per_capita = TRUE) divides
by the per-year F-33 population (Broward 2019: 1,935,878; 2020: 1,952,778)
rather than the static ACS value (1,940,907). Uses absolute-tolerance
expect_true(abs(...) < 1) instead of expect_equal(tolerance=1) because
testthat 3 treats the tolerance argument as relative.
2026-04-29 16:20:41 -04:00
jared df3fe3731b test(views): document hardcoded fixture-pop origin in gov_population_yearly test 2026-04-29 15:37:05 -04:00
jared ed9658d267 feat(sql): add gov_population_yearly view
Exposes one row per (year, canonical_govid) drawn from long.population.
Used by per-capita denominators and peer matching.
2026-04-29 15:12:53 -04:00
jared a28fb2e19b chore(fixture): regenerate against cog_pipeline Layer 1 + Layer 2 fixes
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:

- Salt Lake County now resolves to 20 distinct cities/townships instead
  of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
  townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
  reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
  a separate gap, documented in cog_pipeline).

manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.

All 283 uscogdata tests pass against the new fixture.
2026-04-29 15:05:38 -04:00
jared 24e4449be7 docs(plan): per-year population denominator implementation plan
15-task TDD plan covering: gov_population_yearly view, .attach_per_capita
per-year join, pop_source column + multi-note concatenation, cog_find_peers
year arg, cog_peer_compare cohort_year, rollup unavailable-pop exclusion,
provenance updates, cog_explain rendering, vignette, cog_pipeline data
dictionary, and NEWS entry.

Also corrects spec to match existing rollup semantics (side-by-side, not
summed) and adds /plans to .Rbuildignore.
2026-04-29 09:40:05 -04:00
jared cfcda04e0c docs(spec): correct cog_peer_compare signature in per-year-pop spec
Existing function takes a peers argument (caller supplies cohort) rather
than building one internally. Cohort year flows through via an attr on
the peers tibble produced by cog_find_peers.
2026-04-29 09:17:24 -04:00
jared a25ba5f348 docs: spec for per-year population denominators
Design doc for switching cog_spending / cog_revenue / cog_geographic_rollup
per-capita calculations from a static ACS 2018-2022 population to per-year
F-33 population already present in long.population. Covers shifting peer
matching to a user-selectable cohort year (defaulting to most recent
observed year), type-4/5 NA policy, provenance updates, and a new vignette
enumerating denominator sources for future extensibility.

Specs live in /specs (added to .Rbuildignore) since docs/ is reserved for
pkgdown output.
2026-04-29 09:12:32 -04:00
jared e7fa51eec7 Merge pull request 'feat(search): basket mode for cog_gov_search()' (#1) from feat/cog-gov-search-basket-mode into main
R-CMD-check / check (push) Successful in 1m34s
2026-04-28 15:09:44 -04:00
jared efc0bd16b1 fix(search): soft-fail on per-row excluded type and malformed regex name
R-CMD-check / check (push) Successful in 1m37s
R-CMD-check / check (pull_request) Successful in 1m37s
Cross-task review found two edge cases that violated the basket-mode
soft-fail contract:

- Per-row excluded type (e.g. type = c(NA, "special_district")) hit
  .coerce_type()'s abort inside the per-row resolver, killing the
  whole basket call. Now treated as no_match in the sidecar.
- Malformed regex in the substring fallback (e.g. name = "San(Diego")
  propagated DuckDB engine errors. .escape_regex() now backslash-
  escapes meta characters before the regexp_matches call. Utility-
  mode regex behavior is unchanged.

Plus a new public-surface test for the all-no-match case.
2026-04-28 14:16:37 -04:00
jared 24e67791a0 docs: roxygen, NEWS, and pkgdown for basket mode
Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.
2026-04-28 12:51:52 -04:00
jared 7475696853 test: smoke test cog_gov_search basket -> cog_spending pipe
Confirms a basket result pipes cleanly into the existing query verb
without any input-coercion friction.
2026-04-28 12:10:28 -04:00
jared 56fdd8e5b9 feat: export cog_basket_resolution() and cog_basket_unresolved()
Sidecar accessors for basket-mode results. cog_basket_resolution()
returns the full resolution tibble (drops candidates list-col by
default for readable printing). cog_basket_unresolved() filters to
ambiguous/no_match rows for iterative refinement.
2026-04-28 11:58:05 -04:00
jared ae04d4f48f refactor(search): split sidecar helpers into R/basket.R
Moves .build_sidecar(), .type_to_label(), .basket_summary_message()
out of R/search.R and into a new R/basket.R. No behavior change;
brings R/search.R back under its 350-line budget. R/basket.R will
also host the public sidecar accessors added in the next commit.
2026-04-28 11:53:40 -04:00
jared d9bf0552b7 feat(search): post-resolution summary message in basket mode
Emits a single cli::cli_inform when any input was ambiguous, missed,
or fell back to largest-population. Silent on clean baskets.
2026-04-28 11:41:28 -04:00
jared 98ed6318a6 feat(search): basket mode for cog_gov_search()
Vector name + state + type arguments dispatch to a per-row resolver
that produces a basket tibble with a 'resolution' sidecar attribute.
Utility mode (length-1 name) is unchanged.
2026-04-28 11:39:53 -04:00
jared 2f47f7ae67 feat(search): add disambiguation — largest_pop and ambiguous branches
Multi-row matches within a single govs_type pick the largest-population
row (status=largest_pop). Multi-row matches spanning >=2 types return no
basket row (status=ambiguous) with all candidates preserved for the
sidecar.
2026-04-28 11:08:26 -04:00
jared 392e818f52 feat(search): add substring fallback and no_match handling
.resolve_basket_row() now falls back to case-insensitive substring
match when no exact match is found, and short-circuits empty/whitespace
input to no_match. Disambiguation stub raises pending Task 5.
2026-04-28 11:06:30 -04:00
jared 4dcf1c72b2 feat(search): add per-row resolver — exact match branch
.resolve_basket_row() handles the exact-match case. Substring fallback
and disambiguation branches follow in subsequent commits.
2026-04-28 11:05:14 -04:00
jared 7fd29c919d feat(search): add basket-mode argument validator
Internal .validate_basket_args() handles length validation and
recycling of state/type from length 1. Foundation for basket mode.
2026-04-28 11:02:03 -04:00
jared 68b72faeae refactor(search): rename cog_gov_search() first argument to name
Pre-rename in preparation for basket mode. All existing callers in this
package and cog_explorer/ pass the first argument positionally, so this
rename is non-breaking. No deprecation alias added per design spec
(no external consumers; package is pre-release v0.1.0).

@param roxygen also updated to match the new formal; man/cog_gov_search.Rd
regenerated via devtools::document().
2026-04-28 10:47:56 -04:00
jared b3e544ebd8 docs: add CLAUDE.md agent instructions with current state + v0.1 roadmap
R-CMD-check / check (push) Successful in 1m29s
2026-04-27 13:25:41 -04:00
jared d65e9feefa ci: use testthat::test_local() instead of devtools::test()
R-CMD-check / check (push) Successful in 1m36s
2026-04-27 12:57:17 -04:00
jared a5f1767caf ci: add Gitea Actions workflow for R CMD check + testthat
R-CMD-check / check (push) Failing after 1m21s
2026-04-27 12:52:14 -04:00
jared af6f18f9d7 docs: add Developer notes section to README (testing + live-corpus release) 2026-04-27 12:50:40 -04:00
jared a640c9cc21 test: bundle fixture corpus + wire local-path test helpers
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.

setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).

helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.

test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).

Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
2026-04-27 12:48:58 -04:00
jared 488d03d74b feat: cog_categories
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().

Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:

  type    = NULL | "spending" | "revenue"
  pattern = regex matched case-insensitively on category

The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.

Also: fix @noRd placement in session.R so devtools::document() stops
warning.

Tests: +16 new / 181 total pass. check 0E/0W/0N.
2026-04-24 18:15:26 -04:00
jared a778d790d8 fix: govid input ergonomics + clearer missing-govid message
Two UX fixes surfaced by first real-user use:

1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
   a character vector OR a data.frame with a canonical_govid column
   (e.g. output of cog_gov_search() or cog_find_peers()). Shared
   .coerce_govid_input() helper in session.R. This lets the natural
   pipe work:

     cog_gov_search('MIAMI', state='FL', type='city') |>
       cog_spending(years=2022, category='Police')

   cog_peer_compare already accepted a data.frame for the peer arg;
   behavior there is unchanged.

2. .check_govids_in_scope() message reworded. The old text led with
   'v0.1 covers gov_types 0-3' which falsely implied the missing govids
   were scope-excluded types when the more common real cause is a typo
   or a guessed value. New message leads with typo + pre-2017 PID,
   mentions scope exclusion as one possibility, and points at
   cog_gov_search() as the recovery path.

Tests: 165 pass / 0 fail. check 0E/0W/0N.
2026-04-24 17:49:08 -04:00
jared 377eed1240 feat: cog_gov_search + cog_mirror + scope-aware verb behavior
Three pieces:

1. cog_gov_search: name/state/type search over canonical_fips_xwalk
   for resolving human-readable place names into canonical_govids.
   Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
   0-3 or name ('state','county','city','township') for type. Types
   4/5 emit an explanatory cli message and return an empty tibble
   (v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
   50 states + DC + territories; FIPS 66 = GU (not GA).

2. cog_mirror: downloads manifest-listed files to a local directory
   with SHA-256 idempotency (files with matching hash return status
   'cached'). Supports HTTP and local-path fixture URLs. Round-trip
   test: mirror + re-open against the mirror + query Broward 2020
   returns identical results.

3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
   queries canonical_fips_xwalk for the requested govids, emits a
   cli_inform listing any missing ones, and records the found/missing
   sets under provenance$scope. Wired into cog_spending (and
   transitively into cog_revenue, cog_geographic_rollup,
   cog_peer_compare via their cog_spending calls).

Also: dropped dbplyr from Imports (unused).

Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
2026-04-24 11:38:32 -04:00
jared cc4d21ec82 feat: cog_find_peers + cog_peer_compare
cog_find_peers selects peers from canonical_fips_xwalk by same-type,
same-state, and population-range (ratio or absolute) criteria,
ordered by |log(pop_ratio)| ascending.

cog_peer_compare accepts the find_peers result (or a plain character
vector of govids), pulls spending for target + peers via cog_spending,
and appends summary rows (summary_p25/p50/p75) so the whole result
can be faceted by `role` in a single ggplot call. Summary rows honor
per_capita + adjust_to_year by picking the right value column.
target_rank reports the target's rank among target+peers at max(years).

Provenance is rewritten with verb = cog_peer_compare and peer_count.

Also: globalVariables('.data') in zzz.R to silence R CMD check on
tidy-eval pronouns.

Tests: 21 new / 120 total pass. devtools::check() 0E/0W/2N.
2026-04-24 11:17:46 -04:00
jared a6b53de2a7 feat: cog_geographic_rollup
Wraps cog_spending across a named list of state/county/city layers,
tagging each row with its `layer` and attaching a scope_note that
documents geographic-scope caveats (state totals are statewide, county
totals include areas outside a listed city, city proper excludes
special districts). Per-capita uses each layer's own population from
the canonical_fips_xwalk.

Provenance is inherited from cog_spending but rewritten to reflect
the outer verb (verb, call, layers).

Tests: 19 new / 99 total pass. devtools::check() 0E/0W/2N.
2026-04-24 11:00:20 -04:00
jared c682e6547d feat: cog_spending + cog_revenue + cog_explain
Three core query verbs over the spending_annotated / revenue_annotated
DuckDB views. Each verb accepts vector govid, vector years, optional
category filter, per_capita flag, and adjust_to_year for CPI-U
real-dollar conversion (bundled index).

Amounts are returned in full USD (SUM(amt) * 1000) so callers can
freely rescale to millions/billions. The $1,000s -> $USD conversion
is recorded in provenance$transformations$units_conversion.

Every result carries an attr(., "provenance") list matching
inst/schemas/provenance-v1.json. cog_explain() prints the structured
form via cli or returns the raw list for MCP/JSON consumers.

Also: .fetch_or_cache_manifest() now handles local fixture paths so
tests can point USCOGDATA_FIXTURE_URL at the pipeline publish_cache/
without a working HTTP server.

Tests: 80 pass / 0 fail. devtools::check() 0E/0W/2N (both notes
pre-existing / environmental).
2026-04-24 09:49:09 -04:00
jared cea7a241c5 chore: require corpus schema_version = 3
Tracks the cog_pipeline corpus bump from 2 → 3 (year=2012 partition is
now modern-only; v2 shipped a broken mixed-source partition). session
open aborts with a clear message if a reader hits a legacy v2 corpus.
2026-04-23 14:35:55 -04:00
jared 2b36f860f9 feat: bundle CPIAUCSL + .inflate() helper
Adds R/sysdata.rda with the annual-average CPIAUCSL index (1947-2026,
80 years, sourced from FRED) and an internal .inflate() helper that
converts nominal amounts between two years via the ratio of CPI values.

Bundling CPI in the package (rather than publishing a cpi_annual.parquet
in the corpus) matches the reader-specification intent: real-dollar
conversion is a verb-level option, not a corpus-level artifact, so the
target_year stays flexible at query time.

data-raw/cpi_annual.R carries the one-shot FRED refresh used to build
sysdata.rda. Re-run when the CPI series needs to roll forward.

Verified: CPI(2000)/CPI(2021) ≈ 0.635, matching the ~0.63 sanity
anchor in cog_explorer's existing inflation logic.

Tests: tests/testthat/test-adjust.R covers the known 2000→2021 ≈ 1.574x
anchor, vectorized from_year, error cases for out-of-range years, and
NA-amount preservation. 14 pass / 0 fail.
2026-04-23 13:54:40 -04:00
jared 8d2a8d341f feat: view registration from inst/sql/
Seven DuckDB views register on session open: long (raw), spending_long
and revenue_long (prefix-filtered, NOT is_aggregate per reader-spec §4),
canonical_fips_xwalk and summary_categories (identity), spending_annotated
and revenue_annotated (LEFT JOIN xwalk + categories for verb composition).

File prefix `NN-` enforces creation order so *_annotated views resolve
their *_long dependencies.

Deviation from plan: series_breaks/cpi_annual/legacy_aggregate_map and
*_with_transforms views are deferred — their backing parquets are not
in the v0.1 corpus (manifest.files.metadata only lists canonical_fips_xwalk
and summary_categories). The verbs will compute CPI adjustment verb-side
against a bundled cpi table in a later task.
2026-04-23 10:21:36 -04:00
jared f7035049e7 feat: package skeleton — DESCRIPTION, NAMESPACE, session/manifest/cache/views
Minimal skeleton for uscogdata v0.1. Internal session layer with lazy
cog_open(), manifest fetch+validate+cache, view registration placeholder.
Depends on DuckDB >=1.0, httr2, jsonlite. inst/schemas/provenance-v1.json
ships the structured provenance JSON Schema.
2026-04-23 09:05:54 -04:00