107 Commits
Author SHA1 Message Date
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00
jared 1d553a788f fix: surface ALL-scoped series breaks in provenance (#19)
R-CMD-check / check (push) Successful in 3m1s
R-CMD-check / check (pull_request) Successful in 3m1s
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:

  SB085  1977  dollar precision across the 1976/1977 boundary
  SB087  2002  imputation exclusion FY2002-2006
  SB194  2012  dense -> sparse representation change
  SB086  2017  government id scheme change

SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.

Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.

On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.

Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.

Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
2026-07-30 10:27:48 -04:00
jared c375c55da7 fix: regenerate the bundled fixture against the sparsified corpus (#18)
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m51s
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.

Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.

Three test repairs, each a real consequence of sparsification rather than
a number to bump:

  test-categories.R          "assistance" joined the spending subtype
                             vocabulary with the J-prefix codes.

  test-spending.R            The harmonization block counts rows that
                             exist. Broward's E21/F21/G21 were zero-pads
                             and are gone, so the anchor moves to FL state,
                             whose three NA-mapped rows carry $2.83B --
                             the amount accounting was previously asserted
                             only against 0 and could not have caught a
                             bug. Broward keeps a test of its own, now
                             asserting the zero-pads are absent.

  test-expenditure-concept.R Coverage-gap suggestions are presence-based.
                             AL state's only FY2011 B47 cell was an
                             explicit zero, so ig_federal_b47_wide stopped
                             being a candidate there; FL state carries a
                             real amount, so the counterpart guard is
                             exercised against a suggestion that fires.

test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.

Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
2026-07-30 10:20:09 -04:00
jared 9233c3d18e test: add failing tests for Madison walkthrough findings
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m58s
Six skipped tests, one per issue opened from the Madison walkthrough audit
(docs/walkthroughs/FINDINGS.md in cog_explorer). Each asserts the desired
behaviour, so it fails today and goes green when the fix lands; each is
guarded by a single skip() naming its issue and finding IDs, so the suite
stays green and activating a test is a one-line deletion.

  test-expenditure-concepts.R             #11  F-012, F-017, F-018
  test-revenue-concept-insurance-trust.R  #12  F-014
  test-coverage-disclosure.R              #13  F-020, F-023
  test-peer-summary-scope.R               #14  F-021
  test-amount-units-documented.R          #15  F-004
  test-gov-search-literal-match.R         #16  F-025

helper-walkthrough-raw.R reads the corpus's long parquet partitions directly,
bypassing uscogdata's SQL views. Every expected amount comes from there rather
than from the verb under test - verifying an absence through the filter that
creates it proves nothing, which was the most common defect in the audit itself.

Verified: with the skips removed all six fail (or error) against the bundled
fixture; with them in place the full suite is 576 pass / 0 fail / 6 skip.
2026-07-29 00:14:11 -04:00
jaredandClaude Opus 5 d258cef8c5 fix: gate direct-suppressed flag/note on an actually-covering recipe
R-CMD-check / check (push) Successful in 2m53s
R-CMD-check / check (pull_request) Successful in 2m53s
.detect_direct_suppressed() equated "no Direct sibling row" with "Direct
was suppressed", but the dominant real cause is a government with
genuinely no direct spending in that category (e.g. a state funding K-12
entirely through school districts) -- correct, ordinary data, not
suppression. Measured: 32 of 50 states false-flagged on a clean FY2019
category = NULL total query, and all 141 flagged rows across 50 states x
{2011, 2019} fell back to "no covering recipe found" instead of naming one
-- including AL Corrections, which names corrections_combined correctly
when category is supplied explicitly.

Both the flag and its row note are now gated on a harmonization recipe
actually covering that exact (year, canonical_govid, category) triple, via
a new .covering_recipes() helper that runs the same generic recipe join
per-row regardless of whether the caller supplied a category filter.
.notes_column() takes the precomputed note vector directly instead of
searching a category-gated suggestions list; .direct_suppressed_note() is
removed (its "no recipe found" fallback no longer applies -- if no recipe
covers a triple, it isn't suppression).

Also recomputes two total-spending.Rmd figures the prior wave never
actually reconciled with its own "measured against the fixture" caption:
State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M
(flat 188.3%, now 144%-189% varying by year).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 13:05:58 -04:00
jared a4eb80d823 feat: cog_explain() prints the expenditure concept and direct-suppressed flag (I1)
.print_provenance() printed "Basis:" but nothing about Direct vs Total
-- the most consequential switch this branch adds to cog_spending() was
invisible in the package's designated "what am I looking at" verb. Add
a "Concept: direct|total (<note>)" line next to Basis, and surface a
cli warning when provenance$expenditure_concept_direct_suppressed is
TRUE (see the C1 fix), so the suppressed-Direct case is visible in the
human-readable explain output, not just in the structured provenance.
2026-07-27 12:06:28 -04:00
jared aba7ffbac2 test: cover C1 direct-suppressed handling, C2 corpus guard, and I2 candidate filter
Adds regression coverage for the three preceding fixes:
- "total" on a legacy aggregate-only family (AL Corrections 2011) now
  fires recipe suggestions, flags provenance$expenditure_concept_direct_
  suppressed, and names a recovering recipe in the affected row's notes,
  plus a contrast test confirming the flag stays FALSE when the Direct
  leg is present.
- expenditure_concept = "total" aborts with class
  uscogdata_ig_categories_unsupported against a corpus whose
  summary_categories carries no M/L rows (new
  with_corpus_missing_ig_categories() fixture helper), is unaffected for
  "direct" on the same corpus, and still works on a corpus that does
  carry M/L rows.
- an M/L recipe (corrections_ig_local_combined) no longer appears as a
  raw suggestion for a Direct-flavored cog_spending() call.
2026-07-27 12:06:22 -04:00
jared c7260cb20c fix: scope total's coverage-gap detection to the Direct leg; require IG category rows (C1, C2)
C1: spending_long/spending_long_harmonized filter NOT is_aggregate but
ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a
legacy aggregate-only family (e.g. Corrections pre-2012) can survive on
the IG leg while Direct is suppressed. expenditure_concept = "total"
then UNIONs an IG-only figure that reads as a plausible Total, and the
coverage-gap suggestion machinery -- fed the UNION'd result -- saw the
surviving IG row as coverage and stayed silent.

  (a) .build_suggestions() is now fed a Direct-leg-only view of the
      result (IG rows filtered out before the gap-years computation),
      so the recipe hints fire for "total" exactly as they do for
      "direct".
  (b) Any row where IG has dollars but Direct has none for the same
      (year, canonical_govid, category) is now flagged: the row's
      `notes` name the recovering recipe (drawn from the Direct-leg
      suggestions), and provenance gains an explicit
      `expenditure_concept_direct_suppressed` boolean plus an appended
      warning on `expenditure_concept_note` -- both cheap for a
      downstream consumer (cog-api passes provenance through verbatim)
      to test, rather than silently asserting Direct + IG when that
      arithmetic didn't happen.

Measured before/after on AL state government, Corrections, 2011:
"total" already correctly returns the corpus's actual IG-only figure
($31,358,000, vs. true Direct of $521,651,000 via recipe =
"corrections_combined"), but before this fix it did so with 0
suggestions and an unqualified "Total = Direct + IG" note; after, it
fires 3 recipe hints and both the row notes and provenance say plainly
that Direct is unavailable through this basis.

C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59
with no schema_version bump, so schema_version can't gate "total" --
a pre-#59 corpus can report any supported schema_version and still
have zero M/L category rows, in which case ig_annotated's LEFT JOIN
silently produces NA category/spend_subtype (0 rows for a specific
category, or one invisible NA-subtype group for category = NULL). New
.require_ig_categories() checks summary_categories directly and aborts
with class uscogdata_ig_categories_unsupported, naming PR #59 and
directing the user to a newer corpus.

Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose
synthetic summary_categories carries only one E36 row) by asserting
the new guard fires against that same connection, rather than leaving
the two silently contradictory.
2026-07-27 12:06:03 -04:00
jared 54ece11867 Merge origin/main (PR #8: corpus URL trailing-slash normalization) into feat/expenditure-concept
Local main was stale at fa40266 when this branch was created, so it was
missing 748ca4a. Merging rather than rebasing to preserve the reviewed
commit SHAs recorded in the SDD ledger.
2026-07-27 11:26:17 -04:00
jared 24b2ff7d8c fix: gate IG-counterpart matching to the direct-expenditure flow family
Review found the suffix-set match alone is unsafe: revenue-side recipes
(ig_federal_b47_wide, ig_state_c47_wide, ig_local_d47_wide, and their *_89
siblings) coincidentally share exact suffix sets with M/L expenditure
recipes despite representing a different flow direction. Reachable today via
a mis-scoped cog_spending(category = "IG Federal") call, not just
cog_revenue(). Thread flow_prefixes (same parameter .build_harmonization_block
already uses) through .build_suggestions()/.attach_ig_counterparts() and
require a firing recipe's own prefixes to be both in the calling verb's flow
family and within {E,F,G} before searching the M/L catalog.
2026-07-27 11:08:57 -04:00
jared c28712f62f feat: name the intergovernmental counterpart in firing recipe suggestions
Closes uscogdata #6 item 4. Only extends suggestions that already fire -- a
concept hint on every healthy call would be noise.
2026-07-27 10:46:22 -04:00
jared 7913b0f664 feat: record expenditure_concept in provenance and its JSON schema
Always populated, never implicit, so a downstream artifact says which concept
produced it. cog-api passes provenance through verbatim.
2026-07-27 10:27:12 -04:00
jared 887acf7e81 test: add missing cog_peer_compare coverage to expenditure_concept tests
- 'both cross-government verbs still accept the direct default' now tests both verbs
- 'the refusal message names the fix and the reason' now asserts both functions name
  themselves correctly in their error messages (cog_geographic_rollup vs cog_peer_compare)

Addresses coordinator feedback to prevent test coverage gaps and ensure the helper's
verb name argument is pinned correctly.
2026-07-27 10:19:55 -04:00
jared 81fd1a5279 feat: refuse expenditure_concept = total in the cross-government verbs
Owner ruling R1. Combining Census Total across governments counts
intergovernmental transfers twice, and these results land in Tableau where a
warning would be invisible -- so this is a hard error whose message names the
fix and the reason.
2026-07-27 10:14:16 -04:00
jared e2088458e1 fix: address Task 3 code review (bool_or, invariant tests, guards, docs)
Nine review items on the expenditure_concept = direct|total feature:

- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
  bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
  dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
  wide-era data puts a $0 leaf row in the same group as the real aggregate
  row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
  across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
  aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
  scoped to the aggregate flag rather than bare code presence (M89/L89
  continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
  pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
  retained, NULL harmonized_code coalesced, L-- excluded), rather than
  relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
  necessary (not just incidental) against a corpus whose `long` genuinely
  lacks a harmonized_code column, and rewrote the misleading "v5-only
  parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
  test, guarded .verb_spendrev() against total on a non-spending
  view_base, added a roxygen caveat against summing total across levels
  of government, and replaced an uncheckable corpus-wide SQL comment
  figure with a fixture-verifiable one.

Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
2026-07-27 10:00:07 -04:00
jared fefd4fe969 feat: expenditure_concept = direct|total in cog_spending()
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
2026-07-27 09:28:17 -04:00
jared 7ed1da9b79 fix: drop the inert K prefix from the spending flow prefixes
K matches zero rows corpus-wide (audited pipeline-side). Numerically inert;
removed so the code stops implying a prefix the data never had.
2026-07-27 09:16:41 -04:00
jared 9240a18ea3 chore: regenerate fixture corpus with the intergovernmental category rows
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).

cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.

Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
2026-07-27 09:11:10 -04:00
jared 748ca4a56e fix(#3): normalize the corpus URL's trailing slash at resolution
R-CMD-check / check (pull_request) Successful in 2m33s
R-CMD-check / check (push) Successful in 2m42s
Investigating "gov search doesn't work" (#3) turned up two separate things.

THE REPORTED SYMPTOM IS ALREADY FIXED.
#3 reported `cog_gov_search("Orange")` dying in jsonlite with
`lexical error: invalid char in json text. <html> <head>`. That was fixed the
same day the issue was filed, by 8743472 "fix(manifest): actionable errors when
USCOGDATA_URL is unset or returns non-JSON" (issue filed 2026-05-27 11:30;
commit 2026-05-27). The issue was simply never closed. Verified now: injecting
an HTML manifest.json raises a typed `uscogdata_invalid_manifest` condition
naming the likely causes, with the raw parse error demoted to a footnote, and
`cog_gov_search("Orange")` returns 62 rows against the live corpus.

THE ROOT CAUSE OF THAT HTML WAS STILL LIVE -- and is what this commit fixes.

Every consumer builds locations by CONCATENATION:
  manifest.R:95   paste0(url, "manifest.json")
  mirror.R:48,125 paste0(url, e$path)
  views.R         the parquet glob
and mirror.R:104 documents the invariant outright ('url ends in "/"'). The
error messages tell users to set `"<url-or-local-path>/"`. But `.resolve_url()`
was a bare `.cfg("url")` passthrough -- the invariant was assumed everywhere and
enforced nowhere.

So a URL entered without the slash failed silently and misleadingly:
  HTTPS -> ".../downloadmanifest.json"; the host answers with an HTML 404 page,
           which lands in the JSON parser as EXACTLY the #3 symptom -- and the
           guard then blames "login page / 404 / wrong share" when the real
           cause was one missing character.
  local -> ".../corpusdata/long/**/*.parquet" and a DuckDB "No files found".

Reproduced both: pointing USCOGDATA_URL at the bundled fixture without a
trailing slash gave
  No files found that match ".../fixture_corpusdata/long/**/*.parquet"

Normalizing once at resolution fixes every consumer at the same time, rather
than having each call site re-derive the invariant. An empty setting passes
through untouched so manifest.R's "not configured" guard still fires instead of
the value degrading into a bare "/" filesystem root.

RED->GREEN: 4 tests added, 2 failed first (append-missing-slash, local-path
normalization); the already-correct cases (slash present, empty setting) passed
throughout and pin them against regression. Same fixture path that produced the
DuckDB error above now returns 62 rows.

Suite: FAIL 0 | WARN 0 | SKIP 0 | PASS 471 (was 463; +8 = the new tests).
2026-07-25 18:37:42 -04:00
jaredandClaude Opus 4.8 e813ffd3aa feat: accept corpus schema v6 (FIPS geography harmonization); v6 fixture
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.

Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
  * test-manifest.R: v6 accepted; boundary rejection moves to v7.
  * test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
    extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
    Education NEC local, SB184-186). Broward's 2011 partition zero-pads
    exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
    excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 13:16:11 -04:00
jared 77f48047b1 fix: exercise real harmonized-view SQL in tests; unambiguous recipe provenance
R-CMD-check / check (pull_request) Successful in 3m5s
R-CMD-check / check (push) Successful in 3m1s
test-views.R's harmonized-view test previously ran a hand-rolled REPLACE
query with no WHERE clause, so a regression in any of
inst/sql/22-spending_long_harmonized.sql / 23-revenue_long_harmonized.sql's
three predicates (NOT is_aggregate, harmonized_code IS NOT NULL, the
E/F/G/K or T/A/U/B/C/D prefix filter) would go uncaught. Replaced it with a
test that reads the real SQL files off disk, substitutes {url} exactly as
.register_views() does, and executes them (plus their 10-long.sql
dependency) against a synthetic hive-partitioned parquet tree written via
DuckDB's own COPY ... TO (FORMAT PARQUET) (no arrow dependency, matching
this package's existing convention). Ten rows are crafted so each predicate
is independently falsifiable by a specific row; manually broke each
predicate in turn to confirm the test fails exactly as expected, then
restored the SQL files (see the task report for the RED-phase transcript).

Also fixes a provenance ambiguity: a recipe= query bypasses
spending_annotated(_harmonized)/revenue_annotated(_harmonized) entirely
(.run_recipe() joins `long` directly), so basis= has no effect on it, but
provenance was still reporting basis = "harmonized"/"raw" (whatever the
argument resolved to) with harmonization$applied = FALSE alongside it --
misleading, since it looks like harmonization was evaluated and found
nothing to exclude rather than "not applicable here." Recipe results now
report basis = "recipe" with an inert harmonization block carrying an
explicit note, regardless of what basis= was passed.
2026-07-18 23:55:48 -04:00
jared 4de915b557 feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.

Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).

Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
2026-07-18 23:33:02 -04:00
jared 7818cd2b1a feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.

Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.

Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
2026-07-18 23:19:17 -04:00
jared 4d61692f05 feat: export cog_manifest() accessor + pin CPI coverage 1967-present
R-CMD-check / check (push) Successful in 2m47s
R-CMD-check / check (pull_request) Successful in 2m57s
2026-07-18 13:40:54 -04:00
jared 92c9a7382e test: re-baseline canonical_govid literals to 12-char namespace
Swaps every hardcoded 9-char canonical_govid literal (Broward County,
Fort Lauderdale City, Florida/Alabama state govts, Bexar/Tarrant/Wayne
counties, San Diego/Oakland/Miami/Austin cities) for its 12-char Phase P
equivalent, resolved by name+type+state against the regenerated fixture
xwalk. Also updates two gov_name search patterns that no longer match
under Phase P canonical naming ("FLORIDA STATE GOVT" -> "FLORIDA"; the
"Miami" substring test now pins type = "city" since MIAMI-DADE COUNTY's
canonical name now also contains "Miami", which would otherwise make the
match ambiguous across govs_types instead of resolving via largest-pop).
Underlying per-year population figures for Broward County and Alabama
are unchanged, so no expected data-value literals needed recomputation.
Suite: 126 test blocks / 336 expectations, 0 FAIL / 0 WARN / 0 SKIP.
2026-07-11 09:32:02 -04:00
jared 874347242b fix(manifest): actionable errors when USCOGDATA_URL is unset or returns non-JSON
R-CMD-check / check (push) Successful in 2m5s
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:

1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
   when the resolved URL is empty or still contains the
   `REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
   `Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
   remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
   it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
   Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
   existing caches with non-JSON content are silently refetched instead
   of returning a parse error to the caller.

Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
2026-05-27 11:57:31 -04:00
jared 919548685b polish(per-year-pop): expand popyear in cog_explain + propagate pop_range
R-CMD-check / check (push) Failing after 1m30s
Final-review followups:

1. cog_explain rendered the popyear range as raw 2-digit values
   ("popyear range: 19-20"), which a user could read as the years 19-20.
   Added .expand_popyear() helper to format as 4-digit calendar years
   (2019-2020). Pivot at 70 to handle pre-2000 vintages if the corpus
   ever extends backward.

2. cog_peer_compare provenance was missing pop_range and is_ratio,
   omitted from the spec-required reproducibility metadata.
   cog_find_peers now stamps both as tibble attributes; cog_peer_compare
   reads them through to provenance$pop_range and provenance$is_ratio.

Test coverage extended: explain test asserts the 4-digit format and
rejects the old 2-digit form; peer-compare test asserts pop_range +
is_ratio propagate end-to-end.

326 PASS / 0 FAIL.
2026-04-29 19:16:48 -04:00
jared 916212c327 feat(explain): render new per-capita provenance fields
cog_explain() now prints denominator_source, popyear_range, and
pop_source_counts under the Transformations section.
2026-04-29 19:02:37 -04:00
jared a2ced368f5 feat(provenance): record per-year denominator metadata
Updates transformations\$per_capita with the new denominator_source string,
popyear_range, and pop_source_counts. .attach_per_capita stashes
popyear_range on the result; .verb_spendrev strips the helper attr after
provenance is built.
2026-04-29 19:00:16 -04:00
jared c334be7706 test(rollup): per-year denominator + provenance expectations (failing) 2026-04-29 17:50:07 -04:00
jared a92450ff76 feat(peers): stamp cohort_year on cog_peer_compare results
Reads attr(peers, 'cohort_year') when the caller passed a cog_find_peers()
tibble; NA when the caller passed a bare character vector. Stamped as a
constant column on the result and recorded in provenance alongside the
cohort govids.
2026-04-29 17:44:25 -04:00
jared 54dd40a61d feat(peers): cog_find_peers uses per-year population
Adds optional 'year' argument (defaults to most recent observed year for
the target). Filters and ranks candidates by gov_population_yearly.population
at that year. Returned column renamed population_acs -> population.
Cohort year attached as attr(x, 'cohort_year').

Adds .resolve_cohort_year() helper. Updates test assertions to use
'population' column name. Regenerates man/cog_find_peers.Rd.
2026-04-29 17:26:32 -04:00
jared 807ed35cb7 test(peers): per-year cohort expectations (failing)
Three RED tests that drive Task 6's cog_find_peers() rewrite:
- defaults year to most recent observed (expects cohort_year attr + population column)
- honors explicit year= argument (expects cohort_year attr)
- errors with "no observed population" for unobserved year
2026-04-29 17:17:52 -04:00
jared 9238b04b69 feat(notes): concatenate notes; flag unavailable population
.notes_column now joins multiple per-row notes with '; '. Adds the
'No population denominator available for this gov type' note when
pop_source is 'unavailable'. Gracefully handles absent pop_source
(per_capita = FALSE). Two new tests: one corpus-level (census_f33
branch) and one synthetic unit test covering multi-note concatenation.
2026-04-29 17:03:05 -04:00
jared e4a105013e test(spending): document fixture-pop origin in per-year-denominator test 2026-04-29 16:29:24 -04:00
jared 21b3d66c0e test(spending): per-year denominator expectation (failing)
Adds a RED test asserting that cog_spending(per_capita = TRUE) divides
by the per-year F-33 population (Broward 2019: 1,935,878; 2020: 1,952,778)
rather than the static ACS value (1,940,907). Uses absolute-tolerance
expect_true(abs(...) < 1) instead of expect_equal(tolerance=1) because
testthat 3 treats the tolerance argument as relative.
2026-04-29 16:20:41 -04:00
jared df3fe3731b test(views): document hardcoded fixture-pop origin in gov_population_yearly test 2026-04-29 15:37:05 -04:00
jared ed9658d267 feat(sql): add gov_population_yearly view
Exposes one row per (year, canonical_govid) drawn from long.population.
Used by per-capita denominators and peer matching.
2026-04-29 15:12:53 -04:00
jared efc0bd16b1 fix(search): soft-fail on per-row excluded type and malformed regex name
R-CMD-check / check (push) Successful in 1m37s
R-CMD-check / check (pull_request) Successful in 1m37s
Cross-task review found two edge cases that violated the basket-mode
soft-fail contract:

- Per-row excluded type (e.g. type = c(NA, "special_district")) hit
  .coerce_type()'s abort inside the per-row resolver, killing the
  whole basket call. Now treated as no_match in the sidecar.
- Malformed regex in the substring fallback (e.g. name = "San(Diego")
  propagated DuckDB engine errors. .escape_regex() now backslash-
  escapes meta characters before the regexp_matches call. Utility-
  mode regex behavior is unchanged.

Plus a new public-surface test for the all-no-match case.
2026-04-28 14:16:37 -04:00
jared 7475696853 test: smoke test cog_gov_search basket -> cog_spending pipe
Confirms a basket result pipes cleanly into the existing query verb
without any input-coercion friction.
2026-04-28 12:10:28 -04:00
jared 56fdd8e5b9 feat: export cog_basket_resolution() and cog_basket_unresolved()
Sidecar accessors for basket-mode results. cog_basket_resolution()
returns the full resolution tibble (drops candidates list-col by
default for readable printing). cog_basket_unresolved() filters to
ambiguous/no_match rows for iterative refinement.
2026-04-28 11:58:05 -04:00
jared d9bf0552b7 feat(search): post-resolution summary message in basket mode
Emits a single cli::cli_inform when any input was ambiguous, missed,
or fell back to largest-population. Silent on clean baskets.
2026-04-28 11:41:28 -04:00
jared 98ed6318a6 feat(search): basket mode for cog_gov_search()
Vector name + state + type arguments dispatch to a per-row resolver
that produces a basket tibble with a 'resolution' sidecar attribute.
Utility mode (length-1 name) is unchanged.
2026-04-28 11:39:53 -04:00
jared 2f47f7ae67 feat(search): add disambiguation — largest_pop and ambiguous branches
Multi-row matches within a single govs_type pick the largest-population
row (status=largest_pop). Multi-row matches spanning >=2 types return no
basket row (status=ambiguous) with all candidates preserved for the
sidecar.
2026-04-28 11:08:26 -04:00
jared 392e818f52 feat(search): add substring fallback and no_match handling
.resolve_basket_row() now falls back to case-insensitive substring
match when no exact match is found, and short-circuits empty/whitespace
input to no_match. Disambiguation stub raises pending Task 5.
2026-04-28 11:06:30 -04:00
jared 4dcf1c72b2 feat(search): add per-row resolver — exact match branch
.resolve_basket_row() handles the exact-match case. Substring fallback
and disambiguation branches follow in subsequent commits.
2026-04-28 11:05:14 -04:00
jared 7fd29c919d feat(search): add basket-mode argument validator
Internal .validate_basket_args() handles length validation and
recycling of state/type from length 1. Foundation for basket mode.
2026-04-28 11:02:03 -04:00
jared a640c9cc21 test: bundle fixture corpus + wire local-path test helpers
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.

setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).

helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.

test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).

Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
2026-04-27 12:48:58 -04:00
jared 488d03d74b feat: cog_categories
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().

Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:

  type    = NULL | "spending" | "revenue"
  pattern = regex matched case-insensitively on category

The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.

Also: fix @noRd placement in session.R so devtools::document() stops
warning.

Tests: +16 new / 181 total pass. check 0E/0W/0N.
2026-04-24 18:15:26 -04:00
jared a778d790d8 fix: govid input ergonomics + clearer missing-govid message
Two UX fixes surfaced by first real-user use:

1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
   a character vector OR a data.frame with a canonical_govid column
   (e.g. output of cog_gov_search() or cog_find_peers()). Shared
   .coerce_govid_input() helper in session.R. This lets the natural
   pipe work:

     cog_gov_search('MIAMI', state='FL', type='city') |>
       cog_spending(years=2022, category='Police')

   cog_peer_compare already accepted a data.frame for the peer arg;
   behavior there is unchanged.

2. .check_govids_in_scope() message reworded. The old text led with
   'v0.1 covers gov_types 0-3' which falsely implied the missing govids
   were scope-excluded types when the more common real cause is a typo
   or a guessed value. New message leads with typo + pre-2017 PID,
   mentions scope exclusion as one possibility, and points at
   cog_gov_search() as the recovery path.

Tests: 165 pass / 0 fail. check 0E/0W/0N.
2026-04-24 17:49:08 -04:00