Commit Graph
106 Commits
Author SHA1 Message Date
jaredandClaude Sonnet 5 d0d724c4c8 fix(suggestions): guard empty candidates and restore year typing in .query_covered_years()
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m15s
Addresses roborev jobs 205/208 (reviewing pre-merge draft/original commits
of #33/#34, now split and merged as #69/#70):

- Restore `res$year <- as.integer(res$year)` in .query_covered_years(),
  dropped during the #33 rebuild as apparently-dead code. Its own roxygen
  promises an integer `year` column; without the coercion the
  implementation no longer matches that documented contract.
- Guard .query_covered_years() against empty `candidates`, mirroring the
  existing `gap_years` guard. Not reachable via .build_suggestions() today
  (candidates is checked non-empty before this is called), but the
  extracted helper is independently callable and previously built a
  malformed `WHERE r.recipe_id IN ()` clause for a hypothetical direct
  caller with no candidates.
- Add a direct fixture-only unit test of .query_candidate_recipes()'s
  category_type filtering (#34) in both flow directions -- queries only
  summary_categories/harmonization_recipes metadata, so it runs without
  skip_if_no_corpus(), unlike the two existing end-to-end tests.

1080 tests pass (2 skipped live-corpus).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:28:01 -04:00
jaredandClaude Sonnet 5 392643bd74 fix(suggestions): scope candidate recipes by category_type (#34)
R-CMD-check / check (push) Successful in 4m15s
R-CMD-check / check (pull_request) Successful in 4m25s
.query_candidate_recipes() (extracted in #33) now filters candidates by
category_type ('expenditure' vs 'revenue'), derived from the calling
verb's own flow_prefixes (E/F/G -> 'expenditure', else 'revenue').

Without this, a category shared across both flow families in
summary_categories leaked cross-family recipes: cog_revenue(category =
"Corrections") surfaced the expenditure-only corrections_combined recipe
(E04/E05) merely because "Corrections" is also a spending category name,
and cog_spending(category = "IG Federal") surfaced the revenue-only
ig_federal_b47_wide recipe. Both are wrong: following either hint would
attribute dollars to the wrong flow, or (IG Federal) fire the
coverage-gap machinery for a category the calling verb structurally
cannot report on at all.

Updates the two tests this changes the expected behavior of:
- "a mis-scoped cog_spending() call never attaches an M/L counterpart to
  a revenue-flavored recipe" (test-expenditure-concept.R): IG Federal is
  revenue-only, so a spending call now finds zero candidates outright
  rather than firing the suggestion and then blocking its M/L
  counterpart as a second-order check.
- "cog_revenue never suggests expenditure-only recipes"
  (test-recipes.R, was "I1: ... never fabricates suppressed dollars"):
  corrections_combined is expenditure-only, so a revenue call now never
  considers it as a candidate, rather than considering it and reporting
  zero suppressed dollars.

All 1078 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree stacked on the #33 refactor.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:11:16 -04:00
jared 2bb9d41d72 feat: limit/offset on cog_gov_search() and cog_balances() (#57)
R-CMD-check / check (pull_request) Successful in 4m22s
R-CMD-check / check (push) Successful in 4m28s
verbs were left materializing everything and slicing in R -- the pattern behind
the 2026-08-06 production incident. cog_gov_search() had no LIMIT at all, so an
unfiltered call returns the entire 40,336-row crosswalk.

Extracted the #39 machinery into R/pagination.R first (.validate_pagination(),
.paginate_sql(), .take_pagination_total()) rather than growing a third inline
copy: three definitions of what total_rows means is three places for it to
drift. Conflict refusals stay at the call sites because each verb's conflict
set differs. .verb_spendrev() now uses the shared helpers and is unchanged in
behaviour.

The empty-page fallback query is now passed as a thunk, so the unpaginated SQL
is only BUILT when an offset actually lands past the end instead of on every
paged call.

Two things #57 did not anticipate:

- cog_gov_search()'s ORDER BY was not a total order. population_acs DESC NULLS
  LAST leaves ties -- and the whole NULL block -- in scan order, so two requests
  can order them differently and a paged sweep duplicates one row while dropping
  another. Added canonical_govid as tiebreaker. Unpaginated output changes only
  in the relative order of already-tied rows.

- Basket mode returns one resolved row per requested name plus a sidecar
  covering all of them, so a page of it is not a page of anything the caller
  asked for. Refused with uscogdata_basket_pagination_conflict rather than
  silently ignoring the arguments.

Both default to NULL, so cog-api adopts them behind its existing formals()
probe with no lockstep deploy.

Suite: 1067 passed, 0 failed, 0 warnings (2 pre-existing live-corpus skips).
2026-08-10 19:27:11 -04:00
jared f4ab9b6d90 feat: cog_open() honours a DuckDB thread and memory budget (#60)
R-CMD-check / check (push) Successful in 4m11s
R-CMD-check / check (pull_request) Successful in 4m11s
cog_open() connected with a bare dbConnect() and set no resource pragmas, so
DuckDB claimed every visible core. Right for one interactive session on a
dedicated machine; wrong for a server, where cog-api runs two replicas on an
8-core host budgeted 4 and each replica independently claims all 8.

USCOGDATA_DUCKDB_THREADS and USCOGDATA_DUCKDB_MEMORY_LIMIT now resolve through
.cfg() -- inheriting the env var > option > default precedence USCOGDATA_URL
already had -- and are applied as pragmas when the connection is created.

Unset issues NO pragma, so an unconfigured session is byte-identical to before.
That negative property is asserted directly against a connection opened the
pre-change way rather than against a hardcoded core count.

.cfg() returns an env var as character, so both resolvers coerce and validate
rather than trusting the type: sprintf("SET threads TO %d", "4") would
otherwise abort inside the connection path with an error naming the pragma
instead of the setting the operator got wrong.

Replaces cog-api's getFromNamespace(".ensure_session", "uscogdata") workaround,
which depended on a private name and on the session already being open.
2026-08-10 18:51:56 -04:00
jared 0fbae00e27 feat: name a cohort by state/type predicate instead of a 40k-id IN list
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00
jared 72b2cc3a26 fix: local corpus paths were unreadable on Windows (backslashes eaten)
R-CMD-check / check (push) Successful in 3m35s
R-CMD-check / check (pull_request) Successful in 3m49s
gsub() in regex mode treats backslashes in the REPLACEMENT string as
escape sequences and silently drops them. A Windows corpus path is full of
them, so C:\Users\RUNNER\AppData\... was substituted into the view SQL
as C:UsersRUNNERAppData... and every DuckDB read failed with 'No files
found that match the pattern'.

Effect: uscogdata could not read a LOCAL corpus on Windows at all -- the
bundled fixture included, so the entire test suite failed there, and any
cog_mirror() copy was unusable. Remote https URLs were unaffected, having
no backslashes, which is part of why it stayed hidden.

The bug predates the {long_files} token; it lived in the original {url}
substitution since that code was written. Nothing ever ran on Windows
until the GitHub mirror's check matrix existed, which found it on its
first run: 14 of 15 Windows failures were this, the 15th a downstream
consequence of view registration failing.

fixed = TRUE treats pattern and replacement as literal text. The
regression test reproduces on any platform -- it is string handling, not
filesystem behaviour, so it needs no Windows runner.
2026-08-09 09:03:44 -04:00
jared 44e4953f94 fix: keep doc/ and Meta/ out of the build
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m52s
Removing ^vignettes$ was right; removing ^doc$ and ^Meta$ with it was
not. Those are devtools::build_vignettes() artefacts, not sources -- R CMD
build regenerates inst/doc/ from vignettes/ by itself, and shipping the
local copies earned a 'non-standard file/directory found at top level'
NOTE.

R CMD check --as-cran is now 0 errors, 0 warnings, 0 notes.
2026-08-08 18:25:00 -04:00
jared 331399ab86 docs: rewrite README for a stranger
Reordered around a new user: what the data is, where it comes from,
install, a quickstart that runs with no configuration, then the
full-dollars warning and the concepts that decide whether a published
number is right.

Adds a 'Where the data comes from' section linking the API documentation
site, the live API, the Hugging Face corpus and the Census source, so
attribution and provenance are reachable from the top rather than implied.

Drops the sibling-repo path, the commented-out install line, the Status
block, and the release advice telling you to strip the fixture -- which
would break the vignette and leave public CI unable to check without
credentials.

The quickstart passes years=; cog_spending() has no full-history default,
so the obvious one-liner errors on a reader's first call.
2026-08-08 18:20:24 -04:00
jared 5582c6cb57 docs: index all 14 exports in pkgdown, set the site url
The reference index covered 6 of 14 exports, so pkgdown errored on the
eight missing topics and the docs site did not build at all. Adds a
Comparison & aggregation section and a Corpus metadata section, lists both
vignettes as articles, and sets url so canonical links and search resolve.

build_site() now completes clean: reference metadata ok, no problems.
2026-08-08 17:49:10 -04:00
jared 013af5b0d2 fix: ship the vignettes
.Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed
uscogdata had no vignettes at all -- while the README instructed users to
run vignette("total-spending"), which failed for every one of them.

Both build offline: total-spending points USCOGDATA_URL at the bundled
fixture, population-denominators is eval = FALSE. Confirmed present in the
built tarball as both source and rendered inst/doc/.

The test also pins the fixture as never-excluded -- it is what lets
R CMD check run with no credentials on r-universe and GitHub Actions.
2026-08-08 17:48:07 -04:00
jared 4a92f36d44 chore: add full MIT text, name the copyright holder properly
LICENSE held only the two-line stub and no LICENSE.md existed, so the
repo carried no license text for a human browsing it or for GitHub's
license detector.

usethis::use_mit_license() writes LICENSE.md but leaves an existing
LICENSE alone, so the stub kept saying 'Civilytics' while the full text
said 'Civilytics Consulting LLC'. Corrected by hand, with a test pinning
the two together.
2026-08-08 17:47:18 -04:00
jared a67735f121 chore: release metadata -- author of record, URLs, schema ceiling, 0.3.0
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.

MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.

0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
2026-08-08 17:46:39 -04:00
jared b9f7f7d8d3 test: exercise the public corpus end to end, unconfigured
The remote-read defect survived because every test path used a local
corpus, and so did the API in production. This is the only test that runs
the package the way a new user does: no USCOGDATA_URL, no option, no
fixture -- just install and call a verb.

Gated on USCOGDATA_LIVE_TEST so offline CI skips rather than fails.

Measured against the live corpus while writing this: cog_spending for one
government is 3.9s for a single year and 5.9s across 2000-2022. Well above
the 1.5-2.8s raw parquet scan, because the verbs also join crosswalks,
resolve categories and assemble provenance.
2026-08-08 17:21:53 -04:00
jared 4300b636b1 feat: default to the public corpus so the package works unconfigured
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in
the package supplied a working URL, so a new user installing uscogdata
had no path to a session at all -- just an actionable-looking error with
nothing actionable behind it.

The default is now the public HuggingFace mirror: CC-BY-4.0, no
credential, CDN-backed, and it keeps the origin's uplink out of the read
path. USCOGDATA_URL and options(uscogdata.url=) still override, so
Nextcloud and cog_mirror() copies are unaffected.

The sentinel guard stays for half-edited configs; the two tests covering
it set the URL explicitly, so they only needed renaming to stop calling
it 'the default'.
2026-08-08 17:18:56 -04:00
jared 99e1e86e37 fix: enumerate long partitions from the manifest instead of globbing
DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.

Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.

Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
2026-08-08 17:11:44 -04:00
jared 6a06302036 fix: push limit/offset into the query instead of materializing then slicing
R-CMD-check / check (push) Successful in 4m6s
R-CMD-check / check (pull_request) Successful in 3m33s
cog-api's paginate() sliced an already-fully-materialized result: every
page of a deep sweep re-ran the whole cog_spending()/cog_revenue() query
and re-listified every row, just to keep up to 1000 and discard the
rest. A 193,105-row/194-page fleet-wide sweep (cog_explorer's Southern
guide, corpus summary build) repeated that full cost 194 times and
wedged the production server for hours on 2026-08-06 -- single request,
CPU-bound, single-threaded plumber process, no other request could get
through, not even /health.

cog_spending()/cog_revenue() gain optional limit/offset, pushed into
.build_verb_sql() as SQL LIMIT/OFFSET behind the existing (already
deterministic) ORDER BY. The full unpaginated row count rides along via
COUNT(*) OVER() in the same scan -- exposed as a total_rows attribute --
so a caller walking pages never needs a second round trip to ask how
many there are. A page now costs O(limit), not O(full result).

Mutually exclusive with complete = TRUE (which fills a grid over the
FULL requested (year, category) space -- pagination over a partial slice
of already-grouped rows has no defined meaning for the cells it would
fill) and with recipe (whose result comes from a separate, not-yet-wired
query path). Both abort with a clear classed condition rather than
silently ignoring the parameter.

limit/offset default to NULL; every existing call site is unaffected.
2026-08-06 13:56:27 -04:00
jared 498950afa6 fix: scope all-categories suggestion candidates by subtype, not category (finding 6)
R-CMD-check / check (push) Successful in 3m41s
R-CMD-check / check (pull_request) Successful in 4m52s
.build_suggestions()'s recipe-candidate sub-select was keyed on
`WHERE category IN (<category>)`. The reserved pseudo-category
"All Categories" is never itself a row in summary_categories.category, so
in all-categories mode `candidates` always came back empty and coverage
signposting (uscogdata#9) was structurally impossible for the one mode
whose entire premise is "you cannot sum the wrong scope" -- measured on
Los Angeles County FY2011: category = "Public Welfare" reports 2
suggestions (incl. $271,589,000 excluded E68), category = "All Categories"
reported 0, silently losing that same signal.

Apply the branch's own design principle: the concept boundary is subtype,
not category. .build_suggestions() now accepts all_categories/subtype_col/
subtype_scope (all optional, default off, so no other caller's behaviour
changes) and, when all-categories mode is active, scopes the candidate
sub-select by `<subtype_col> IN (<subtype_scope>)` instead -- symmetric
with .build_verb_sql()'s own WHERE predicate. The M/L recipe exclusion and
the is.null(category) early return are unchanged.

After the fix, LA County FY2011 "All Categories" reports 5 suggestions,
including welfare_cash_e68_wide for the exact $271,589,000 gap.

Adds two covering tests to test-all-categories.R using the bundled fixture
(AL state gov, FY2011, "Corrections"): one end-to-end (per-category and
all-categories both signpost the same recipe) and one direct on
.build_suggestions() proving the subtype-vs-category branch is what
changes the query. Updates the 0.2.0 NEWS entry.
2026-08-05 12:38:12 -04:00
jared a5f86d87b3 fix: close five final-review gaps in all-categories mode
- .detect_direct_suppressed() keys on (year, canonical_govid, category);
  all-categories mode collapses category to one literal value, so the key
  collides and the detector silently reports FALSE instead of "unknown".
  Report NA there instead, and stop isTRUE() in .build_provenance() from
  collapsing that NA back to FALSE. Schema widened to allow null.
- Refuse complete = TRUE + category = "All Categories": the completion grid
  has no per-category cells left to fill once categories are collapsed,
  so the prior silent 0-rows-filled result was never actually checked.
- cog_balances(category = "All Categories") returned zero rows with no
  error. .validate_verb_inputs() gains allow_all_categories (default
  FALSE); .verb_spendrev() passes TRUE, cog_balances() does not, so the
  three verbs share one place to reject it instead of drifting again.
- Fix the false `subtype = "operations"` argument claim (no such argument
  exists) in NEWS.md and an internal spending.R comment.

Adds three covering tests to test-all-categories.R for the three
behaviour changes above.
2026-08-05 12:30:23 -04:00
jared 44e9b40b86 docs: n_units_reporting is category-conditional, not a response rate
Closes uscogdata#36. It counts governments with rows for the requested
category, so a surveyed government that genuinely spends nothing there
is indistinguishable from one never surveyed. In FY2022, a complete
census year, Georgia reports 393 of 567 cities for Police -- the gap is
cities that contract to the sheriff.

Documents the comparison that IS valid: same category, census year vs
sample year.
2026-08-05 11:58:25 -04:00
jared f1e9aa383a test: prove 'All Categories' passes through the geographic rollup
Geographic totals are the expensive case cog-api#37 was filed about --
without this a caller issues one rollup per category and sums them.
The pass-through was expected to work by construction; this asserts it
rather than assuming it, including under per_capita and inflation
adjustment.
2026-08-05 11:53:27 -04:00
jaredandClaude Opus 5 12a9be110f feat: advertise 'All Categories' from cog_categories()
A reserved value nobody can discover is a trap, and this is the view
the API's /categories endpoint is built from. Emitted for the two flow
vocabularies only -- cog_balances() returns a stock and has no concept
to sum within.

Also fix test-categories.R to exclude pseudo-category rows from
crosswalk-specific assertions (one row per (category, subtype) pair,
non-empty item_codes, valid subtypes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:43:02 -04:00
jared 11ae99c382 feat: accept category = 'All Categories' on cog_spending/cog_revenue
Returns one summed row per (year, govid, subtype) across every
category in the requested concept's subtype scope, so a caller never
sums categories client-side and cannot sum the wrong scope.

Combining it with other category names is an error rather than a
silent partial sum.
2026-08-05 11:32:47 -04:00
jared 5e22e940e7 feat: all-categories mode in .build_verb_sql()
The concept boundary in this package is subtype, not category, so a
total is the existing query with the category dimension collapsed and
no category predicate applied. subtype is deliberately kept in the
grouping: subtype=operations plus all-categories is 'operating
expenditure', which is the measure a fiscal comparison wants.

Named 'All Categories' rather than 'Total' because category='Total'
would sit one argument from expenditure_concept='total' and mean
something different.
2026-08-05 11:21:35 -04:00
jaredandClaude Opus 5 77074621d8 revert: drop the I3(b) suppression pre-check gate (#9)
Scoped re-review measured .needs_suppression_query() against the fixture
and found it doesn't pay for itself: it skips the round trip on ~3% of
healthy candidate-bearing calls, ~0% of the multi-govid batch shape
(cog_geographic_rollup()/cog_peer_compare()) it was meant to help, and
reaching the gate costs an unconditional metadata query that on its own
roughly cancels the expected saving -- net slower on the fixture. The
gate was also correct (0 unsound skips) but left an untested exactness
invariant (result$codes_included and the anti-join sharing the harmonized
item_code space) whose silent violation would kill signposting, which is
the exact failure class uscogdata#9 exists to prevent.

Owner's call: revert it and keep the code simple. A batch-aware
optimization, if warranted, is a separate issue.

Removes .needs_suppression_query() entirely (function, roxygen, call
site, comp_rows/flow_components), restoring .build_suggestions() to call
.suppressed_components() directly -- unchanged from f77adb6 except that
it still threads flow_prefixes through (I1, kept). Also removes the two
tests that existed solely to exercise the gate (the five-branch synthetic
test and the local_mocked_bindings call-counter test); no test asserting
real signposting behavior was touched.

I1 (flow_prefixes filter), I2 (schema wording + ig_recipe_id required),
I3(a) (restated NOT EXISTS literals for partition pruning), and M4
(reworded scope claims) are all untouched by this revert.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 07:59:19 -04:00
jaredandClaude Opus 5 4b749205a5 fix: scope suppressed-dollar measurement to the calling verb's own flow family (#9)
Final whole-branch review fix wave for the partial-coverage signposting
feature:

- I1: .suppressed_components() now filters measured recipe components to
  the calling verb's own flow_prefixes. Without this, a candidate recipe
  from the OTHER flow family was always absent from the verb's own view by
  construction and so was always reported as "suppressed" -- fabricating a
  dollar claim across flow families (cog_revenue(category = "Corrections")
  claimed $3.63B excluded that cog_spending() actually reports in full).
- I2: reworded provenance-v1.json's trigger/suppressed_amount descriptions
  to describe what the code actually measures (the verb's underlying long
  view, not "the result"), and to note suppressed_amount can be negative.
  Added ig_recipe_id to the suggestions items' required list, matching the
  key's always-set/nullable runtime behavior.
- I3(a): restated the year/govid literals inside .suppressed_components()'s
  NOT EXISTS subquery so DuckDB can partition-prune that side too (verified
  via EXPLAIN: Scanning Files 1/4 instead of an unfiltered full scan;
  all.equal(old, new) results confirmed unchanged).
- I3(b): added .needs_suppression_query(), a free, exact pre-check reusing
  the verb's own already-computed result$codes_included to skip the anti-
  join round trip on the common fully-covered path, without weakening the
  "suppression can fire with zero gap years" guarantee.
- M4: corrected the overbroad "confines every fire to 2011" scope claim in
  R/suggestions.R and NEWS.md -- the suppressed-dollar measurement is now
  flow-scoped (post-I1), but the empty_year trigger itself is not, and can
  still fire in modern years for a mis-scoped cross-flow-family category.

Added a regression test for I1 plus direct unit-test coverage for the new
flow-family filter and the I3(b) pre-check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 23:22:52 -04:00
jared 230f3401c4 docs: document the suggestion trigger and suppressed-dollar fields (#9) 2026-08-04 22:11:11 -04:00
jared 7522b48a08 feat: report suppressed component dollars in the signpost message (#9)
.inform_suggestions() and cog_explain() now render suppressed_amount /
suppressed_years / suppressed_codes as a continuation line on each
suggestion bullet whenever suppressed_amount > 0 (an empty_year fire can
carry them too, so this keys off the amount, not trigger). Also renames
the cli header from "Coverage gap detected" to "Incomplete coverage" --
a partial-coverage fire is not a gap, the year has rows, they're just
short.
2026-08-04 22:02:01 -04:00
jared 693f8d81a6 fix: signpost aggregate-suppressed components in a category that still has rows (#9) 2026-08-04 21:49:23 -04:00
jared db35fa9058 feat: measure structurally-suppressed recipe component dollars (#9) 2026-08-04 21:38:22 -04:00
jared e067a5930f test: accept schema_version 7, and keep the upper bound enforced
R-CMD-check / check (pull_request) Successful in 6m3s
R-CMD-check / check (push) Has been cancelled
The merged schema-v7 fix (b59b79b) widened .validate_schema()'s allow-list but
left this test asserting that 7 is REJECTED, so main went red. CI had been
hanging on the apt step before ever reaching the suite, which is why the
failure only surfaced once the HTTPS fix let the job get that far.

Flips 7L from expect_error to expect_silent, and ADDS an 8L rejection case.
That second part is the point: simply deleting the 7L expectation would have
left the test unable to prove any upper bound is enforced at all, so a future
v8 corpus with a genuinely breaking change would pass validation silently. The
test should assert the boundary moved, not that it disappeared.

Suite: 796 PASS, 0 FAIL, 0 WARN, 0 SKIP.
2026-08-04 12:07:59 -04:00
jared da726a61f6 fix: cog_categories() surfaces balance subtypes and accepts type = "balance"
R-CMD-check / check (pull_request) Successful in 3m39s
R-CMD-check / check (push) Successful in 3m39s
The balance work added category_type = "balance" rows to the corpus and
cog_balances() to read them, but left cog_categories() -- the discovery
surface -- unable to describe them:

- subtype COALESCEd only spend_subtype and revenue_subtype, so every balance
  row came back with subtype = NA
- type rejected "balance", so there was no way to ask for the holdings
  taxonomy at all

Both matter downstream: cog-api derives its subtype vocabulary from
cog_categories(), so an NA subtype becomes an unusable API parameter. Found
while implementing cog-api#26.

Note cog_balances() itself still takes no subtype argument -- for holdings
category is a strict coarsening of balance_subtype -- but the value belongs
in the discovery surface regardless.

Tests read the expected subtype set independently from the crosswalk parquet
rather than from the function under test.
2026-08-03 12:02:26 -04:00
jared fde62eb6cc test(balances): pin the behaviours the final review found untested or weakly asserted (#25)
Findings F-1..F-8. Every assertion below was verified to FAIL before its
fix (or under mutation, where the behaviour already worked) and pass after.

F-3: 'an unknown recipe id is rejected' used a bare expect_error(). Deleting
.validate_recipe_id() leaves .recipe_components() returning 0 rows and
comps$label[[1]] throwing 'subscript out of bounds' -- still an error, so
the test passed on the regression while the user lost the curated message.
Now asserts class = 'uscogdata_unknown_recipe'. Mutation-checked.

F-4: no test ever set per_capita and adjust_to_year together, so the
load-bearing ordering comment at R/balances.R was unverified. Reversing
those two calls silently drops amt_per_capita_real (.attach_real_dollars()
no-ops when amt_per_capita_nominal does not exist yet). New test asserts
presence AND that the per-capita column is deflated by the same factor as
the level column; mutation-checked by reversing the order (2 failures).

F-5: the spec's 'Gating' requirement had no test -- nothing ever called
cog_balances() on a corpus without balance_subtype. Extended the existing
with_corpus_missing_balance_subtype() block to assert class =
'uscogdata_no_balance_support'; mutation-checked by dropping the guard.

F-1/F-2: added mutual-exclusivity and four-argument validation tests, each
pinned to the message or class (all four inputs already produced *some*
error or *some* quiet wrong answer, so bare expect_error() was useless
here). Plus an ordering guard: a data-frame govid must still work, which
is what fails if validation is put before .coerce_govid_input().

F-6: asserts on the RENDERED cog_explain() text (both streams -- cli
routes through conditions that land on stderr), with a negative case
proving money-verb output is unaffected and that the capture is not vacuous.

F-7: pins that coverage_window is corpus-scoped while truncated is
query-scoped; mutation-checked by scoping the windows to observed subtypes.

F-8: pins the memo slot is populated on first call and cleared by
cog_close().
2026-08-03 11:36:32 -04:00
jared b03f095e49 docs: document cog_balances() and correct stale CLAUDE.md claims (#25)
Adds the NEWS entry, a Financial data pkgdown reference section (none
existed for cog_spending/cog_revenue), and corrects CLAUDE.md's SQL-layer
claim, view count, test count and fixture-year description against
measured values. Also documents balance_caveats in
inst/schemas/provenance-v1.json (test-first: added a schema-documentation
test to test-balances.R, confirmed it failed, then fixed the schema) and
fleshes out cog_balances()'s @return roxygen to enumerate its conditional
columns, regenerating man/cog_balances.Rd.
2026-08-03 11:02:16 -04:00
jared 724b6bd58b docs: disambiguate live-corpus vs fixture year claim in balances test comment (#25) 2026-08-03 10:50:46 -04:00
jared 82e4face4e feat: balance_caveats provenance + once-per-session disclosure (#25) 2026-08-03 10:48:11 -04:00
jared 6c5bdb3048 test: clarify why 2002 must stay in the SB195 recipe test's year vector 2026-08-03 10:40:23 -04:00
jared 90d2e6019e feat: recipe= bridges the wide-era holdings series (#25) 2026-08-03 10:37:54 -04:00
jared 769164c824 feat: per_capita and adjust_to_year for cog_balances() (#25) 2026-08-03 10:24:59 -04:00
jared de3a58d105 fix: attach govids_found/govids_missing to cog_balances() provenance
Mirrors R/spending.R:465-466 -- .check_govids_in_scope()'s return was
previously captured only for its message side effect. Also drops a
redundant duplicate assertion in the flow-code guard test.
2026-08-03 10:19:32 -04:00
jared cdb574d3d0 test: drop arrow dependency from cog_balances tests, use direct DuckDB reads
Also add explicit non-empty assertion to the flow-code guard test so it
cannot pass vacuously on a zero-row result.
2026-08-03 10:12:57 -04:00
jared a281a9621f feat: cog_balances() core verb (#25) 2026-08-03 10:01:40 -04:00
jared 825ac394f2 test: replace vacuous is_aggregate assertion with synthetic-parquet coverage
The bundled fixture has no balance item_code with is_aggregate = TRUE, so
asserting COUNT(*) FROM balance_long WHERE is_aggregate = 0 passed whether
or not the view's AND NOT is_aggregate predicate existed. Follows the
synthetic hive-partitioned parquet pattern already used for the 22-/23-
and 24-/25- view predicates in test-views.R: reads the real
inst/sql/26-balance_long.sql text off disk and executes it against a
synthetic corpus containing both an aggregate and non-aggregate row under
a real balance item_code (W01).
2026-08-03 09:54:53 -04:00
jared d09bfd6aef feat: register balance_long / balance_annotated behind a column gate (#25) 2026-08-03 09:46:36 -04:00
jared 4b23dbd9f4 feat: revenue_concept = c("general", "total") off the crosswalk (#12)
R-CMD-check / check (push) Successful in 3m47s
R-CMD-check / check (pull_request) Successful in 3m27s
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.

`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:

  general = own_source + federal + state + local_aid   (the default)
  total   = general + utility + liquor_store + insurance_trust

The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.

- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
  subtype; the concept narrows in R via the existing subtype_scope
  machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
  printed `expenditure_concept` unconditionally, so a cog_revenue()
  caller was told "Concept: primary" -- a spending concept their result
  has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).

Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).

Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.

Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.

Closes #12
2026-07-30 20:49:49 -04:00
jared 93300ae0c1 feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00
jared 7d798b9937 chore: regenerate fixture corpus at pipeline_commit e64a046
R-CMD-check / check (pull_request) Successful in 3m18s
R-CMD-check / check (push) Successful in 4m6s
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.

Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.

test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.

Suite: 0 failures, 2 skips (the #11 and #12 blocks).
2026-07-30 16:04:35 -04:00
jared d95c9032c5 feat: coverage argument + always-on reporting-coverage metadata (#13)
R-CMD-check / check (pull_request) Successful in 3m13s
R-CMD-check / check (push) Successful in 3m18s
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".

On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.

Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.

Decisions worth stating:

  - n_units_expected is the universe the CALLER named, not the national one.
    That is what makes the ratio mean something: "597 of the 608 Wisconsin
    cities you asked about". For peers it is the cohort size, counted over
    peer rows only -- including the target would inflate every count by one
    and make a cohort that has entirely stopped reporting look non-empty.

  - The coverage table is built from the REQUESTED years, not the years
    present in the result, so a year in which nothing reported still appears
    with n_units_reporting = 0. A year that vanishes silently is precisely
    the disclosure failure at issue.

  - "census" filters years BEFORE the query, and aborts when the range holds
    no census year rather than returning an empty result for a query the
    caller believes they made.

  - "consistent" exempts the peer-comparison target: it is the subject of the
    comparison, not a member of the cohort being balanced, and dropping it
    would leave nothing to compare. The summary_* quantiles are computed
    AFTER the filter so they describe the cohort actually returned.

  - is_census_year is documented as a statement about the survey CALENDAR,
    never a claim of completeness -- FY1967 is a census year in which only 97
    of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
    that tells the truth.

On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().

One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.

Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
2026-07-30 11:57:11 -04:00
jared af85a23ea7 feat: complete = TRUE fills absent cells with their meaning (#18)
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
2026-07-30 11:47:51 -04:00
jared 2e8383b098 fix: let the doc-content tests survive R CMD check
R-CMD-check / check (push) Successful in 3m5s
R-CMD-check / check (pull_request) Successful in 3m11s
CI failed on the previous commit. testthat::test_local() from a checkout was
green, but rcmdcheck was not: under R CMD check the suite runs against the
INSTALLED package, where README.md, vignettes/ and man/ do not exist. Both
newly-activated tests read them through test_path("..", "..", ...) and died
on `cannot open the connection`.

The defect was latent in the committed tests, not introduced here -- they
shipped skip()ped, so CI had never executed either one. Removing the skips
is what exposed it, which is the mechanism working as intended.

Guarded with skip_if_no_source_tree(), so they skip in the installed-package
context that structurally cannot satisfy them. They are NOT thereby unchecked
in CI: the workflow runs testthat::test_local() from the checkout as its own
step before rcmdcheck, and there the paths resolve and the assertions run.

Deliberately not split: test-peer-summary-scope.R's numeric pin needs only
the corpus and would survive check on its own, but it exists to protect the
sentence above it. Separating them would let the prose drift while the pin
kept passing.

Verified locally: test_local 629 pass / 0 fail / 3 skip; rcmdcheck
0 errors / 0 warnings / 0 notes.
2026-07-30 11:31:28 -04:00
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00