Compare commits

...
Author SHA1 Message Date
jared 56f610ea3f Merge pull request 'fix(suggestions): scope candidate recipes by category_type (#34)' (#70) from fix/category-type-34 into main
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m36s
2026-09-09 11:18:52 -04:00
jared 62741343ee Merge pull request 'refactor(suggestions): decompose .build_suggestions() into named helpers (#33)' (#69) from issue-33 into main
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m33s
2026-09-09 11:18:01 -04:00
jaredandClaude Sonnet 5 392643bd74 fix(suggestions): scope candidate recipes by category_type (#34)
R-CMD-check / check (push) Successful in 4m15s
R-CMD-check / check (pull_request) Successful in 4m25s
.query_candidate_recipes() (extracted in #33) now filters candidates by
category_type ('expenditure' vs 'revenue'), derived from the calling
verb's own flow_prefixes (E/F/G -> 'expenditure', else 'revenue').

Without this, a category shared across both flow families in
summary_categories leaked cross-family recipes: cog_revenue(category =
"Corrections") surfaced the expenditure-only corrections_combined recipe
(E04/E05) merely because "Corrections" is also a spending category name,
and cog_spending(category = "IG Federal") surfaced the revenue-only
ig_federal_b47_wide recipe. Both are wrong: following either hint would
attribute dollars to the wrong flow, or (IG Federal) fire the
coverage-gap machinery for a category the calling verb structurally
cannot report on at all.

Updates the two tests this changes the expected behavior of:
- "a mis-scoped cog_spending() call never attaches an M/L counterpart to
  a revenue-flavored recipe" (test-expenditure-concept.R): IG Federal is
  revenue-only, so a spending call now finds zero candidates outright
  rather than firing the suggestion and then blocking its M/L
  counterpart as a second-order check.
- "cog_revenue never suggests expenditure-only recipes"
  (test-recipes.R, was "I1: ... never fabricates suppressed dollars"):
  corrections_combined is expenditure-only, so a revenue call now never
  considers it as a candidate, rather than considering it and reporting
  zero suppressed dollars.

All 1078 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree stacked on the #33 refactor.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:11:16 -04:00
jaredandClaude Sonnet 5 0c7c7eb299 refactor(suggestions): decompose .build_suggestions() into named helpers (#33)
R-CMD-check / check (pull_request) Successful in 4m12s
R-CMD-check / check (push) Successful in 4m20s
Extract three functions from the ~140-line .build_suggestions()
orchestrator to comply with the 'functions under 50 lines' convention:

- .query_candidate_recipes(): candidate recipe lookup by category/subtype
  scope, plus M/L self-exclusion
- .query_recipe_meta(): metadata lookup for labels and year spans
- .query_covered_years(): Path 1 gap-year coverage query via the recipe's
  own generic join; returns empty data frame when gap_years is empty

Kept inline per design: the for-loop that merges covered-years +
suppressed-components into suggestion objects, the M/L-exclusion comment
block as call-site rationale, and .attach_ig_counterparts() at the end.

Pure extraction, no behavior change -- SQL text is unchanged apart from
whitespace. Restored real multi-line SQL string literals in the two new
helpers (the original candidate/covered-years queries were written that
way; keep it consistent with .query_recipe_meta()) and normal roxygen
'#'' comment-marker spacing throughout, both of which drifted during
extraction in an earlier pass.

All 1084 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:09:01 -04:00
jared 7274ce3bfe chore: correct the roborev exclusions and refresh the guidelines
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m46s
roborev matches excluded_commit_patterns as substrings, and compass's cadence rule
requires `type(ws): subject (#N)`, so `chore:` never matched `chore(engine): ...`.
Adds `chore(`.

Does not add `docs(`. Those commits carry the journal entry and the board
narrative, and the plain-language rule exists to check exactly that prose -- it has
no other commit to fire on.

The guidelines were also stale: they were composed before base.md gained the
plain-language rule, and nothing re-composes them on its own. Refreshed, which is
what put that rule in this repository for the first time.
2026-08-23 23:57:05 -04:00
jared 24e86ed598 chore: retire the three plans whose work has shipped
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m21s
186 unticked checkboxes across three plans, none of them outstanding work.
Superpowers-style plans are execution transcripts: nobody ticks the boxes, and
the plan is abandoned at the point the work is done. Left in place they are
indistinguishable from a live backlog -- the old compass retrofit rule, "open
checkboxes become issues", would have filed 186 issues for finished work.

Evidence, from scripts/plans.py plus a check by hand:

  2026-08-04-partial-coverage-signposting  NEWS: "Coverage signposting ..."
  2026-08-08-public-release                NEWS: "First public release."
  2026-04-29-per-year-population-denominator
      no NEWS line matched and 19 of the 21 files it names exist, so the
      classifier called it ambiguous. Confirmed shipped by hand:
      33c0274 docs(news): per-year population denominators, plus the feat
      commits behind it.

specs/ is untouched. A spec explains why the design is what it is and stays
useful; a plan is scaffolding, and once the building stands it is in the way.
All three remain in git history.
2026-08-23 23:32:11 -04:00
jared 1ec20174b7 chore: move compass out of docs/, which pkgdown deletes
Mirror to GitHub / mirror (push) Successful in 7s
R-CMD-check / check (push) Successful in 3m28s
compass.toml, JOURNAL.md, STATUS.md and decisions/ were sitting inside pkgdown's
output directory. Asked directly, pkgdown listed docs/pm and docs/decisions among
the 28 top-level entries clean_site() would delete, and the guard that would have
refused -- check_dest_is_pkgdown() -- was satisfied by docs/pkgdown.yml. After the
move it lists 26 and none of them are compass's.

Nothing was lost. The journal had no entries and there were no decision records
yet, so this was the cheapest moment to move.

The .gitignore workaround goes with it. Re-including two children of an excluded
docs/ forced the rule to be written as /docs/* plus two negations, which changed
the anchoring and made the fixture corpus's own docs/ need a separate rule. A bare
docs/ matches at any depth again, so both are unnecessary.

.Rbuildignore gains ^pm$ -- R CMD check flags a non-standard top-level directory.

Compass reads both layouts, so this repository worked either way; the point is
that docs/ is a directory another tool empties.
2026-08-23 23:26:33 -04:00
Jared Knowles 2e317d3a0f chore: adopt compass for project tracking
Mirror to GitHub / mirror (push) Successful in 10s
R-CMD-check / check (push) Successful in 3m55s
Three workstreams -- the query verbs, the corpus and how it is mounted, and the
docs that explain both. Each can go stale independently, which is what the
workstream boundary is for: a change to a verb obligates the vignettes, a change
to the corpus obligates NEWS.

All six open issues now carry ws/ and type/ labels, applied additively so #36
kept its existing kodor/, severity/, south-guide and verdict/ labels. The
generated board is pinned as issue #68.

.roborev.toml carries review guidelines composed from a shared baseline, the R
package overlay, and this project's own conventions read out of CLAUDE.md: the
ensure_session-then-dbGetQuery order, the tbl_df-with-provenance return
contract, coerce_govid_input at the boundary, the two SQL layers, no arrow, and
withr as tests-only.

The .gitignore change is load-bearing. pkgdown output made docs/ ignored, which
would have left every compass file untracked and unable to travel to another
machine. Git cannot re-include anything beneath an excluded directory, so the
rule had to list children instead. That forced an anchoring change: a bare
"docs/" matches at any depth, "/docs/*" only at the root, so the fixture corpus
docs directory needed an explicit rule to stay excluded as before.
2026-08-23 16:17:20 -04:00
jared f7c606984f Merge pull request 'ci: mirror canonical tag objects, not the lightweight refs checkout builds' (#67) from ci/mirror-canonical-tags into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m14s
Reviewed-on: #67
2026-08-11 10:03:43 -04:00
jared 9617b86a26 ci: mirror canonical tag objects, not the lightweight refs checkout builds
R-CMD-check / check (push) Successful in 3m45s
R-CMD-check / check (pull_request) Successful in 3m29s
v0.4.0 failed to mirror, and it would have failed at every future release.

On a tag-triggered run, checkout materializes refs/tags/<tag> as a LIGHTWEIGHT
tag at the commit SHA; the annotated tag object Gitea holds is never fetched.
Run 2070 therefore pushed a lightweight v0.4.0 to GitHub and reported success.
Run 2076, on main with fetch-depth 0, did fetch the real annotated object and
was rejected with "already exists" trying to correct it -- git will not clobber
an existing tag. Gitea had e138eeb (annotated), GitHub had d2caa6d (the commit).

Re-fetch canonical tag objects from Gitea before pushing. --force rewrites LOCAL
tag refs only; it is not a force push and does not weaken the non-force
guarantee on main. It is required: without it the fetch is rejected with "would
clobber existing tag" and the lightweight ref survives to be mirrored again.

Verified against a scratch clone -- lightweight d2caa6d becomes annotated
e138eeb, peeling back to the same commit; without --force the tag is unchanged.

The GitHub tag was repaired by hand out of band, so the two remotes already
agree; this stops it recurring.
2026-08-11 09:57:14 -04:00
jared 224e5d0530 Merge pull request 'chore: CI badge, mirror PR explainer, and the tag-pin release step (#47)' (#66) from chore/release-47-badges-mirror-pr into main
Mirror to GitHub / mirror (push) Failing after 11s
R-CMD-check / check (push) Successful in 3m15s
Reviewed-on: #66
2026-08-11 09:44:33 -04:00
jared 5f81ae386b chore: CI badge, mirror PR explainer, and the tag-pin release step (#47)
R-CMD-check / check (pull_request) Successful in 3m49s
R-CMD-check / check (push) Successful in 3m52s
Three items #47 absorbed from #46, held back so a dead badge would not sit
beside an unresolved r-universe one. Both resolve now that v0.4.0 is tagged.

- R-CMD-check badge pointing at the GitHub mirror's workflow, where the
  4-platform matrix actually runs.

- A pull_request_target workflow explaining the mirror flow on every incoming
  PR. A PR here is landed on Gitea and syncs back, and because the merge
  preserves the contributor's commits at their original SHAs, GitHub marks the
  PR 'Merged' with nobody visibly clicking Merge. To a first-time contributor
  that reads as rejection. Say so before it happens.

  pull_request_target rather than pull_request because a fork PR's token is
  read-only under the latter -- it could not comment, which is the entire job.
  That is only safe because this never checks out or runs contributor code; the
  file says so and says not to add a checkout.

- CONTRIBUTING's release checklist now spells out that the tag goes on Gitea and
  the mirror carries it, and that r-universe does NOT pick up a release until
  packages.json's branch pin is edited. '*release' would automate it but needs a
  GitHub Release object, and the mirror pushes tags only -- so it would silently
  never update. Learned while doing this release.
2026-08-10 19:47:26 -04:00
jared d2caa6de97 Merge pull request 'docs: re-measure the corpus-access table against the published corpus (#56)' (#65) from docs/readme-perf-remeasure-56 into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m25s
2026-08-10 19:36:42 -04:00
jared 303aa59b07 docs: order the 0.4.0 NEWS sections by user impact, not merge order
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m38s
The three 0.4.0 features landed in the order their PRs merged, which buried the
headline change (cohort predicates, 4.8x) below an operator config knob and a
documentation note. Reordered to: cohorts, pagination, DuckDB budget, docs,
Fixes -- with Fixes last, where it was already.

The stacked PRs each appended their own '## Fixes' heading, so resolving the
conflicts also folded two of them into the single section that belongs there.

Content is byte-identical to what merged; only section order changed. Verified
by diffing the sorted non-blank lines against the previous commit.
2026-08-10 19:29:43 -04:00
jared 81f72321ee docs: re-measure the corpus-access table against the published corpus (#56)
figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09)
and reported the mirrored column as 'local speed' with no number -- hiding the
largest difference available to a user.

Measured 2026-08-10, fresh R session per arm, against the live corpus at
pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation.

Three findings the old table could not express:

- A local mirror is 60-80x faster. A one-off question is ~12 s end to end
  remotely against ~0.15 s mirrored. Stated outright now, because it is a
  bigger and cheaper win for users than anything in the R code.

- Opening the session is the LARGEST remote cost (~7.5 s), bigger than any
  individual query, and it lands on the user's first query rather than on
  library(). The old table accounted for it nowhere, so every per-query figure
  was quietly missing it.

- The remote cost is round-trips, not scanning: a repeat query over
  already-touched partitions is ~1.5 s against ~4 s cold, and a full-history
  query costs ~7 s whether it runs first or last (verified by running the arms
  in both orders). This is why #93's 1.4-1.7x, measured through cog-api against
  a local mount, does not show up on the remote path -- there, network latency
  swamps scan time.

Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was
ambiguous between MB and MiB besides. Measured from the manifest and on disk.
The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release.

Also documented HTTP 429: a burst of remote queries gets rate-limited by the
host. Hit while taking these measurements.
2026-08-10 19:29:05 -04:00
jared 5cb83d8f1d Merge pull request 'feat: limit/offset on cog_gov_search() and cog_balances() (#57)' (#63) from feat/pagination-search-balances-57 into main
Mirror to GitHub / mirror (push) Successful in 12s
R-CMD-check / check (push) Successful in 4m26s
2026-08-10 19:28:09 -04:00
jared 2bb9d41d72 feat: limit/offset on cog_gov_search() and cog_balances() (#57)
R-CMD-check / check (pull_request) Successful in 4m22s
R-CMD-check / check (push) Successful in 4m28s
verbs were left materializing everything and slicing in R -- the pattern behind
the 2026-08-06 production incident. cog_gov_search() had no LIMIT at all, so an
unfiltered call returns the entire 40,336-row crosswalk.

Extracted the #39 machinery into R/pagination.R first (.validate_pagination(),
.paginate_sql(), .take_pagination_total()) rather than growing a third inline
copy: three definitions of what total_rows means is three places for it to
drift. Conflict refusals stay at the call sites because each verb's conflict
set differs. .verb_spendrev() now uses the shared helpers and is unchanged in
behaviour.

The empty-page fallback query is now passed as a thunk, so the unpaginated SQL
is only BUILT when an offset actually lands past the end instead of on every
paged call.

Two things #57 did not anticipate:

- cog_gov_search()'s ORDER BY was not a total order. population_acs DESC NULLS
  LAST leaves ties -- and the whole NULL block -- in scan order, so two requests
  can order them differently and a paged sweep duplicates one row while dropping
  another. Added canonical_govid as tiebreaker. Unpaginated output changes only
  in the relative order of already-tied rows.

- Basket mode returns one resolved row per requested name plus a sidecar
  covering all of them, so a page of it is not a page of anything the caller
  asked for. Refused with uscogdata_basket_pagination_conflict rather than
  silently ignoring the arguments.

Both default to NULL, so cog-api adopts them behind its existing formals()
probe with no lockstep deploy.

Suite: 1067 passed, 0 failed, 0 warnings (2 pre-existing live-corpus skips).
2026-08-10 19:27:11 -04:00
jared 700ae93c9c Merge pull request 'feat: cog_open() honours a DuckDB thread and memory budget (#60)' (#62) from feat/duckdb-threads-60 into main
Mirror to GitHub / mirror (push) Successful in 10s
R-CMD-check / check (push) Successful in 4m6s
Reviewed-on: #62
2026-08-10 19:23:14 -04:00
jared f4ab9b6d90 feat: cog_open() honours a DuckDB thread and memory budget (#60)
R-CMD-check / check (push) Successful in 4m11s
R-CMD-check / check (pull_request) Successful in 4m11s
cog_open() connected with a bare dbConnect() and set no resource pragmas, so
DuckDB claimed every visible core. Right for one interactive session on a
dedicated machine; wrong for a server, where cog-api runs two replicas on an
8-core host budgeted 4 and each replica independently claims all 8.

USCOGDATA_DUCKDB_THREADS and USCOGDATA_DUCKDB_MEMORY_LIMIT now resolve through
.cfg() -- inheriting the env var > option > default precedence USCOGDATA_URL
already had -- and are applied as pragmas when the connection is created.

Unset issues NO pragma, so an unconfigured session is byte-identical to before.
That negative property is asserted directly against a connection opened the
pre-change way rather than against a hardcoded core count.

.cfg() returns an env var as character, so both resolvers coerce and validate
rather than trusting the type: sprintf("SET threads TO %d", "4") would
otherwise abort inside the connection path with an error naming the pragma
instead of the setting the operator got wrong.

Replaces cog-api's getFromNamespace(".ensure_session", "uscogdata") workaround,
which depended on a private name and on the session already being open.
2026-08-10 18:51:56 -04:00
jared 9508b98676 Merge pull request 'feat: cohort predicates instead of a 40k-id IN list (#58)' (#61) from feat/cohort-predicates-58 into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m18s
Reviewed-on: #61
2026-08-09 14:31:53 -04:00
jared 0fbae00e27 feat: name a cohort by state/type predicate instead of a 40k-id IN list
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00
jared 698812a25c Merge pull request 'fix: local corpus paths were unreadable on Windows (backslashes eaten)' (#55) from fix/windows-backslash-paths into main
Mirror to GitHub / mirror (push) Failing after 8s
R-CMD-check / check (push) Successful in 3m34s
2026-08-09 09:08:41 -04:00
jared d74ecdd4a5 Merge pull request 'ci: mirror main and tags to the GitHub mirror' (#54) from ci/mirror-to-github into main
R-CMD-check / check (push) Successful in 3m29s
Mirror to GitHub / mirror (push) Failing after 11s
2026-08-09 09:04:28 -04:00
jared 72b2cc3a26 fix: local corpus paths were unreadable on Windows (backslashes eaten)
R-CMD-check / check (push) Successful in 3m35s
R-CMD-check / check (pull_request) Successful in 3m49s
gsub() in regex mode treats backslashes in the REPLACEMENT string as
escape sequences and silently drops them. A Windows corpus path is full of
them, so C:\Users\RUNNER\AppData\... was substituted into the view SQL
as C:UsersRUNNERAppData... and every DuckDB read failed with 'No files
found that match the pattern'.

Effect: uscogdata could not read a LOCAL corpus on Windows at all -- the
bundled fixture included, so the entire test suite failed there, and any
cog_mirror() copy was unusable. Remote https URLs were unaffected, having
no backslashes, which is part of why it stayed hidden.

The bug predates the {long_files} token; it lived in the original {url}
substitution since that code was written. Nothing ever ran on Windows
until the GitHub mirror's check matrix existed, which found it on its
first run: 14 of 15 Windows failures were this, the 15th a downstream
consequence of view registration failing.

fixed = TRUE treats pattern and replacement as literal text. The
regression test reproduces on any platform -- it is string handling, not
filesystem behaviour, so it needs no Windows runner.
2026-08-09 09:03:44 -04:00
jared 28c500f47a ci: mirror main and tags to the GitHub mirror
R-CMD-check / check (pull_request) Successful in 3m23s
R-CMD-check / check (push) Successful in 4m2s
A plain non-force git push rather than Gitea's built-in push mirror. A
push mirror force-updates the refs it owns, so a Merge clicked on a GitHub
PR would be silently overwritten on the next sync -- the PR still reading
'Merged' while its commit became unreachable. A non-force push is rejected
instead, which turns that into a red CI run.

Secret is PAT_GH, not GITHUB_MIRROR_PAT: Gitea reserves the GITHUB_ and
GITEA_ prefixes for its own injected variables and refuses secrets using
them.
2026-08-08 20:20:56 -04:00
jared fe9238a6ef Merge pull request 'ci: multi-platform R CMD check for the GitHub mirror' (#53) from ci/github-actions-matrix into main
R-CMD-check / check (push) Successful in 3m5s
2026-08-08 19:54:01 -04:00
jared 6392a74013 ci: multi-platform R CMD check for the GitHub mirror
R-CMD-check / check (push) Successful in 3m27s
R-CMD-check / check (pull_request) Successful in 3m19s
The canonical Gitea runner is Linux-only, and this package hard-depends on
duckdb and httr2 -- both compiled, both with real platform variance --
while having never been checked on Windows or macOS. A large share of the
audience is on Windows.

Inert here: Gitea reads .gitea/workflows, GitHub reads .github/workflows,
so both configs coexist and the Gitea CI stays authoritative for deploys.
This goes live the moment the mirror repo exists (#44).

The matrix needs no credentials -- setup.R points at the bundled fixture --
which is why that fixture must never be .Rbuildignore'd.
2026-08-08 19:49:52 -04:00
jared de2ba0cfb9 Merge pull request 'feat: uscogdata 0.3.0 — first public release' (#41) from feat/public-release-0.3.0 into main
R-CMD-check / check (push) Successful in 3m27s
2026-08-08 18:56:14 -04:00
jared e912a926c2 Merge pull request 'chore: regenerate fixture against pipeline e7394a4 (SB203-SB209)' (#40) from chore/fixture-sb203 into main
R-CMD-check / check (push) Successful in 3m29s
2026-08-08 18:55:59 -04:00
jared 44e4953f94 fix: keep doc/ and Meta/ out of the build
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m52s
Removing ^vignettes$ was right; removing ^doc$ and ^Meta$ with it was
not. Those are devtools::build_vignettes() artefacts, not sources -- R CMD
build regenerates inst/doc/ from vignettes/ by itself, and shipping the
local copies earned a 'non-standard file/directory found at top level'
NOTE.

R CMD check --as-cran is now 0 errors, 0 warnings, 0 notes.
2026-08-08 18:25:00 -04:00
jared a5500f0b6a docs: add CONTRIBUTING with the canonical-on-Gitea PR flow
Moves developer, testing and release instructions out of the README,
minus the fixture-stripping advice, which was wrong.

Explains that a GitHub PR closes itself as merged once the mirror syncs,
because the merge preserves the contributor's SHAs -- so a PR closing
without a visible Merge click reads as success rather than rejection.

Documents the cog-api dependency: its CI clones this package at
USCOGDATA_REF, defaulting to main with no pin, so anything merged here
reaches the API's next build. Includes the commands to run its suite
against a branch first.
2026-08-08 18:20:42 -04:00
jared da2839f885 docs: recast NEWS around the first public release
NEWS described changes relative to states no user had ever seen --
'Breaking: corpus schema_version 4', 'the package now requires...' --
across the whole pre-release development. To someone deciding whether to
depend on this, that reads as instability.

0.3.0 is written as an announcement: what it covers, the verbs, that
reading the corpus now works out of the box, four things to know before a
first query, and the known limits. The 0.2.0 changelog is kept verbatim.
The 0.1.0 development log is dropped; that history is in git.

cog_explain() now documents what provenance actually holds, since the
README points readers at it -- in particular why series_break_refs and
corpus_break_refs are separate fields rather than one list.
2026-08-08 18:20:33 -04:00
jared 331399ab86 docs: rewrite README for a stranger
Reordered around a new user: what the data is, where it comes from,
install, a quickstart that runs with no configuration, then the
full-dollars warning and the concepts that decide whether a published
number is right.

Adds a 'Where the data comes from' section linking the API documentation
site, the live API, the Hugging Face corpus and the Census source, so
attribution and provenance are reachable from the top rather than implied.

Drops the sibling-repo path, the commented-out install line, the Status
block, and the release advice telling you to strip the fixture -- which
would break the vignette and leave public CI unable to check without
credentials.

The quickstart passes years=; cog_spending() has no full-history default,
so the obvious one-liner errors on a reader's first call.
2026-08-08 18:20:24 -04:00
jared 5582c6cb57 docs: index all 14 exports in pkgdown, set the site url
The reference index covered 6 of 14 exports, so pkgdown errored on the
eight missing topics and the docs site did not build at all. Adds a
Comparison & aggregation section and a Corpus metadata section, lists both
vignettes as articles, and sets url so canonical links and search resolve.

build_site() now completes clean: reference metadata ok, no problems.
2026-08-08 17:49:10 -04:00
jared 013af5b0d2 fix: ship the vignettes
.Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed
uscogdata had no vignettes at all -- while the README instructed users to
run vignette("total-spending"), which failed for every one of them.

Both build offline: total-spending points USCOGDATA_URL at the bundled
fixture, population-denominators is eval = FALSE. Confirmed present in the
built tarball as both source and rendered inst/doc/.

The test also pins the fixture as never-excluded -- it is what lets
R CMD check run with no credentials on r-universe and GitHub Actions.
2026-08-08 17:48:07 -04:00
jared 4a92f36d44 chore: add full MIT text, name the copyright holder properly
LICENSE held only the two-line stub and no LICENSE.md existed, so the
repo carried no license text for a human browsing it or for GitHub's
license detector.

usethis::use_mit_license() writes LICENSE.md but leaves an existing
LICENSE alone, so the stub kept saying 'Civilytics' while the full text
said 'Civilytics Consulting LLC'. Corrected by hand, with a test pinning
the two together.
2026-08-08 17:47:18 -04:00
jared a67735f121 chore: release metadata -- author of record, URLs, schema ceiling, 0.3.0
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.

MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.

0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
2026-08-08 17:46:39 -04:00
jared b9f7f7d8d3 test: exercise the public corpus end to end, unconfigured
The remote-read defect survived because every test path used a local
corpus, and so did the API in production. This is the only test that runs
the package the way a new user does: no USCOGDATA_URL, no option, no
fixture -- just install and call a verb.

Gated on USCOGDATA_LIVE_TEST so offline CI skips rather than fails.

Measured against the live corpus while writing this: cog_spending for one
government is 3.9s for a single year and 5.9s across 2000-2022. Well above
the 1.5-2.8s raw parquet scan, because the verbs also join crosswalks,
resolve categories and assemble provenance.
2026-08-08 17:21:53 -04:00
jared 4300b636b1 feat: default to the public corpus so the package works unconfigured
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in
the package supplied a working URL, so a new user installing uscogdata
had no path to a session at all -- just an actionable-looking error with
nothing actionable behind it.

The default is now the public HuggingFace mirror: CC-BY-4.0, no
credential, CDN-backed, and it keeps the origin's uplink out of the read
path. USCOGDATA_URL and options(uscogdata.url=) still override, so
Nextcloud and cog_mirror() copies are unaffected.

The sentinel guard stays for half-edited configs; the two tests covering
it set the URL explicitly, so they only needed renaming to stop calling
it 'the default'.
2026-08-08 17:18:56 -04:00
jared 99e1e86e37 fix: enumerate long partitions from the manifest instead of globbing
DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.

Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.

Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
2026-08-08 17:11:44 -04:00
jared c042ee0b90 docs: retarget the release at 0.3.0 and preserve the 0.2.0 changelog
Spec and plan were written against a branch 25 commits behind main, where
the package still read 0.1.0. It is 0.2.0, with a real 0.2.0 changelog in
NEWS that the plan would have deleted.

0.3.0 rather than 0.2.0 because remote corpus reads go from broken to
working and the default URL from placeholder to live -- user-visible
behaviour, so a minor bump. Not 1.0.0: types 4 and 5 remain out of scope.

Task 9 now prepends a 0.3.0 section, keeps 0.2.0 verbatim with a diff
check to prove it, and drops only the 0.1.0 development churn. Task 4
gains the Version bump.
2026-08-08 17:07:19 -04:00
jared 29dc8199e0 docs: implementation plan for the uscogdata 0.1.0 public release
Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest
enumeration, working default URL, and the live-corpus test whose absence
let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite
README, NEWS and CONTRIBUTING.

Distribution mechanics stay out of scope -- r-universe publishes check
results on registration, so it comes after final verification is green.
2026-08-08 17:03:40 -04:00
jared 619b167ab1 docs: design spec for the uscogdata 0.1.0 public release
Covers the P0 finding that the package cannot read the corpus remotely at
all -- no working default URL, and Hive globs are unsupported over generic
HTTP by DuckDB 1.5.5. Fix is manifest-driven file enumeration (measured:
46,148,034 rows over plain https, identical to the hf:// glob) plus a
working public default.

Also: seven release-readiness fixes, a README restructured for a stranger,
NEWS rewritten as an initial release rather than a pre-release churn log,
and the Gitea-canonical/GitHub-mirror/r-universe distribution mechanics.
2026-08-08 17:03:40 -04:00
jared dfda39051e docs: point agents at the Civilytics values file before they start
Mirrors the same block in cog_explorer/CLAUDE.md. A project-level @ import
does not preload, so the instruction to read it is the mechanism.
2026-08-08 17:03:39 -04:00
jared 785f3af16d chore: regenerate fixture against pipeline e7394a4 (SB203-SB209)
R-CMD-check / check (pull_request) Successful in 3m50s
R-CMD-check / check (push) Successful in 4m9s
Picks up seven new catalogued series breaks, all break_year 2017,
recording that the employee-retirement X-codes (X21, X30, X42, X44, X47,
Z77, Z78) were last collected in the annual finance file at FY2016 before
those systems moved to the Annual Survey of Public Pensions.

series_breaks.parquet 202 -> 209 rows; nothing removed. manifest built_at
and pipeline_commit updated, and the series_breaks sha256 with them.
2026-08-08 17:02:14 -04:00
jared c587c8ba87 Merge pull request 'fix: push limit/offset into the query instead of materializing then slicing' (#39) from fix/pushdown-pagination into main
R-CMD-check / check (push) Successful in 3m35s
Reviewed-on: #39
2026-08-06 14:19:42 -04:00
jared 6a06302036 fix: push limit/offset into the query instead of materializing then slicing
R-CMD-check / check (push) Successful in 4m6s
R-CMD-check / check (pull_request) Successful in 3m33s
cog-api's paginate() sliced an already-fully-materialized result: every
page of a deep sweep re-ran the whole cog_spending()/cog_revenue() query
and re-listified every row, just to keep up to 1000 and discard the
rest. A 193,105-row/194-page fleet-wide sweep (cog_explorer's Southern
guide, corpus summary build) repeated that full cost 194 times and
wedged the production server for hours on 2026-08-06 -- single request,
CPU-bound, single-threaded plumber process, no other request could get
through, not even /health.

cog_spending()/cog_revenue() gain optional limit/offset, pushed into
.build_verb_sql() as SQL LIMIT/OFFSET behind the existing (already
deterministic) ORDER BY. The full unpaginated row count rides along via
COUNT(*) OVER() in the same scan -- exposed as a total_rows attribute --
so a caller walking pages never needs a second round trip to ask how
many there are. A page now costs O(limit), not O(full result).

Mutually exclusive with complete = TRUE (which fills a grid over the
FULL requested (year, category) space -- pagination over a partial slice
of already-grouped rows has no defined meaning for the cells it would
fill) and with recipe (whose result comes from a separate, not-yet-wired
query path). Both abort with a clear classed condition rather than
silently ignoring the parameter.

limit/offset default to NULL; every existing call site is unaffected.
2026-08-06 13:56:27 -04:00
jared e3ab26c3e6 Merge pull request 'feat: 'All Categories' pseudo-category + n_units_reporting semantics' (#37) from feat/all-categories-37 into main
R-CMD-check / check (push) Successful in 3m50s
Reviewed-on: #37
2026-08-05 12:50:46 -04:00
jared 498950afa6 fix: scope all-categories suggestion candidates by subtype, not category (finding 6)
R-CMD-check / check (push) Successful in 3m41s
R-CMD-check / check (pull_request) Successful in 4m52s
.build_suggestions()'s recipe-candidate sub-select was keyed on
`WHERE category IN (<category>)`. The reserved pseudo-category
"All Categories" is never itself a row in summary_categories.category, so
in all-categories mode `candidates` always came back empty and coverage
signposting (uscogdata#9) was structurally impossible for the one mode
whose entire premise is "you cannot sum the wrong scope" -- measured on
Los Angeles County FY2011: category = "Public Welfare" reports 2
suggestions (incl. $271,589,000 excluded E68), category = "All Categories"
reported 0, silently losing that same signal.

Apply the branch's own design principle: the concept boundary is subtype,
not category. .build_suggestions() now accepts all_categories/subtype_col/
subtype_scope (all optional, default off, so no other caller's behaviour
changes) and, when all-categories mode is active, scopes the candidate
sub-select by `<subtype_col> IN (<subtype_scope>)` instead -- symmetric
with .build_verb_sql()'s own WHERE predicate. The M/L recipe exclusion and
the is.null(category) early return are unchanged.

After the fix, LA County FY2011 "All Categories" reports 5 suggestions,
including welfare_cash_e68_wide for the exact $271,589,000 gap.

Adds two covering tests to test-all-categories.R using the bundled fixture
(AL state gov, FY2011, "Corrections"): one end-to-end (per-category and
all-categories both signpost the same recipe) and one direct on
.build_suggestions() proving the subtype-vs-category branch is what
changes the query. Updates the 0.2.0 NEWS entry.
2026-08-05 12:38:12 -04:00
jared a5f86d87b3 fix: close five final-review gaps in all-categories mode
- .detect_direct_suppressed() keys on (year, canonical_govid, category);
  all-categories mode collapses category to one literal value, so the key
  collides and the detector silently reports FALSE instead of "unknown".
  Report NA there instead, and stop isTRUE() in .build_provenance() from
  collapsing that NA back to FALSE. Schema widened to allow null.
- Refuse complete = TRUE + category = "All Categories": the completion grid
  has no per-category cells left to fill once categories are collapsed,
  so the prior silent 0-rows-filled result was never actually checked.
- cog_balances(category = "All Categories") returned zero rows with no
  error. .validate_verb_inputs() gains allow_all_categories (default
  FALSE); .verb_spendrev() passes TRUE, cog_balances() does not, so the
  three verbs share one place to reject it instead of drifting again.
- Fix the false `subtype = "operations"` argument claim (no such argument
  exists) in NEWS.md and an internal spending.R comment.

Adds three covering tests to test-all-categories.R for the three
behaviour changes above.
2026-08-05 12:30:23 -04:00
jared 61b9c95731 chore: release 0.2.0
R-CMD-check / check (push) Successful in 3m34s
R-CMD-check / check (pull_request) Successful in 6m30s
Bumps the minor version because 'All Categories' adds public surface
without breaking any existing call.

The bump is load-bearing, not cosmetic: cog-api installs this package
with install_local(), which no-ops when the version already matches.
Without it, Phase 1 would silently test against the 0.1.0 reader and
pass while proving nothing.
2026-08-05 12:04:24 -04:00
jared 44e9b40b86 docs: n_units_reporting is category-conditional, not a response rate
Closes uscogdata#36. It counts governments with rows for the requested
category, so a surveyed government that genuinely spends nothing there
is indistinguishable from one never surveyed. In FY2022, a complete
census year, Georgia reports 393 of 567 cities for Police -- the gap is
cities that contract to the sheriff.

Documents the comparison that IS valid: same category, census year vs
sample year.
2026-08-05 11:58:25 -04:00
jared f1e9aa383a test: prove 'All Categories' passes through the geographic rollup
Geographic totals are the expensive case cog-api#37 was filed about --
without this a caller issues one rollup per category and sums them.
The pass-through was expected to work by construction; this asserts it
rather than assuming it, including under per_capita and inflation
adjustment.
2026-08-05 11:53:27 -04:00
jaredandClaude Opus 5 12a9be110f feat: advertise 'All Categories' from cog_categories()
A reserved value nobody can discover is a trap, and this is the view
the API's /categories endpoint is built from. Emitted for the two flow
vocabularies only -- cog_balances() returns a stock and has no concept
to sum within.

Also fix test-categories.R to exclude pseudo-category rows from
crosswalk-specific assertions (one row per (category, subtype) pair,
non-empty item_codes, valid subtypes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:43:02 -04:00
jared 503fa6562f docs: fix false subtype= claim in All Categories roxygen (F1)
The @param category text on cog_spending()/cog_revenue() told users to
"Combine with subtype = ..." but neither verb has a subtype argument.
Replace with accurate guidance: filter the returned frame's
spend_subtype/revenue_subtype column.
2026-08-05 11:35:47 -04:00
jared 11ae99c382 feat: accept category = 'All Categories' on cog_spending/cog_revenue
Returns one summed row per (year, govid, subtype) across every
category in the requested concept's subtype scope, so a caller never
sums categories client-side and cannot sum the wrong scope.

Combining it with other category names is an error rather than a
silent partial sum.
2026-08-05 11:32:47 -04:00
jared 5e22e940e7 feat: all-categories mode in .build_verb_sql()
The concept boundary in this package is subtype, not category, so a
total is the existing query with the category dimension collapsed and
no category predicate applied. subtype is deliberately kept in the
grouping: subtype=operations plus all-categories is 'operating
expenditure', which is the measure a fiscal comparison wants.

Named 'All Categories' rather than 'Total' because category='Total'
would sit one argument from expenditure_concept='total' and mean
something different.
2026-08-05 11:21:35 -04:00
jared 2fc9e7585b Merge pull request 'fix: signpost partially-suppressed categories (#9)' (#32) from fix/partial-coverage-signposting-9 into main
R-CMD-check / check (push) Successful in 3m24s
2026-08-05 08:15:46 -04:00
jared 8bf9c4ccc1 refactor: split .suppressed_components() into R/suppression.R (#9)
R-CMD-check / check (pull_request) Successful in 3m40s
R-CMD-check / check (push) Successful in 3m46s
R/suggestions.R crossed the project's 400-line limit (424 lines).
Pure move of .suppressed_components() and its roxygen block per the
plan's Task 5 Step 2 remedy; no logic, SQL, or wording changed.
2026-08-05 08:11:15 -04:00
jaredandClaude Opus 5 77074621d8 revert: drop the I3(b) suppression pre-check gate (#9)
Scoped re-review measured .needs_suppression_query() against the fixture
and found it doesn't pay for itself: it skips the round trip on ~3% of
healthy candidate-bearing calls, ~0% of the multi-govid batch shape
(cog_geographic_rollup()/cog_peer_compare()) it was meant to help, and
reaching the gate costs an unconditional metadata query that on its own
roughly cancels the expected saving -- net slower on the fixture. The
gate was also correct (0 unsound skips) but left an untested exactness
invariant (result$codes_included and the anti-join sharing the harmonized
item_code space) whose silent violation would kill signposting, which is
the exact failure class uscogdata#9 exists to prevent.

Owner's call: revert it and keep the code simple. A batch-aware
optimization, if warranted, is a separate issue.

Removes .needs_suppression_query() entirely (function, roxygen, call
site, comp_rows/flow_components), restoring .build_suggestions() to call
.suppressed_components() directly -- unchanged from f77adb6 except that
it still threads flow_prefixes through (I1, kept). Also removes the two
tests that existed solely to exercise the gate (the five-branch synthetic
test and the local_mocked_bindings call-counter test); no test asserting
real signposting behavior was touched.

I1 (flow_prefixes filter), I2 (schema wording + ig_recipe_id required),
I3(a) (restated NOT EXISTS literals for partition pruning), and M4
(reworded scope claims) are all untouched by this revert.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 07:59:19 -04:00
jaredandClaude Opus 5 4b749205a5 fix: scope suppressed-dollar measurement to the calling verb's own flow family (#9)
Final whole-branch review fix wave for the partial-coverage signposting
feature:

- I1: .suppressed_components() now filters measured recipe components to
  the calling verb's own flow_prefixes. Without this, a candidate recipe
  from the OTHER flow family was always absent from the verb's own view by
  construction and so was always reported as "suppressed" -- fabricating a
  dollar claim across flow families (cog_revenue(category = "Corrections")
  claimed $3.63B excluded that cog_spending() actually reports in full).
- I2: reworded provenance-v1.json's trigger/suppressed_amount descriptions
  to describe what the code actually measures (the verb's underlying long
  view, not "the result"), and to note suppressed_amount can be negative.
  Added ig_recipe_id to the suggestions items' required list, matching the
  key's always-set/nullable runtime behavior.
- I3(a): restated the year/govid literals inside .suppressed_components()'s
  NOT EXISTS subquery so DuckDB can partition-prune that side too (verified
  via EXPLAIN: Scanning Files 1/4 instead of an unfiltered full scan;
  all.equal(old, new) results confirmed unchanged).
- I3(b): added .needs_suppression_query(), a free, exact pre-check reusing
  the verb's own already-computed result$codes_included to skip the anti-
  join round trip on the common fully-covered path, without weakening the
  "suppression can fire with zero gap years" guarantee.
- M4: corrected the overbroad "confines every fire to 2011" scope claim in
  R/suggestions.R and NEWS.md -- the suppressed-dollar measurement is now
  flow-scoped (post-I1), but the empty_year trigger itself is not, and can
  still fire in modern years for a mis-scoped cross-flow-family category.

Added a regression test for I1 plus direct unit-test coverage for the new
flow-family filter and the I3(b) pre-check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 23:22:52 -04:00
jared f77adb6c83 fix: revert DESCRIPTION to original roxygen2 7.3.3 config
R-CMD-check / check (push) Successful in 3m37s
R-CMD-check / check (pull_request) Successful in 3m34s
devtools::document() should not change DESCRIPTION when only JSON/Markdown/test files are edited. Restore the original RoxygenNote: 7.3.3 and remove the Config/roxygen2/version auto-generated line that resulted from running devtools::document() locally.
2026-08-04 22:16:56 -04:00
jared 230f3401c4 docs: document the suggestion trigger and suppressed-dollar fields (#9) 2026-08-04 22:11:11 -04:00
jared 7522b48a08 feat: report suppressed component dollars in the signpost message (#9)
.inform_suggestions() and cog_explain() now render suppressed_amount /
suppressed_years / suppressed_codes as a continuation line on each
suggestion bullet whenever suppressed_amount > 0 (an empty_year fire can
carry them too, so this keys off the amount, not trigger). Also renames
the cli header from "Coverage gap detected" to "Incomplete coverage" --
a partial-coverage fire is not a gap, the year has rows, they're just
short.
2026-08-04 22:02:01 -04:00
jared 693f8d81a6 fix: signpost aggregate-suppressed components in a category that still has rows (#9) 2026-08-04 21:49:23 -04:00
jared db35fa9058 feat: measure structurally-suppressed recipe component dollars (#9) 2026-08-04 21:38:22 -04:00
jared cabe2e2799 docs: plan for partial-coverage signposting (#9) 2026-08-04 21:33:09 -04:00
jared 6cd219a291 Merge pull request 'ci: fetch apt indexes over HTTPS so the install step stops hanging' (#31) from ci/apt-https into main
R-CMD-check / check (push) Successful in 5m2s
Reviewed-on: #31
2026-08-04 12:28:49 -04:00
jared e067a5930f test: accept schema_version 7, and keep the upper bound enforced
R-CMD-check / check (pull_request) Successful in 6m3s
R-CMD-check / check (push) Has been cancelled
The merged schema-v7 fix (b59b79b) widened .validate_schema()'s allow-list but
left this test asserting that 7 is REJECTED, so main went red. CI had been
hanging on the apt step before ever reaching the suite, which is why the
failure only surfaced once the HTTPS fix let the job get that far.

Flips 7L from expect_error to expect_silent, and ADDS an 8L rejection case.
That second part is the point: simply deleting the 7L expectation would have
left the test unable to prove any upper bound is enforced at all, so a future
v8 corpus with a genuinely breaking change would pass validation silently. The
test should assert the boundary moved, not that it disappeared.

Suite: 796 PASS, 0 FAIL, 0 WARN, 0 SKIP.
2026-08-04 12:07:59 -04:00
jared 342debaefa ci: fetch apt indexes over HTTPS so the install step stops hanging
R-CMD-check / check (push) Has been cancelled
R-CMD-check / check (pull_request) Has been cancelled
The "Install system libraries" step was stalling indefinitely. It was not
deadlocked on a config prompt and not slow-but-progressing: measured inside the
live runner container, /var/cache/apt/archives stayed at 0 .deb files after 3+
minutes, with apt's http workers parked in S state waiting on the network.

Root cause is the http:// mirror path being pathologically slow from this
runner, not broken. Measured 2026-08-04 from inside the CI container, same
index file, back to back:

  http://archive.ubuntu.com/ubuntu/dists/noble/Release   20.1s
  https://archive.ubuntu.com/ubuntu/dists/noble/Release    3.1s

apt fetches many indexes serially, so ~20s apiece compounds into what looks
like a hang. Rewriting the deb822 sources to https makes the step complete.

Verified before committing, in the running CI container (rocker/r-ver:4.4):
- ca-certificates present and apt 2.8.3 ships the https method built in, so
  nothing has to be installed over http first to bootstrap TLS
- the sed rewrites both URIs (archive + security); the only remaining http://
  is an inert comment line
- '#' is used as the sed delimiter deliberately: '|' collides with the
  alternation and fails with "unknown option to `s'"
- the regex survives YAML block-scalar parsing with backslashes intact

`|| true` guards each sed because the step runs under `sh -e`, so a
missing-sources-file on some other base image must not kill the job.
2026-08-04 11:57:49 -04:00
jared b59b79b2d5 Merge pull request 'fix: accept corpus schema_version 7 (#80)' (#30) from fix/schema-v7 into main
R-CMD-check / check (push) Has been cancelled
Reviewed-on: #30
2026-08-04 11:38:45 -04:00
jared 5668d6b102 fix: accept corpus schema_version 7 (#80)
R-CMD-check / check (pull_request) Has been cancelled
R-CMD-check / check (push) Has been cancelled
2026-08-04 11:22:28 -04:00
jared 0a6d878a36 Merge pull request 'fix: cog_categories() surfaces balance subtypes and accepts type = "balance"' (#29) from fix/cog-categories-balance-subtype into main
R-CMD-check / check (push) Successful in 3m28s
Reviewed-on: #29
2026-08-03 12:20:10 -04:00
jared da726a61f6 fix: cog_categories() surfaces balance subtypes and accepts type = "balance"
R-CMD-check / check (pull_request) Successful in 3m39s
R-CMD-check / check (push) Successful in 3m39s
The balance work added category_type = "balance" rows to the corpus and
cog_balances() to read them, but left cog_categories() -- the discovery
surface -- unable to describe them:

- subtype COALESCEd only spend_subtype and revenue_subtype, so every balance
  row came back with subtype = NA
- type rejected "balance", so there was no way to ask for the holdings
  taxonomy at all

Both matter downstream: cog-api derives its subtype vocabulary from
cog_categories(), so an NA subtype becomes an unusable API parameter. Found
while implementing cog-api#26.

Note cog_balances() itself still takes no subtype argument -- for holdings
category is a strict coarsening of balance_subtype -- but the value belongs
in the discovery surface regardless.

Tests read the expected subtype set independently from the crosswalk parquet
rather than from the function under test.
2026-08-03 12:02:26 -04:00
jared 03c313b46d Merge pull request 'feat: cog_balances(), a reader surface for cash and security holdings (#25)' (#28) from feat/cog-balances-25 into main
R-CMD-check / check (push) Successful in 3m18s
Reviewed-on: #28
2026-08-03 11:52:13 -04:00
jared 2c532bde19 docs: record the two balance_caveats contract facts cog-api#26 must carry
R-CMD-check / check (push) Successful in 3m38s
R-CMD-check / check (pull_request) Successful in 3m35s
Both were settled during implementation and are easy to get wrong from
outside the package:

- coverage_window is corpus-scoped, not result-scoped. It reports the observed
  year extent of every balance subtype, not only those a query returned. The
  sibling field `truncated` is the result-scoped one.
- balance_caveats is present only on cog_balances() results; an API layer that
  assumes it is universal will read NULL from the money verbs.
2026-08-03 11:45:43 -04:00
jared a9e80858d4 docs: correct coverage_window scope and the stale CLAUDE.md Current State block (#25)
F-7: inst/schemas/provenance-v1.json described coverage_window as mapping
each *observed* balance_subtype, but the query at R/balance_caveats.R has
no predicate tied to the query's codes and always returns every subtype in
the mounted corpus. Took option (b) of the two the review offered -- change
the doc, not the code. Reporting all windows is the better product
behaviour (it answers 'is there a family I missed?'), it is what cog-api#26
already forwards verbatim, and option (a) would make the block empty for a
0-row result. Reworded to say the windows are corpus-wide and that
'truncated' is the query-scoped field. Pinned by a new test either way.

F-10: the 'Current State' block was self-contradictory after a partial
update -- headed 2026-04-27, claiming branch main @ d65e9fe, with a
2026-08-03 test count measured on feat/cog-balances-25 underneath it, and
listing README.md / _pkgdown.yml as outstanding when both exist and
_pkgdown.yml was edited by this branch. All numbers below re-measured on
the final tree after every other fix in this wave, not before:
788 tests (testthat::test_local()), 14 exports (NAMESPACE), 14 man/*.Rd,
2 vignettes, no docs/ (pkgdown::build_site() genuinely still outstanding,
as is the .Rbuildignore fixture entry -- both kept in the list).

The related deferred README.md item is closed with no change, per the
review's ruling: README.md enumerates no verbs at all, so naming
cog_balances would make it the only non-cog_spending verb mentioned.
2026-08-03 11:37:15 -04:00
jared fde62eb6cc test(balances): pin the behaviours the final review found untested or weakly asserted (#25)
Findings F-1..F-8. Every assertion below was verified to FAIL before its
fix (or under mutation, where the behaviour already worked) and pass after.

F-3: 'an unknown recipe id is rejected' used a bare expect_error(). Deleting
.validate_recipe_id() leaves .recipe_components() returning 0 rows and
comps$label[[1]] throwing 'subscript out of bounds' -- still an error, so
the test passed on the regression while the user lost the curated message.
Now asserts class = 'uscogdata_unknown_recipe'. Mutation-checked.

F-4: no test ever set per_capita and adjust_to_year together, so the
load-bearing ordering comment at R/balances.R was unverified. Reversing
those two calls silently drops amt_per_capita_real (.attach_real_dollars()
no-ops when amt_per_capita_nominal does not exist yet). New test asserts
presence AND that the per-capita column is deflated by the same factor as
the level column; mutation-checked by reversing the order (2 failures).

F-5: the spec's 'Gating' requirement had no test -- nothing ever called
cog_balances() on a corpus without balance_subtype. Extended the existing
with_corpus_missing_balance_subtype() block to assert class =
'uscogdata_no_balance_support'; mutation-checked by dropping the guard.

F-1/F-2: added mutual-exclusivity and four-argument validation tests, each
pinned to the message or class (all four inputs already produced *some*
error or *some* quiet wrong answer, so bare expect_error() was useless
here). Plus an ordering guard: a data-frame govid must still work, which
is what fails if validation is put before .coerce_govid_input().

F-6: asserts on the RENDERED cog_explain() text (both streams -- cli
routes through conditions that land on stderr), with a negative case
proving money-verb output is unaffected and that the capture is not vacuous.

F-7: pins that coverage_window is corpus-scoped while truncated is
query-scoped; mutation-checked by scoping the windows to observed subtypes.

F-8: pins the memo slot is populated on first call and cleared by
cog_close().
2026-08-03 11:36:32 -04:00
jared 22c2478634 fix(balances): validate the full signature, surface caveats in cog_explain, memoise coverage windows (#25)
Final-review findings F-1, F-2, F-6, F-8 (plus the F-9 @return reword,
which shares R/balances.R).

F-2: .validate_balance_inputs() checked 2 of cog_balances()' 7 arguments.
years = integer(0) leaked a raw DuckDB 'Parser Error ... AND year IN ()'
with the generated SQL echoed back; govid = character(0) and a non-character
category returned 0 rows with no error at all; recipe = c("a","b") threw
'the condition has length > 1' from inside .validate_recipe_id(). Replaced
with a call to the money verbs' own .validate_verb_inputs() (R/spending.R),
which validates the exact superset needed. Deleted the local copy rather
than extending it -- two validators is how they drift. Placed AFTER
.coerce_govid_input(), because .validate_verb_inputs() asserts
is.character(govid) and a data-frame govid is not unwrapped before that.
This is helper reuse of the same kind as .build_verb_sql()/.attach_per_capita();
the verb still does NOT route through .verb_spendrev().

F-1: falls out of F-2 for free -- the recipe/category mutual-exclusivity
guard lives inside .validate_verb_inputs(). Previously recipe silently
discarded category AND overwrote provenance$category with the recipe label,
so a caller asking for Fund Balances got X40/Z77 insurance-trust holdings
with no trace of the dropped filter.

F-6: cog_explain() rendered every provenance caveat block except
balance_caveats. Since .emit_balance_caveats() fires at most once per
session -- and is routinely consumed by a suppressMessages() call or an
unread knitr chunk -- cog_explain() is the only surface left for a caller
who deliberately audits the result. Added a 'Holdings caveats' section
guarded on !is.null(prov$balance_caveats). Also relabels the cosmetic
'Concept: NA' line on balance results as 'not applicable (holdings are a
stock, not a flow)'.

F-8: the coverage-window query has no govid and no year predicate -- its
answer depends only on the mounted corpus -- yet it scanned all of
balance_long on every call (35% of verb runtime on the fixture, and a
per-request throughput ceiling for cog-api#26). Memoised in
.uscogdata_env$balance_coverage_windows, invalidated by cog_close(), the
same pattern as .uscogdata_env$manifest.
2026-08-03 11:36:18 -04:00
jared 225cd60968 docs: fix stale test count and phantom notes column in cog_balances docs (#25)
Re-measured CLAUDE.md's test count on the final tree (764, not 763 --
the earlier number predated the balance_caveats schema test). Removed
notes from cog_balances()'s @return block: it was copied from
cog_spending()'s @return style without checking cog_balances() never
calls .verb_spendrev(), the only place that sets notes. Verified the
remaining documented columns against colnames() observed across every
argument combination (bare, per_capita, adjust_to_year, both, recipe,
category filter).
2026-08-03 11:10:46 -04:00
jared b03f095e49 docs: document cog_balances() and correct stale CLAUDE.md claims (#25)
Adds the NEWS entry, a Financial data pkgdown reference section (none
existed for cog_spending/cog_revenue), and corrects CLAUDE.md's SQL-layer
claim, view count, test count and fixture-year description against
measured values. Also documents balance_caveats in
inst/schemas/provenance-v1.json (test-first: added a schema-documentation
test to test-balances.R, confirmed it failed, then fixed the schema) and
fleshes out cog_balances()'s @return roxygen to enumerate its conditional
columns, regenerating man/cog_balances.Rd.
2026-08-03 11:02:16 -04:00
jared 724b6bd58b docs: disambiguate live-corpus vs fixture year claim in balances test comment (#25) 2026-08-03 10:50:46 -04:00
jared 82e4face4e feat: balance_caveats provenance + once-per-session disclosure (#25) 2026-08-03 10:48:11 -04:00
jared 6c5bdb3048 test: clarify why 2002 must stay in the SB195 recipe test's year vector 2026-08-03 10:40:23 -04:00
jared b8189aeb7f docs: caveat 4 needs a year span crossing FY2002, not just a recipe query
Task 4's implementer found that SB195 does not surface for a
recipe query spanning only 2011-2012. .build_series_break_refs() matches
break_year BETWEEN min(years) AND max(years), and SB195's break_year is 2002.

That is correct behaviour rather than a gap: a series lying entirely after the
book -> market change sits on one consistent basis, so disclosing a break it
never crosses would be noise. .build_corpus_break_refs() applies the same rule
deliberately.

The spec's caveat table overclaimed by omitting the span condition. Corrected.
2026-08-03 10:39:09 -04:00
jared 90d2e6019e feat: recipe= bridges the wide-era holdings series (#25) 2026-08-03 10:37:54 -04:00
jared 769164c824 feat: per_capita and adjust_to_year for cog_balances() (#25) 2026-08-03 10:24:59 -04:00
jared de3a58d105 fix: attach govids_found/govids_missing to cog_balances() provenance
Mirrors R/spending.R:465-466 -- .check_govids_in_scope()'s return was
previously captured only for its message side effect. Also drops a
redundant duplicate assertion in the flow-code guard test.
2026-08-03 10:19:32 -04:00
jared cdb574d3d0 test: drop arrow dependency from cog_balances tests, use direct DuckDB reads
Also add explicit non-empty assertion to the flow-code guard test so it
cannot pass vacuously on a zero-row result.
2026-08-03 10:12:57 -04:00
jared a281a9621f feat: cog_balances() core verb (#25) 2026-08-03 10:01:40 -04:00
jared 825ac394f2 test: replace vacuous is_aggregate assertion with synthetic-parquet coverage
The bundled fixture has no balance item_code with is_aggregate = TRUE, so
asserting COUNT(*) FROM balance_long WHERE is_aggregate = 0 passed whether
or not the view's AND NOT is_aggregate predicate existed. Follows the
synthetic hive-partitioned parquet pattern already used for the 22-/23-
and 24-/25- view predicates in test-views.R: reads the real
inst/sql/26-balance_long.sql text off disk and executes it against a
synthetic corpus containing both an aggregate and non-aggregate row under
a real balance item_code (W01).
2026-08-03 09:54:53 -04:00
jared d09bfd6aef feat: register balance_long / balance_annotated behind a column gate (#25) 2026-08-03 09:46:36 -04:00
jared a11e29a0e0 docs: use Wisconsin state govt (550000227544) as the cog_balances test government
Standardises on the identifier other agents use for state governments, which
is stable across corpus vintages and is the same id used against the live API.

Verified in the bundled fixture, and it is strictly better coverage than the
previous pick: Wisconsin reaches four of the five balance subtypes (adds
workers_comp_trust via Y21) and carries BOTH wide->modern recipe bridges
(X40->Z77 and X41->Z78), so a second recipe test is added. Y61
(other_insurance_trust) is absent for Wisconsin; no test depends on it.

Also notes not to assert on gov_name -- the fixture carries both "WISCONSIN"
and "WISCONSIN STATE GOVT" and the verb COALESCEs them.
2026-08-03 09:42:47 -04:00
jared 7ac4dc6882 docs: implementation plan for cog_balances() (#25)
Six TDD tasks: the two views + registration gate, the core verb, per_capita
and adjust_to_year, recipe=, balance_caveats provenance, docs.

Every internal the plan calls was verified to exist with the signature used
(.build_verb_sql, .shape_recipe_result, .attach_per_capita, .run_recipe,
.require_schema_v5, ...), so the tasks reuse the shared machinery rather than
reimplementing it. The verb deliberately does not route through
.verb_spendrev(), whose concept scoping, IG leg and complete= grid are all
flow-specific.

Test government is ALABAMA STATE GOVT (010000226085), which covers every case
in the bundled fixture: W01/W31/W61 in 2012/2019/2020, X21+Z77 in 2012,
Y07/Y08 throughout, and X40 in 2011 -- so the wide-era recipe bridge is
testable offline.
2026-08-03 09:36:42 -04:00
jared 57212e3399 docs: restore recipe= to cog_balances(); the pipeline was right
Corrects this spec. The earlier draft deferred recipe= and proposed adding
summary_categories rows for X40/X41. Both were wrong, and the pipeline state
they were meant to fix is correct and documented.

cog_pipeline/docs/phase_r_harmonization_review.md records the decisions:

- Sec 0.2: the wide era exposes these split families ONLY as aggregates, so
  the recipe join deliberately does NOT filter is_aggregate. Safe by
  construction -- wide rows are aggregate-only, modern rows leaf-only, every
  component year-scoped.
- Sec 1: the planned X40->Z77 harmonization MAP rows were dropped on purpose;
  continuity ships as recipes instead. That is why harmonization_map carries
  no balance-code rows.

The reader already implements this (R/recipes.R, R/spending.R). Verified
against the live corpus rather than trusting the comment: corrections_combined
FY2007, whose wide leg E05 is likewise aggregate-only, returns $906,743,000.

Also withdraws the claim that SB155/156 and SB195/196 contradict each other.
X40 rows after FY2002 are the wide-era SAS column persisting through the era
boundary; they say nothing about a classification-level rename. The two sets
describe different layers.

What survives is one narrow, non-blocking gap: no series_breaks row exists at
2016/2017 for Z77/Z78/X30, though review doc Sec 2 recommended exactly that.
Recorded as out-of-scope item 1 with the SB197-SB202 precedent.
2026-08-03 09:28:52 -04:00
jared 9f9d40e1c3 docs: drop the subtype argument from cog_balances()
balance is the only category_type whose subtype column is not orthogonal to
category. Measured against the crosswalk: 5 of 6 expenditure subtypes and 1 of
7 revenue subtypes span more than one category, but 0 of 5 balance subtypes do.
Balance is a strict tree -- Fund Balances = {general}, Retirement System
Holdings = {employee_retirement}, Insurance Trust Balances = the three trust
subtypes.

Exposing both arguments would admit no useful combination: of the 15 pairs, 3
are redundant and 12 are guaranteed empty for every government in every year,
failing as an empty tibble that reads as "holds none" rather than as a
contradiction.

Dropping it also keeps the verb aligned -- no uscogdata verb exposes a subtype
argument; the API layers its own subtype row filter on top, which cog-api#26
can do for /balances. #25's one-filter requirement is still met, since
category = "Fund Balances" is exactly W01/W31/W61.

Adds two tests: that one-filter equivalence, and an assertion that the
subtype -> category tree holds, so an upstream change making category lossy
fails here rather than in a user's analysis.
2026-08-03 09:15:50 -04:00
jared d7e14156ff docs: design spec for cog_balances(), the uscogdata#25 holdings surface
Requirement 1 of #25 shipped with #11/#12. This specs requirement 2 only.

Records three upstream gaps found while measuring the corpus, which change
the shipping scope:

- X40/X41 carry ~42.7K rows (1967-2011) but have no summary_categories row,
  so they cannot appear in a category_type='balance' view. Both holdings
  recipes span X40/X41 + Z77/Z78, so recipe= would silently return only the
  2012-2016 leg. recipe= is therefore deferred to v2.
- SB195/SB196 attach to fin_code X40/X41, outside the balance view.
- SB197-SB202 attach to flow codes, not the holdings codes, so the FY2016
  termination of X21/X30/X42/X44/X47/Z77/Z78 has no catalogued break.

Caveats 2-4 are therefore surfaced reader-side via a computed coverage_window
rather than through the existing series-break builders.
2026-08-03 08:44:18 -04:00
jared de7ccbebc7 Merge pull request 'feat: revenue_concept = c("general", "total") off the crosswalk (#12)' (#27) from feat/revenue-concepts-12 into main
R-CMD-check / check (push) Successful in 3m9s
Reviewed-on: #27
2026-07-30 22:09:44 -04:00
jared 4b23dbd9f4 feat: revenue_concept = c("general", "total") off the crosswalk (#12)
R-CMD-check / check (push) Successful in 3m47s
R-CMD-check / check (pull_request) Successful in 3m27s
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.

`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:

  general = own_source + federal + state + local_aid   (the default)
  total   = general + utility + liquor_store + insurance_trust

The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.

- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
  subtype; the concept narrows in R via the existing subtype_scope
  machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
  printed `expenditure_concept` unconditionally, so a cog_revenue()
  caller was told "Concept: primary" -- a spending concept their result
  has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).

Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).

Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.

Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.

Closes #12
2026-07-30 20:49:49 -04:00
jared 93300ae0c1 feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00
jared 5d77d39711 Merge pull request 'chore: regenerate fixture corpus at pipeline_commit e64a046 (#11 groundwork)' (#26) from feat/expenditure-concepts-11 into main
R-CMD-check / check (push) Successful in 3m3s
Reviewed-on: #26
2026-07-30 16:22:19 -04:00
jared 7d798b9937 chore: regenerate fixture corpus at pipeline_commit e64a046
R-CMD-check / check (pull_request) Successful in 3m18s
R-CMD-check / check (push) Successful in 4m6s
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.

Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.

test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.

Suite: 0 failures, 2 skips (the #11 and #12 blocks).
2026-07-30 16:04:35 -04:00
jared 915a4d0678 Merge pull request 'feat: coverage argument + always-on reporting-coverage metadata (#13)' (#24) from feat/coverage-disclosure-13 into main
R-CMD-check / check (push) Successful in 3m30s
2026-07-30 12:06:53 -04:00
jared 6f98d061a9 Merge pull request 'feat: complete = TRUE fills absent cells with their meaning (#18)' (#23) from feat/complete-argument-18 into main
R-CMD-check / check (push) Successful in 4m14s
Reviewed-on: #23
2026-07-30 12:04:27 -04:00
jared d95c9032c5 feat: coverage argument + always-on reporting-coverage metadata (#13)
R-CMD-check / check (pull_request) Successful in 3m13s
R-CMD-check / check (push) Successful in 3m18s
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".

On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.

Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.

Decisions worth stating:

  - n_units_expected is the universe the CALLER named, not the national one.
    That is what makes the ratio mean something: "597 of the 608 Wisconsin
    cities you asked about". For peers it is the cohort size, counted over
    peer rows only -- including the target would inflate every count by one
    and make a cohort that has entirely stopped reporting look non-empty.

  - The coverage table is built from the REQUESTED years, not the years
    present in the result, so a year in which nothing reported still appears
    with n_units_reporting = 0. A year that vanishes silently is precisely
    the disclosure failure at issue.

  - "census" filters years BEFORE the query, and aborts when the range holds
    no census year rather than returning an empty result for a query the
    caller believes they made.

  - "consistent" exempts the peer-comparison target: it is the subject of the
    comparison, not a member of the cohort being balanced, and dropping it
    would leave nothing to compare. The summary_* quantiles are computed
    AFTER the filter so they describe the cohort actually returned.

  - is_census_year is documented as a statement about the survey CALENDAR,
    never a claim of completeness -- FY1967 is a census year in which only 97
    of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
    that tells the truth.

On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().

One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.

Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
2026-07-30 11:57:11 -04:00
jared af85a23ea7 feat: complete = TRUE fills absent cells with their meaning (#18)
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
2026-07-30 11:47:51 -04:00
jared 8db944e4a0 Merge pull request 'fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)' (#22) from fix/kodor-batch-14-15-16 into main
R-CMD-check / check (push) Successful in 3m11s
Reviewed-on: #22
2026-07-30 11:37:08 -04:00
jared 2e8383b098 fix: let the doc-content tests survive R CMD check
R-CMD-check / check (push) Successful in 3m5s
R-CMD-check / check (pull_request) Successful in 3m11s
CI failed on the previous commit. testthat::test_local() from a checkout was
green, but rcmdcheck was not: under R CMD check the suite runs against the
INSTALLED package, where README.md, vignettes/ and man/ do not exist. Both
newly-activated tests read them through test_path("..", "..", ...) and died
on `cannot open the connection`.

The defect was latent in the committed tests, not introduced here -- they
shipped skip()ped, so CI had never executed either one. Removing the skips
is what exposed it, which is the mechanism working as intended.

Guarded with skip_if_no_source_tree(), so they skip in the installed-package
context that structurally cannot satisfy them. They are NOT thereby unchecked
in CI: the workflow runs testthat::test_local() from the checkout as its own
step before rcmdcheck, and there the paths resolve and the assertions run.

Deliberately not split: test-peer-summary-scope.R's numeric pin needs only
the corpus and would survive check on its own, but it exists to protect the
sentence above it. Separating them would let the prose drift while the pin
kept passing.

Verified locally: test_local 629 pass / 0 fail / 3 skip; rcmdcheck
0 errors / 0 warnings / 0 notes.
2026-07-30 11:31:28 -04:00
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00
jared ebac39e6de Merge pull request 'fix: surface ALL-scoped series breaks in provenance (#19)' (#21) from fix/all-scoped-series-breaks-19 into main
R-CMD-check / check (push) Successful in 3m23s
2026-07-30 10:33:24 -04:00
jared 47dc08c4b0 Merge pull request 'fix: regenerate the bundled fixture against the sparsified corpus (#18)' (#20) from fix/regen-fixture-corpus-18 into main
R-CMD-check / check (push) Successful in 3m7s
2026-07-30 10:32:40 -04:00
jared 1d553a788f fix: surface ALL-scoped series breaks in provenance (#19)
R-CMD-check / check (push) Successful in 3m1s
R-CMD-check / check (pull_request) Successful in 3m1s
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:

  SB085  1977  dollar precision across the 1976/1977 boundary
  SB087  2002  imputation exclusion FY2002-2006
  SB194  2012  dense -> sparse representation change
  SB086  2017  government id scheme change

SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.

Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.

On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.

Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.

Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
2026-07-30 10:27:48 -04:00
jared c375c55da7 fix: regenerate the bundled fixture against the sparsified corpus (#18)
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m51s
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.

Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.

Three test repairs, each a real consequence of sparsification rather than
a number to bump:

  test-categories.R          "assistance" joined the spending subtype
                             vocabulary with the J-prefix codes.

  test-spending.R            The harmonization block counts rows that
                             exist. Broward's E21/F21/G21 were zero-pads
                             and are gone, so the anchor moves to FL state,
                             whose three NA-mapped rows carry $2.83B --
                             the amount accounting was previously asserted
                             only against 0 and could not have caught a
                             bug. Broward keeps a test of its own, now
                             asserting the zero-pads are absent.

  test-expenditure-concept.R Coverage-gap suggestions are presence-based.
                             AL state's only FY2011 B47 cell was an
                             explicit zero, so ig_federal_b47_wide stopped
                             being a candidate there; FL state carries a
                             real amount, so the counterpart guard is
                             exercised against a suggestion that fires.

test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.

Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
2026-07-30 10:20:09 -04:00
jared 82acda6f93 Merge pull request 'test: add failing tests for Madison walkthrough findings' (#17) from test/walkthrough-findings into main
R-CMD-check / check (push) Successful in 3m6s
2026-07-29 10:35:31 -04:00
111 changed files with 9207 additions and 2075 deletions
+4 -3
View File
@@ -3,17 +3,18 @@
^\.Rproj\.user$ ^\.Rproj\.user$
^_pkgdown\.yml$ ^_pkgdown\.yml$
^docs$ ^docs$
^pm$
^Meta$
^doc$
^pkgdown$ ^pkgdown$
^\.github$ ^\.github$
^LICENSE\.md$ ^LICENSE\.md$
^\.git$ ^\.git$
^\.gitignore$ ^\.gitignore$
\.gitkeep$ \.gitkeep$
^vignettes$
^specs$ ^specs$
^plans$ ^plans$
^doc$
^Meta$
^\.gitea$ ^\.gitea$
^CLAUDE\.md$ ^CLAUDE\.md$
^\.superpowers$ ^\.superpowers$
^CONTRIBUTING\.md$
+15
View File
@@ -11,6 +11,21 @@ jobs:
steps: steps:
- name: Install system libraries and Node.js (required by actions/checkout) - name: Install system libraries and Node.js (required by actions/checkout)
run: | run: |
# Switch apt to HTTPS mirrors. Measured from this runner on
# 2026-08-04: the SAME index file takes 20.1s over http:// and 3.1s
# over https://. apt fetches many indexes serially, so http:// does
# not read as "slow" -- it reads as a hang (zero bytes in
# /var/cache/apt/archives after 3+ minutes, apt's http workers parked
# in S state). rocker/r-ver:4.4 already ships ca-certificates and
# apt 2.8.3 has the https method built in, so nothing needs to be
# installed over http first to bootstrap this.
# `|| true` because the step runs under `sh -e`: on an image whose
# sources live in the other location, the missing-file sed must not
# kill the job.
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
/etc/apt/sources.list.d/ubuntu.sources 2>/dev/null || true
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
/etc/apt/sources.list 2>/dev/null || true
apt-get update -qq apt-get update -qq
apt-get install -y --no-install-recommends \ apt-get install -y --no-install-recommends \
nodejs git \ nodejs git \
+55
View File
@@ -0,0 +1,55 @@
# Mirror the canonical Gitea repo to the public GitHub mirror.
#
# Deliberately a plain `git push`, NOT Gitea's built-in push mirror. A push
# mirror force-updates the refs it owns: if anyone ever clicks Merge on a
# GitHub PR, the next sync silently overwrites main, the PR still displays
# "Merged", the commit becomes unreachable, and nothing anywhere says so.
# A non-force push is REJECTED as non-fast-forward the moment that happens,
# turning a silent data-loss trap into a red CI run in a place we already look.
#
# Do NOT add --force here, and do NOT add GitHub branch protection to the
# mirror: protection rules block the mirror's legitimate pushes too, breaking
# normal syncing to catch an abnormal case.
#
# PAT_GH is a GitHub personal access token (repo + workflow scope; workflow is
# required because this pushes .github/workflows/). It is stored as a Gitea
# Actions secret. The name cannot begin with GITHUB_ or GITEA_ -- Gitea
# reserves both prefixes for its own injected variables and rejects the secret.
name: Mirror to GitHub
on:
push:
branches: [main]
tags: ['v*']
jobs:
mirror:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Push main and tags to the GitHub mirror
env:
PAT_GH: ${{ secrets.PAT_GH }}
run: |
set -eu
if [ -z "${PAT_GH:-}" ]; then
echo "PAT_GH is unset -- add it under Settings > Actions > Secrets." >&2
exit 1
fi
# On a tag-triggered run, checkout materializes refs/tags/<tag> as a
# LIGHTWEIGHT tag at the commit SHA -- the annotated tag object Gitea
# holds is never fetched. Mirroring that strips the annotation, and the
# NEXT run on main (which does fetch the real object) is then rejected
# with "already exists" trying to correct it, because git will not
# clobber an existing tag. That is why v0.4.0 failed to mirror.
#
# Re-fetch canonical tag objects from Gitea first. --force here rewrites
# LOCAL tag refs only; it is not a force push and does not weaken the
# non-force guarantee on main documented above.
git fetch --tags --force origin
git push "https://x-access-token:${PAT_GH}@github.com/civilytics/uscogdata.git" \
HEAD:refs/heads/main --tags
+61
View File
@@ -0,0 +1,61 @@
# Multi-platform R CMD check, running on the GitHub mirror.
#
# This exists because the canonical Gitea runner is Linux-only, and this
# package hard-depends on duckdb and httr2 -- both compiled, both with real
# platform variance -- while having never been checked on Windows or macOS.
# A large share of the audience is on Windows.
#
# Gitea reads .gitea/workflows and GitHub reads .github/workflows, so this
# file is inert on the canonical repo and coexists with the Gitea CI that
# remains authoritative for deploys.
#
# The suite needs NO credentials: tests/testthat/setup.R points USCOGDATA_URL
# at the bundled fixture corpus. That is exactly why inst/extdata/fixture_corpus
# must never be added to .Rbuildignore.
on:
push:
branches: [main]
pull_request:
name: R-CMD-check
permissions: read-all
jobs:
R-CMD-check:
runs-on: ${{ matrix.config.os }}
name: ${{ matrix.config.os }} (${{ matrix.config.r }})
strategy:
fail-fast: false
matrix:
config:
- {os: macos-latest, r: 'release'}
- {os: windows-latest, r: 'release'}
- {os: ubuntu-latest, r: 'devel', http-user-agent: 'release'}
- {os: ubuntu-latest, r: 'release'}
env:
GITHUB_PAT: ${{ secrets.GITHUB_TOKEN }}
R_KEEP_PKG_SOURCE: yes
steps:
- uses: actions/checkout@v4
- uses: r-lib/actions/setup-pandoc@v2
- uses: r-lib/actions/setup-r@v2
with:
r-version: ${{ matrix.config.r }}
http-user-agent: ${{ matrix.config.http-user-agent }}
use-public-rspm: true
- uses: r-lib/actions/setup-r-dependencies@v2
with:
extra-packages: any::rcmdcheck
needs: check
- uses: r-lib/actions/check-r-package@v2
with:
upload-snapshots: true
build_args: 'c("--no-manual")'
+64
View File
@@ -0,0 +1,64 @@
# Explain the mirror contribution flow on every incoming pull request.
#
# This repository is a MIRROR. A PR opened here is landed on the canonical Gitea
# repository and syncs back; because the merge preserves the contributor's
# commits at their original SHAs, GitHub marks the PR "Merged" on its own as
# soon as the mirror syncs -- with nobody visibly clicking Merge.
#
# Without this comment, that reads as a rejection: the contributor sees their PR
# close with no review, no merge button pressed, and no explanation. It is
# actually the successful outcome. Say so up front, before it happens.
#
# WHY pull_request_target AND NOT pull_request:
# a `pull_request` run from a fork gets a read-only token, so it cannot post a
# comment -- which is exactly the case this workflow exists to serve.
# `pull_request_target` runs in the context of the BASE repo and gets a writable
# token. That is only safe because this job never checks out or executes the
# contributor's code; it posts a fixed string. Do not add a checkout of
# `github.event.pull_request.head.sha` here -- that combination is the standard
# pull_request_target privilege-escalation hole.
name: Explain the mirror flow
on:
pull_request_target:
types: [opened]
permissions:
pull-requests: write
jobs:
comment:
runs-on: ubuntu-latest
steps:
- name: Post the contribution-flow explainer
uses: actions/github-script@v7
with:
script: |
const body = [
"Thanks for this — and one thing worth knowing before it happens.",
"",
"**This repository is a mirror.** Development happens on Gitea at",
"`gitea.civilytics.org/Civilytics/uscogdata`. Your pull request will be fetched",
"from here, landed there, and synced back.",
"",
"Because that merge preserves your commits at their original SHAs, **GitHub will",
"mark this pull request \"Merged\" on its own** — without anyone visibly clicking",
"the Merge button, and possibly without a review comment on this page first.",
"",
"> If your pull request closes as \"Merged\" and nobody appears to have merged it,",
"> that is the normal, successful outcome — not a rejection.",
"",
"If it is *not* going to be merged, you will get an actual reply saying so.",
"",
"Substantial contributions get a `ctb` entry in `DESCRIPTION`, which surfaces in",
"`citation(\"uscogdata\")`. There is no CLA and no DCO sign-off.",
"",
"Full details: [CONTRIBUTING.md](https://github.com/civilytics/uscogdata/blob/main/CONTRIBUTING.md).",
].join("\n");
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: context.payload.pull_request.number,
body,
});
+7
View File
@@ -4,6 +4,10 @@
.Ruserdata .Ruserdata
*.Rproj *.Rproj
inst/doc inst/doc
# pkgdown output. Compass used to keep its files in docs/pm/ and
# docs/decisions/, which forced this to be written as children with two
# re-includes -- git cannot re-include anything beneath an excluded directory.
# Compass lives in pm/ now, so the whole directory can be excluded again.
docs/ docs/
/doc/ /doc/
/Meta/ /Meta/
@@ -12,3 +16,6 @@ docs/
# SDD working artifacts (ledger, briefs, review packages) — plans/ stays tracked # SDD working artifacts (ledger, briefs, review packages) — plans/ stays tracked
.superpowers/sdd/ .superpowers/sdd/
.compass-cache/
# roborev snapshots
/.roborev/
+60
View File
@@ -0,0 +1,60 @@
# roborev configuration, initialised by compass.
# Reviews are queued to a background daemon -- they never block a commit.
post_commit_review = 'commit'
excluded_commit_patterns = ['WIP', 'chore:', 'chore(', 'docs:', 'Merge ']
review_guidelines = '''
# --- compass:begin (generated -- edit the sources, not this) ---
- Prefer returning new values to mutating arguments in place. A function that edits
its caller's object is a bug waiting for a second caller.
- Validate at system boundaries -- user input, API responses, file contents, config.
Fail fast with a message naming the field and the file.
- Never swallow an error. Handle it or let it propagate; a bare catch that continues
is worse than a crash.
- No hardcoded secrets, tokens, or credentials, and no secrets in log output or error
messages.
- Parameterise every query. String-built SQL is a defect even when the input looks safe.
- Keep functions under roughly 50 lines and files under roughly 400. Flag nesting
deeper than four levels.
- No magic numbers or hardcoded paths -- name them as constants or read them from config.
- New behaviour needs a test. A bug fix needs a test that fails without the fix.
- Prose a person reads -- an issue title or body, a journal entry, a decision record,
the narrative on the status board -- names the action or the thing, not the shape of
the machinery. Flag "gate", "seam", "surface area", "load-bearing", "first-class",
"primitive", "blast radius". A project's own defined vocabulary is not the target.
- Use the native pipe `|>`, not magrittr `%>%`.
- snake_case for objects and functions; UPPER_SNAKE for constants. Never use `.` as a
word separator in a function name -- it collides with S3 dispatch.
- Validate arguments at the top of exported functions with `stopifnot()` or an explicit
check, and say which argument was wrong.
- Never `setDT()`, `set()`, or otherwise modify by reference a data.table the caller
still owns. `as.data.table()` copies; use it.
- Prefer `vapply()` to `sapply()` -- `sapply()` silently returns a list when the type
varies, which turns a type error into a downstream mystery.
- Use `seq_len(n)` / `seq_along(x)`, never `1:n`, which iterates backwards when n is 0.
- Compare strings with `==` only after checking for NA; use `identical()` for scalars
where NA would be wrong.
- Do not call `library()` inside package or module files; attach packages in scripts and
test helpers only.
- Namespace-qualify calls into other packages (`stats::sd`) in code that is sourced.
- Every exported function needs roxygen with `@param` for each argument (type, meaning,
and why the default is what it is) and `@return`. Add `@examples` for exported API.
- Declare dependencies in DESCRIPTION. Prefer base R or an existing dependency over
adding a new one; a package with zero hard deps is worth keeping that way.
- Signal errors with `stop()` carrying a condition class, so callers can catch the kind
rather than matching on message text.
- Keep internals internal. Export only what a user needs; an accidentally exported
helper becomes an API you have to keep.
- Tests use testthat edition 3. Each test is self-sufficient -- no reliance on state
left by an earlier test or on a fixture built elsewhere in the file.
- Prefer duplication in tests over a helper that hides what is being asserted.
- Every verb calls .ensure_session() first, then queries via DBI::dbGetQuery().
- A verb's return value is always a tbl_df carrying a provenance attribute.
- govid inputs always go through .coerce_govid_input(); it accepts a character vector or a data frame.
- SQL has two layers: view definitions are numbered .sql files in inst/sql/ registered by .register_views(); query construction is inline sprintf() in R. Add a view as a file; build a query in R.
- No arrow dependency -- DuckDB reads parquet natively.
- withr is Suggests-only and must appear in tests alone.
- Tests must pass offline against the bundled fixture; tests/testthat/setup.R sets USCOGDATA_URL for that.
# --- compass:end ---
'''
File diff suppressed because it is too large Load Diff
+43 -14
View File
@@ -28,8 +28,23 @@ USCOGDATA_URL (local path or https://)
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()` - `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache) - `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`) - `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated` - `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
`revenue_annotated`/`ig_annotated`), and metadata views
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
`representation`, `code_set`)
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()` - `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
**stock** (a balance at a point in time) rather than a **flow** (activity
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
disclosure + measured per-subtype coverage windows).
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer) - `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()` - `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters) - `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
@@ -45,28 +60,31 @@ USCOGDATA_URL (local path or https://)
Any value without `://` is treated as a local path by `.is_local_path()` and reads Any value without `://` is treated as a local path by `.is_local_path()` and reads
`manifest.json` directly from disk (no HTTP, no TTL cache). `manifest.json` directly from disk (no HTTP, no TTL cache).
## Current State (2026-04-27) ## Current State (2026-08-03)
**Version:** 0.1.0 (pre-release) **Version:** 0.1.0 (pre-release)
**Branch:** `main`, commit `d65e9fe` **Branch:** `feat/cog-balances-25`, commit `fde62eb`
**Tests:** 181 PASS / 0 FAIL / 0 SKIP **Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`) **CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
### Completed (Tasks 2.1–2.7) ### Completed (Tasks 2.1–2.7)
All 8 exported verbs implemented and tested: All **14** exports implemented and tested (measured from `NAMESPACE`):
`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`, `cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`, `cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
plus `cog_categories`. `cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
2019+2020, all 50 states). Tests run fully offline — no credentials needed. 2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
### Remaining to v0.1 release ### Remaining to v0.1 release
1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports; 1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`. `README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
Vignettes can be stubbed for v0.1. `total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
`pkgdown::build_site()` has never been run (no `docs/`).
2. **Phase 3 — cog_explorer bridge:** create 2. **Phase 3 — cog_explorer bridge:** create
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs `cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
@@ -99,6 +117,17 @@ devtools::test()
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()` - All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
- Return value is always a `tbl_df` with a `provenance` attribute - Return value is always a `tbl_df` with a `provenance` attribute
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame) - govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
- SQL lives in `inst/sql/` — never inline SQL strings in R files - SQL has two layers. **View definitions** live in `inst/sql/` and are
registered by `.register_views()`, which globs the directory in sorted order
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
a numbered `.sql` file; build a query in R.
- No arrow dependency — DuckDB reads parquet natively - No arrow dependency — DuckDB reads parquet natively
- `withr` is a Suggests-only dep; only used in tests - `withr` is a Suggests-only dep; only used in tests
## Domain context — read this first
**Before doing any work in this repo, read `~/.claude/memory/values/civilytics.md`.**
It carries the purpose, direction, and constraints for this domain. It is not optional
context — read it before planning or writing code, not after. (An `@` import will not
work here; project-level imports don't preload. The read is the mechanism.)
+113
View File
@@ -0,0 +1,113 @@
# Contributing to uscogdata
Thanks for reading this — a package like this gets better mostly through people
noticing that a number looks wrong.
## Where the code lives
Development happens on **Gitea**, at
`gitea.civilytics.org/Civilytics/uscogdata`. The repository at
`github.com/civilytics/uscogdata` is a **mirror** that accepts issues and pull
requests.
## What happens to a GitHub pull request
Open it normally. Behind the scenes it is fetched and landed on the canonical
Gitea repository, then syncs back:
```sh
git fetch github refs/pull/42/head:pr-42
git switch main && git merge --no-ff pr-42
git push origin main # Gitea -> mirror -> GitHub
```
Because the merge preserves your commits at their original SHAs, **GitHub marks
your PR merged on its own** as soon as the mirror syncs. So:
> If your pull request closes as "Merged" without anyone visibly clicking
> Merge, that is the normal, successful outcome — not a rejection.
Substantial contributions get a `ctb` entry in `DESCRIPTION`, which surfaces in
`citation("uscogdata")`.
There is no CLA and no DCO sign-off requirement.
## Running the tests
```r
devtools::test() # bundled fixture; no network, no credentials
```
`tests/testthat/setup.R` points `USCOGDATA_URL` at
`inst/extdata/fixture_corpus/` automatically — a four-year slice (2011, 2012,
2019, 2020) covering all 50 states. That is the whole data setup.
## Testing against the live corpus
```sh
USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test(filter = "live-corpus")'
```
This is worth understanding rather than skipping. Until 0.3.0 the package
**could not read a remote corpus at all** — the partitioned view used a glob,
and DuckDB cannot expand a glob over generic HTTP. It went unnoticed for months
because every test path used a local corpus (the bundled fixture), and so did
the production API (a host mount). Nothing exercised the package the way a new
user does.
`test-live-corpus.R` is the only test that runs with no `USCOGDATA_URL`, no
option, and no fixture. If you change anything touching view registration,
manifest handling, or configuration, run it.
## Do not exclude the fixture from the build
There is a temptation to add `^inst/extdata/fixture_corpus$` to
`.Rbuildignore` because 15 MB feels large for a package. Don't:
- `vignette("total-spending")` reads from it and would fail to build.
- `R CMD check` on r-universe and GitHub Actions would have no corpus, so the
suite could not run without credentials.
This package is not going to CRAN, so its 5 MB guidance does not apply. A
package-size NOTE in `R CMD check` is expected and acceptable.
## Downstream consumers
`cog-api` depends on this package and its CI clones uscogdata at
`USCOGDATA_REF`, **defaulting to `main`**. There is no pin. Anything merged
here reaches the API's next build, so before merging a change to the reader,
run the API suite against your branch:
```sh
Rscript -e "remotes::install_local('/path/to/uscogdata', upgrade = 'never')"
cd /path/to/cog-api/api/tests/testthat
Rscript -e 'testthat::test_dir(".", stop_on_failure = TRUE)'
```
The API calls only exported verbs, so internal refactors are usually safe —
but "usually" is not a release gate.
## Release checklist
1. `devtools::test()` — green against the bundled fixture, offline.
2. `USCOGDATA_LIVE_TEST=true devtools::test()` — green against the live corpus.
3. cog-api suite green against this branch (above).
4. `devtools::check(args = "--as-cran")` — 0 errors, 0 warnings.
5. `pkgdown::build_site()` completes.
6. Vignettes resolve from an installed copy:
`vignette("total-spending", package = "uscogdata")`.
7. **Cold-start check**: on a machine that has never had this package,
install it and run the README quickstart verbatim with no environment
variables set. This is the only check that catches a
corpus-unreachable defect, and its absence is why 0.3.0 needed fixing.
8. Bump `Version` and add a `NEWS.md` section.
9. Tag on **Gitea** (`git tag -a vX.Y.Z && git push origin vX.Y.Z`). The mirror
workflow carries tags to GitHub on its own — confirm the tag appears at
`github.com/civilytics/uscogdata/tags` before continuing.
10. Update the r-universe registry pin at
`github.com/civilytics/civilytics.r-universe.dev` — edit `packages.json`'s
`branch` to the new tag. **r-universe will not pick up a release until this
is edited**: the pin is a tag, deliberately, so a mid-refactor `main` is
never published as a release. `"branch": "*release"` would track releases
automatically, but it needs a GitHub *Release* object and the mirror pushes
tags only — so it would silently never update.
+11 -4
View File
@@ -1,14 +1,21 @@
Package: uscogdata Package: uscogdata
Type: Package Type: Package
Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus
Version: 0.1.0 Version: 0.4.0
Authors@R: Authors@R: c(
person("Civilytics", , , "jknowles@gmail.com", role = c("aut", "cre")) person(c("Jared", "E."), "Knowles",
email = "jared@civilytics.com",
role = c("aut", "cre"),
comment = c(ORCID = "0000-0003-0005-9478")),
person("Civilytics Consulting LLC", role = c("cph", "fnd")))
Description: Curated R verbs over the Civilytics US Census of Governments Description: Curated R verbs over the Civilytics US Census of Governments
finance corpus. Provides unit-level financial profiles, geographic finance corpus. Provides unit-level financial profiles, geographic
rollups, and peer comparisons with auditable provenance and built-in rollups, and peer comparisons with auditable provenance and built-in
cross-vintage correctness. cross-vintage correctness.
License: MIT + file LICENSE License: MIT + file LICENSE
URL: https://github.com/civilytics/uscogdata,
https://civilytics.r-universe.dev/uscogdata
BugReports: https://github.com/civilytics/uscogdata/issues
Encoding: UTF-8 Encoding: UTF-8
LazyData: false LazyData: false
Depends: R (>= 4.1) Depends: R (>= 4.1)
@@ -32,4 +39,4 @@ Config/testthat/edition: 3
VignetteBuilder: knitr VignetteBuilder: knitr
RoxygenNote: 7.3.3 RoxygenNote: 7.3.3
MinCorpusSchema: 4 MinCorpusSchema: 4
MaxCorpusSchema: 5 MaxCorpusSchema: 7
+1 -1
View File
@@ -1,2 +1,2 @@
YEAR: 2026 YEAR: 2026
COPYRIGHT HOLDER: Civilytics COPYRIGHT HOLDER: Civilytics Consulting LLC
+21
View File
@@ -0,0 +1,21 @@
# MIT License
Copyright (c) 2026 Civilytics Consulting LLC
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+1
View File
@@ -1,5 +1,6 @@
# Generated by roxygen2: do not edit by hand # Generated by roxygen2: do not edit by hand
export(cog_balances)
export(cog_basket_resolution) export(cog_basket_resolution)
export(cog_basket_unresolved) export(cog_basket_unresolved)
export(cog_categories) export(cog_categories)
+213 -83
View File
@@ -1,100 +1,230 @@
# uscogdata 0.1.0 (development) # uscogdata 0.4.0
## Breaking: corpus schema_version 4 (Phase P canonical ids) ## Cohorts can be named by predicate, not just by id
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` / `cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state`
`MaxCorpusSchema` in `DESCRIPTION` are both `4`); older corpora built and `type` arguments. Both default to `NULL`, so every existing call behaves
against schema 3 are rejected by `cog_open()` with a clear version-mismatch exactly as before.
error. `canonical_govid` is now uniformly 12 characters across every
vintage the corpus covers (previously a mix of 9-char legacy ids and
12-char FIPS ids depending on source year) — **every hardcoded
`canonical_govid` literal from a pre-Phase-P corpus is now invalid** and
must be re-resolved via `cog_gov_search()` or the new `canonical_alias`
lookup table. `canonical_fips_xwalk` gains four columns
(`legacy_govs_id`, `census_geoid`, `id_source`; `confidence` is renamed to
`pop_confidence`) and a companion `canonical_alias` table ships in the
corpus for mapping legacy/alternate ids onto the current canonical
namespace. The bundled fixture corpus (`inst/extdata/fixture_corpus/`) has
been regenerated against the Phase P publish tree, now ships the full
`canonical_fips_xwalk` and `canonical_alias` master tables alongside the
2019-2020 long partitions, and is reproducible via
`data-raw/regenerate_fixture_corpus.R`.
## Clearer errors when `USCOGDATA_URL` is unconfigured or returns non-JSON Passing them expresses the cohort as a subquery against `canonical_fips_xwalk`
inside each statement, instead of round-tripping the ids through R and
rendering them back into a literal `IN` list:
* `cog_open()` now aborts with the `uscogdata_url_not_configured` error ```r
class when the resolved corpus URL still contains the placeholder # before: resolve 20,106 ids in R, then embed them in every statement
`REPLACE_WITH_SHARE_TOKEN` sentinel (or is empty). The message lists both ids <- cog_gov_search(NULL, state = "CA", type = "city")$canonical_govid
remediation paths (`Sys.setenv(USCOGDATA_URL = ...)` and cog_spending(ids, years = 2022)
`options(uscogdata.url = ...)`) and points at the bundled fixture for
offline testing. Previously the package proceeded to fetch the placeholder
URL, cached the resulting HTML welcome page, and failed downstream with a
cryptic `jsonlite` lexical-error.
* `.fetch_or_cache_manifest()` now parses the HTTP response body before
persisting it. Non-JSON responses (login pages, 404 HTML) raise
`uscogdata_invalid_manifest` with the URL, Content-Type, and underlying
parse error — and never write to the on-disk cache.
* Manifest cache writes are now atomic (write to `manifest.json.tmp.<pid>`
in `cache_dir`, then `file.rename` over the target), so an interrupted
fetch cannot replace a previously-good cache.
* Existing caches with non-JSON content (poisoned by the prior code path)
are silently refetched instead of returning a parse error to the caller.
* Local `USCOGDATA_URL` paths whose `manifest.json` is not valid JSON now
surface the same `uscogdata_invalid_manifest` class with file context.
## Per-capita denominators now use per-year Census F-33 population # now: the cohort never leaves the database
cog_spending(years = 2022, state = "CA", type = "city")
```
* `cog_spending()` and `cog_revenue()` previously divided all years' amounts Measured against the production corpus, same FY2022 aggregate over the
by a single ACS 2018-2022 estimate (`canonical_fips_xwalk.population_acs`), 20,106-government `type = "city"` cohort:
producing biased per-capita values for time-series analysis. They now
divide by the F-33 `population` recorded on each gov-year via the new
`gov_population_yearly` view. Result tibbles gain a `pop_source` column
with values `"census_f33"` or `"unavailable"`. `notes` is updated to
concatenate multiple notes with `"; "`.
## Peer cohorts can be set to a chosen year | cohort expressed as | time |
|---|---:|
| `IN (20,106 literals)` | 449 ms |
| join against a temp cohort table | 99 ms |
| predicate on `canonical_fips_xwalk` | **94 ms** |
| no cohort filter at all (the floor) | 88 ms |
* `cog_find_peers()` adds a `year` argument (default: most recent year for **4.8x, within 7% of the floor.** The rendered `IN` list was 301,591
which the target has an observed population in `gov_population_yearly`). characters and was re-parsed in 5-8 separate statements per call, so the cost
The returned column previously named `population_acs` is now `population` was paid repeatedly; the predicate's size is constant in the cohort.
and reflects the cohort year's vintage. The cohort year is attached to the
returned tibble as `attr(x, "cohort_year")`.
* `cog_peer_compare()` now stamps a `cohort_year` column on its result (read
from the peers tibble's attribute) and records `cohort_year` plus
`cohort_govids` in provenance. When the caller supplies a bare character
vector instead of a `cog_find_peers()` result, `cohort_year` is `NA`.
## Rollups exclude govs missing population `state` and `type` use the same vocabulary and the same internal coercion as
`cog_gov_search()` -- `state` is a postal abbreviation (`"WI"`) even though the
crosswalk column holds a FIPS code (`"55"`).
* `cog_geographic_rollup(per_capita = TRUE)` drops rows whose government has Supplying `govid` **and** `state`/`type` intersects them: the governments in
`pop_source == "unavailable"` and records the dropped govids in `govid` that also match the predicate. Naming no cohort at all now aborts with
`provenance$rollup$excluded_govids`. This excludes special districts class `uscogdata_no_cohort` rather than R's "argument is missing" error.
(type 4) and school districts (type 5) from per-capita rollups by design.
## New: vignette and provenance metadata When the cohort is named by predicate there is no id list to report, so
`provenance$scope$govids_found`/`govids_missing` are empty and
`provenance$scope$cohort` carries `state`, `type` and `n_governments` instead.
A `govid`-named cohort's provenance is unchanged.
* New vignette `population-denominators` covers the four population sources, ## `cog_gov_search()` and `cog_balances()` gain `limit`/`offset`
the type-4/5 coverage gap, the popyear quirk, and how to build moving-window
peer cohorts manually. Pagination arrived on `cog_spending()`/`cog_revenue()` in 0.3.0; the other two
* Provenance gains `transformations$per_capita$popyear_range` and verbs were left materializing everything and slicing in R. Both now take
`pop_source_counts`. `cog_explain()` renders both. `limit`/`offset` with the same semantics: `NULL` default, the page applied in
SQL behind a deterministic `ORDER BY`, and the unpaginated count returned as a
`total_rows` attribute computed by `COUNT(*) OVER()` in the same scan rather
than a second query.
`cog_gov_search()` had no `LIMIT` at all, which made it the one verb that
returns the entire 40,336-row crosswalk when called with no filter.
Two refusals rather than silent surprises:
* `cog_balances(recipe = , limit = )` aborts with class
`uscogdata_recipe_pagination_conflict` -- a recipe's result comes from a
separate query that pagination is not wired into.
* `cog_gov_search()` in basket mode (`length(name) > 1`) aborts with class
`uscogdata_basket_pagination_conflict`. Basket mode returns one resolved row
per requested name with a sidecar covering all of them; a page of that is not
a page of anything the caller asked for.
## DuckDB's resource budget is configurable
`USCOGDATA_DUCKDB_THREADS` and `USCOGDATA_DUCKDB_MEMORY_LIMIT` (with matching
`options(uscogdata.duckdb_threads = )` / `options(uscogdata.duckdb_memory_limit = )`
spellings) cap the DuckDB connection the package opens. Both follow the same
env-var > option > default precedence as `USCOGDATA_URL`.
Unset, **no pragma is issued at all** and DuckDB's own defaults apply exactly as
before -- every visible core. That is right for one interactive session on a
dedicated machine and wrong for a server: where several readers share a host, each
otherwise claims the whole machine and they contend. Capping measured ~5% on a
single-government all-years query (502 ms at 2 threads vs 475 ms uncapped on 16
cores), which is cheap enough that a server should always cap.
This replaces a workaround in which a consumer reached into the package namespace
at boot -- `getFromNamespace(".ensure_session", "uscogdata")()` followed by a manual
`SET threads` -- depending both on a private name and on the session already being
open.
## Documentation: the corpus-access table is re-measured and honest
The README's "two ways to read the corpus" table carried figures taken before
the corpus was re-chunked into row groups (cog_pipeline#93, published
2026-08-09) and reported the mirrored column as "local speed" with no number at
all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit
3d28ddd`), fresh R session per arm:
* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s
end to end remotely against ~0.15 s mirrored. That is the largest single
difference available to a user and it is now stated outright rather than left
as "local speed".
* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch
plus 23 view registrations over HTTPS), larger than any individual query, and
it lands on the first query rather than on `library(uscogdata)`. The old table
did not account for it anywhere.
* **The remote cost is round-trips, not scanning.** A repeat query over
already-touched partitions is ~1.5 s against ~4 s cold, and a full-history
query costs ~7 s whether it runs first or last.
* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4%
and the old figure was ambiguous between MB and MiB besides.
* Documented that a burst of remote queries can be rate-limited by the host
(`HTTP 429`), which is another reason to mirror for real work.
## Fixes
* `cog_gov_search()` now orders by `population_acs DESC NULLS LAST,
canonical_govid`. **`population_acs` alone is not a total order** -- ties, and
the entire `NULLS LAST` block, came back in whatever order the scan produced.
That was invisible while every call returned the full result set, but it makes
a paged sweep unsound: two requests can order tied rows differently, so a row
is duplicated on one page and missing from the next. Unpaginated results are
unchanged except for the relative order of rows that were already tied.
* An unknown `state` abbreviation now aborts with "Unknown state abbreviation"
(class `uscogdata_unknown_state`) instead of base R's "subscript out of
bounds". `.state_abbrev_to_fips` is a named character vector, so `[[` on an
absent name threw before the curated message could be reached -- making that
message unreachable dead code in every verb that takes a `state`.
# uscogdata 0.3.0
First public release.
`uscogdata` provides curated R verbs over the Civilytics US Census of
Governments finance corpus: unit-level financial profiles, geographic rollups
and peer comparisons, with auditable provenance on every result.
## What it covers
Government types 0-3 (state, county, municipality, township), FY1967-FY2024 --
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
or FY1969. Special districts (type 4) and school districts (type 5) are out of
scope pending validation.
## The verbs
`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and holdings;
`cog_gov_search()` to resolve place names (including basket mode for many at
once); `cog_find_peers()` and `cog_peer_compare()` for cohorts;
`cog_geographic_rollup()` for aggregates; `cog_categories()`, `cog_recipes()`,
`cog_manifest()` and `cog_explain()` for metadata and provenance; and
`cog_mirror()` for a local copy of the corpus.
## Reading the corpus now works out of the box
* The package reads the published corpus over HTTPS **with no configuration**.
Previously the default was a placeholder sentinel and no document in the
package supplied a working URL, so a new user had no path to a session.
* Remote reads work at all. The partitioned view used a glob, and DuckDB
cannot expand a glob over generic HTTP -- there is no directory listing to
expand against. Partition paths are now enumerated from the corpus manifest,
which is host-agnostic: an HTTPS mirror, a Nextcloud share and a local
`cog_mirror()` copy all take the same path.
* Nothing is written to disk in remote mode; DuckDB fetches only the row
groups a query needs.
## Four things to know before your first query
* **Amounts are in full US dollars.** The raw Census files report thousands;
the verbs multiply by 1000 on the way out. Do not multiply again.
* **Multi-government aggregates disclose their coverage.** The Census is a
complete enumeration only in years ending in 2 and 7; every other year is a
sample. Every such result carries `provenance$coverage` with per-year
`n_units_reporting`.
* **Absence means two different things.** Before FY2012 an absent cell means
Census published $0; from FY2012 it means not reported. `complete = TRUE`
labels which.
* **Series breaks reach you unasked.** Catalogued breaks intersecting your
query appear in provenance and in `cog_explain()`.
## Known limits
* Special districts (type 4) and school districts (type 5) are out of scope.
* Per-capita rollups exclude governments with no F-33 population, which is by
design but does silently narrow a rollup.
* `n_units_reporting` is category-conditional and is not a response rate.
* Employee-retirement (`X`) codes stop at FY2016, when those systems moved to
the Annual Survey of Public Pensions.
# uscogdata 0.2.0
## New features ## New features
* `cog_gov_search()` gains a **basket mode**: passing vector `name` * `cog_spending()` and `cog_revenue()` accept the reserved category
/ `state` / `type` arguments resolves multiple place names in one `"All Categories"`, returning one summed row per
call and returns a tibble of canonical rows in input order, ready `(year, canonical_govid, subtype)` across every category inside the
to pipe into `cog_spending()` / `cog_revenue()`. Per-row resolution requested concept's subtype scope. Filtering the result to
follows an exact-then-substring matching algorithm with deterministic `spend_subtype == "operations"` gives an operating-expenditure total.
disambiguation; ambiguous and missing entries are surfaced via a `cog_geographic_rollup()` inherits it,
sidecar audit tibble plus a single console summary message. which is the efficient way to build a geographic total — previously a
* New exports `cog_basket_resolution()` and `cog_basket_unresolved()` caller had to issue one rollup per category and sum the results
expose the basket sidecar for iterative query refinement. (cog-api#37).
## Breaking changes `"All Categories"` is not the same thing as `expenditure_concept = "total"`.
The concept chooses which subtypes are in scope; `"All Categories"` chooses
whether the rows inside that scope are broken out or summed.
* The first formal of `cog_gov_search()` was renamed from `pattern` * `cog_categories()` advertises `"All Categories"` for the expenditure and
to `name`. All existing call sites in `cog_explorer/` and the revenue vocabularies, so the reserved value is discoverable.
package itself use positional first-arg, so this rename is
non-breaking in practice. Callers that pass `pattern = ...` by name * Coverage signposting (see "Signposting now catches partially-suppressed
must update to `name = ...`. categories" below) now also works in `category = "All Categories"` mode.
The recipe-suggestion candidate query used to be scoped by `category`,
which is never a match for the reserved `"All Categories"` value, so
`provenance$suggestions` always came back empty there — the one mode whose
whole point is "you cannot sum the wrong scope" was silently unable to
signal a wrong scope. The candidate query is now scoped by the concept's
subtype allowlist instead, symmetric with how `.build_verb_sql()` itself
scopes the summed total: Los Angeles County FY2011, `category = "All
Categories"` still excludes $271,589,000 of aggregate-published Public
Welfare (`E68`), but now names `recipe = "welfare_cash_e68_wide"` to
recover it instead of reporting zero suggestions.
## Documentation
* `cog_geographic_rollup()` and `cog_peer_compare()` now document that
`provenance$coverage`'s `n_units_reporting` is **category-conditional** and
is not a response rate: a government that was surveyed and genuinely spends
nothing in the requested category is indistinguishable from one never
surveyed (uscogdata#36).
+122
View File
@@ -0,0 +1,122 @@
# R/balance_caveats.R
#
# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and
# security holdings. Each one silently invalidates an obvious analysis, so
# they travel in provenance (machine-readable, for cog-api#26) rather than
# living only in prose.
#
# Two of the four are already carried by the code-driven series-break
# builders and are deliberately NOT duplicated here:
# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via
# series_break_refs on the recipe path, the only path that observes those
# codes.
# What remains is the GAAP distinction (a constant) and the coverage windows
# (measured, never hardcoded, so they stay correct as the corpus grows).
#' Per-subtype observed year extents, plus which requested families are
#' truncated relative to the requested span.
#' @noRd
.balance_caveats <- function(con, codes_observed, years) {
cw <- .balance_coverage_windows(con)
observed_subtypes <- if (length(codes_observed) == 0L) {
character(0)
} else {
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT balance_subtype FROM summary_categories
WHERE item_code IN (%s) AND balance_subtype IS NOT NULL",
.sql_lit_chr(codes_observed)
))$balance_subtype
}
# A family is "truncated" when the caller asked for years outside the span
# that family actually covers -- the FY2016 employee-retirement termination
# and the FY2021 end of the W family are both this shape.
truncated <- character(0)
if (length(years) > 0L) {
for (s in observed_subtypes) {
w <- cw[[s]]
if (is.null(w)) next
if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s)
}
}
list(
not_gaap = TRUE,
not_gaap_note = paste0(
"Census holdings are gross -- no liabilities are netted -- and are NOT ",
"GAAP fund balance. A reserve ratio built from them overstates what is ",
"actually available."
),
coverage_window = cw,
truncated = sort(unique(truncated))
)
}
#' Per-subtype [min year, max year] extents for EVERY balance subtype in the
#' mounted corpus, memoised for the session.
#'
#' The query carries no govid and no year predicate -- its answer is a property
#' of the mounted corpus alone and cannot change between calls -- but it scans
#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime
#' on the bundled fixture and would be a per-request throughput ceiling once
#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and
#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`.
#'
#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking
#' "is there a family I missed?" needs every window. The observed-scoped field
#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json.
#' @noRd
.balance_coverage_windows <- function(con) {
cached <- .uscogdata_env$balance_coverage_windows
if (!is.null(cached)) return(cached)
windows <- DBI::dbGetQuery(con,
"SELECT c.balance_subtype AS subtype,
MIN(l.year) AS year_min,
MAX(l.year) AS year_max
FROM balance_long l
JOIN summary_categories c USING (item_code)
WHERE c.balance_subtype IS NOT NULL
GROUP BY 1
ORDER BY 1"
)
cw <- stats::setNames(
lapply(seq_len(nrow(windows)),
function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))),
windows$subtype
)
.uscogdata_env$balance_coverage_windows <- cw
cw
}
#' TRUE the first time `key` is seen this session, FALSE thereafter.
#' Reset by cog_close().
#' @noRd
.balance_caveat_once <- function(key) {
seen <- .uscogdata_env$balance_caveats_shown
if (is.null(seen)) seen <- character(0)
if (key %in% seen) return(FALSE)
.uscogdata_env$balance_caveats_shown <- c(seen, key)
TRUE
}
#' Emit at most one message per caveat class per session.
#' @noRd
.emit_balance_caveats <- function(caveats) {
if (.balance_caveat_once("not_gaap")) {
cli::cli_inform(c(
"!" = "Census holdings are gross and are {.strong not} GAAP fund balance.",
"i" = "No liabilities are netted; a reserve ratio built from them overstates available funds."
))
}
if (length(caveats$truncated) > 0L &&
.balance_caveat_once("coverage_window")) {
cli::cli_inform(c(
"!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.",
"i" = "See {.code provenance$balance_caveats$coverage_window}."
))
}
invisible(NULL)
}
+216
View File
@@ -0,0 +1,216 @@
# R/balances.R
#
# Cash and security holdings. A third verb rather than an argument on a money
# verb because holdings are a STOCK -- a balance at a point in time -- while
# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money
# verbs' whole argument vocabulary (expenditure_concept, revenue_concept,
# complete=) describes flows and is meaningless here, so this deliberately
# does NOT route through .verb_spendrev().
#' Cash and security holdings for one or more governments
#'
#' Returns Census cash-and-security holdings (`category_type = "balance"`):
#' fund balances, retirement system holdings and insurance trust balances.
#'
#' @section Holdings are not GAAP fund balance:
#' Census holdings are **gross** -- no liabilities are netted -- so a reserve
#' ratio built from them overstates what is actually available. They are not
#' comparable to a GAAP fund balance from an ACFR.
#'
#' @param govid Canonical govid(s): a character vector, or a data frame with a
#' `canonical_govid` column (e.g. from [cog_gov_search()]). `NULL` to name
#' the cohort by `state`/`type` instead.
#' @inheritParams cog_spending
#' @param years Integer vector of fiscal years.
#' @param category Optional character vector of categories to keep. One of
#' `"Fund Balances"`, `"Insurance Trust Balances"`,
#' `"Retirement System Holdings"`. There is deliberately no `subtype`
#' argument: for holdings, `category` is a strict coarsening of
#' `balance_subtype` (unlike the money verbs, where the two axes cross), so
#' every combination would be either redundant or empty.
#' `category = "Fund Balances"` is exactly the `general` family
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
#' one `dplyr::filter()` away. The reserved pseudo-category
#' `"All Categories"` (see [cog_spending()]) is **not** supported here and
#' errors with class `uscogdata_all_categories_unsupported`: it sums a
#' concept's subtype scope, and holdings are a stock with no concept
#' vocabulary to sum across. Omit `category` to get every category broken
#' out instead.
#' @param per_capita Divide holdings by population. Note this is a **stock per
#' resident** (reserves per person), which is *not* comparable to
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
#' @param adjust_to_year Deflate to this year's dollars (CPI-U).
#' @param basis Accepted for uniformity with the money verbs, but currently a
#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized
#' and raw space are identical for holdings. Reported in
#' `provenance$basis_note`.
#' @param recipe Optional harmonization recipe id (see [cog_recipes()]).
#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
#' wide era to the modern one.
#' @param limit Maximum number of result rows to return, pushed into the SQL
#' rather than applied after materializing every row. `NULL` (default)
#' returns everything. Cannot be combined with `recipe` -- see `offset` and
#' `total_rows`.
#' @param offset Rows to skip before `limit` starts counting (0-based).
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
#'
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`,
#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and
#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when
#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
#' nominal per-capita column to deflate otherwise). Amounts are full US
#' dollars.
#'
#' Carries a `provenance` attribute matching
#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
#' every balance subtype in the mounted corpus, not only the observed ones)
#' and `truncated` (the observed subtypes whose coverage falls short of the
#' requested years). `expenditure_concept`/`revenue_concept` are `NA` --
#' holdings are a stock, not a flow, so neither concept vocabulary applies.
#'
#' When `limit` is set, also carries a `total_rows` attribute: the full
#' unpaginated row count, computed by the same query (`COUNT(*) OVER()`)
#' rather than a second scan.
#' @export
cog_balances <- function(govid = NULL, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
state = NULL, type = NULL,
limit = NULL, offset = NULL) {
call <- match.call()
basis <- match.arg(basis, c("harmonized", "raw"))
# Coerce FIRST, validate second: .validate_verb_inputs() asserts
# is.character(govid), and a data-frame govid (cog_gov_search() output) has
# not been unwrapped yet at this point.
govid <- if (is.null(govid)) NULL else .coerce_govid_input(govid)
# The money verbs' validator, reused rather than re-implemented (R/spending.R).
# It covers the exact superset cog_balances() needs -- including the
# recipe/category mutual-exclusivity guard -- so a second local copy would
# only be a place for the two to drift apart. This is the same kind of
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
# route the verb through .verb_spendrev(), which stays deliberately unused
# here because its flow vocabulary is meaningless for a stock.
#
# allow_all_categories is left at its FALSE default (contrast
# .verb_spendrev(), which passes TRUE): the all-categories mode's "sum"
# only means something in terms of a concept's subtype scope, and holdings
# have no concept vocabulary. The reuse above is exactly why this can be a
# one-line default rather than a second bespoke check -- see the
# validator's own doc comment for the incident that made that matter.
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
# Same semantics as the money verbs (R/pagination.R). Only the `recipe`
# conflict applies here: cog_balances() has no `complete` argument, and a
# recipe's result comes from .run_recipe()'s own query, which pagination is
# not wired into.
paging <- .validate_pagination(limit, offset)
limit <- paging$limit
offset <- paging$offset
if (!is.null(limit) && !is.null(recipe)) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with `recipe`.",
"i" = "A recipe's result comes from a separate query (`.run_recipe()`) that pagination is not wired into yet.",
"*" = "Drop `limit`/`offset`, or drop `recipe`."
), class = "uscogdata_recipe_pagination_conflict")
}
years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
cohort <- .make_cohort(govid, state, type)
con <- .ensure_session()
.require_balance_support(con)
scope <- .check_govids_in_scope(govid)
basis_note <- paste0(
"`basis` has no effect on holdings: harmonization_map carries no ",
"balance-code rows, so harmonized and raw space are identical here."
)
manifest <- .uscogdata_env$manifest
recipe_block <- NULL
category_for_prov <- category
total_rows <- NULL # set below only when limit is non-NULL (non-recipe path)
if (!is.null(recipe)) {
.require_schema_v5(con, manifest, "recipe =")
.validate_recipe_id(con, recipe)
comps <- .recipe_components(con, recipe)
recipe_label <- comps$label[[1]]
result <- .run_recipe(con, recipe, cohort, years)
sql <- attr(result, "sql_query")
result <- .shape_recipe_result(result, "balance_subtype", recipe_label)
recipe_block <- list(
recipe_id = recipe, label = recipe_label,
components = .df_to_row_list(comps)
)
category_for_prov <- recipe_label
} else {
sql <- .build_verb_sql("balance_annotated", "balance_subtype",
cohort, years, category,
ig_view = NULL, subtype_scope = NULL,
limit = limit, offset = offset)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
if (!is.null(limit)) {
paged <- .take_pagination_total(result, con, function() {
.build_verb_sql("balance_annotated", "balance_subtype",
cohort, years, category,
ig_view = NULL, subtype_scope = NULL)
})
result <- paged$result
total_rows <- paged$total_rows
}
}
# Order matters (matches .verb_spendrev()): per-capita first, so
# .attach_real_dollars() deflates the nominal per-capita column into
# amt_per_capita_real rather than needing amt_per_capita_nominal recomputed.
if (isTRUE(per_capita)) result <- .attach_per_capita(result, con)
if (!is.null(adjust_to_year)) {
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
}
prov <- .build_provenance(
verb = "cog_balances", call = call, govid = govid, years = years,
category = category_for_prov, per_capita = per_capita,
adjust_to_year = adjust_to_year, result = result, sql = sql,
subtype_col = "balance_subtype",
basis = basis, basis_note = basis_note,
# Neither concept vocabulary applies to a stock.
expenditure_concept = NA_character_,
revenue_concept = NA_character_,
recipe = recipe_block
)
prov$scope$govids_found <- scope$found
prov$scope$govids_missing <- scope$missing
prov$scope$cohort <- .cohort_provenance(con, cohort)
prov$balance_caveats <- .balance_caveats(
con, prov$codes_summed$observed, years
)
.emit_balance_caveats(prov$balance_caveats)
attr(result, "provenance") <- prov
if (!is.null(limit)) attr(result, "total_rows") <- total_rows
result
}
#' Abort unless the mounted corpus classifies balance codes.
#'
#' `balance_subtype` arrived with cog_pipeline #76/#77 without a
#' schema_version bump, so the check is on the column, not the version.
#' @noRd
.require_balance_support <- function(con) {
if (.corpus_has_balance_subtype(con)) return(invisible(TRUE))
cli::cli_abort(
c("This corpus does not classify cash and security holdings.",
i = "`summary_categories` has no {.field balance_subtype} column.",
i = "Republish from cog_pipeline at #76/#77 or later."),
class = "uscogdata_no_balance_support"
)
}
+17 -8
View File
@@ -44,11 +44,18 @@
#' Count + sum item-level rows that basis="harmonized" excludes because they #' Count + sum item-level rows that basis="harmonized" excludes because they
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the #' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
#' requested flow type (spending or revenue), govids, and years. Only #' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
#' meaningful when the resolved basis is "harmonized"; returns an #' the same subtype-membership classification the verb SQL uses, never
#' applied = FALSE stub otherwise (raw basis never excludes rows this way). #' item-code prefixes), govids, and years. Only meaningful when the resolved
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
#' basis never excludes rows this way).
#'
#' The intergovernmental leg is deliberately outside this count even for
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
#' @noRd #' @noRd
.build_harmonization_block <- function(con, govid, years, resolved, flow_prefixes) { .build_harmonization_block <- function(con, cohort, years, resolved,
subtype_col, subtype_scope) {
if (!identical(resolved$basis, "harmonized")) { if (!identical(resolved$basis, "harmonized")) {
return(list( return(list(
applied = FALSE, applied = FALSE,
@@ -61,11 +68,13 @@
sql <- sprintf( sql <- sprintf(
"SELECT COUNT(*) AS n, COALESCE(SUM(amt), 0) * 1000.0 AS amt "SELECT COUNT(*) AS n, COALESCE(SUM(amt), 0) * 1000.0 AS amt
FROM long FROM long
WHERE canonical_govid IN (%s) AND year IN (%s) WHERE %s AND year IN (%s)
AND NOT is_aggregate AND harmonized_code IS NULL AND NOT is_aggregate AND harmonized_code IS NULL
AND LEFT(item_code, 1) IN (%s)", AND item_code IN (
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","), SELECT item_code FROM summary_categories WHERE %s IN (%s)
.sql_lit_chr(flow_prefixes) )",
.cohort_sql(cohort), paste(as.integer(years), collapse = ","),
subtype_col, .sql_lit_chr(subtype_scope)
) )
na <- DBI::dbGetQuery(con, sql) na <- DBI::dbGetQuery(con, sql)
+41 -8
View File
@@ -5,23 +5,33 @@
#' Returns the category taxonomy exposed by the corpus's #' Returns the category taxonomy exposed by the corpus's
#' `summary_categories` view, grouped to one row per #' `summary_categories` view, grouped to one row per
#' `(category, subtype)` pair. Use this to discover valid `category` #' `(category, subtype)` pair. Use this to discover valid `category`
#' values for [cog_spending()] / [cog_revenue()] / #' values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
#' [cog_geographic_rollup()] and to audit which Census item codes feed #' [cog_geographic_rollup()] and to audit which Census item codes feed
#' each category. #' each category.
#' #'
#' @param type Either `NULL` (default, return both spending and revenue #' `subtype` COALESCEs the crosswalk's three subtype columns, so it carries
#' rows), `"spending"`, or `"revenue"`. #' `spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
#' `balance_subtype` on balance rows. Note that [cog_balances()] itself takes
#' no `subtype` argument — for holdings, `category` is a strict coarsening of
#' `balance_subtype` — but the value is surfaced here because it is the
#' discovery surface downstream consumers build their vocabulary from.
#'
#' @param type Either `NULL` (default, every row: expenditure, revenue and
#' balance), `"spending"`, `"revenue"`, or `"balance"`.
#' @param pattern Optional regex matched case-insensitively against the #' @param pattern Optional regex matched case-insensitively against the
#' `category` column (e.g. `"Police"` or `"Tax"`). #' `category` column (e.g. `"Police"` or `"Tax"`).
#' @return Tibble with columns `category`, `category_type`, `subtype`, #' @return Tibble with columns `category`, `category_type`, `subtype`,
#' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by #' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
#' `category_type`, `category`, `subtype`. #' `category_type`, `category`, `subtype`. Includes one row per flow for the
#' reserved pseudo-category `"All Categories"`, which carries `NA` for
#' `subtype`, `n_codes` and `item_codes` because it is a query mode rather
#' than a crosswalk entry — see [cog_spending()]'s `category` argument.
#' @export #' @export
cog_categories <- function(type = NULL, pattern = NULL) { cog_categories <- function(type = NULL, pattern = NULL) {
if (!is.null(type)) { if (!is.null(type)) {
if (!is.character(type) || length(type) != 1L || if (!is.character(type) || length(type) != 1L ||
!type %in% c("spending", "revenue")) { !type %in% c("spending", "revenue", "balance")) {
cli::cli_abort('`type` must be NULL, "spending", or "revenue".') cli::cli_abort('`type` must be NULL, "spending", "revenue", or "balance".')
} }
} }
if (!is.null(pattern) && if (!is.null(pattern) &&
@@ -48,7 +58,7 @@ cog_categories <- function(type = NULL, pattern = NULL) {
sql <- paste( sql <- paste(
"SELECT category, category_type, "SELECT category, category_type,
COALESCE(spend_subtype, revenue_subtype) AS subtype, COALESCE(spend_subtype, revenue_subtype, balance_subtype) AS subtype,
COUNT(DISTINCT item_code) AS n_codes, COUNT(DISTINCT item_code) AS n_codes,
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS item_codes string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS item_codes
FROM summary_categories", FROM summary_categories",
@@ -56,5 +66,28 @@ cog_categories <- function(type = NULL, pattern = NULL) {
"GROUP BY category, category_type, subtype "GROUP BY category, category_type, subtype
ORDER BY category_type, category, subtype" ORDER BY category_type, category, subtype"
) )
tibble::as_tibble(DBI::dbGetQuery(con, sql)) out <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
# The reserved pseudo-category is a query mode, not a crosswalk row, so it
# has no item codes to report -- hence NA rather than 0 for n_codes. It is
# emitted for the two FLOW vocabularies only: cog_balances() returns a stock
# and has no concept argument to sum within.
pseudo <- tibble::tibble(
category = .ALL_CATEGORIES,
category_type = c("expenditure", "revenue"),
subtype = NA_character_,
n_codes = NA_integer_,
item_codes = NA_character_
)
if (!is.null(type)) {
db_type <- if (type == "spending") "expenditure" else type
pseudo <- pseudo[pseudo$category_type == db_type, , drop = FALSE]
}
if (!is.null(pattern) && nrow(pseudo) > 0L) {
keep <- grepl(pattern, pseudo$category, ignore.case = TRUE)
pseudo <- pseudo[keep, , drop = FALSE]
}
if (nrow(pseudo) == 0L) return(out)
out <- rbind(out, pseudo)
out[order(out$category_type, out$category, out$subtype), , drop = FALSE]
} }
+127
View File
@@ -0,0 +1,127 @@
# How a verb names the set of governments it queries.
#
# Historically there was one way: a `govid` character vector, rendered by
# .sql_lit_chr() into a quoted IN list. That is fine for a handful of
# governments and pathological for a fleet. Measured against the production
# corpus, the same FY2022 aggregate over the 20,106-government `type = "city"`
# cohort:
#
# cohort expressed as time
# IN (20,106 literals) 449 ms
# join against a temp cohort table 99 ms
# predicate on canonical_fips_xwalk 94 ms
# no cohort filter at all (the floor) 88 ms
#
# 4.8x, and within 7% of the no-filter floor. The rendered IN list is 301,591
# characters and .verb_spendrev() embeds it in 5-8 separate statements per
# call, so the parse-and-plan cost is paid over and over (uscogdata#58).
#
# A cohort therefore has two independent halves, and a query can carry either
# or both:
#
# ids an explicit canonical_govid vector -> literal IN list
# predicate state/type over canonical_fips_xwalk -> IN (SELECT ...)
#
# Both together is an INTERSECTION -- "these ids, narrowed to that state/type"
# -- never a precedence rule where one silently wins.
#' Build the internal cohort object shared by every query verb.
#'
#' `state` and `type` are coerced with the SAME helpers `cog_gov_search()`
#' uses. That is load-bearing, not tidiness: the public argument is a postal
#' abbreviation (`"WI"`) while `canonical_fips_xwalk.fips_state` holds a FIPS
#' code (`"55"`), and `type` is a label (`"city"`) against an integer
#' `govs_type`. A predicate written against the raw parameter matches nothing
#' and returns an empty result indistinguishable from "this government
#' reported nothing" -- cog-api hit exactly that trap optimizing this path.
#' One definition of the translation, not two.
#'
#' @param govid Already-coerced character vector of canonical_govids, or NULL.
#' @param state Postal abbreviation or FIPS code, or NULL.
#' @param type Type label or integer code, or NULL.
#' @noRd
.make_cohort <- function(govid = NULL, state = NULL, type = NULL) {
if (is.null(govid) && is.null(state) && is.null(type)) {
cli::cli_abort(c(
"A cohort must be named.",
"*" = "Pass {.arg govid} for specific governments, or {.arg state}/{.arg type} for every government matching a predicate.",
"i" = "Passing both intersects them: the governments in {.arg govid} that also match {.arg state}/{.arg type}."
), class = "uscogdata_no_cohort")
}
structure(
list(
ids = govid,
state = state,
type = type,
state_fips = if (is.null(state)) NULL else .coerce_state_to_fips(state),
type_int = if (is.null(type)) NULL else .coerce_type(type)
),
class = "uscogdata_cohort"
)
}
#' Is any part of this cohort expressed as an xwalk predicate?
#' @noRd
.cohort_by_predicate <- function(cohort) {
!is.null(cohort$state_fips) || !is.null(cohort$type_int)
}
#' Render the cohort as a SQL boolean expression over `col`.
#'
#' `col` may be qualified (`"l.canonical_govid"`, `"x.canonical_govid"`) --
#' several call sites join the xwalk under an alias. The subquery's own
#' projected column stays unqualified: it selects from canonical_fips_xwalk,
#' not from the outer relation.
#' @noRd
.cohort_sql <- function(cohort, col = "canonical_govid") {
preds <- character(0)
if (!is.null(cohort$ids)) {
preds <- c(preds, sprintf("%s IN (%s)", col, .sql_lit_chr(cohort$ids)))
}
if (.cohort_by_predicate(cohort)) {
xwalk_preds <- character(0)
if (!is.null(cohort$state_fips)) {
xwalk_preds <- c(xwalk_preds,
sprintf("fips_state = %s", .sql_lit_chr(cohort$state_fips)))
}
if (!is.null(cohort$type_int)) {
xwalk_preds <- c(xwalk_preds, sprintf("govs_type = %d", cohort$type_int))
}
preds <- c(preds, sprintf(
"%s IN (SELECT canonical_govid FROM canonical_fips_xwalk WHERE %s)",
col, paste(xwalk_preds, collapse = " AND ")
))
}
paste(preds, collapse = " AND ")
}
#' How many governments the cohort covers.
#'
#' One COUNT against the crosswalk, used only to populate the provenance
#' `scope$cohort` block. Deliberately a count rather than the id list: a
#' fleet-scale cohort would otherwise put 20,000 ids into every response body,
#' which is the cost this issue exists to remove.
#' @noRd
.cohort_count <- function(con, cohort) {
sql <- sprintf(
"SELECT COUNT(*) AS n FROM canonical_fips_xwalk WHERE %s",
.cohort_sql(cohort)
)
as.integer(DBI::dbGetQuery(con, sql)$n[[1]])
}
#' The provenance `scope$cohort` block for a predicate cohort, or NULL when
#' the cohort was named by id alone (in which case `govids_found`/
#' `govids_missing` already describe it exactly).
#' @noRd
.cohort_provenance <- function(con, cohort) {
if (!.cohort_by_predicate(cohort)) return(NULL)
list(
state = if (is.null(cohort$state)) NA_character_ else as.character(cohort$state),
type = if (is.null(cohort$type)) NA_character_ else as.character(cohort$type),
n_governments = .cohort_count(con, cohort)
)
}
+149
View File
@@ -0,0 +1,149 @@
# R/complete.R
#
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
# cell the corpus does not carry still appears, labelled with WHY it is
# missing.
#
# The corpus stopped storing the wide era's explicit zeros
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
#
# <= FY2011 dense_source absent => Census published $0 (census_zero)
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
#
# Before sparsification a wide-era query whose cells were all $0 came back as
# explicit $0 rows; afterwards it came back empty, with nothing to say which
# of the two meanings applied. This restores that -- and improves on it,
# because the pre-sparsification corpus could not distinguish the two either.
#
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
# That difference is the entire point: writing 0 into a modern absence would
# invent data, which is the error the representation contract exists to stop.
#' @noRd
.abort_complete_unsupported <- function(reason, alternative) {
cli::cli_abort(c(
"{.code complete = TRUE} is not supported for this query.",
x = reason,
i = alternative
), class = "uscogdata_complete_unsupported")
}
#' @noRd
.require_representation <- function(con, manifest) {
needed <- c("representation.parquet", "code_set.parquet")
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
logical(1))]
if (length(missing) == 0L) return(invisible(TRUE))
cli::cli_abort(c(
"This corpus does not publish the representation contract.",
x = "Missing: {.file {missing}}.",
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
), class = "uscogdata_representation_unavailable")
}
#' The cells a government-year COULD carry: every code in force for that
#' government's own type, mapped through `summary_categories`, restricted to
#' the calling verb's crosswalk subtype scope (the same subtype-membership
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
#' operations/capital/assistance) and (when given) its category filter.
#'
#' Scoped by `govs_type` deliberately. Filling against the union of all types
#' would invent cells that the government can never report -- a county row for
#' "state IG transfer to school districts" -- and those inventions would then
#' be indistinguishable from real census zeros.
#'
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
#' aggregate rows. Without it the grid would offer cells the verb structurally
#' never returns, so every one of them would fill as a phantom $0.
#' @noRd
.completion_grid_sql <- function(subtype_col, cohort, years, category,
subtype_scope) {
category_pred <- if (is.null(category)) {
""
} else {
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
}
sprintf(
"SELECT DISTINCT
cs.year,
x.canonical_govid,
x.gov_name,
c.%1$s AS subtype_value,
c.category,
r.absence_means
FROM code_set cs
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
JOIN summary_categories c ON c.item_code = cs.item_code
JOIN representation r ON r.year = cs.year
WHERE %2$s
AND cs.year IN (%3$s)
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%1$s IN (%4$s)
%5$s",
subtype_col, .cohort_sql(cohort, "x.canonical_govid"),
paste(as.integer(years), collapse = ","),
.sql_lit_chr(subtype_scope), category_pred
)
}
#' Fill `result` out to the full grid, stamping `value_source` on every row.
#'
#' Returns the completed tibble with a `.completion` attribute carrying the
#' provenance block. Reported rows are passed through untouched -- filling
#' must never alter or drop what the corpus actually published.
#' @noRd
.complete_result <- function(result, con, subtype_col, cohort, years, category,
subtype_scope) {
grid <- tibble::as_tibble(DBI::dbGetQuery(
con, .completion_grid_sql(subtype_col, cohort, years, category, subtype_scope)
))
result$value_source <- rep("reported", nrow(result))
if (nrow(grid) == 0L) {
attr(result, ".completion") <- list(
applied = TRUE, rows_filled = 0L, absence_means = list()
)
return(result)
}
names(grid)[names(grid) == "subtype_value"] <- subtype_col
key <- function(d) {
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
}
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
if (nrow(missing) > 0L) {
filled <- tibble::tibble(
year = as.integer(missing$year),
canonical_govid = as.character(missing$canonical_govid),
gov_name = as.character(missing$gov_name),
category = as.character(missing$category),
# census_zero is a value Census published; not_reported is unknown and
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
amt_nominal = ifelse(missing$absence_means == "census_zero",
0, NA_real_),
codes_included = NA_character_,
aggregate_fallback = NA,
value_source = as.character(missing$absence_means)
)
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
if ("notes" %in% names(result)) filled$notes <- NA_character_
result <- dplyr::bind_rows(result, filled)
result <- result[order(result$year, result$canonical_govid,
result[[subtype_col]], result$category), ,
drop = FALSE]
}
rules <- unique(grid[, c("year", "absence_means")])
attr(result, ".completion") <- list(
applied = TRUE,
rows_filled = nrow(missing),
absence_means = stats::setNames(
as.list(as.character(rules$absence_means)), as.character(rules$year)
)
)
result
}
+72 -2
View File
@@ -5,9 +5,28 @@
.uscogdata_env <- new.env(parent = emptyenv()) .uscogdata_env <- new.env(parent = emptyenv())
.uscogdata_defaults <- list( .uscogdata_defaults <- list(
url = "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/", # Public HuggingFace mirror of the published corpus: CC-BY-4.0, no
# credential, CDN-backed. This is the default so `library(uscogdata)`
# followed by a verb works with zero configuration -- previously the
# default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in the
# package supplied a working URL, so a new user had no path to a session.
#
# The trailing slash is required: every consumer concatenates onto this
# (see .resolve_url(), which enforces it anyway).
#
# Override with USCOGDATA_URL or options(uscogdata.url=) to read a
# Nextcloud share or a local copy made by cog_mirror().
url = "https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/",
cache_dir = NULL, cache_dir = NULL,
manifest_ttl_secs = 3600L manifest_ttl_secs = 3600L,
# NULL means "emit no pragma", which leaves DuckDB's own defaults intact:
# every visible core, and 80% of RAM. That is right for one interactive
# session on a dedicated machine and wrong for a server, where several
# readers share a box and each would otherwise claim all of it. See
# .resolve_duckdb_threads() for why this is a supported option rather than
# something a consumer reaches into the namespace to set.
duckdb_threads = NULL,
duckdb_memory_limit = NULL
) )
#' Resolve a config value: env var > option > default #' Resolve a config value: env var > option > default
@@ -50,3 +69,54 @@
v <- .cfg("cache_dir") v <- .cfg("cache_dir")
if (is.null(v)) tools::R_user_dir("uscogdata", "cache") else v if (is.null(v)) tools::R_user_dir("uscogdata", "cache") else v
} }
#' Resolve the DuckDB thread cap, or NULL to leave DuckDB's default alone.
#'
#' `cog_open()` used to connect with a bare `dbConnect()` and set no `threads`
#' pragma, so DuckDB claimed every core it could see. cog-api works around that
#' by reaching into this namespace at boot --
#' `getFromNamespace(".ensure_session", "uscogdata")()` followed by a manual
#' `SET threads` -- which depends on a private name and on the session already
#' being open. Making it a resolved option removes the reason to do that.
#'
#' `.cfg()` returns an environment variable as CHARACTER, so this coerces
#' rather than trusting the type: `USCOGDATA_DUCKDB_THREADS=4` arrives as "4",
#' and `sprintf("SET threads TO %d", "4")` would abort inside the connection
#' path with an error about the pragma rather than about the setting.
#' @noRd
.resolve_duckdb_threads <- function() {
v <- .cfg("duckdb_threads")
if (is.null(v) || (is.character(v) && !nzchar(v))) return(NULL)
n <- suppressWarnings(as.integer(v))
if (length(n) != 1L || is.na(n) || n < 1L) {
cli::cli_abort(c(
"{.envvar USCOGDATA_DUCKDB_THREADS} must be a single positive integer.",
x = "Got {.val {v}}.",
i = "Unset it (or {.code options(uscogdata.duckdb_threads = NULL)}) to use DuckDB's default of every visible core."
), class = "uscogdata_invalid_duckdb_threads")
}
n
}
#' Resolve the DuckDB memory limit, or NULL to leave DuckDB's default alone.
#'
#' The value is a DuckDB size string (`"4GB"`, `"512MB"`). Only its SHAPE is
#' checked here -- DuckDB owns the unit vocabulary, and re-implementing that
#' parse would be a second definition free to drift from the engine's. An
#' unrecognised unit therefore surfaces as DuckDB's own error at `SET` time,
#' which names the setting correctly; the check here exists to reject the
#' inputs that would otherwise reach the connection as a SQL fragment.
#' @noRd
.resolve_duckdb_memory_limit <- function() {
v <- .cfg("duckdb_memory_limit")
if (is.null(v) || (is.character(v) && !nzchar(v))) return(NULL)
if (length(v) != 1L || !is.character(v) ||
!grepl("^[0-9]+(\\.[0-9]+)?\\s*[A-Za-z]{0,3}$", v)) {
cli::cli_abort(c(
"{.envvar USCOGDATA_DUCKDB_MEMORY_LIMIT} must be a single DuckDB size string.",
x = "Got {.val {v}}.",
i = "Examples: {.val 4GB}, {.val 512MB}, {.val 1.5GB}."
), class = "uscogdata_invalid_duckdb_memory_limit")
}
trimws(v)
}
+107
View File
@@ -0,0 +1,107 @@
# R/coverage.R
#
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
# findings F-020 and F-023).
#
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
# 7. Every other year is a sample, and the sample varies enormously: on the
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
# what the verbs have always done -- correctly -- but the return value said
# nothing about it, so a statewide total resting on 18% of the universe looked
# exactly like one resting on 98%.
#
# Owner's settled design: a `coverage` argument selecting WHICH units to
# include, plus always-on metadata saying how many there were either way. The
# principle behind it: using these verbs correctly must not require the caller
# to know the survey calendar.
# Years ending in 2 or 7 are full censuses of every government; all others are
# samples.
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
#' @noRd
.is_census_year <- function(years) {
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
}
#' @noRd
.validate_coverage <- function(coverage) {
tryCatch(
match.arg(coverage, c("all", "census", "consistent")),
error = function(e) {
cli::cli_abort(
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
class = "uscogdata_invalid_coverage", parent = e
)
}
)
}
#' Restrict `years` to census years for `coverage = "census"`.
#'
#' Aborts rather than returning an empty result when the requested range holds
#' no census year: silently handing back zero rows for a query the caller
#' believes they made is the failure mode this whole issue is about.
#' @noRd
.apply_census_years <- function(years, coverage, verb) {
if (!identical(coverage, "census")) return(as.integer(years))
keep <- as.integer(years)[.is_census_year(years)]
if (length(keep) == 0L) {
cli::cli_abort(c(
"{.code coverage = \"census\"} leaves no years to query.",
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
), class = "uscogdata_no_census_years")
}
sort(keep)
}
#' Keep only units that report in EVERY requested year (a balanced panel).
#'
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
#' filter (the peer-comparison target, which is the subject of the comparison
#' rather than a member of the cohort being balanced).
#' @noRd
.filter_consistent <- function(result, years, id_col = "canonical_govid",
keep_ids = character(0)) {
years <- unique(as.integer(years))
if (nrow(result) == 0L || length(years) <= 1L) return(result)
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
present <- vapply(ids, function(g) {
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
}, logical(1))
consistent <- c(ids[present], keep_ids)
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
drop = FALSE]
}
#' Per-year coverage metadata, always attached regardless of mode.
#'
#' Built from the REQUESTED years rather than the years present in the result,
#' so a year in which nothing reported still appears -- with
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
#' missing year fails to make.
#'
#' `n_units_reporting` describes the result the caller actually received, so
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
#' `n_units_reporting` is the number that tells the truth.
#' @noRd
.coverage_table <- function(result, years, n_expected,
id_col = "canonical_govid", rows = NULL) {
years <- sort(unique(as.integer(years)))
src <- if (is.null(rows)) result else rows
reporting <- vapply(years, function(y) {
ids <- src[[id_col]][as.integer(src$year) == y]
length(unique(ids[!is.na(ids)]))
}, integer(1))
tibble::tibble(
year = years,
n_units_reporting = as.integer(reporting),
n_units_expected = rep(as.integer(n_expected), length(years)),
is_census_year = .is_census_year(years)
)
}
+114 -2
View File
@@ -11,6 +11,30 @@
#' returns `result` invisibly for chaining. `"list"` returns the raw #' returns `result` invisibly for chaining. `"list"` returns the raw
#' provenance list (identical to `attr(result, "provenance")`). #' provenance list (identical to `attr(result, "provenance")`).
#' @return Either `result` (invisibly) or the provenance list. #' @return Either `result` (invisibly) or the provenance list.
#' @section Two kinds of series break:
#' Catalogued breaks reach you without being asked for, in two disjoint
#' fields, because a caveat about one series and a caveat about the whole
#' corpus are different claims:
#'
#' * **`series_break_refs`** — breaks matched against the item codes actually
#' present in this result. A break in one code you queried.
#' * **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
#' which are statements about the corpus rather than about any one code:
#' dollar precision across the 1976/1977 boundary (`SB085`), imputation
#' exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
#' representation change (`SB194`), and the FY2017 government-identifier
#' change (`SB086`). These are selected on the break-year window alone.
#'
#' `SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
#' crosses the boundary where an absent cell stops meaning "Census published
#' $0" and starts meaning "not reported".
#' @section Other provenance blocks:
#' `transformations$units_conversion` records the `$1,000s`-to-dollars
#' multiply that every amount column has already had applied.
#' `transformations$per_capita` records the population denominator and its
#' year range. `coverage` and `coverage_mode` appear on multi-government
#' results (see [cog_geographic_rollup()]). `completion` appears when
#' `complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
#' @export #' @export
cog_explain <- function(result, format = c("print", "list")) { cog_explain <- function(result, format = c("print", "list")) {
format <- match.arg(format) format <- match.arg(format)
@@ -60,7 +84,20 @@ cog_explain <- function(result, format = c("print", "list")) {
cli::cli_text("Basis: {prov$basis}{note}") cli::cli_text("Basis: {prov$basis}{note}")
} }
if (!is.null(prov$expenditure_concept)) { # Each verb reports its OWN concept. Both fields are always present (each
# defaults to its concept's default), so printing `expenditure_concept`
# unconditionally would tell a cog_revenue() caller "Concept: primary",
# which names a spending concept their result has nothing to do with.
if (identical(prov$verb, "cog_revenue")) {
if (!is.null(prov$revenue_concept)) {
cli::cli_text("Concept: {prov$revenue_concept} revenue")
}
} else if (identical(prov$verb, "cog_balances")) {
# Both concept fields are deliberately NA here (a stock has no flow
# concept). Printing the raw NA reads as a missing value rather than an
# intentional one, so say what it means instead.
cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)")
} else if (!is.null(prov$expenditure_concept)) {
concept_note <- if (!is.null(prov$expenditure_concept_note) && concept_note <- if (!is.null(prov$expenditure_concept_note) &&
!is.na(prov$expenditure_concept_note)) { !is.na(prov$expenditure_concept_note)) {
sprintf(" (%s)", prov$expenditure_concept_note) sprintf(" (%s)", prov$expenditure_concept_note)
@@ -112,17 +149,92 @@ cog_explain <- function(result, format = c("print", "list")) {
if (length(prov$suggestions) > 0L) { if (length(prov$suggestions) > 0L) {
cli::cli_h2("Suggestions") cli::cli_h2("Suggestions")
sugg_lines <- vapply(prov$suggestions, function(s) { sugg_lines <- vapply(prov$suggestions, function(s) {
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label, line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
s$available_years[1], s$available_years[2], s$hint) s$available_years[1], s$available_years[2], s$hint)
if (isTRUE(s$suppressed_amount > 0)) {
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
line
}, character(1)) }, character(1))
cli::cli_ul(sugg_lines) cli::cli_ul(sugg_lines)
} }
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
cli::cli_h2("Reporting coverage")
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
cov <- prov$coverage
cli::cli_ul(sprintf(
"%d: %d of %d units reporting (%.0f%%) -- %s year",
cov$year, cov$n_units_reporting, cov$n_units_expected,
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
ifelse(cov$is_census_year, "census", "sample")
))
if (any(!cov$is_census_year)) {
cli::cli_text(
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
)
}
}
if (isTRUE(prov$completion$applied)) {
cli::cli_h2("Completion")
cli::cli_text(
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
)
rules <- prov$completion$absence_means
if (length(rules) > 0L) {
cli::cli_ul(vapply(names(rules), function(y) {
sprintf("%s: an absent cell means %s", y,
if (identical(rules[[y]], "census_zero")) {
"Census published $0 (filled as 0)"
} else {
"the government did not report (filled as NA, not 0)"
})
}, character(1)))
}
}
if (length(prov$series_break_refs) > 0L) { if (length(prov$series_break_refs) > 0L) {
cli::cli_h2("Series breaks") cli::cli_h2("Series breaks")
cli::cli_ul(.series_break_story_lines(prov$series_break_refs)) cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
} }
# Kept in a section of its own: these qualify the whole result, so folding
# them in with the per-code breaks above would invite reading them as a
# caveat about one series.
if (length(prov$corpus_break_refs) > 0L) {
cli::cli_h2("Corpus-wide caveats")
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
}
# Balance results only (NULL on money-verb provenance, so they are
# unaffected). This is the ONLY on-demand surface for the GAAP disclosure:
# .emit_balance_caveats() fires at most once per session, and is routinely
# consumed by a suppressMessages() call or by a knitted chunk nobody reads,
# so a caller who deliberately audits a result with cog_explain() must still
# be told.
bc <- prov$balance_caveats
if (!is.null(bc)) {
cli::cli_h2("Holdings caveats")
if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note)
if (length(bc$truncated) > 0L) {
cli::cli_text(
"Requested years extend beyond what these families actually cover:"
)
cli::cli_ul(vapply(bc$truncated, function(s) {
w <- bc$coverage_window[[s]]
if (length(w) == 2L) {
sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2])
} else {
s
}
}, character(1)))
}
}
cli::cli_h2("Transformations") cli::cli_h2("Transformations")
uc <- prov$transformations$units_conversion uc <- prov$transformations$units_conversion
if (isTRUE(uc$applied)) { if (isTRUE(uc$applied)) {
+2 -2
View File
@@ -28,7 +28,7 @@
"*" = "{.code Sys.setenv(USCOGDATA_URL = \"<url-or-local-path>/\")}", "*" = "{.code Sys.setenv(USCOGDATA_URL = \"<url-or-local-path>/\")}",
"*" = "{.code options(uscogdata.url = \"<url-or-local-path>/\")}", "*" = "{.code options(uscogdata.url = \"<url-or-local-path>/\")}",
i = "For an offline smoke test, use the bundled fixture: {.code system.file(\"extdata/fixture_corpus\", package = \"uscogdata\")}.", i = "For an offline smoke test, use the bundled fixture: {.code system.file(\"extdata/fixture_corpus\", package = \"uscogdata\")}.",
i = "For the live Civilytics corpus, request the Nextcloud share URL from the package maintainer." i = "The public corpus is the default: unset USCOGDATA_URL to use it, or point it at a local copy made by {.code cog_mirror()}."
), class = "uscogdata_url_not_configured") ), class = "uscogdata_url_not_configured")
} }
invisible(url) invisible(url)
@@ -138,7 +138,7 @@
#' year, matching canonical_fips_xwalk) rather than as-of-year; as-of-year #' year, matching canonical_fips_xwalk) rather than as-of-year; as-of-year
#' moved to the *_asof columns. This package's own geography always came from #' moved to the *_asof columns. This package's own geography always came from
#' the xwalk (already present-based), so behaviour is unchanged. #' the xwalk (already present-based), so behaviour is unchanged.
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L)) { .validate_schema <- function(manifest, supported = c(4L, 5L, 6L, 7L)) {
if (!manifest$schema_version %in% supported) { if (!manifest$schema_version %in% supported) {
cli::cli_abort(c( cli::cli_abort(c(
"Corpus schema version mismatch.", "Corpus schema version mismatch.",
+81
View File
@@ -0,0 +1,81 @@
# R/pagination.R
#
# Shared limit/offset machinery. #39 established the semantics inside
# .verb_spendrev(); #57 extends them to cog_gov_search() and cog_balances(),
# which is what made a single definition worth having: three inline copies of
# "coerce, refuse, unwrap the count" would be three places for the meaning of
# `total_rows` to drift.
#
# The SQL side stays in .build_verb_sql() (R/spending.R) -- it already wraps
# the aggregate in an outer SELECT so COUNT(*) OVER() sees the post-GROUP-BY
# row count rather than the pre-aggregation one, and that is the subtle part
# worth not duplicating either.
#' Coerce and check a limit/offset pair.
#'
#' Returns the coerced pair, or NULL for `limit` when no page was requested.
#' `offset` defaults to 0 whenever `limit` is set, so a caller can supply just
#' `limit` and get the first page.
#'
#' Conflicts with other arguments are deliberately NOT checked here: they
#' differ per verb (`complete`/`recipe` for the money verbs, basket mode for
#' `cog_gov_search()`, `recipe` alone for `cog_balances()`), and a shared
#' function taking a list of conflict flags would be harder to read than the
#' three explicit refusals at the call sites.
#' @noRd
.validate_pagination <- function(limit, offset) {
if (is.null(limit)) {
return(list(limit = NULL, offset = NULL))
}
limit <- as.integer(limit)
if (length(limit) != 1L || is.na(limit) || limit < 0L) {
cli::cli_abort("`limit` must be a single non-negative integer.",
class = "uscogdata_invalid_pagination")
}
offset <- if (is.null(offset)) 0L else as.integer(offset)
if (length(offset) != 1L || is.na(offset) || offset < 0L) {
cli::cli_abort("`offset` must be a single non-negative integer.",
class = "uscogdata_invalid_pagination")
}
list(limit = limit, offset = offset)
}
#' Wrap a query so one page comes back carrying the unpaginated total.
#'
#' `COUNT(*) OVER()` rides along as an ordinary column, so the caller gets the
#' true total from the SAME scan instead of a second round trip. The outer
#' `SELECT *` matters: appending LIMIT/OFFSET directly to a grouped query would
#' have the window function count pre-aggregation rows.
#' @noRd
.paginate_sql <- function(base_sql, limit, offset) {
if (is.null(limit)) return(base_sql)
sprintf(
"SELECT *, COUNT(*) OVER() AS pagination_total_rows
FROM (%s) AS _paged
LIMIT %d OFFSET %d",
base_sql, limit, offset
)
}
#' Strip the count column back out and report the unpaginated total.
#'
#' Returns `list(result = , total_rows = )`.
#'
#' An empty page -- an offset past the end -- carries no row to read the window
#' function off, so that one case falls back to a second, unpaginated
#' `COUNT(*)` rather than reporting a wrong zero. `unpaged_sql` is passed as a
#' function so the fallback query is only BUILT when it is actually needed;
#' every caller's unpaginated SQL is otherwise constructed on every paged call
#' and thrown away.
#' @noRd
.take_pagination_total <- function(result, con, unpaged_sql) {
if (nrow(result) > 0L) {
total <- result$pagination_total_rows[[1]]
result$pagination_total_rows <- NULL
return(list(result = result, total_rows = as.integer(total)))
}
count_sql <- sprintf("SELECT COUNT(*) AS n FROM (%s) AS _uncounted",
if (is.function(unpaged_sql)) unpaged_sql() else unpaged_sql)
list(result = result,
total_rows = as.integer(DBI::dbGetQuery(con, count_sql)$n[[1]]))
}
+134 -10
View File
@@ -19,6 +19,13 @@
#' target's population at `year` to produce absolute bounds. If `FALSE`, #' target's population at `year` to produce absolute bounds. If `FALSE`,
#' `pop_range` is interpreted as absolute population counts. #' `pop_range` is interpreted as absolute population counts.
#' @param max_peers Integer cap on the number of peers returned. #' @param max_peers Integer cap on the number of peers returned.
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
#' most recent census year with an observed population, so a cohort is not
#' built from a sample year in which most of the candidate universe is
#' absent. `"consistent"` needs a year range, which cohort selection does not
#' have, so it selects like `"all"` and is carried on the result as
#' `attr(x, "coverage")` for [cog_peer_compare()].
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`, #' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as #' `population`, `pop_ratio`, `rank`. The cohort year is attached as
#' `attr(x, "cohort_year")`. #' `attr(x, "cohort_year")`.
@@ -29,7 +36,9 @@ cog_find_peers <- function(target_govid,
same_state = FALSE, same_state = FALSE,
pop_range = c(0.7, 1.3), pop_range = c(0.7, 1.3),
is_ratio = TRUE, is_ratio = TRUE,
max_peers = 10L) { max_peers = 10L,
coverage = c("all", "census", "consistent")) {
coverage <- .validate_coverage(coverage)
if (!is.character(target_govid) || length(target_govid) != 1L) { if (!is.character(target_govid) || length(target_govid) != 1L) {
cli::cli_abort("`target_govid` must be a length-1 character string.") cli::cli_abort("`target_govid` must be a length-1 character string.")
} }
@@ -59,7 +68,7 @@ cog_find_peers <- function(target_govid,
)) ))
} }
cohort_year <- .resolve_cohort_year(con, target_govid, year) cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
pop_sql <- sprintf( pop_sql <- sprintf(
"SELECT population FROM gov_population_yearly "SELECT population FROM gov_population_yearly
@@ -107,12 +116,34 @@ cog_find_peers <- function(target_govid,
attr(peers, "cohort_year") <- as.integer(cohort_year) attr(peers, "cohort_year") <- as.integer(cohort_year)
attr(peers, "pop_range") <- as.numeric(pop_range) attr(peers, "pop_range") <- as.numeric(pop_range)
attr(peers, "is_ratio") <- isTRUE(is_ratio) attr(peers, "is_ratio") <- isTRUE(is_ratio)
attr(peers, "coverage") <- coverage
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
peers peers
} }
# `coverage` picks the cohort vintage when the caller did not name one.
# "census" snaps to the most recent CENSUS year with an observed population,
# so a cohort is not silently built from a sample year in which most of the
# candidate universe is absent. "consistent" is a comparison-time concept --
# it needs a year RANGE, which cohort selection does not have -- so it selects
# like "all" here and is carried on the result for cog_peer_compare().
#' @noRd #' @noRd
.resolve_cohort_year <- function(con, target_govid, year) { .resolve_cohort_year <- function(con, target_govid, year,
coverage = "all") {
if (!is.null(year)) return(as.integer(year)) if (!is.null(year)) return(as.integer(year))
if (identical(coverage, "census")) {
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
.sql_lit_chr(target_govid)
)
y <- DBI::dbGetQuery(con, sql)$y
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
cli::cli_abort(c(
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
), class = "uscogdata_no_census_years")
}
sql <- sprintf( sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly "SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s", WHERE canonical_govid = %s",
@@ -133,7 +164,9 @@ cog_find_peers <- function(target_govid,
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and #' [cog_find_peers()] result or a character vector of `canonical_govid`) and
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`, #' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot #' `summary_p75`) so the result can be faceted by `role` in a single ggplot
#' call. #' call. Those summary rows are quantiles **within each category**, not
#' quantiles of each peer's total — see the `@return` section before summing
#' them.
#' #'
#' @param target_govid Character scalar. #' @param target_govid Character scalar.
#' @param peers A tibble from [cog_find_peers()] or a character vector of #' @param peers A tibble from [cog_find_peers()] or a character vector of
@@ -143,10 +176,36 @@ cog_find_peers <- function(target_govid,
#' @param per_capita Default `TRUE` — peer compare usually normalizes by #' @param per_capita Default `TRUE` — peer compare usually normalizes by
#' population. #' population.
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`. #' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only #' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for #' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' single-government queries but cannot be used here because combining Total #' refused here because combining Total across peer sets counts
#' across peer sets counts intergovernmental transfers twice. #' intergovernmental transfers twice; `"primary"` and `"direct"` combine
#' safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#'
#' The comparison target is exempt from `"consistent"` balancing -- it is the
#' subject of the comparison, not a member of the cohort -- and the
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
#' cohort actually returned. `n_units_reporting` counts peers only, against
#' the cohort size: "3 of your 15 peers reported in FY2019".
#' @return Tibble matching [cog_spending()]'s columns, plus a `role` #' @return Tibble matching [cog_spending()]'s columns, plus a `role`
#' column taking values `"target"`, `"peer"`, `"summary_p25"`, #' column taking values `"target"`, `"peer"`, `"summary_p25"`,
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank #' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
@@ -155,12 +214,57 @@ cog_find_peers <- function(target_govid,
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character #' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`, #' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
#' `cohort_year`, and `cohort_govids`. #' `cohort_year`, and `cohort_govids`.
#'
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
#' Each one is computed **within each `(year, spend_subtype,
#' category)` cell** across the peer set, so a `summary_p50` row is *the
#' median peer's value in that one category*, not *the value of the median
#' peer's total*. The median peer for Police and the median peer for Fire
#' are usually different governments, so summing `summary_*` rows across
#' categories does not give any peer's total and misstates the band it
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
#' one cohort, with a sign flip at FY2012.
#'
#' Facet by `role` **and** `category` (the documented use, and what the
#' rows are built for). For a genuine "median peer's total spending" line,
#' sum each peer's own categories first and take the quantile of those
#' per-government totals:
#'
#' ```r
#' library(dplyr)
#' cmp |>
#' filter(role %in% c("target", "peer")) |>
#' group_by(year, role, canonical_govid) |>
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
#' filter(role == "peer") |>
#' group_by(year) |>
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
#' ```
#' @section Reading `coverage`:
#' `provenance$coverage` reports `n_units_reporting` against
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
#' it counts cohort members with rows for the category you asked for, not
#' cohort members collected that year.** A government that was surveyed and
#' genuinely spends nothing in that category is indistinguishable here from one
#' that was never surveyed.
#'
#' The ratio is therefore **not a response rate** and must not be used as one.
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
#' contract policing to the county sheriff, not non-response.
#'
#' The comparison that *is* valid is the same category across a census year
#' (ending in 2 or 7) and a sample year, where the real-zero component is
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
#' marks which is which.
#' @export #' @export
cog_peer_compare <- function(target_govid, peers, category, years, cog_peer_compare <- function(target_govid, peers, category, years,
per_capita = TRUE, adjust_to_year = NULL, per_capita = TRUE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) { expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call() call <- match.call()
expenditure_concept <- match.arg(expenditure_concept) expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) { if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_peer_compare") .abort_concept_not_aggregatable("cog_peer_compare")
} }
@@ -183,9 +287,21 @@ cog_peer_compare <- function(target_govid, peers, category, years,
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)] peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
all_govids <- unique(c(target_govid, peer_govids)) all_govids <- unique(c(target_govid, peer_govids))
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) years <- .apply_census_years(years, coverage, "cog_peer_compare")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer") r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
# The target is exempt from balancing: it is the subject of the comparison,
# not a member of the cohort being balanced, and dropping it would leave a
# peer comparison with nothing to compare. Filtering happens BEFORE the
# quantiles below, so a "consistent" cohort's summary rows describe that
# cohort rather than the unbalanced one.
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years, keep_ids = target_govid)
}
value_col <- .peer_value_col(per_capita, adjust_to_year) value_col <- .peer_value_col(per_capita, adjust_to_year)
summary_rows <- .peer_summary_rows(r, value_col) summary_rows <- .peer_summary_rows(r, value_col)
@@ -206,6 +322,14 @@ cog_peer_compare <- function(target_govid, peers, category, years,
canonical_govid = target_govid, canonical_govid = target_govid,
gov_name = unique(r$gov_name[r$role == "target"]) gov_name = unique(r$gov_name[r$role == "target"])
) )
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
# reported in FY2019". Including the target would inflate every count by one
# and make a cohort that has entirely stopped reporting look non-empty.
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(
out, years, length(peer_govids),
rows = r[r$role == "peer", , drop = FALSE]
)
attr(out, "provenance") <- prov attr(out, "provenance") <- prov
out out
} }
+32 -4
View File
@@ -6,11 +6,13 @@
per_capita, adjust_to_year, result, sql, per_capita, adjust_to_year, result, sql,
subtype_col, basis = NA_character_, subtype_col, basis = NA_character_,
basis_note = NA_character_, basis_note = NA_character_,
expenditure_concept = "direct", expenditure_concept = "primary",
expenditure_concept_note = NA_character_, expenditure_concept_note = NA_character_,
expenditure_concept_direct_suppressed = FALSE, expenditure_concept_direct_suppressed = FALSE,
revenue_concept = "general",
harmonization = NULL, recipe = NULL, harmonization = NULL, recipe = NULL,
suggestions = list()) { suggestions = list(),
completion = NULL) {
manifest <- .uscogdata_env$manifest manifest <- .uscogdata_env$manifest
codes <- result[["codes_included"]] codes <- result[["codes_included"]]
@@ -37,11 +39,20 @@
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L)) schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
con <- .uscogdata_env$con con <- .uscogdata_env$con
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) { have_con <- !is.null(con) && DBI::dbIsValid(con)
break_refs <- if (have_con) {
.build_series_break_refs(con, codes_observed, years, schema_version) .build_series_break_refs(con, codes_observed, years, schema_version)
} else { } else {
character(0) character(0)
} }
# Corpus-wide caveats travel separately: they qualify the whole result
# rather than one series, and they do not depend on codes_observed (see
# .build_corpus_break_refs()).
corpus_refs <- if (have_con) {
.build_corpus_break_refs(con, years, schema_version)
} else {
character(0)
}
list( list(
verb = verb, verb = verb,
@@ -56,7 +67,16 @@
basis_note = basis_note, basis_note = basis_note,
expenditure_concept = expenditure_concept, expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note, expenditure_concept_note = expenditure_concept_note,
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed), # isTRUE() alone would collapse a deliberate NA (all-categories mode,
# where suppression detection cannot run -- see .verb_spendrev()) down to
# FALSE, turning "we don't know" back into the false claim this field
# exists to avoid. Preserve NA; otherwise normalize to a strict logical.
expenditure_concept_direct_suppressed = if (isTRUE(is.na(expenditure_concept_direct_suppressed))) {
NA
} else {
isTRUE(expenditure_concept_direct_suppressed)
},
revenue_concept = revenue_concept,
harmonization = harmonization %||% list( harmonization = harmonization %||% list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0, applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
note = NA_character_ note = NA_character_
@@ -116,6 +136,14 @@
) )
), ),
series_break_refs = break_refs, series_break_refs = break_refs,
corpus_break_refs = corpus_refs,
# What `complete = TRUE` filled, and the rule it filled by. Always
# present so a consumer can read `completion$applied` without testing
# for the key -- an absent block and applied = FALSE would otherwise be
# indistinguishable from an older reader version.
completion = completion %||% list(
applied = FALSE, rows_filled = 0L, absence_means = list()
),
manifest = list( manifest = list(
schema_version = as.integer(manifest$schema_version), schema_version = as.integer(manifest$schema_version),
pipeline_commit = manifest$pipeline_commit %||% NA_character_, pipeline_commit = manifest$pipeline_commit %||% NA_character_,
+3 -3
View File
@@ -113,7 +113,7 @@ cog_recipes <- function(pattern = NULL) {
#' `year_min`/`year_max` -- so there is no double-counting. (Checkpoint #' `year_min`/`year_max` -- so there is no double-counting. (Checkpoint
#' review docs/phase_r_harmonization_review.md § 0.2.) #' review docs/phase_r_harmonization_review.md § 0.2.)
#' @noRd #' @noRd
.run_recipe <- function(con, recipe_id, govid, years) { .run_recipe <- function(con, recipe_id, cohort, years) {
sql <- sprintf( sql <- sprintf(
"SELECT l.year, l.canonical_govid, "SELECT l.year, l.canonical_govid,
COALESCE(x.gov_name, l.gov_name) AS gov_name, COALESCE(x.gov_name, l.gov_name) AS gov_name,
@@ -128,11 +128,11 @@ cog_recipes <- function(pattern = NULL) {
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
WHERE r.recipe_id = %1$s WHERE r.recipe_id = %1$s
AND l.canonical_govid IN (%2$s) AND %2$s
AND l.year IN (%3$s) AND l.year IN (%3$s)
GROUP BY 1, 2, 3 GROUP BY 1, 2, 3
ORDER BY 1, 2", ORDER BY 1, 2",
.sql_lit_chr(recipe_id), .sql_lit_chr(govid), .sql_lit_chr(recipe_id), .cohort_sql(cohort, "l.canonical_govid"),
paste(as.integer(years), collapse = ",") paste(as.integer(years), collapse = ",")
) )
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
+54 -4
View File
@@ -8,14 +8,58 @@
#' multiplies by 1000 and records the conversion in `provenance`). #' multiplies by 1000 and records the conversion in `provenance`).
#' #'
#' @inheritParams cog_spending #' @inheritParams cog_spending
#' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories broken out
#' one row each. The reserved value `"All Categories"` instead returns a
#' single summed row per `(year, canonical_govid, subtype)`, covering every
#' category inside the requested concept's subtype scope. It cannot be
#' combined with other category names, and it is not the same thing as
#' `revenue_concept = "total"`: the concept chooses which subtypes are in
#' scope, `"All Categories"` chooses whether rows inside that scope are
#' broken out or summed. Because the result keeps one row per
#' `revenue_subtype`, filtering the returned frame to
#' `revenue_subtype == "own_source"` gives an own-source revenue total.
#' @param revenue_concept Which of Census's two published revenue concepts to
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
#' values -- never as item-code first letters, which cannot classify
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
#' prefix `X` does the same):
#'
#' * `"general"` (default) -- Census General Revenue: `own_source` +
#' `federal` + `state` + `local_aid`. The manual defines this concept by
#' subtraction (section 4.3: *"General revenue comprises all revenue
#' except that classified as liquor store, utility, or insurance trust
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
#' insurance trust revenue are all excluded.
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
#' `general` plus utility, liquor store, and insurance trust revenue
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
#' `X01`/`X02`/`X05`/`X08`).
#'
#' The two are related by Census's own identity, `Total Revenue = General +
#' Utility + Liquor Store + Insurance Trust`.
#'
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
#' systems moved out of the annual finance file into the separate Annual
#' Survey of Public Pensions, so a `"total"` series steps down at the
#' FY2016/FY2017 seam for reasons that are about collection scope rather
#' than revenue (series breaks `SB197`-`SB202`).
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`, #' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, #' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`. #' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' and `value_source` when `complete = TRUE`.
#' @export #' @export
cog_revenue <- function(govid, years, category = NULL, cog_revenue <- function(govid = NULL, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL, per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL) { basis = c("harmonized", "raw"), recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE, limit = NULL, offset = NULL,
state = NULL, type = NULL) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
# this verb's recipe families (see R/suggestions.R).
.verb_spendrev( .verb_spendrev(
verb = "cog_revenue", verb = "cog_revenue",
view_base = "revenue_annotated", view_base = "revenue_annotated",
@@ -28,6 +72,12 @@ cog_revenue <- function(govid, years, category = NULL,
per_capita = per_capita, per_capita = per_capita,
adjust_to_year = adjust_to_year, adjust_to_year = adjust_to_year,
basis = basis, basis = basis,
recipe = recipe recipe = recipe,
revenue_concept = revenue_concept,
complete = complete,
limit = limit,
offset = offset,
state = state,
type = type
) )
} }
+66 -9
View File
@@ -19,30 +19,72 @@
#' `state`, `county`, `city`. Each element is a character vector of #' `state`, `county`, `city`. Each element is a character vector of
#' `canonical_govid` values. At least one layer required. #' `canonical_govid` values. At least one layer required.
#' @param category Single category name or character vector (passed through #' @param category Single category name or character vector (passed through
#' to [cog_spending()]). #' to [cog_spending()]), or the reserved `"All Categories"` for one summed
#' row per `(year, canonical_govid, subtype)` covering every category in the
#' concept's scope. `"All Categories"` is the efficient way to build a
#' geographic total: without it a caller must issue one rollup per category
#' and sum the results themselves.
#' @param years Integer vector of years. #' @param years Integer vector of years.
#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year #' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
#' population from `gov_population_yearly`. Govs with missing population #' population from `gov_population_yearly`. Govs with missing population
#' are excluded from the result. #' are excluded from the result.
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`. #' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only #' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for #' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' single-government queries but cannot be used here because combining Total #' refused here because combining Total across multiple layers of
#' across multiple layers of government double-counts intergovernmental #' government double-counts intergovernmental transfers (a state's payment
#' transfers (a state's payment to a school district is the same dollar the #' to a school district is the same dollar the district reports as its own
#' district reports as its own Direct spending). #' Direct spending); `"primary"` and `"direct"` combine safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` / #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`, #' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a #' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`, #' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
#' and `rollup$included_govids` / `rollup$excluded_govids`. #' and `rollup$included_govids` / `rollup$excluded_govids`.
#' @section Reading `coverage`:
#' `provenance$coverage` reports `n_units_reporting` against
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
#' it counts governments with rows for the category you asked for, not
#' governments collected that year.** A government that was surveyed and
#' genuinely spends nothing in that category is indistinguishable here from one
#' that was never surveyed.
#'
#' The ratio is therefore **not a response rate** and must not be used as one.
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
#' contract policing to the county sheriff, not non-response.
#'
#' The comparison that *is* valid is the same category across a census year
#' (ending in 2 or 7) and a sample year, where the real-zero component is
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
#' marks which is which.
#' @export #' @export
cog_geographic_rollup <- function(govids, category, years, cog_geographic_rollup <- function(govids, category, years,
per_capita = FALSE, adjust_to_year = NULL, per_capita = FALSE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) { expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call() call <- match.call()
expenditure_concept <- match.arg(expenditure_concept) expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) { if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_geographic_rollup") .abort_concept_not_aggregatable("cog_geographic_rollup")
} }
@@ -59,11 +101,21 @@ cog_geographic_rollup <- function(govids, category, years,
layer = rep(layer_names, lengths(govids)) layer = rep(layer_names, lengths(govids))
) )
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) # coverage = "census" drops non-census years BEFORE the query rather than
# after: a sample year's rows are not wanted at all, and fetching them only
# to discard them would also let them into the coverage table.
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r <- dplyr::left_join(r, layer_map, by = "canonical_govid", r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
relationship = "many-to-many") relationship = "many-to-many")
r$scope_note <- .rollup_scope_note(r$layer) r$scope_note <- .rollup_scope_note(r$layer)
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years)
}
excluded <- character(0) excluded <- character(0)
if (isTRUE(per_capita) && "pop_source" %in% names(r)) { if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
drop <- r$pop_source == "unavailable" drop <- r$pop_source == "unavailable"
@@ -82,6 +134,11 @@ cog_geographic_rollup <- function(govids, category, years,
included_govids = included, included_govids = included,
excluded_govids = excluded excluded_govids = excluded
) )
# n_units_expected is the universe the CALLER named -- the govids passed in
# -- not the national universe. That is what makes the ratio meaningful:
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
attr(r, "provenance") <- prov attr(r, "provenance") <- prov
r r
+80 -18
View File
@@ -6,8 +6,11 @@
#' the cross-vintage canonical-government registry. Operates in two modes: #' the cross-vintage canonical-government registry. Operates in two modes:
#' #'
#' * **Utility mode** (single `name`, the original behavior): returns all #' * **Utility mode** (single `name`, the original behavior): returns all
#' rows whose `gov_name` matches the regex case-insensitively, sorted by #' rows whose `gov_name` contains `name` as a **literal, case-insensitive
#' `population_acs` descending. Useful for exploratory lookups. #' substring**, sorted by `population_acs` descending. Useful for
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
#' government is findable by its own complete name even when that name
#' contains parentheses or a period.
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a #' * **Basket mode** (`length(name) > 1`): resolves each input row to a
#' single canonical govid and returns a tibble in input order, suitable #' single canonical govid and returns a tibble in input order, suitable
#' for piping straight into [cog_spending()] / [cog_revenue()] / #' for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -19,7 +22,8 @@
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`. #' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
#' 2. **Exact pass:** case-insensitive equality against `gov_name`. #' 2. **Exact pass:** case-insensitive equality against `gov_name`.
#' Single hit -> resolved. Multiple -> step 4. #' Single hit -> resolved. Multiple -> step 4.
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`. #' 3. **Substring fallback:** case-insensitive literal substring against
#' `gov_name` (metacharacters escaped).
#' Single hit -> resolved (`match_method = "substring"`). Zero hits -> #' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
#' `status = "no_match"`. Multiple hits -> step 4. #' `status = "no_match"`. Multiple hits -> step 4.
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the #' 4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -40,15 +44,27 @@
#' in basket mode (recycles from length 1). Excluded types `4`/`5` (or #' in basket mode (recycles from length 1). Excluded types `4`/`5` (or
#' `"special_district"` / `"school_district"`) trigger an explanatory #' `"special_district"` / `"school_district"`) trigger an explanatory
#' message and an empty result. #' message and an empty result.
#' @param limit Maximum number of rows to return, applied in SQL. `NULL`
#' (default) returns every match -- which, with no other filter, is the
#' entire crosswalk. Utility mode only: pagination has no meaning in basket
#' mode, where the result is one resolved row per requested name in input
#' order, and is refused there with class
#' `uscogdata_basket_pagination_conflict`.
#' @param offset Rows to skip before `limit` starts counting (0-based).
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
#' @return A tibble of `canonical_fips_xwalk` rows. In utility mode, all #' @return A tibble of `canonical_fips_xwalk` rows. In utility mode, all
#' matches sorted by `population_acs` desc. In basket mode, resolved #' matches sorted by `population_acs` desc, ties broken by
#' rows in input order, with `attr(., "resolution")` set to the #' `canonical_govid`. In basket mode, resolved rows in input order, with
#' sidecar tibble. #' `attr(., "resolution")` set to the sidecar tibble.
#'
#' When `limit` is set, carries a `total_rows` attribute: the full
#' unpaginated match count, computed by the same query (`COUNT(*) OVER()`)
#' rather than a second scan.
#' @seealso [cog_basket_resolution()], [cog_basket_unresolved()], #' @seealso [cog_basket_resolution()], [cog_basket_unresolved()],
#' [cog_spending()], [cog_revenue()]. #' [cog_spending()], [cog_revenue()].
#' @examples #' @examples
#' \dontrun{ #' \dontrun{
#' # Utility mode — exploratory regex lookup #' # Utility mode — exploratory substring lookup
#' cog_gov_search("broward", state = "FL") #' cog_gov_search("broward", state = "FL")
#' #'
#' # Basket mode — resolve a known cohort #' # Basket mode — resolve a known cohort
@@ -78,7 +94,12 @@
#' ) #' )
#' } #' }
#' @export #' @export
cog_gov_search <- function(name = NULL, state = NULL, type = NULL) { cog_gov_search <- function(name = NULL, state = NULL, type = NULL,
limit = NULL, offset = NULL) {
paging <- .validate_pagination(limit, offset)
limit <- paging$limit
offset <- paging$offset
if (!is.null(type) && length(type) == 1L && .is_excluded_type(type)) { if (!is.null(type) && length(type) == 1L && .is_excluded_type(type)) {
cli::cli_inform(c( cli::cli_inform(c(
i = "v0.1 covers gov_types 0-3 (state/county/city/township) only.", i = "v0.1 covers gov_types 0-3 (state/county/city/township) only.",
@@ -90,6 +111,18 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
con <- .ensure_session() con <- .ensure_session()
if (length(name) > 1L) { if (length(name) > 1L) {
# Basket mode returns one resolved row per requested name, in input order,
# with a resolution sidecar describing how each was matched. A page of that
# is not a page of anything the caller asked for -- the sidecar would still
# describe every name -- so refuse rather than silently ignoring the
# arguments. Same shape as the recipe/complete refusals in .verb_spendrev().
if (!is.null(limit)) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with basket mode.",
"i" = "Basket mode ({.code length(name) > 1}) returns one resolved row per requested name, in input order, with a resolution sidecar covering all of them.",
"*" = "Drop `limit`/`offset`, or search one name at a time."
), class = "uscogdata_basket_pagination_conflict")
}
return(.resolve_basket(name = name, state = state, type = type, con = con)) return(.resolve_basket(name = name, state = state, type = type, con = con))
} }
@@ -98,9 +131,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
if (!is.character(name) || length(name) != 1L) { if (!is.character(name) || length(name) != 1L) {
cli::cli_abort("`name` must be a length-1 character string.") cli::cli_abort("`name` must be a length-1 character string.")
} }
# Escaped, so `name` is a literal case-insensitive substring -- the same
# treatment basket mode has always given it. Interpolating it raw made a
# government unfindable by its own name whenever that name contains a
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
# match-everything wildcard, and let malformed pattern text reach the
# engine as an error -- which cog-api surfaced as a 500, reachable by
# typing a real name one character at a time (uscogdata#16, F-025).
preds <- c(preds, preds <- c(preds,
sprintf("regexp_matches(gov_name, %s, 'i')", sprintf("regexp_matches(gov_name, %s, 'i')",
.sql_lit_chr(name))) .sql_lit_chr(.escape_regex(name))))
} }
if (!is.null(state)) { if (!is.null(state)) {
st_fips <- .coerce_state_to_fips(state) st_fips <- .coerce_state_to_fips(state)
@@ -112,12 +152,26 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
} }
where <- if (length(preds) == 0L) "" else paste("WHERE", paste(preds, collapse = " AND ")) where <- if (length(preds) == 0L) "" else paste("WHERE", paste(preds, collapse = " AND "))
sql <- paste( # canonical_govid breaks ties. population_acs alone is NOT a total order --
# governments sharing a population, and the whole NULLS LAST block, came back
# in whatever order the scan produced. That was invisible while every call
# returned the full result set, but it makes a paged sweep unsound: two
# requests can order the tied rows differently, so a row is duplicated on one
# page and missing from the next. Any pagination has to sit on a total order.
base_sql <- paste(
"SELECT * FROM canonical_fips_xwalk", "SELECT * FROM canonical_fips_xwalk",
where, where,
"ORDER BY population_acs DESC NULLS LAST" "ORDER BY population_acs DESC NULLS LAST, canonical_govid"
) )
tibble::as_tibble(DBI::dbGetQuery(con, sql)) result <- tibble::as_tibble(
DBI::dbGetQuery(con, .paginate_sql(base_sql, limit, offset))
)
if (is.null(limit)) return(result)
paged <- .take_pagination_total(result, con, base_sql)
out <- paged$result
attr(out, "total_rows") <- paged$total_rows
out
} }
#' @noRd #' @noRd
@@ -136,8 +190,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
#' @noRd #' @noRd
.escape_regex <- function(x) { .escape_regex <- function(x) {
# Backslash-escape POSIX regex metacharacters so `name` is treated as a # Backslash-escape POSIX regex metacharacters so `name` is treated as a
# literal substring in the DuckDB regexp_matches call (substring fallback # literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
# only; utility-mode intentionally preserves regex behavior). # utility mode used to interpolate raw, which was a defect rather than a
# feature -- see the call site and uscogdata#16.
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE) gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
} }
@@ -172,11 +227,18 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
if (!is.character(state) || length(state) != 1L) { if (!is.character(state) || length(state) != 1L) {
cli::cli_abort("`state` must be a 2-letter USPS abbrev or a FIPS integer.") cli::cli_abort("`state` must be a 2-letter USPS abbrev or a FIPS integer.")
} }
fips <- .state_abbrev_to_fips[[toupper(state)]] # Membership tested before the lookup, not after: `.state_abbrev_to_fips` is
if (is.null(fips)) { # a named CHARACTER vector, and `[[` on a name it does not carry throws
cli::cli_abort("Unknown state abbreviation: {state}.") # base R's "subscript out of bounds" rather than returning NULL -- which
# made the curated message below unreachable dead code. Reported as a bare
# subscript error, `cog_gov_search(state = "ZZ")` gave no hint that the
# argument wants a postal abbreviation.
key <- toupper(state)
if (!key %in% names(.state_abbrev_to_fips)) {
cli::cli_abort("Unknown state abbreviation: {state}.",
class = "uscogdata_unknown_state")
} }
fips .state_abbrev_to_fips[[key]]
} }
# USPS state / territory abbreviation -> 2-digit FIPS code. # USPS state / territory abbreviation -> 2-digit FIPS code.
+33 -1
View File
@@ -14,9 +14,41 @@
sql <- sprintf( sql <- sprintf(
"SELECT DISTINCT break_id "SELECT DISTINCT break_id
FROM series_breaks_pq FROM series_breaks_pq
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d WHERE fin_code IN (%s) AND fin_code <> 'ALL'
AND break_year BETWEEN %d AND %d
ORDER BY break_id", ORDER BY break_id",
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years)) .sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
) )
DBI::dbGetQuery(con, sql)$break_id DBI::dbGetQuery(con, sql)$break_id
} }
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
#' `"ALL"` rather than an item code. They qualify the whole result, so they
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
#' before uscogdata#19. Selection is on the break_year window alone: which
#' codes a result happens to contain is irrelevant to a caveat about the
#' corpus.
#'
#' All four catalogued entries are *boundary* caveats (dollar precision
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
#' representation change at 2012, the id scheme change at 2017), so the same
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
#' path uses is the right one -- a request that never crosses the boundary
#' is not affected by it.
#'
#' Returned separately from `series_break_refs` so a consumer can tell a
#' whole-result caveat from a break in one series; the two are disjoint by
#' construction.
#' @noRd
.build_corpus_break_refs <- function(con, years, schema_version) {
if (schema_version < 5L || length(years) == 0L) return(character(0))
sql <- sprintf(
"SELECT DISTINCT break_id
FROM series_breaks_pq
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
ORDER BY break_id",
min(as.integer(years)), max(as.integer(years))
)
DBI::dbGetQuery(con, sql)$break_id
}
+35 -2
View File
@@ -2,17 +2,27 @@
#' Internal: open session, register views, cache manifest. #' Internal: open session, register views, cache manifest.
#' Not exported. Called lazily by verbs via .ensure_session(). #' Not exported. Called lazily by verbs via .ensure_session().
#'
#' `threads` and `memory_limit` default to the resolved configuration and are
#' applied as pragmas on the new connection. When both resolve to NULL -- which
#' is the case unless the operator sets one -- NO pragma is issued at all, so an
#' unconfigured session connects exactly as it did before this argument existed.
#' @noRd #' @noRd
cog_open <- function(url = .resolve_url(), cog_open <- function(url = .resolve_url(),
cache_dir = .resolve_cache_dir()) { cache_dir = .resolve_cache_dir(),
threads = .resolve_duckdb_threads(),
memory_limit = .resolve_duckdb_memory_limit()) {
.check_url_configured(url) .check_url_configured(url)
if (!dir.exists(cache_dir)) dir.create(cache_dir, recursive = TRUE) if (!dir.exists(cache_dir)) dir.create(cache_dir, recursive = TRUE)
con <- DBI::dbConnect(duckdb::duckdb()) con <- DBI::dbConnect(duckdb::duckdb())
# Before anything else touches the connection: httpfs reads the corpus, and
# a remote read should already be bound by whatever budget the operator set.
.apply_duckdb_limits(con, threads, memory_limit)
DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;") DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;")
manifest <- .fetch_or_cache_manifest(url, cache_dir) manifest <- .fetch_or_cache_manifest(url, cache_dir)
.validate_schema(manifest, supported = c(4L, 5L, 6L)) .validate_schema(manifest, supported = c(4L, 5L, 6L, 7L))
.validate_scope(manifest) .validate_scope(manifest)
.register_views(con, url, manifest) .register_views(con, url, manifest)
@@ -25,6 +35,26 @@ cog_open <- function(url = .resolve_url(),
invisible(con) invisible(con)
} }
#' Apply the operator's DuckDB resource budget to a fresh connection.
#'
#' Split out from cog_open() so the "unset changes nothing" property is one
#' readable branch rather than two conditionals buried in the connection path.
#' Both settings are session-scoped in DuckDB, so this must run per connection;
#' cog_close() discards the connection and the next cog_open() re-resolves,
#' which is what makes a changed option take effect on the next session.
#' @noRd
.apply_duckdb_limits <- function(con, threads, memory_limit) {
if (!is.null(threads)) {
DBI::dbExecute(con, sprintf("SET threads TO %d", threads))
}
if (!is.null(memory_limit)) {
# Quoted as a string literal: DuckDB's memory_limit takes '4GB', not 4GB.
DBI::dbExecute(con, sprintf("SET memory_limit TO %s",
.sql_lit_chr(memory_limit)))
}
invisible(con)
}
#' @noRd #' @noRd
.ensure_session <- function() { .ensure_session <- function() {
if (is.null(.uscogdata_env$con) || if (is.null(.uscogdata_env$con) ||
@@ -95,4 +125,7 @@ cog_close <- function() {
} }
.uscogdata_env$con <- NULL .uscogdata_env$con <- NULL
.uscogdata_env$manifest <- NULL .uscogdata_env$manifest <- NULL
.uscogdata_env$balance_caveats_shown <- NULL
# Memoised corpus-constant; a different corpus may be mounted next.
.uscogdata_env$balance_coverage_windows <- NULL
} }
+474 -60
View File
@@ -1,5 +1,65 @@
# R/spending.R # R/spending.R
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
# `spend_subtype` values. Classification is crosswalk membership, never
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
# route it (finding F-018).
#
# primary = operations + capital + assistance (the default)
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
# total = direct + intergovernmental (via the ig_* views)
#
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
# than intergovernmental -- including payments to retirees, i.e. insurance
# trust benefits. Verified against Census's own published FY2020 state
# aggregates (20statetypepu.txt): `total` reproduces the published
# expenditure sum to the dollar; omitting insurance benefits understates
# California's Direct by 10.9%.
.spend_subtypes_primary <- c("operations", "capital", "assistance")
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
# The reserved pseudo-category. Deliberately NOT "Total": `category = "Total"`
# would sit one argument away from `expenditure_concept = "total"` and mean
# something different -- the concept selects WHICH SUBTYPES are in scope, this
# selects whether the rows inside that scope are broken out by category or
# summed. "All Categories" states the operation and cannot be misread as the
# concept.
.ALL_CATEGORIES <- "All Categories"
#' @noRd
.expenditure_concept_subtypes <- function(concept) {
switch(concept,
primary = .spend_subtypes_primary,
# "total" = the direct subtypes here PLUS the intergovernmental leg,
# which travels through the ig_* views rather than this scope (see
# .build_verb_sql()).
direct = ,
total = .spend_subtypes_direct
)
}
# The two revenue concepts (uscogdata#12), again as crosswalk subtype sets.
# Census's manual section 4.3 defines the first by SUBTRACTING from the second
# -- "General revenue comprises all revenue except that classified as liquor
# store, utility, or insurance trust revenue" -- giving the identity
#
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
#
# Verified against Census's own computed concept fields (IndFin FY2012,
# Wisconsin state): 31,410,686 + 0 + 0 + 4,469,906 = 35,880,592, exact.
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
.revenue_subtypes_total <- c(.revenue_subtypes_general, "utility",
"liquor_store", "insurance_trust")
#' @noRd
.revenue_concept_subtypes <- function(concept) {
switch(concept,
general = .revenue_subtypes_general,
total = .revenue_subtypes_total
)
}
#' Summarized spending by category #' Summarized spending by category
#' #'
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are #' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
@@ -8,10 +68,21 @@
#' millions/billions). The conversion is recorded in the provenance attribute #' millions/billions). The conversion is recorded in the provenance attribute
#' under `transformations$units_conversion`. #' under `transformations$units_conversion`.
#' #'
#' @param govid Character vector of `canonical_govid` values. #' @param govid Character vector of `canonical_govid` values, or `NULL` to name
#' the cohort by `state`/`type` instead. One of `govid`, `state`, or `type`
#' is required.
#' @param years Integer vector of years. #' @param years Integer vector of years.
#' @param category Character vector of category names (from #' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories. #' `summary_categories.category`), or `NULL` for all categories broken out
#' one row each. The reserved value `"All Categories"` instead returns a
#' single summed row per `(year, canonical_govid, subtype)`, covering every
#' category inside the requested concept's subtype scope. It cannot be
#' combined with other category names, and it is not the same thing as
#' `expenditure_concept = "total"`: the concept chooses which subtypes are in
#' scope, `"All Categories"` chooses whether rows inside that scope are
#' broken out or summed. Because the result keeps one row per
#' `spend_subtype`, filtering the returned frame to
#' `spend_subtype == "operations"` gives an operating-expenditure total.
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and #' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
#' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year #' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
#' Census F-33 population from `gov_population_yearly`. Result also gains #' Census F-33 population from `gov_population_yearly`. Result also gains
@@ -42,22 +113,34 @@
#' `basis = "recipe"` with an inert `harmonization` block (`applied = #' `basis = "recipe"` with an inert `harmonization` block (`applied =
#' FALSE`, pointing at the `recipe` block instead) rather than a #' FALSE`, pointing at the `recipe` block instead) rather than a
#' possibly-misleading `"harmonized"`/`"raw"` value. #' possibly-misleading `"harmonized"`/`"raw"` value.
#' @param expenditure_concept `"direct"` (default) returns only the #' @param expenditure_concept Which spending concept to return. Concepts are
#' government's own direct spending (item codes `E`/`F`/`G`), unchanged #' defined as sets of the crosswalk's `spend_subtype` values -- never as
#' from prior releases. `"total"` additionally UNIONs in the #' item-code first letters, which cannot classify correctly (prefix `Y`
#' intergovernmental leg -- payments to local governments (`M` codes) and #' alone spans revenue, expenditure, and balance codes):
#' to the state government (`L` codes, excluding the `L--` family-total #'
#' rollup) -- so results gain rows with `spend_subtype == #' * `"primary"` (default) -- the government's own service provision:
#' "intergovernmental"`. Requires the active corpus's `summary_categories` #' `operations` + `capital` + `assistance` subtypes.
#' to carry M/L rows (added by cog_pipeline PR #59); aborts with class #' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
#' `uscogdata_ig_categories_unsupported` on an older corpus rather than #' `interest` (interest on debt) and `insurance_benefits` (insurance
#' silently under-reporting. Mutually exclusive with `recipe` (a recipe #' trust benefit payments, e.g. pensions -- Census manual section
#' already defines its own component codes). **Do not sum `"total"` #' 5.2.2.1 includes payments to retirees in Direct).
#' results across levels of government** (e.g. state + county + city): #' * `"total"` -- `direct` plus the intergovernmental leg: payments to
#' a state's `M12` payment to a school district is the same dollar the #' local governments (`M` codes), to the state government (`L` codes,
#' district reports as its own direct `E12`, so summing both double-counts #' excluding the `L--` family-total rollup), and state payments to
#' it. This matters in particular with [cog_geographic_rollup()], which #' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
#' sums across exactly that kind of multi-layer government set. #' `spend_subtype == "intergovernmental"`. Requires the active corpus's
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
#' older corpus rather than silently under-reporting. Mutually
#' exclusive with `recipe` (a recipe already defines its own component
#' codes).
#'
#' **Do not sum `"total"` results across levels of government** (e.g.
#' state + county + city): a state's `M12` payment to a school district is
#' the same dollar the district reports as its own direct `E12`, so
#' summing both double-counts it. This matters in particular with
#' [cog_geographic_rollup()], which sums across exactly that kind of
#' multi-layer government set.
#' #'
#' In the legacy wide era (<= FY2011), some functions are published ONLY #' In the legacy wide era (<= FY2011), some functions are published ONLY
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05` #' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
@@ -69,17 +152,87 @@
#' component (when one exists), and #' component (when one exists), and
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the #' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
#' figure in those rows is the intergovernmental leg alone, not Direct + #' figure in those rows is the intergovernmental leg alone, not Direct +
#' IG. #' IG. When `category = "All Categories"` is combined with
#' `expenditure_concept = "total"`, this detection cannot run (it keys on
#' per-category rows, which all-categories mode collapses to one literal
#' value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
#' possibly-false `FALSE`; query an explicit `category` to get a real
#' answer.
#' @param complete If `TRUE`, fill the requested grid so that a cell the
#' corpus does not carry still appears, labelled with **why** it is
#' missing, and add a `value_source` column to every row:
#'
#' * `"reported"` — the corpus carries this cell.
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
#' Census published `$0`. `amt_nominal` is `0`.
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
#' government did not report, and the value is unknown. `amt_nominal` is
#' `NA`, **not** `0` — writing a zero there would invent data.
#'
#' The grid comes from the corpus's `code_set` table, scoped to each
#' government's own type, so a county is never filled with cells only a
#' state can report. Reported rows are passed through untouched.
#'
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
#' when `representation`/`code_set` began shipping; aborts with class
#' `uscogdata_representation_unavailable` otherwise. Not available with
#' `recipe` or with `expenditure_concept = "total"` (class
#' `uscogdata_complete_unsupported`) — neither draws its cells from
#' `code_set`.
#' @param limit Maximum number of result rows to return, pushed into the SQL
#' query itself (`LIMIT`/`OFFSET`) rather than applied after the full
#' result is materialized. `NULL` (the default) returns every matching row,
#' exactly as before this parameter existed. Mutually exclusive with
#' `recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.
#' @param offset Rows to skip before `limit` starts counting (0-based).
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
#' @param state,type Name the cohort by predicate instead of by id: `state` is
#' a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
#' `"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
#' the same vocabulary, and the same internal coercion, as
#' [cog_gov_search()]. Both default to `NULL`.
#'
#' The cohort is then expressed as a subquery against `canonical_fips_xwalk`
#' inside each statement rather than round-tripped through R as a literal id
#' list. For a fleet-scale cohort that is the difference between a
#' 301,591-character `IN` list re-parsed in 5--8 statements per call and a
#' constant-size predicate: measured at **94 ms versus 449 ms** for the same
#' FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
#' 7% of the no-filter floor.
#'
#' Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
#' governments in `govid` that also match the predicate -- rather than one
#' silently taking precedence. Naming no cohort at all (`govid`, `state` and
#' `type` all `NULL`) aborts with class `uscogdata_no_cohort`.
#'
#' When the cohort is named by predicate, `provenance$scope$govids_found`
#' and `govids_missing` are empty -- there is no id list to report against --
#' and `provenance$scope$cohort` carries `state`, `type` and
#' `n_governments` instead. A `govid`-named cohort reports exactly as before.
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`, #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, #' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`. #' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`. #' and `value_source` when `complete = TRUE`.
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
#' whose `completion` block reports `applied`, `rows_filled`, and the
#' per-year `absence_means` rule that was applied. When `limit` is set,
#' also carries a `total_rows` attribute: the full unpaginated row count,
#' computed by the same query (`COUNT(*) OVER()`) rather than a second
#' round trip -- so a caller walking pages never has to ask "how many are
#' there" separately.
#' @export #' @export
cog_spending <- function(govid, years, category = NULL, cog_spending <- function(govid = NULL, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL, per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL, basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total")) { expenditure_concept = c("primary", "direct", "total"),
complete = FALSE, limit = NULL, offset = NULL,
state = NULL, type = NULL) {
# flow_prefixes no longer classifies rows (crosswalk subtype membership
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
# machinery to this verb's recipe families (see R/suggestions.R; the
# catalog only has E/F/G-component direct-expenditure recipes).
.verb_spendrev( .verb_spendrev(
verb = "cog_spending", verb = "cog_spending",
view_base = "spending_annotated", view_base = "spending_annotated",
@@ -93,7 +246,12 @@ cog_spending <- function(govid, years, category = NULL,
adjust_to_year = adjust_to_year, adjust_to_year = adjust_to_year,
basis = basis, basis = basis,
recipe = recipe, recipe = recipe,
expenditure_concept = expenditure_concept expenditure_concept = expenditure_concept,
complete = complete,
limit = limit,
offset = offset,
state = state,
type = type
) )
} }
@@ -101,8 +259,9 @@ cog_spending <- function(govid, years, category = NULL,
.abort_concept_not_aggregatable <- function(verb) { .abort_concept_not_aggregatable <- function(verb) {
cli::cli_abort(c( cli::cli_abort(c(
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.", "{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
"*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\ "*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
comparison or sum that spans more than one government.", {.code \"direct\"} for any comparison or sum that spans more than \\
one government.",
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\ "i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
money it hands to other governments. The receiving government reports \\ money it hands to other governments. The receiving government reports \\
that same dollar again as its own Direct when it actually spends it, \\ that same dollar again as its own Direct when it actually spends it, \\
@@ -118,26 +277,75 @@ cog_spending <- function(govid, years, category = NULL,
govid, years, category, govid, years, category,
per_capita, adjust_to_year, per_capita, adjust_to_year,
basis = c("harmonized", "raw"), recipe = NULL, basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total")) { expenditure_concept = c("primary", "direct", "total"),
revenue_concept = c("general", "total"),
complete = FALSE, limit = NULL, offset = NULL,
state = NULL, type = NULL) {
basis_explicit <- length(basis) == 1L basis_explicit <- length(basis) == 1L
basis <- match.arg(basis, c("harmonized", "raw")) basis <- match.arg(basis, c("harmonized", "raw"))
# match.arg() itself throws a base `simpleError`, not an rlang-classed # match.arg() itself throws a base `simpleError`, not an rlang-classed
# condition; wrap it so an invalid expenditure_concept aborts consistently # condition; wrap it so an invalid expenditure_concept aborts consistently
# with the rest of this package's validation (cli::cli_abort -> rlang_error). # with the rest of this package's validation (cli::cli_abort -> rlang_error).
expenditure_concept <- tryCatch( expenditure_concept <- tryCatch(
match.arg(expenditure_concept, c("direct", "total")), match.arg(expenditure_concept, c("primary", "direct", "total")),
error = function(e) { error = function(e) {
cli::cli_abort( cli::cli_abort(
"`expenditure_concept` must be one of {.val direct} or {.val total}.", "`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
class = "uscogdata_invalid_expenditure_concept", class = "uscogdata_invalid_expenditure_concept",
parent = e parent = e
) )
} }
) )
govid <- .coerce_govid_input(govid, arg = "govid") revenue_concept <- tryCatch(
match.arg(revenue_concept, c("general", "total")),
error = function(e) {
cli::cli_abort(
"`revenue_concept` must be one of {.val general} or {.val total}.",
class = "uscogdata_invalid_revenue_concept",
parent = e
)
}
)
# The concept's subtype scope. Every code path below -- the verb SQL, the
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
# by crosswalk subtype membership, never by item-code prefix. The
# expenditure "total" concept's extra intergovernmental leg is the one
# exception: it travels through the ig_* views rather than this scope,
# because its legacy rows are aggregate-flagged.
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
.expenditure_concept_subtypes(expenditure_concept)
} else {
.revenue_concept_subtypes(revenue_concept)
}
# NULL `govid` means "the cohort is named by predicate"; anything else is
# coerced and validated exactly as before, so an empty or wrong-typed vector
# still fails with its original message rather than being read as absent.
govid <- if (is.null(govid)) NULL else .coerce_govid_input(govid, arg = "govid")
# allow_all_categories = TRUE: cog_spending()/cog_revenue() are the two
# verbs the reserved pseudo-category is defined for. cog_balances() shares
# this validator but leaves the argument at its FALSE default, so it
# rejects "All Categories" instead of silently returning zero rows
# (finding 3, all-categories review).
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year, .validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe) recipe, allow_all_categories = TRUE)
# Built after validation so the argument-shape errors above keep firing
# first, and before .ensure_session() so a bad state/type costs no I/O.
cohort <- .make_cohort(govid, state, type)
# Recognize the reserved pseudo-category. Detected after type validation so a
# non-character `category` still fails with the ordinary type error.
all_categories <- !is.null(category) && .ALL_CATEGORIES %in% category
if (all_categories && length(category) > 1L) {
cli::cli_abort(c(
"{.val {(.ALL_CATEGORIES)}} cannot be combined with other categories.",
"i" = "It already sums every category in the requested concept's scope.",
"*" = "Ask for it alone, or list the specific categories you want."
), class = "uscogdata_all_categories_not_combinable")
}
if (!is.null(recipe) && identical(expenditure_concept, "total")) { if (!is.null(recipe) && identical(expenditure_concept, "total")) {
cli::cli_abort(c( cli::cli_abort(c(
@@ -148,10 +356,10 @@ cog_spending <- function(govid, years, category = NULL,
} }
# .verb_spendrev() is shared with cog_revenue(), which never exposes # .verb_spendrev() is shared with cog_revenue(), which never exposes
# expenditure_concept and always resolves it to "direct" -- so nothing on # expenditure_concept and always resolves it to the default -- so nothing
# the public API can reach this today. But it's a cheap guard against a # on the public API can reach this today. But it's a cheap guard against a
# future call (direct or via a modified cog_revenue()) that would UNION # future call (direct or via a modified cog_revenue()) that would UNION
# the IG leg's expenditure M/L rows into a revenue result, which has no # the IG leg's expenditure M/L/Q rows into a revenue result, which has no
# matching IG view and no sensible meaning. # matching IG view and no sensible meaning.
if (identical(expenditure_concept, "total") && if (identical(expenditure_concept, "total") &&
!identical(view_base, "spending_annotated")) { !identical(view_base, "spending_annotated")) {
@@ -165,23 +373,70 @@ cog_spending <- function(govid, years, category = NULL,
) )
} }
complete <- isTRUE(complete)
if (complete && !is.null(recipe)) {
.abort_complete_unsupported(
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
)
}
if (complete && identical(expenditure_concept, "total")) {
.abort_complete_unsupported(
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
)
}
if (complete && all_categories) {
.abort_complete_unsupported(
"`category = \"All Categories\"` collapses the category dimension that `code_set` grids over (see `.completion_grid_sql()`), so there is no per-category grid left to fill -- filling a summed row has no defined semantics.",
"Drop `complete`, or use `complete = TRUE` with an explicit `category` (or `category = NULL` for every category)."
)
}
# limit/offset push the page into the SQL itself (see .build_verb_sql()),
# so the two things that would make "a page of what" ambiguous are refused
# up front rather than silently ignored: complete = TRUE fills a grid over
# the FULL requested (year, category) space, and a recipe's result comes
# from .run_recipe()'s own query, which this function does not touch.
paging <- .validate_pagination(limit, offset)
limit <- paging$limit
offset <- paging$offset
if (!is.null(limit)) {
if (complete) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with `complete = TRUE`.",
"i" = "`complete` fills a grid over the FULL requested (year, category) space; paginating a slice of already-grouped rows has no defined meaning for the cells it would fill.",
"*" = "Drop `limit`/`offset`, or drop `complete`."
), class = "uscogdata_complete_pagination_conflict")
}
if (!is.null(recipe)) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with `recipe`.",
"i" = "A recipe's result comes from a separate query (`.run_recipe()`) that pagination is not wired into yet.",
"*" = "Drop `limit`/`offset`, or drop `recipe`."
), class = "uscogdata_recipe_pagination_conflict")
}
}
years <- as.integer(years) years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
con <- .ensure_session() con <- .ensure_session()
manifest <- .uscogdata_env$manifest manifest <- .uscogdata_env$manifest
scope <- .check_govids_in_scope(govid) scope <- .check_govids_in_scope(govid)
if (complete) .require_representation(con, manifest)
resolved <- .resolve_basis(basis, basis_explicit, manifest) resolved <- .resolve_basis(basis, basis_explicit, manifest)
recipe_block <- NULL recipe_block <- NULL
category_for_prov <- category category_for_prov <- category
total_rows <- NULL # set below only when limit is non-NULL (non-recipe path)
if (!is.null(recipe)) { if (!is.null(recipe)) {
.require_schema_v5(con, manifest, "recipe =") .require_schema_v5(con, manifest, "recipe =")
.validate_recipe_id(con, recipe) .validate_recipe_id(con, recipe)
comps <- .recipe_components(con, recipe) comps <- .recipe_components(con, recipe)
recipe_label <- comps$label[[1]] recipe_label <- comps$label[[1]]
result <- .run_recipe(con, recipe, govid, years) result <- .run_recipe(con, recipe, cohort, years)
sql <- attr(result, "sql_query") sql <- attr(result, "sql_query")
result <- .shape_recipe_result(result, subtype_col, recipe_label) result <- .shape_recipe_result(result, subtype_col, recipe_label)
recipe_block <- list( recipe_block <- list(
@@ -197,11 +452,37 @@ cog_spending <- function(govid, years, category = NULL,
} else { } else {
NULL NULL
} }
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view) sql <- .build_verb_sql(view, subtype_col, cohort, years,
if (all_categories) NULL else category,
ig_view, subtype_scope,
all_categories = all_categories,
limit = limit, offset = offset)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
if (!is.null(limit)) {
paged <- .take_pagination_total(result, con, function() {
.build_verb_sql(view, subtype_col, cohort, years,
if (all_categories) NULL else category,
ig_view, subtype_scope,
all_categories = all_categories)
})
result <- paged$result
total_rows <- paged$total_rows
}
} }
if (per_capita) result <- .attach_per_capita(result, con, govid) # Fill BEFORE per_capita / inflation so the added cells get the same
# treatment as reported ones: a census_zero stays $0 per capita and in real
# dollars, and a not_reported stays NA through both rather than becoming a
# spurious 0.
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
if (complete) {
result <- .complete_result(result, con, subtype_col, cohort, years,
category, subtype_scope)
completion <- attr(result, ".completion")
attr(result, ".completion") <- NULL
}
if (per_capita) result <- .attach_per_capita(result, con)
if (!is.null(adjust_to_year)) { if (!is.null(adjust_to_year)) {
result <- .attach_real_dollars(result, adjust_to_year, per_capita) result <- .attach_real_dollars(result, adjust_to_year, per_capita)
} }
@@ -229,7 +510,7 @@ cog_spending <- function(govid, years, category = NULL,
basis_for_prov <- resolved$basis basis_for_prov <- resolved$basis
basis_note_for_prov <- resolved$note basis_note_for_prov <- resolved$note
harmonization <- .build_harmonization_block( harmonization <- .build_harmonization_block(
con, govid, years, resolved, flow_prefixes con, cohort, years, resolved, subtype_col, subtype_scope
) )
# C1(a): gap detection must run against the Direct leg alone. `result` # C1(a): gap detection must run against the Direct leg alone. `result`
# can also carry UNION'd intergovernmental rows (expenditure_concept = # can also carry UNION'd intergovernmental rows (expenditure_concept =
@@ -244,9 +525,13 @@ cog_spending <- function(govid, years, category = NULL,
} else { } else {
result result
} }
suggestions <- .build_suggestions(con, govid, years, category, suggestions <- .build_suggestions(con, cohort, years, category,
direct_leg_result, direct_leg_result,
resolved$basis, flow_prefixes) resolved$basis, flow_prefixes,
.select_long_view(view_base, resolved$basis),
all_categories = all_categories,
subtype_col = subtype_col,
subtype_scope = subtype_scope)
} }
# C1(b): when expenditure_concept = "total", flag any row where the IG # C1(b): when expenditure_concept = "total", flag any row where the IG
@@ -258,13 +543,32 @@ cog_spending <- function(govid, years, category = NULL,
# direct spending in that category, which is correct, ordinary data). When # direct spending in that category, which is correct, ordinary data). When
# a covering recipe is found, both the row-level notes and the provenance # a covering recipe is found, both the row-level notes and the provenance
# say so rather than pass silently as a plausible Total. # say so rather than pass silently as a plausible Total.
direct_suppressed_info <- if (identical(expenditure_concept, "total")) { #
# In all-categories mode this cannot run at all: .detect_direct_suppressed()
# keys on (year, canonical_govid, category), and every row shares the same
# literal "All Categories" value, so the key collides across every real
# category for that (year, govid) -- an IG-only row for a suppressed
# category becomes indistinguishable from one sharing a key with an
# unrelated category's ordinary Direct row. `has_direct` would then read
# TRUE whenever the government has ANY direct spending at all, and the
# detector could never fire. Rather than run it and report a false FALSE,
# skip it and record NA -- the provenance must stop making a claim it
# cannot support (finding 1, all-categories review).
suppression_unavailable <- all_categories &&
identical(expenditure_concept, "total")
direct_suppressed_info <- if (suppression_unavailable) {
list(flag = rep(NA, nrow(result)), notes = rep(NA_character_, nrow(result)))
} else if (identical(expenditure_concept, "total")) {
.detect_direct_suppressed(con, result, subtype_col) .detect_direct_suppressed(con, result, subtype_col)
} else { } else {
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result))) list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
} }
direct_suppressed <- direct_suppressed_info$flag direct_suppressed <- direct_suppressed_info$flag
direct_suppressed_flag <- isTRUE(any(direct_suppressed)) direct_suppressed_flag <- if (suppression_unavailable) {
NA
} else {
isTRUE(any(direct_suppressed))
}
result$notes <- .notes_column(result, direct_suppressed_info$notes) result$notes <- .notes_column(result, direct_suppressed_info$notes)
@@ -273,9 +577,20 @@ cog_spending <- function(govid, years, category = NULL,
# leg is suppressed for at least one requested (year, category), append an # leg is suppressed for at least one requested (year, category), append an
# explicit warning rather than let the base note's "Total = Direct + IG" # explicit warning rather than let the base note's "Total = Direct + IG"
# framing stand unqualified for rows where that arithmetic didn't happen. # framing stand unqualified for rows where that arithmetic didn't happen.
# When suppression detection itself is unavailable (all-categories mode),
# say so instead of silently reusing the unqualified base note.
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) { expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded." base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
if (direct_suppressed_flag) { if (suppression_unavailable) {
paste0(
base_note,
" NOTE: direct-leg-suppression detection is unavailable when ",
"`category = \"All Categories\"` -- it keys on per-category rows, ",
"which this mode collapses. `expenditure_concept_direct_suppressed` ",
"is NA here rather than a possibly-false FALSE; query an explicit ",
"`category` (or `category = NULL`) to get a real answer."
)
} else if (isTRUE(direct_suppressed_flag)) {
paste0( paste0(
base_note, base_note,
" NOTE: for at least one requested (year, category) the Direct leg ", " NOTE: for at least one requested (year, category) the Direct leg ",
@@ -307,24 +622,53 @@ cog_spending <- function(govid, years, category = NULL,
expenditure_concept = expenditure_concept, expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note_for_prov, expenditure_concept_note = expenditure_concept_note_for_prov,
expenditure_concept_direct_suppressed = direct_suppressed_flag, expenditure_concept_direct_suppressed = direct_suppressed_flag,
revenue_concept = revenue_concept,
harmonization = harmonization, harmonization = harmonization,
recipe = recipe_block, recipe = recipe_block,
suggestions = suggestions suggestions = suggestions,
completion = completion
) )
prov$scope$govids_found <- scope$found prov$scope$govids_found <- scope$found
prov$scope$govids_missing <- scope$missing prov$scope$govids_missing <- scope$missing
# A predicate-named cohort has no id list to report found/missing against
# (both stay empty), so it describes itself instead. Deliberately a COUNT
# rather than the resolved ids: enumerating them would put 20,000 govids in
# every fleet-scale response body, which is the cost this path exists to
# remove. NULL for a govid-named cohort, so that output is untouched.
prov$scope$cohort <- .cohort_provenance(con, cohort)
attr(result, "provenance") <- prov attr(result, "provenance") <- prov
attr(result, ".popyear_range") <- NULL attr(result, ".popyear_range") <- NULL
# Attached here, after every downstream transform (per_capita/real-dollar
# joins, notes, subtype filtering), the same way provenance is -- an
# attribute set before those runs is not guaranteed to survive them.
if (!is.null(limit)) attr(result, "total_rows") <- total_rows
if (length(suggestions) > 0L) .inform_suggestions(suggestions) if (length(suggestions) > 0L) .inform_suggestions(suggestions)
result result
} }
#' Shared input validation for the money/holdings verbs.
#'
#' `allow_all_categories` gates the reserved pseudo-category
#' `.ALL_CATEGORIES` ("All Categories"). It is meaningful only where a
#' concept's subtype scope defines what "all" sums over --
#' `cog_spending()`/`cog_revenue()`, via `.verb_spendrev()`, pass `TRUE`.
#' `cog_balances()` leaves it at the `FALSE` default: holdings are a stock
#' with no concept vocabulary to sum across (see R/balances.R), and before
#' this guard existed `cog_balances(category = "All Categories")` silently
#' matched zero crosswalk rows and returned an empty result with no error
#' (finding 3, all-categories review). This validator is shared specifically
#' so the three verbs cannot drift apart on this again.
#' @noRd #' @noRd
.validate_verb_inputs <- function(govid, years, category, .validate_verb_inputs <- function(govid, years, category,
per_capita, adjust_to_year, recipe = NULL) { per_capita, adjust_to_year, recipe = NULL,
if (!is.character(govid) || length(govid) == 0L) { allow_all_categories = FALSE) {
# NULL is allowed only because the caller has already established that the
# cohort is named some other way (`state`/`type`); .make_cohort() is what
# refuses a call that names no cohort at all. A supplied-but-empty `govid`
# still fails here, exactly as before.
if (!is.null(govid) && (!is.character(govid) || length(govid) == 0L)) {
cli::cli_abort("`govid` must be a non-empty character vector.") cli::cli_abort("`govid` must be a non-empty character vector.")
} }
if (!(is.integer(years) || is.numeric(years)) || length(years) == 0L) { if (!(is.integer(years) || is.numeric(years)) || length(years) == 0L) {
@@ -333,6 +677,14 @@ cog_spending <- function(govid, years, category = NULL,
if (!is.null(category) && !is.character(category)) { if (!is.null(category) && !is.character(category)) {
cli::cli_abort("`category` must be character or NULL.") cli::cli_abort("`category` must be character or NULL.")
} }
if (!allow_all_categories && !is.null(category) &&
.ALL_CATEGORIES %in% category) {
cli::cli_abort(c(
"{.val {(.ALL_CATEGORIES)}} is not supported here.",
i = "It sums a spending or revenue concept's subtype scope; this verb has no concept vocabulary to sum across.",
i = "Use {.fn cog_spending} or {.fn cog_revenue} for an all-categories total."
), class = "uscogdata_all_categories_unsupported")
}
if (!is.logical(per_capita) || length(per_capita) != 1L) { if (!is.logical(per_capita) || length(per_capita) != 1L) {
cli::cli_abort("`per_capita` must be a length-1 logical.") cli::cli_abort("`per_capita` must be a length-1 logical.")
} }
@@ -361,6 +713,18 @@ cog_spending <- function(govid, years, category = NULL,
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
} }
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
#' anti-joins the LONG view rather than the annotated one: they have identical
#' row membership (the annotated views are the long views plus LEFT JOINs, see
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
#' the suppression test is asking about.
#' @noRd
.select_long_view <- function(view_base, basis) {
.select_view(sub("_annotated$", "_long", view_base), basis)
}
#' @noRd #' @noRd
.select_ig_view <- function(basis) { .select_ig_view <- function(basis) {
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated" if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
@@ -405,19 +769,38 @@ cog_spending <- function(govid, years, category = NULL,
} }
#' @noRd #' @noRd
.build_verb_sql <- function(view, subtype_col, govid, years, category, .build_verb_sql <- function(view, subtype_col, cohort, years, category,
ig_view = NULL) { ig_view = NULL, subtype_scope = NULL,
govid_lit <- .sql_lit_chr(govid) all_categories = FALSE, limit = NULL, offset = NULL) {
cohort_pred <- .cohort_sql(cohort)
years_lit <- paste(as.integer(years), collapse = ",") years_lit <- paste(as.integer(years), collapse = ",")
category_pred <- if (is.null(category)) { # In all-categories mode there is no category filter: the sum is defined by
# the concept's SUBTYPE allowlist (subtype_pred below), which is the real
# concept boundary. Filtering by category as well would be a no-op at best
# and, if the crosswalk ever gained an uncategorized code, a silent
# under-count of the very total this mode exists to guarantee.
category_pred <- if (all_categories || is.null(category)) {
"" ""
} else { } else {
sprintf("AND category IN (%s)", .sql_lit_chr(category)) sprintf("AND category IN (%s)", .sql_lit_chr(category))
} }
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
# The base views carry every subtype of their flow (spending_annotated has
# all five non-IG expenditure subtypes); the concept narrows here. For
# "total", the IG leg's rows are 'intergovernmental', so that value joins
# the allowlist exactly when ig_view is present.
subtype_pred <- if (is.null(subtype_scope)) {
""
} else {
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
}
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL, # expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
# never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L), # never UNION: the two legs are disjoint by crosswalk subtype (the direct
# so de-duplication would be pure cost, and a silent row-drop if two # view excludes 'intergovernmental'; the IG view is only that), so
# de-duplication would be pure cost, and a silent row-drop if two
# governments ever reported identical values. # governments ever reported identical values.
source_expr <- if (is.null(ig_view)) { source_expr <- if (is.null(ig_view)) {
view view
@@ -435,28 +818,59 @@ cog_spending <- function(govid, years, category = NULL,
# though its dollars came entirely from an aggregate row, silently # though its dollars came entirely from an aggregate row, silently
# suppressing the "Aggregate fallback applied" note on exactly the rows # suppressing the "Aggregate fallback applied" note on exactly the rows
# this feature exists to surface. # this feature exists to surface.
sprintf(
# Collapse the category dimension. subtype is deliberately KEPT: it is what
# lets a caller filter the result to `spend_subtype == "operations"` and
# get an operating-expenditure total, the measure a fiscal comparison
# actually wants. (There is no `subtype` argument -- this is a post-hoc
# filter on the returned column, not a query parameter.)
category_select <- if (all_categories) {
sprintf("%s AS category", .sql_lit_chr(.ALL_CATEGORIES))
} else {
"category"
}
category_group <- if (all_categories) "" else ", category"
base_sql <- sprintf(
"SELECT "SELECT
year, year,
canonical_govid, canonical_govid,
COALESCE(xwalk_gov_name, gov_name) AS gov_name, COALESCE(xwalk_gov_name, gov_name) AS gov_name,
%1$s, %1$s,
category, %7$s,
SUM(amt) * 1000.0 AS amt_nominal, SUM(amt) * 1000.0 AS amt_nominal,
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included, string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
bool_or(is_aggregate) AS aggregate_fallback bool_or(is_aggregate) AS aggregate_fallback
FROM %2$s FROM %2$s
WHERE canonical_govid IN (%3$s) WHERE %3$s
AND year IN (%4$s) AND year IN (%4$s)
%5$s %5$s
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category %6$s
ORDER BY year, canonical_govid, %1$s, category", GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s%8$s
subtype_col, source_expr, govid_lit, years_lit, category_pred ORDER BY year, canonical_govid, %1$s%8$s",
subtype_col, source_expr, cohort_pred, years_lit, category_pred, subtype_pred,
category_select, category_group
) )
# limit/offset push the page into the query itself instead of pulling every
# matching row across the network only to slice and discard most of it
# afterward (the pattern behind the 2026-08-06 production incident: a
# 193,105-row/194-page sweep re-ran the full query and re-listified every
# row on EVERY page). See .paginate_sql() in R/pagination.R for why the
# wrapping is an outer SELECT rather than a bare LIMIT on base_sql.
.paginate_sql(base_sql, limit, offset)
} }
#' Join population onto a result and derive the per-capita columns.
#'
#' The population lookup is keyed on the govids PRESENT IN `result`, not on the
#' cohort that produced it. Those are the only ones the LEFT JOIN below can
#' match, so the joined output is identical either way -- but it means this
#' works unchanged for a cohort named by predicate (where no id list exists in
#' R at all), and on a paginated call it looks up one page's governments
#' instead of the whole fleet's.
#' @noRd #' @noRd
.attach_per_capita <- function(result, con, govid) { .attach_per_capita <- function(result, con) {
if (nrow(result) == 0L) { if (nrow(result) == 0L) {
result$amt_per_capita_nominal <- numeric(0) result$amt_per_capita_nominal <- numeric(0)
result$pop_source <- character(0) result$pop_source <- character(0)
@@ -469,7 +883,7 @@ cog_spending <- function(govid, years, category = NULL,
FROM gov_population_yearly FROM gov_population_yearly
WHERE canonical_govid IN (%s) WHERE canonical_govid IN (%s)
AND year IN (%s)", AND year IN (%s)",
.sql_lit_chr(govid), years_lit .sql_lit_chr(unique(result$canonical_govid)), years_lit
) )
pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
result <- dplyr::left_join(result, pops, result <- dplyr::left_join(result, pops,
+276 -55
View File
@@ -1,8 +1,16 @@
# R/suggestions.R # R/suggestions.R
# Recipe-component-driven signposting: when a basis = "harmonized" query for # Recipe-component-driven signposting. When a basis = "harmonized" query for
# a category comes back with a coverage gap in some requested years (the # a category comes back incomplete in some requested year -- and a
# result has no rows at all in that year) that a harmonization recipe would # harmonization recipe would actually fill it for this government -- surface
# actually fill for this government, surface that recipe as a suggestion. # that recipe as a suggestion. "Incomplete" has two forms, and a recipe
# qualifies on either:
# 1. empty_year -- the result has no rows at all in that year.
# 2. suppressed_component -- the result HAS rows, but a component code
# carries dollars the verb's own long view structurally excludes
# (aggregate-published, or absent from summary_categories). This is
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
# aggregate-only E67/E68, so form 1 never fired and the caller got a
# number a third too low with no signpost at all.
# #
# This is deliberately keyed off the recipe catalog's component codes, not # This is deliberately keyed off the recipe catalog's component codes, not
# off harmonization_map rows: no live map row carries a non-blank # off harmonization_map rows: no live map row carries a non-blank
@@ -34,8 +42,8 @@
#' verb call: recipes whose generic join would fill a real gap in `result`. #' verb call: recipes whose generic join would fill a real gap in `result`.
#' #'
#' @param con Active DuckDB connection. #' @param con Active DuckDB connection.
#' @param govid Character vector of canonical_govid values (the verb's raw #' @param cohort The verb's cohort object (see `.make_cohort()`), naming the
#' `govid`). #' governments by id, by state/type predicate, or both.
#' @param years Integer vector of requested years. #' @param years Integer vector of requested years.
#' @param category `category` argument as passed to the verb (character #' @param category `category` argument as passed to the verb (character
#' vector or `NULL`; suggestions are only computed when non-NULL). #' vector or `NULL`; suggestions are only computed when non-NULL).
@@ -48,11 +56,54 @@
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to #' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart #' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
#' lookup scoped to the calling verb's own flow family. #' lookup scoped to the calling verb's own flow family.
#' @param long_view Name of the verb's own long view (from
#' `.select_long_view()`), passed through to `.suppressed_components()` to
#' measure the second qualifying path (uscogdata#9).
#' @param all_categories `TRUE` when the caller's `category` is the reserved
#' pseudo-category (`.ALL_CATEGORIES`). Defaults to `FALSE` so no other
#' caller's behaviour changes. When `TRUE`, the candidate-recipe sub-select
#' is scoped by `subtype_col`/`subtype_scope` instead of by `category` --
#' symmetric with `.build_verb_sql()`'s own all-categories branch (see
#' R/spending.R): the concept's subtype allowlist is the real scope
#' boundary, not any literal category value, and
#' `.ALL_CATEGORIES` ("All Categories") is never itself a row in
#' `summary_categories.category`, so leaving the category-keyed sub-select
#' in place here always returned zero candidates and silently disabled
#' signposting in all-categories mode (final whole-branch review, finding
#' 6).
#' @param subtype_col Name of the `summary_categories` subtype column to
#' scope by when `all_categories = TRUE` (`"spend_subtype"` or
#' `"revenue_subtype"` -- the same value `.build_verb_sql()` already
#' receives as its own `subtype_col`). Ignored when `all_categories =
#' FALSE`. `NULL` by default.
#' @param subtype_scope Character vector of subtype values to scope by when
#' `all_categories = TRUE` (the same value `.build_verb_sql()` already
#' receives as its own `subtype_scope` -- the concept's subtype allowlist,
#' e.g. `.expenditure_concept_subtypes(expenditure_concept)`). Ignored when
#' `all_categories = FALSE`. `NULL` by default.
#' @return List of `list(recipe_id, label, available_years, hint, #' @return List of `list(recipe_id, label, available_years, hint,
#' ig_recipe_id)`, possibly empty. #' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
#' suppressed_codes)`, possibly empty.
#'
#' Decomposed (Issue #33) into three extracted helpers to stay within the
#' project's "functions under 50 lines" convention:
#' \itemize{
#' \item `.query_candidate_recipes()` -- candidate recipe lookup by
#' category/subtype scope + `category_type` filter (#34) + M/L exclusion.
#' \item `.query_recipe_meta()` -- metadata (label, year spans).
#' \item `.query_covered_years()` -- Path 1 gap-year coverage via the
#' recipe's own generic join.
#' }
#' The for-loop that merges covered-years + suppressed-components into
#' suggestion objects stays inline here because it interleaves
#' empty_hit/supp_hit precedence with field assembly. Likewise kept inline:
#' the M/L-exclusion design-comment block and the final
#' `.attach_ig_counterparts()` call.
#' @noRd #' @noRd
.build_suggestions <- function(con, govid, years, category, result, basis, .build_suggestions <- function(con, cohort, years, category, result, basis,
flow_prefixes) { flow_prefixes, long_view,
all_categories = FALSE,
subtype_col = NULL, subtype_scope = NULL) {
if (!identical(basis, "harmonized") || is.null(category)) return(list()) if (!identical(basis, "harmonized") || is.null(category)) return(list())
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe -- # Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
@@ -67,17 +118,24 @@
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe # flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
# should never be suggested as a coverage-gap filler for EITHER verb, not # should never be suggested as a coverage-gap filler for EITHER verb, not
# just kept from being named as the *counterpart* of another suggestion. # just kept from being named as the *counterpart* of another suggestion.
candidates <- DBI::dbGetQuery(con, sprintf( #
"SELECT DISTINCT recipe_id FROM harmonization_recipes # The inner sub-select is the concept boundary (finding 6, final
WHERE component_code IN ( # whole-branch review): in all-categories mode it is scoped by
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s) # `subtype_col`/`subtype_scope` -- the same allowlist `.build_verb_sql()`
) # applies as a WHERE predicate to make the summed result a *concept*, not
AND recipe_id NOT IN ( # by `category` (`.ALL_CATEGORIES` is never a row in
SELECT DISTINCT recipe_id FROM harmonization_recipes # `summary_categories.category`, so a category-keyed sub-select always
WHERE LEFT(component_code, 1) IN ('M', 'L') # came back empty here). The M/L exclusion below is unchanged either way.
)", #
.sql_lit_chr(category) # Issue #34: scope the candidate query by `category_type` ('expenditure'
))$recipe_id # vs 'revenue') to prevent cross-flow-family leakage -- e.g.
# `cog_revenue(category = "Corrections")` must not surface
# expenditure-only recipes (E04/E05) merely because they share the same
# category name in summary_categories. The type is derived from
# flow_prefixes: E/F/G -> 'expenditure', anything else -> 'revenue'.
candidates <- .query_candidate_recipes(con, category, flow_prefixes,
all_categories, subtype_col,
subtype_scope)
if (length(candidates) == 0L) return(list()) if (length(candidates) == 0L) return(list())
result_years <- if (is.null(result) || nrow(result) == 0L) { result_years <- if (is.null(result) || nrow(result) == 0L) {
@@ -86,22 +144,170 @@
unique(as.integer(result$year)) unique(as.integer(result$year))
} }
gap_years <- setdiff(as.integer(years), result_years) gap_years <- setdiff(as.integer(years), result_years)
if (length(gap_years) == 0L) return(list())
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf( # Path 2 (uscogdata#9): component dollars this government holds that the
"SELECT recipe_id, any_value(label) AS label, # verb's own view structurally excludes. Measured across ALL requested
MIN(year_min) AS year_min, MAX(year_max) AS year_max # years, not just gap years -- the whole point is that a year with rows can
FROM harmonization_recipes # still be missing dollars. Scoped to the calling verb's own flow_prefixes
WHERE recipe_id IN (%s) # (I1) -- see `.suppressed_components()`'s own roxygen for why.
GROUP BY recipe_id", #
.sql_lit_chr(candidates) # This runs unconditionally whenever there are candidates -- an earlier
))) # revision of this fix wave tried a free, in-memory pre-check
# (`.needs_suppression_query()`) to skip the round trip on an already-
# covered path, but a scoped re-review measured it against the fixture and
# found it didn't pay for itself (it skipped ~3% of healthy calls, ~0% of
# the multi-govid batch shape it was meant to help, at a net cost increase
# once its own always-run metadata query was counted) while adding an
# untested exactness invariant -- that `result$codes_included` and this
# anti-join share the harmonized `item_code` space -- whose silent
# violation would kill signposting, the exact failure class uscogdata#9
# exists to prevent. Owner's call: keep this simple; a batch-aware
# optimization, if one is worth building, is a separate issue.
supp <- .suppressed_components(con, candidates, cohort, years, long_view, flow_prefixes)
# Which (recipe_id, year) pairs the recipe's own generic join actually if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
# covers for this government, restricted to the gap years -- the same
# join .run_recipe() uses (component year_min/year_max + gov_type_scope, meta <- .query_recipe_meta(con, candidates)
# no is_aggregate filter), just checking existence instead of summing.
covered <- DBI::dbGetQuery(con, sprintf( # Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
# covers for this government, restricted to the gap years.
covered <- .query_covered_years(con, candidates, cohort, gap_years)
suggestions <- list()
for (rid in candidates) {
empty_hit <- rid %in% covered$recipe_id
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
supp_hit <- nrow(s_rows) > 0L
if (!empty_hit && !supp_hit) next
m <- meta[meta$recipe_id == rid, ]
suggestions[[length(suggestions) + 1L]] <- list(
recipe_id = rid,
label = m$label[[1]],
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
hint = sprintf("re-run with recipe = '%s'", rid),
# An empty year is the stronger claim -- the category returned nothing
# at all -- so it wins when both paths qualify. The suppressed_* fields
# are still populated, so an empty_year fire also reports its dollars.
trigger = if (empty_hit) "empty_year" else "suppressed_component",
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
suppressed_years = if (supp_hit) {
sort(unique(as.integer(s_rows$year)))
} else {
integer(0)
},
suppressed_codes = if (supp_hit) {
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
} else {
character(0)
}
)
}
.attach_ig_counterparts(con, suggestions, flow_prefixes)
}
#' Query candidate harmonization recipe IDs for a coverage-gap suggestion.
#'
#' Selects recipes whose component codes fall within the requested scope
#' (category or subtype allowlist), excluding any recipe that is ITSELF an
#' intergovernmental (M/L) recipe -- i.e. every one of its own component
#' codes is M/L-prefixed. Without this exclusion, a category whose
#' summary_categories rows span both a Direct family (e.g. E04/E05,
#' "Corrections") and its M/L counterpart (M04/M05) makes the M/L recipe
#' itself a raw top-level candidate for a plain `cog_spending()` call --
#' following that hint would silently return intergovernmental dollars
#' under `expenditure_concept = "direct"` provenance.
#'
#' In all-categories mode (`all_categories = TRUE`) the inner sub-select is
#' scoped by `subtype_col`/`subtype_scope` -- the same allowlist
#' `.build_verb_sql()` applies as a WHERE predicate to make the summed
#' result a *concept* (see R/spending.R), not by `category`.
#' `.ALL_CATEGORIES` ("All Categories") is never itself a row in
#' `summary_categories.category`, so a category-keyed sub-select always
#' returns zero candidates and silently disables signposting.
#'
#' Scope is also by `category_type` ('expenditure' vs 'revenue', Issue #34)
#' to prevent cross-flow-family leakage: `cog_revenue(category =
#' "Corrections")` must not surface expenditure-only recipes (E04/E05)
#' merely because they share the same category name in summary_categories.
#' The type is derived from flow_prefixes: E/F/G -> 'expenditure', anything
#' else -> 'revenue'.
#'
#' @param con Active DuckDB connection.
#' @param category Category name, or `NULL`.
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
#' `.build_suggestions()`). Used to derive `category_type` (#34).
#' @param all_categories `TRUE` when the caller used `.ALL_CATEGORIES`.
#' @param subtype_col Name of the summary_categories subtype column to
#' scope by when `all_categories = TRUE`; ignored otherwise.
#' @param subtype_scope Character vector of subtype values to scope by
#' when `all_categories = TRUE`; ignored otherwise.
#' @return Character vector of recipe IDs (possibly empty).
#' @noRd
.query_candidate_recipes <- function(con, category, flow_prefixes,
all_categories = FALSE,
subtype_col = NULL,
subtype_scope = NULL) {
# Issue #34: derive category_type from flow_prefixes to prevent
# cross-flow-family leakage -- e.g. cog_revenue(category = "Corrections")
# must not surface expenditure-only recipes merely because they share the
# same category name in summary_categories.
category_type <- if (all(flow_prefixes %in% c("E", "F", "G"))) {
"expenditure"
} else {
"revenue"
}
candidate_scope_sql <- if (isTRUE(all_categories)) {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories
WHERE %s IN (%s) AND category_type = '%s'",
subtype_col, .sql_lit_chr(subtype_scope), category_type
)
} else {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories
WHERE category IN (%s) AND category_type = '%s'",
.sql_lit_chr(category), category_type
)
}
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE component_code IN (
%s
)
AND recipe_id NOT IN (
SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE LEFT(component_code, 1) IN ('M', 'L')
)",
candidate_scope_sql
))$recipe_id
}
#' Query gap-year coverage: which (recipe, year) pairs the recipe's own
#' generic join covers for this government, restricted to `gap_years`.
#'
#' This is Path 1 of a suggestion (unchanged): it finds recipes whose
#' component codes' generic join produces at least one row for this
#' government in each gap year -- i.e. the category returned nothing in
#' that year but a recipe would fill it.
#'
#' @param con Active DuckDB connection.
#' @param candidates Character vector of recipe IDs to check coverage for.
#' @param cohort The verb's cohort object (see `.make_cohort()`), rendered
#' into the govid predicate on the joined `long` scan via `.cohort_sql()`.
#' @param gap_years Integer vector of requested years absent from the
#' result.
#' @return Data frame with columns `recipe_id` (character) and `year`
#' (integer). Returns an empty data frame (`recipe_id = character(0)`,
#' `year = integer(0)`) when `gap_years` is empty, so callers can safely
#' reference `$recipe_id`.
#' @noRd
.query_covered_years <- function(con, candidates, cohort, gap_years) {
if (length(gap_years) == 0L) {
return(data.frame(recipe_id = character(0), year = integer(0)))
}
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.year "SELECT DISTINCT r.recipe_id, l.year
FROM long l FROM long l
JOIN harmonization_recipes r JOIN harmonization_recipes r
@@ -111,24 +317,30 @@
OR (r.gov_type_scope = 'state' AND l.type = 0) OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%s) WHERE r.recipe_id IN (%s)
AND l.canonical_govid IN (%s) AND %s
AND l.year IN (%s)", AND l.year IN (%s)",
.sql_lit_chr(candidates), .sql_lit_chr(govid), .sql_lit_chr(candidates), .cohort_sql(cohort, "l.canonical_govid"),
paste(gap_years, collapse = ",") paste(gap_years, collapse = ",")
)) ))
}
suggestions <- list() #' Query recipe metadata: labels and year spans for a set of candidate
for (rid in candidates) { #' recipes.
if (!rid %in% covered$recipe_id) next #'
m <- meta[meta$recipe_id == rid, ] #' @param con Active DuckDB connection.
suggestions[[length(suggestions) + 1L]] <- list( #' @param candidates Character vector of recipe IDs to look up.
recipe_id = rid, #' @return Tibble with columns `recipe_id`, `label`, `year_min` (int), and
label = m$label[[1]], #' `year_max` (int).
available_years = c(as.integer(m$year_min), as.integer(m$year_max)), #' @noRd
hint = sprintf("re-run with recipe = '%s'", rid) .query_recipe_meta <- function(con, candidates) {
) tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
} "SELECT recipe_id, any_value(label) AS label,
.attach_ig_counterparts(con, suggestions, flow_prefixes) MIN(year_min) AS year_min, MAX(year_max) AS year_max
FROM harmonization_recipes
WHERE recipe_id IN (%s)
GROUP BY recipe_id",
.sql_lit_chr(candidates)
)))
} }
#' Attach `ig_recipe_id` to each suggestion: the intergovernmental-expenditure #' Attach `ig_recipe_id` to each suggestion: the intergovernmental-expenditure
@@ -176,7 +388,7 @@
#' `R/basis.R`). This blocks a recipe surfaced through a mis-scoped #' `R/basis.R`). This blocks a recipe surfaced through a mis-scoped
#' category from ever reaching the M/L search, e.g. `cog_spending()`'s #' category from ever reaching the M/L search, e.g. `cog_spending()`'s
#' flow_prefixes are `c("E","F","G")`, which `ig_federal_b47_wide`'s own #' flow_prefixes are `c("E","F","G")`, which `ig_federal_b47_wide`'s own
#' `"B"` is not part of. #' "B" is not part of.
#' 2. `own_prefix %in% c("E","F","G")`: M/L only ever pairs with the #' 2. `own_prefix %in% c("E","F","G")`: M/L only ever pairs with the
#' DIRECT-expenditure family, never with revenue (`cog_revenue()`'s #' DIRECT-expenditure family, never with revenue (`cog_revenue()`'s
#' flow_prefixes already fold B/C/D in as ordinary revenue -- there is #' flow_prefixes already fold B/C/D in as ordinary revenue -- there is
@@ -184,10 +396,6 @@
#' adds one for spending) and never with ANOTHER M/L recipe (without #' adds one for spending) and never with ANOTHER M/L recipe (without
#' this check, `ige_local_m47_wide` would wrongly match sibling #' this check, `ige_local_m47_wide` would wrongly match sibling
#' `ige_state_l47_wide` on their shared {"47","94"} suffix set). #' `ige_state_l47_wide` on their shared {"47","94"} suffix set).
#' Condition 1 alone does not catch this: under `cog_revenue()`,
#' `ig_federal_b47_wide`'s own `"B"` IS inside revenue's own
#' `flow_prefixes`, so only this second, family-specific check blocks
#' the search.
#' @noRd #' @noRd
.attach_ig_counterparts <- function(con, suggestions, flow_prefixes) { .attach_ig_counterparts <- function(con, suggestions, flow_prefixes) {
if (length(suggestions) == 0L) return(suggestions) if (length(suggestions) == 0L) return(suggestions)
@@ -232,12 +440,25 @@
#' expressions. When a suggestion has an `ig_recipe_id`, one indented #' expressions. When a suggestion has an `ig_recipe_id`, one indented
#' continuation line is appended naming the intergovernmental counterpart #' continuation line is appended naming the intergovernmental counterpart
#' recipe (embedded `\n` renders as a hanging-indent continuation of the #' recipe (embedded `\n` renders as a hanging-indent continuation of the
#' same bullet under cli, not a new bullet). #' same bullet under cli, not a new bullet). Same treatment for
#' `suppressed_amount` (uscogdata#9): only present when dollars were
#' actually measured as excluded (an `empty_year` fire can carry them too --
#' see `.build_suggestions()` -- so this keys off the amount, not `trigger`).
#' @noRd #' @noRd
.inform_suggestions <- function(suggestions) { .inform_suggestions <- function(suggestions) {
bullets <- vapply(suggestions, function(s) { bullets <- vapply(suggestions, function(s) {
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id, bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
s$available_years[1], s$available_years[2], s$hint) s$available_years[1], s$available_years[2], s$hint)
# Only present when dollars were actually measured as excluded. An
# empty_year fire can carry them too -- the year had no rows AND the
# component was suppressed -- which is strictly more informative.
if (isTRUE(s$suppressed_amount > 0)) {
bullet <- paste0(bullet, sprintf(
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
if (!is.null(s$ig_recipe_id)) { if (!is.null(s$ig_recipe_id)) {
bullet <- paste0(bullet, sprintf( bullet <- paste0(bullet, sprintf(
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id)) "\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
@@ -245,7 +466,7 @@
bullet bullet
}, character(1)) }, character(1))
cli::cli_inform(c( cli::cli_inform(c(
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:", i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
stats::setNames(bullets, rep("*", length(bullets))) stats::setNames(bullets, rep("*", length(bullets)))
)) ))
} }
+117
View File
@@ -0,0 +1,117 @@
# R/suppression.R
# Split out of R/suggestions.R (2026-08-05) to keep files under the project's
# 400-line limit. Owns the second qualifying path for coverage signposting
# (uscogdata#9): measuring, per government, the component dollars the
# calling verb's own long view structurally excludes (aggregate-published,
# or absent from summary_categories). See R/suggestions.R for the
# orchestrator (`.build_suggestions()`) that calls this and the full
# uscogdata#9 background.
#' Measure, per (recipe, year), the component dollars this government holds
#' that the calling verb's own long view structurally excludes.
#'
#' This is the second qualifying path for a suggestion (uscogdata#9). The
#' first -- row absence -- only fires when a category returns NOTHING in a
#' requested year, which is how Corrections behaves in the wide era. Public
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
#' 2011, and absent from `summary_categories` entirely) are dropped. The
#' caller gets a plausible number a third too low, silently.
#'
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
#' view rather than restating its WHERE clause, so this stays correct if
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
#' anti-join is keyed on `item_code`, which is sound only because
#' harmonization never renames a recipe component -- asserted by the "no
#' recipe component is ever renamed by harmonization" test in
#' tests/testthat/test-recipes.R.
#'
#' Note what this deliberately does NOT count as suppressed: a component
#' excluded from the RESULT for scoping reasons -- because it belongs to a
#' different `category`, or because `expenditure_concept` narrowed the
#' subtypes -- is still present in the view, so it never fires. Suggesting a
#' recipe is a coverage fix, not a category redefinition.
#'
#' `flow_prefixes` (uscogdata#9 review, finding I1) restricts the measured
#' components to the CALLING VERB's own flow family (`c("E","F","G")` for
#' spending, `c("T","A","U","B","C","D")` for revenue). Without this, a
#' candidate recipe belonging to the OTHER flow family is always absent from
#' this verb's view (by construction -- `cog_revenue()`'s view never carries
#' an E-coded row) and so was always reported as "suppressed", fabricating a
#' dollar claim across flow families (`cog_revenue(category = "Corrections")`
#' claimed $3.63B excluded that `cog_spending()` reports and fully accounts
#' for). Filtering on `LEFT(r.component_code, 1)` also drops M/L-prefixed
#' components from measurement under `cog_spending()` (`flow_prefixes` never
#' includes "M"/"L") -- harmless today, because a recipe's own M/L components
#' (e.g. `corrections_ig_local_combined`'s M04/M05) are present in the view
#' in every year they exist and so never fired as suppressed anyway, but
#' worth recording since this filter is now the thing relied on to prevent
#' it.
#'
#' @param con Active DuckDB connection.
#' @param candidates Character vector of recipe ids to measure.
#' @param cohort The verb's cohort object (see `.make_cohort()`), rendered
#' into the govid predicate on both the outer scan and the restated
#' NOT EXISTS filter.
#' @param years Integer vector of requested years.
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
#' `.build_suggestions()`). Only recipe components whose first character is
#' in this set are measured.
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
#' nothing is suppressed.
#' @noRd
.suppressed_components <- function(con, candidates, cohort, years, long_view,
flow_prefixes) {
empty <- tibble::tibble(
recipe_id = character(0), year = numeric(0),
suppressed_amount = numeric(0), suppressed_codes = character(0)
)
if (length(candidates) == 0L) return(empty)
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
# never be quoted safely. It is always internally derived from a fixed
# view_base, so an off-allowlist value is a programming error, not input.
if (!long_view %in% c("spending_long", "spending_long_harmonized",
"revenue_long", "revenue_long_harmonized")) {
cli::cli_abort(
"Internal error: unexpected `long_view` {.val {long_view}}.",
class = "uscogdata_internal_error"
)
}
sql <- sprintf(
"SELECT r.recipe_id,
l.year,
SUM(l.amt) * 1000.0 AS suppressed_amount,
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
AS suppressed_codes
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%1$s)
AND %2$s
AND l.year IN (%3$s)
AND l.amt <> 0
AND LEFT(r.component_code, 1) IN (%5$s)
AND NOT EXISTS (
SELECT 1 FROM %4$s v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year
AND v.item_code = l.item_code
AND v.year IN (%3$s) -- restated: enables partition pruning (I3a)
AND %6$s -- restated: pushes the cohort filter (I3a)
)
GROUP BY 1, 2
ORDER BY 1, 2",
.sql_lit_chr(candidates), .cohort_sql(cohort, "l.canonical_govid"),
paste(as.integer(years), collapse = ","), long_view,
.sql_lit_chr(flow_prefixes), .cohort_sql(cohort, "v.canonical_govid")
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
+117 -2
View File
@@ -32,6 +32,117 @@
"45-ig_annotated_harmonized.sql" "45-ig_annotated_harmonized.sql"
) )
# The representation contract (cog_pipeline#64): two parquet tables that say
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
# on schema_version, because the sparsification that introduced them did not
# bump the version -- the pre-sparsification corpus this package shipped
# against until 2026-07-30 was already schema v6 and carried neither table.
# Keying off the version number would therefore register a view over a file
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
# check exists to tolerate.
.representation_view_files <- c(
"36-representation.sql" = "representation.parquet",
"37-code_set.sql" = "code_set.parquet"
)
# Cash and security holdings (uscogdata#25). 46- selects
# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and
# WITHOUT a schema_version bump -- so neither existing gate applies:
# .harmonization_view_files keys on schema_version, .representation_view_files
# on the presence of a FILE. Here the discriminator is a COLUMN on a table
# that exists either way. CREATE VIEW resolves its source schema eagerly, so
# on an older corpus 46- would fail at registration with "Binder Error:
# Referenced column balance_subtype not found" rather than at query time.
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
#' Does the mounted corpus's `summary_categories` carry `balance_subtype`?
#' Probed against the live connection rather than the manifest, because the
#' manifest describes files, not columns.
#' @noRd
.corpus_has_balance_subtype <- function(con) {
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM information_schema.columns
WHERE table_name = 'summary_categories'
AND column_name = 'balance_subtype'"
)$n
isTRUE(as.integer(n) > 0L)
}
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
#' works identically for a local fixture and a remote share.
#' @noRd
.corpus_has_table <- function(manifest, file) {
paths <- vapply(manifest$files$metadata %||% list(),
function(f) as.character(f$path %||% ""), character(1))
file %in% basename(paths)
}
#' Build the SQL path expression for the partitioned `long` table.
#'
#' DuckDB cannot expand a glob over generic HTTP: there is no directory
#' listing to expand against, and `allow_asterisks_in_http_paths` only
#' forwards the literal `**/*` as a filename, which 404s. Measured against
#' the published corpus on 2026-08-08, an explicit file list returns the
#' same 46,148,034 rows the (working) `hf://` glob does, and
#' `hive_partitioning = true` still recovers `year` from the paths.
#'
#' The manifest already enumerates every partition, so we build the list
#' from it. This is host-agnostic -- Nextcloud, HuggingFace and a local
#' fixture take the same path -- where an `hf://` URL would tie the reader
#' to one vendor's protocol and still need special-casing, since manifest
#' fetching goes through httr2, which cannot speak `hf://`.
#'
#' Falls back to the glob when the manifest carries no partition list: a
#' hand-built manifest in a test (see test-views.R) or a corpus predating
#' the field. Both are local, where globbing works.
#' @noRd
.long_files_sql <- function(url, manifest) {
parts <- manifest$files$long_partitions %||% list()
if (length(parts) == 0L) {
return(.sql_lit_chr(paste0(url, "data/long/**/*.parquet")))
}
paths <- vapply(parts, function(p) as.character(p$path), character(1))
paste0("[", .sql_lit_chr(paste0(url, paths)), "]")
}
#' Substitute the corpus-location tokens in a view's SQL text.
#'
#' One place knows the token vocabulary. `.register_views()` and the tests
#' that execute a view file directly both route through here. This exists
#' because four test sites had hand-rolled the `{url}` substitution -- one
#' of them commented as doing it "exactly as .register_views() does" -- and
#' every one of them broke the moment a second token was introduced.
#'
#' `{long_files}` must be substituted BEFORE `{url}`: it expands to a string
#' that itself contains the url, so the reverse order leaves the token in
#' place and DuckDB's parser fails on the brace.
#'
#' `manifest` defaults to empty, which routes `.long_files_sql()` to its glob
#' fallback -- correct for the local temp corpora the direct-execution tests
#' build.
#' @noRd
#' `fixed = TRUE` is load-bearing, not a style choice.
#'
#' In regex mode, `gsub()` interprets backslashes in the REPLACEMENT string as
#' escape sequences and silently drops them. A Windows corpus path is full of
#' them, so `C:\Users\RUNNER\AppData\...` was substituted in as
#' `C:UsersRUNNERAppData...` and every DuckDB read failed with "No files found
#' that match the pattern". `fixed = TRUE` treats pattern and replacement as
#' literal text, which is what a filesystem path needs.
#'
#' This is why the package could not read a LOCAL corpus on Windows at all --
#' including the test fixture, hence the entire suite, and any `cog_mirror()`
#' copy. Remote https URLs were unaffected, having no backslashes, which is
#' part of why it stayed hidden: the bug predates the `{long_files}` token and
#' lived in the original `{url}` substitution, unnoticed because nothing ever
#' ran on Windows until the mirror's check matrix existed.
#' @noRd
.render_view_sql <- function(sql, url, manifest = list()) {
sql <- gsub("{long_files}", .long_files_sql(url, manifest), sql, fixed = TRUE)
gsub("{url}", url, sql, fixed = TRUE)
}
#' Register DuckDB views from inst/sql/ SQL files #' Register DuckDB views from inst/sql/ SQL files
#' @noRd #' @noRd
.register_views <- function(con, url, manifest) { .register_views <- function(con, url, manifest) {
@@ -39,9 +150,13 @@
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE)) files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L)) schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
for (f in files) { for (f in files) {
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next base <- basename(f)
if (base %in% .harmonization_view_files && schema_version < 5L) next
if (base %in% names(.representation_view_files) &&
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next
sql <- paste(readLines(f, warn = FALSE), collapse = "\n") sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) sql <- .render_view_sql(sql, url, manifest)
DBI::dbExecute(con, sql) DBI::dbExecute(con, sql)
} }
} }
+261 -56
View File
@@ -1,78 +1,283 @@
# uscogdata # uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus. <!-- badges: start -->
[![R-CMD-check](https://github.com/civilytics/uscogdata/actions/workflows/R-CMD-check.yaml/badge.svg)](https://github.com/civilytics/uscogdata/actions/workflows/R-CMD-check.yaml)
[![r-universe](https://civilytics.r-universe.dev/badges/uscogdata)](https://civilytics.r-universe.dev/uscogdata)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE.md)
<!-- badges: end -->
Provides unit-level financial profiles, geographic rollups, and peer comparisons A curated R reader for the Civilytics US Census of Governments finance corpus —
with auditable provenance and built-in cross-vintage correctness. Reads the every dollar that US state, county, municipal and township governments reported
published corpus (Hive-partitioned parquet + manifest.json) directly from raising and spending, from **FY1967 to FY2024**, in one queryable place.
Nextcloud via DuckDB httpfs — no local bulk downloads required.
## Status The Census of Governments is the only nationwide source for local government
finance, and it is hard to use: item codes change meaning across vintages,
government identifiers were renumbered in 2017, and an absent value means
"published zero" in one era and "not reported" in the next. This package
handles each of those problems, and it tells you when it has — every result
carries provenance describing what was converted, what was aggregated, and
which known series breaks intersect your query.
Under active development (Phase 2 of the cog_pipeline project). See **Scope:** government types 0–3 (state, county, municipality, township).
`../cog_pipeline/docs/reader-specification.md` for the reader contract this 56 fiscal years, 46,148,034 rows, ~201 MB. There is no source data for FY1968
package implements. or FY1969. Special districts (type 4) and school districts (type 5) are
excluded pending validation.
## Installation ## Where the data comes from
The corpus is published and documented at the **[US Census of Governments
Finance API](https://pages.civilytics.org/cog-api/)**. Start there for how the
data was built, how the identifier and item-code reconciliation works, and what
the corpus does and does not cover.
- **[API documentation and walkthroughs](https://pages.civilytics.org/cog-api/)**
— reference, data dictionary, and worked examples such as the
[Southern states guide](https://pages.civilytics.org/cog-api/cog-api-south-guide.html)
- **[Live API](https://cog-api.civilytics.org/api/v1/)** — the same corpus over
HTTP, for Tableau, Python, or anything that isn't R
- **[Bulk corpus on Hugging Face](https://huggingface.co/datasets/civilytics/us-cog-finance)**
— CC-BY-4.0; the same parquet files this package reads
- **[Census Bureau source data](https://www.census.gov/programs-surveys/gov-finances.html)**
— the underlying public files
## Install
```r ```r
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata") install.packages("uscogdata",
repos = c("https://civilytics.r-universe.dev",
"https://cloud.r-project.org"))
``` ```
## Configuration Or from source:
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
## Direct vs Total spending
`cog_spending(..., expenditure_concept = c("direct", "total"))` controls
whose spending a result counts. `"direct"` (the default) is a government's
own current operations, capital outlay, and other direct spending. `"total"`
additionally adds in the intergovernmental legs — money it hands to other
governments to spend on its behalf — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
**Rule of thumb: any figure that spans more than one government uses
`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this
by refusing `expenditure_concept = "total"`. See
`vignette("total-spending", package = "uscogdata")` for the full
explanation with worked examples.
## Developer notes
### Testing
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
at this fixture, so the full test suite runs offline with no network
dependency:
```r ```r
devtools::test() # uses bundled fixture, no credentials required pak::pkg_install("git::https://gitea.civilytics.org/Civilytics/uscogdata.git")
``` ```
### Releasing against the live corpus ## Quickstart
Before cutting a release, run the test suite against the published corpus to No configuration, no credentials, no download. The package reads the published
catch any drift between the fixture and the real data: corpus over HTTPS by default.
```r ```r
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>") library(uscogdata)
devtools::test()
# Resolve a place name to a canonical government id
madison <- cog_gov_search(name = "Madison", state = "WI", type = 2)
madison$canonical_govid
#> [1] "552025209777"
# Police spending, inflation-adjusted and per capita
spend <- cog_spending(
madison$canonical_govid,
years = 2012:2022,
category = "Police",
per_capita = TRUE,
adjust_to_year = 2023
)
# What did that result do to the numbers, and what should you know about them?
cog_explain(spend)
``` ```
When the live-corpus run is clean, strip the fixture from the built package by `years` is required — there is no implicit full-history default.
adding this line to `.Rbuildignore`:
``` ## Two ways to read the corpus
^inst/extdata/fixture_corpus$
| | Remote (default) | Mirrored |
|---|---|---|
| Setup | none | `cog_mirror(dest)`, ~201 MB once |
| Disk used | **0 MB** — HTTP range requests only | ~201 MB |
| Opening a session | ~7.5 s | ~0.1 s |
| One government, one year | ~4 s | ~0.05 s |
| One government, full history | ~7 s | ~0.1 s |
| Later queries, same session | ~1.5 s | ~0.05 s |
| Good for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility |
**A local mirror is roughly 60–80x faster, and it is one function call.** That is
by far the largest difference any of these settings makes. If you are going to
ask more than a handful of questions, mirror first.
Measured 2026-08-10 on a 16-core Linux workstation against the published corpus
(schema v7, `pipeline_commit 3d28ddd`), fresh R session per arm. A one-off
question costs about **12 seconds end to end remotely and 0.15 seconds
mirrored**, session setup included.
Two things the per-query rows hide:
- **Opening the session is the single largest remote cost** — larger than any
one query. It fetches the manifest and registers 23 SQL views over HTTPS, and
it lands on your first query, not on `library(uscogdata)`.
- **The cost is network round-trips, not scanning.** A repeat query against
partitions this session has already touched is ~1.5 s rather than ~4 s, and a
full-history query costs ~7 s whether it runs first or last. What you are
paying for is reaching each of the 56 yearly files over HTTPS the first time.
Nothing is written to disk in remote mode: DuckDB fetches the parquet footer,
works out which row groups it needs, and reads only those. Nothing is cached
between sessions either, so every query goes back to the network — and a session
that issues many remote queries in quick succession can be rate-limited by the
host (`HTTP Error: ... 429`). Both are further reasons to mirror for real work.
The default points at a public HuggingFace mirror of the corpus. If you would
rather not depend on a third party — for reproducibility, for an air-gapped
environment, or on principle — **the escape hatch is one function call**:
```r
cog_mirror("~/cog-corpus")
Sys.setenv(USCOGDATA_URL = "~/cog-corpus/")
``` ```
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when After that, nothing in your analysis touches an external service.
the bundled fixture is absent, so no test code changes are needed for the
release run or after stripping the fixture. ### Configuration
- `USCOGDATA_URL` — corpus root: an HTTPS URL or a local path, **trailing slash required**
- `USCOGDATA_CACHE_DIR` — where the manifest is cached (default: user cache dir)
- `USCOGDATA_MANIFEST_TTL_SECS` — manifest re-fetch interval (default 3600)
- `USCOGDATA_DUCKDB_THREADS` — cap DuckDB's thread count (default: every visible core)
- `USCOGDATA_DUCKDB_MEMORY_LIMIT` — cap DuckDB's memory, e.g. `"4GB"` (default: DuckDB's own)
Each also has an `options()` spelling — `uscogdata.url`, `uscogdata.duckdb_threads`,
and so on — and the environment variable wins where both are set.
The two DuckDB caps exist for **servers**, not laptops. Unset, DuckDB claims every
core it can see, which is right for one interactive session on your own machine and
wrong when several readers share a box: each claims the whole machine and they fight.
Capping costs roughly 5% on a single query and is worth it anywhere the process is
sharing hardware.
## Amounts are in full US dollars
Every amount column this package returns — `amt_nominal`, `amt_real`,
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
The raw Census source files report **thousands of dollars**, and the corpus's
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
you never have to. The conversion is recorded in every result:
```r
attr(spend, "provenance")$transformations$units_conversion
#> $applied TRUE
#> $source_unit "$1,000s (raw Census)"
#> $target_unit "$USD"
#> $multiplier 1000
```
**Do not multiply again.** If you have read elsewhere that COG amounts are in
`$1,000s` — which is true of the raw Census files and of the corpus's own `amt`
column — that rule does not apply to anything a `cog_*()` verb hands you.
Applying it twice overstates every figure by 1000x, and the result looks
plausible rather than obviously wrong.
## Concepts worth understanding before you publish a number
### Primary vs Direct vs Total spending
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
controls *whose* spending a result counts. Concepts are defined as sets of the
crosswalk's `spend_subtype` values, never item-code first letters — the letter
`Y` alone spans revenue, expenditure and balance codes.
- **`"primary"`** (default) — the government's own service provision: current
operations, capital outlay, assistance payments.
- **`"direct"`** — Census's published Direct Expenditure: `primary` plus
interest on debt and insurance trust benefits (e.g. pensions).
- **`"total"`** — adds the intergovernmental leg, money handed to other
governments to spend. Meaningful for one government's own budget over time,
but it double-counts when summed across governments: a state's payment to a
county is the same dollar the county reports as its own direct spending.
**Rule of thumb: any figure spanning more than one government uses `primary`
or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce that
by refusing `"total"` outright. Worked examples in
`vignette("total-spending", package = "uscogdata")`.
### General vs Total revenue
`cog_revenue(..., revenue_concept = c("general", "total"))`:
- **`"general"`** (default) — Census General Revenue: own-source taxes,
charges and miscellaneous, plus federal, state and local aid.
- **`"total"`** — General plus utility revenue (`A91`–`A94`), liquor store
revenue (`A90`), and insurance trust revenue.
Census defines these by its own identity:
```
Total Revenue = General + Utility + Liquor Store + Insurance Trust
```
Two things to know before switching to `"total"`. **Utility revenue is large
for cities** — measured on the bundled fixture, utility plus liquor store is
15.9% of city revenue, against 1.2% for states and 1.7% for counties. And the
**employee-retirement (`X`) codes stop at FY2016**, when those systems moved to
the separate Annual Survey of Public Pensions, so a `"total"` series steps down
at the FY2016/FY2017 boundary for reasons of collection scope, not revenue
(series breaks `SB197`–`SB209`).
### Reporting coverage: the Census is only sometimes a census
**The Census of Governments is a complete enumeration only in years ending in
2 and 7.** Every other year is a sample, and the sample varies enormously —
measured on the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019.
A statewide total resting on a fifth of the universe looks exactly like one
resting on all of it, so every multi-government result now says which it is:
```r
attr(rollup, "provenance")$coverage # per-year n_units_reporting, is_census_year
```
`cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` take a
`coverage` argument — `"all"` (default), `"census"` (census years only), or
`"consistent"` (only units reporting in every requested year, a balanced
panel).
`n_units_reporting` is **category-conditional**, and it is not a response rate. A government that was surveyed and genuinely spends
nothing in the requested category is indistinguishable from one never surveyed.
### Absent cells mean two different things
Before FY2012, an absent cell means Census published `$0`. From FY2012 on, it
means not reported. `cog_spending(..., complete = TRUE)` fills the requested
grid and labels every row with which it is, via `value_source`:
| `value_source` | meaning | `amt_nominal` |
|---|---|---|
| `reported` | the corpus carries this cell | as published |
| `census_zero` | dense-source year (≤ FY2011), absent — Census published `$0` | `0` |
| `not_reported` | sparse-source year (≥ FY2012), absent — unknown | `NA` |
That `NA` is deliberate. Filling a modern absence with `0` would invent data.
### Series breaks surface on their own
Catalogued breaks that intersect your query appear in provenance whether or not
you went looking for them — `series_break_refs` for breaks in a specific item code, and
`corpus_break_refs` for caveats about the corpus as a whole (dollar precision
across the 1976/1977 boundary, the FY2017 identifier change, the FY2012
dense→sparse representation change). `cog_explain()` prints both.
## How to cite
```r
citation("uscogdata")
```
The corpus itself is published under CC-BY-4.0. Cite it as:
> Civilytics Consulting. US Census of Governments finance corpus.
> https://huggingface.co/datasets/civilytics/us-cog-finance
## Contributing
Development happens on [Gitea](https://gitea.civilytics.org/Civilytics/uscogdata);
[GitHub](https://github.com/civilytics/uscogdata) is a mirror that accepts
issues and pull requests. See [CONTRIBUTING.md](CONTRIBUTING.md) for how a
patch gets from there to here.
## License
MIT © Civilytics Consulting LLC. See [LICENSE.md](LICENSE.md).
+27 -5
View File
@@ -1,19 +1,41 @@
url: ~ url: https://civilytics.r-universe.dev/uscogdata
template: template:
bootstrap: 5 bootstrap: 5
reference: reference:
- title: Financial data
desc: Spending, revenue and balance-sheet holdings for one or more governments.
contents:
- cog_spending
- cog_revenue
- cog_balances
- title: Search & basket - title: Search & basket
desc: Resolve place names into canonical govids. desc: Resolve place names into canonical govids.
contents: contents:
- cog_gov_search - cog_gov_search
- cog_basket_resolution - cog_basket_resolution
- cog_basket_unresolved - cog_basket_unresolved
- title: Session - title: Comparison & aggregation
desc: Peer cohorts and geographic aggregates.
contents: contents:
- has_keyword("internal") - cog_find_peers
- cog_peer_compare
- cog_geographic_rollup
- title: Corpus metadata
desc: >
What the corpus contains, where a given result came from, and how to
hold a local copy of it.
contents:
- cog_categories
- cog_recipes
- cog_manifest
- cog_explain
- cog_mirror
articles: articles:
- title: Getting started - title: Concepts
navbar: ~ navbar: ~
contents: [] contents:
- total-spending
- population-denominators
+38 -34
View File
@@ -11,12 +11,14 @@
# Each partition is a full year (all states/govs) as published, so # Each partition is a full year (all states/govs) as published, so
# Broward County FL and every other previously-pinned government stay # Broward County FL and every other previously-pinned government stay
# covered without any per-gov slicing logic. # covered without any per-gov slicing logic.
# 2. Copies the full canonical_fips_xwalk.parquet, canonical_alias.parquet, # 2. Copies every metadata parquet the publish tree ships (see
# summary_categories.parquet, harmonization_map.parquet, # .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
# harmonization_recipes.parquet, and series_breaks.parquet metadata # registries, not partitioned by year, so the fixture ships the complete
# tables as-is (these are small cross-vintage registries, not # tables rather than a year-scoped subset. representation.parquet and
# partitioned by year, so the fixture ships the complete tables rather # code_set.parquet are what make the sparse wide era interpretable --
# than a year-scoped subset). # absence means "$0" in a dense_source year and "not reported" in a
# sparse_source one -- so a fixture without them cannot represent the
# published corpus.
# 3. Resyncs the four reference docs (data_dictionary.md, # 3. Resyncs the four reference docs (data_dictionary.md,
# reader-specification.md, README.md, series_breaks.md) from the # reader-specification.md, README.md, series_breaks.md) from the
# publish tree's docs/. # publish tree's docs/.
@@ -38,6 +40,22 @@
# source("data-raw/regenerate_fixture_corpus.R") # source("data-raw/regenerate_fixture_corpus.R")
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache") # regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
# Every metadata parquet the publish tree ships, in the order they appear in
# the corpus manifest. Single source of truth for both the copy step and the
# fixture manifest, so the two can never drift apart.
.FIXTURE_METADATA_FILES <- c(
"canonical_alias.parquet",
"canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet",
"code_set.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"lineage_events.parquet",
"representation.parquet",
"series_breaks.parquet",
"summary_categories.parquet"
)
regenerate_fixture_corpus <- function( regenerate_fixture_corpus <- function(
publish_cache_dir = file.path( publish_cache_dir = file.path(
"..", "cog_pipeline", "_targets", "publish_cache" "..", "cog_pipeline", "_targets", "publish_cache"
@@ -100,20 +118,11 @@ regenerate_fixture_corpus <- function(
invisible(NULL) invisible(NULL)
} }
# Copy the full (not year-scoped) canonical_fips_xwalk, canonical_alias, # Copy the full (not year-scoped) metadata tables listed in
# summary_categories, and (schema v5+) harmonization_map/ # .FIXTURE_METADATA_FILES.
# harmonization_recipes/series_breaks parquet tables.
#' @noRd #' @noRd
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) { .copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
files <- c( for (f in .FIXTURE_METADATA_FILES) {
"canonical_fips_xwalk.parquet",
"canonical_alias.parquet",
"summary_categories.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"series_breaks.parquet"
)
for (f in files) {
src <- file.path(publish_cache_dir, "data", f) src <- file.path(publish_cache_dir, "data", f)
dst <- file.path(fixture_dir, "data", f) dst <- file.path(fixture_dir, "data", f)
if (!file.exists(src)) { if (!file.exists(src)) {
@@ -179,15 +188,7 @@ regenerate_fixture_corpus <- function(
) )
}) })
metadata_files <- c( metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
"canonical_alias.parquet",
"canonical_fips_xwalk.parquet",
"summary_categories.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"series_breaks.parquet"
)
metadata <- lapply(metadata_files, function(f) {
rel <- file.path("data", f) rel <- file.path("data", f)
path <- file.path(fixture_dir, rel) path <- file.path(fixture_dir, rel)
list( list(
@@ -203,13 +204,16 @@ regenerate_fixture_corpus <- function(
pipeline_commit = source_manifest$pipeline_commit, pipeline_commit = source_manifest$pipeline_commit,
fixture_note = paste( fixture_note = paste(
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full", "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
"corpus available via USCOGDATA_URL. Regenerated for Phase R2", "corpus available via USCOGDATA_URL. Regenerated from the sparsified",
"(schema_version 5, harmonization_map/harmonization_recipes/", "schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
"series_breaks parquet tables added). 2011/2012 straddle the", "zeros, so FY2011 absence means Census published $0 while FY2012+",
"wide-aggregate -> modern-leaf format boundary exercised by basis=", "absence means not reported. representation.parquet and",
"\"harmonized\" and recipe= queries; 2019/2020 retain the prior", "code_set.parquet carry that rule and ship in full, as do every other",
"per-capita/CPI regression anchors. Full canonical_fips_xwalk master", "metadata table in the publish tree. 2011/2012 straddle both the",
"and canonical_alias lookup table included via", "wide-aggregate -> modern-leaf format boundary (exercised by",
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
"representation boundary (SB194); 2019/2020 retain the prior",
"per-capita/CPI regression anchors. Regenerated via",
"data-raw/regenerate_fixture_corpus.R." "data-raw/regenerate_fixture_corpus.R."
), ),
data_vintage = source_manifest$data_vintage, data_vintage = source_manifest$data_vintage,
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+30 -10
View File
@@ -1,8 +1,8 @@
{ {
"schema_version": 6, "schema_version": 6,
"built_at": "2026-07-27T13:04:05Z", "built_at": "2026-08-03T16:51:32Z",
"pipeline_commit": "6098baf", "pipeline_commit": "e7394a4",
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.", "fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
"data_vintage": { "data_vintage": {
"source_vintages": { "source_vintages": {
"2012": "10162019", "2012": "10162019",
@@ -36,9 +36,9 @@
{ {
"year": 2011, "year": 2011,
"path": "data/long/year=2011/part-0.parquet", "path": "data/long/year=2011/part-0.parquet",
"sha256": "84302ab364dc9fc3b3fbbc3c3f8b826e3508b4d73ff7c42d094d3863cd1e37b5", "sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
"row_count": 2864212, "row_count": 496004,
"size_bytes": 3845911 "size_bytes": 2202455
}, },
{ {
"year": 2012, "year": 2012,
@@ -74,9 +74,14 @@
"description": "canonical_fips_xwalk.parquet" "description": "canonical_fips_xwalk.parquet"
}, },
{ {
"path": "data/summary_categories.parquet", "path": "data/census_collection_coverage.parquet",
"sha256": "0985b607f3f35a8dff62c0561261ab6922423b81d11c07b03bcb3e3461f85e33", "sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
"description": "summary_categories.parquet" "description": "census_collection_coverage.parquet"
},
{
"path": "data/code_set.parquet",
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
"description": "code_set.parquet"
}, },
{ {
"path": "data/harmonization_map.parquet", "path": "data/harmonization_map.parquet",
@@ -88,10 +93,25 @@
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c", "sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
"description": "harmonization_recipes.parquet" "description": "harmonization_recipes.parquet"
}, },
{
"path": "data/lineage_events.parquet",
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
"description": "lineage_events.parquet"
},
{
"path": "data/representation.parquet",
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
"description": "representation.parquet"
},
{ {
"path": "data/series_breaks.parquet", "path": "data/series_breaks.parquet",
"sha256": "b0b6794b6887a4f300079adfa10029c2a77109faa4952fbff1c5a270793cc02b", "sha256": "731998516cd802f63fcf7fb66053c7a62b7be955ab0794cad4a4979cb7628b87",
"description": "series_breaks.parquet" "description": "series_breaks.parquet"
},
{
"path": "data/summary_categories.parquet",
"sha256": "e3b0efa00ce713b8f45829b89cfde24b55333f26101f0495df82d85997d18d8e",
"description": "summary_categories.parquet"
} }
] ]
}, },
+77 -6
View File
@@ -14,25 +14,96 @@
"basis_note": { "type": ["string", "null"] }, "basis_note": { "type": ["string", "null"] },
"expenditure_concept": { "expenditure_concept": {
"type": "string", "type": "string",
"enum": ["direct", "total"], "enum": ["primary", "direct", "total"],
"description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments." "description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
}, },
"expenditure_concept_note": { "expenditure_concept_note": {
"type": ["string", "null"], "type": ["string", "null"],
"description": "How the intergovernmental leg was assembled; null for 'direct'." "description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
}, },
"expenditure_concept_direct_suppressed": { "expenditure_concept_direct_suppressed": {
"type": "boolean", "type": ["boolean", "null"],
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any." "description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. null (NA) when expenditure_concept = 'total' AND category = 'All Categories': the detector keys on per-category rows, which that mode collapses, so suppression cannot be computed -- see `expenditure_concept_note`. See the affected rows' `notes` for the recovering recipe, if any."
},
"revenue_concept": {
"type": "string",
"enum": ["general", "total"],
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
}, },
"harmonization": { "type": "object" }, "harmonization": { "type": "object" },
"recipe": { "type": ["object", "null"] }, "recipe": { "type": ["object", "null"] },
"suggestions": { "type": "array" }, "suggestions": {
"type": "array",
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
"items": {
"type": "object",
"required": ["recipe_id", "label", "available_years", "hint", "ig_recipe_id",
"trigger", "suppressed_amount", "suppressed_years", "suppressed_codes"],
"properties": {
"recipe_id": { "type": "string" },
"label": { "type": "string" },
"available_years": {
"type": "array",
"items": { "type": "integer" },
"description": "[year_min, year_max] of the recipe's component coverage."
},
"hint": { "type": "string" },
"ig_recipe_id": {
"type": ["string", "null"],
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
},
"trigger": {
"type": "string",
"enum": ["empty_year", "suppressed_component"],
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars this government reports in the requested years that the verb's underlying long view structurally excludes -- aggregate-published, carrying no harmonized code, or absent from summary_categories. This is NOT the same thing as 'excluded from the result': a component present in the view under a different category (a scoping choice, e.g. a different `category` or a narrower `expenditure_concept`) contributes 0 and never fires. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way, using the same underlying-view measurement, and can be 0 even on an 'empty_year' fire."
},
"suppressed_amount": {
"type": "number",
"description": "Full US dollars this government reports, in the recipe's component codes, in the requested years, that the verb's underlying long view structurally excludes (aggregate-published, carrying no harmonized code, or absent from summary_categories) -- summed across those years. This is NOT the same quantity as 'what the result excludes': a component present in the view under a different category or a narrower `expenditure_concept` is scoped out on purpose, counts as 0 here, and is not suppression. 0 does not always mean full coverage -- see 'trigger' and 'empty_year'. May be negative where Census publishes a negative `amt` for the excluded rows."
},
"suppressed_years": {
"type": "array",
"items": { "type": "integer" },
"description": "The requested years contributing to suppressed_amount."
},
"suppressed_codes": {
"type": "array",
"items": { "type": "string" },
"description": "The excluded component item codes, sorted."
}
}
}
},
"scope": { "type": "object" }, "scope": { "type": "object" },
"codes_summed": { "type": "object" }, "codes_summed": { "type": "object" },
"aggregate_fallback": { "type": ["object", "null"] }, "aggregate_fallback": { "type": ["object", "null"] },
"transformations":{ "type": "object" }, "transformations":{ "type": "object" },
"series_break_refs": { "type": "array", "items": { "type": "string" } }, "series_break_refs": { "type": "array", "items": { "type": "string" } },
"completion": {
"type": "object",
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
"properties": {
"applied": { "type": "boolean" },
"rows_filled": { "type": "integer" },
"absence_means": { "type": "object" }
}
},
"corpus_break_refs": {
"type": "array",
"items": { "type": "string" },
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
},
"balance_caveats": {
"type": ["object", "null"],
"description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.",
"properties": {
"not_gaap": { "type": "boolean" },
"not_gaap_note": { "type": "string" },
"coverage_window": { "type": "object" },
"truncated": { "type": "array", "items": { "type": "string" } }
}
},
"manifest": { "type": "object" }, "manifest": { "type": "object" },
"sql_query": { "type": "string" } "sql_query": { "type": "string" }
} }
+5 -1
View File
@@ -1,3 +1,7 @@
CREATE OR REPLACE VIEW long AS CREATE OR REPLACE VIEW long AS
SELECT * SELECT *
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true); -- {long_files} carries its own quoting: a bracketed list of every partition
-- the manifest enumerates, or a single quoted glob on fallback. Do NOT wrap
-- it in quotes. See .long_files_sql() in R/views.R for why a glob alone
-- cannot work over HTTP.
FROM read_parquet({long_files}, hive_partitioning = true);
+7
View File
@@ -0,0 +1,7 @@
-- Category crosswalk. Numbered 11 (not with the other reference tables at
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
-- already exist when they register.
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+18 -1
View File
@@ -1,5 +1,22 @@
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
-- (summary_categories.category_type = 'expenditure'), never by item-code
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
-- every non-intergovernmental expenditure subtype: operations, capital,
-- assistance, interest, insurance_benefits.
--
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
-- 24-ig_long.sql).
CREATE OR REPLACE VIEW spending_long AS CREATE OR REPLACE VIEW spending_long AS
SELECT * SELECT *
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G') WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
)
AND NOT is_aggregate; AND NOT is_aggregate;
+14 -1
View File
@@ -1,5 +1,18 @@
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
--
-- Carries EVERY revenue subtype. Which of Census's two published concepts a
-- query actually returns is decided per revenue_concept in R
-- (.verb_spendrev), exactly as expenditure_concept narrows spending_long:
-- general = own_source + federal + state + local_aid (the default)
-- total = general + utility + liquor_store + insurance_trust
-- Census defines the first by subtracting the other three from the second
-- (manual section 4.3), so both concepts need all four families present here.
CREATE OR REPLACE VIEW revenue_long AS CREATE OR REPLACE VIEW revenue_long AS
SELECT * SELECT *
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D') WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
)
AND NOT is_aggregate; AND NOT is_aggregate;
+11 -1
View File
@@ -1,6 +1,16 @@
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
-- classification, applied to harmonized_code (the code the row is folded
-- onto) rather than the published item_code. Safe because the harmonized
-- space is leaf-only and every harmonized_code in the corpus is a
-- summary_categories member (verified at fixture regen; a code the
-- crosswalk cannot classify would be silently dropped here).
CREATE OR REPLACE VIEW spending_long_harmonized AS CREATE OR REPLACE VIEW spending_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code) SELECT * REPLACE (harmonized_code AS item_code)
FROM long FROM long
WHERE NOT is_aggregate WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G'); AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
);
+7 -1
View File
@@ -1,6 +1,12 @@
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
-- classification (every revenue subtype; the concept narrows in R), applied
-- to harmonized_code rather than the published item_code.
CREATE OR REPLACE VIEW revenue_long_harmonized AS CREATE OR REPLACE VIEW revenue_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code) SELECT * REPLACE (harmonized_code AS item_code)
FROM long FROM long
WHERE NOT is_aggregate WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D'); AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
);
+12 -5
View File
@@ -1,4 +1,6 @@
-- Intergovernmental expenditure rows (M = to local govts, L = to state govts). -- Intergovernmental expenditure rows: crosswalk spend_subtype =
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
-- = state payments to school systems -- uscogdata#11, finding F-017).
-- --
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the -- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published -- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
@@ -9,10 +11,15 @@
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument -- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
-- the pipeline's recipe joins use. -- the pipeline's recipe joins use.
-- --
-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up -- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
-- the L-NN codes, so including it would double-count. -- rolls up the L-NN codes, so including it would double-count. The crosswalk
-- deliberately carries no `--` family-total codes, so membership excludes it
-- (guarded by "the IG leg never includes the L-- family total" in
-- tests/testthat/test-expenditure-concept.R).
CREATE OR REPLACE VIEW ig_long AS CREATE OR REPLACE VIEW ig_long AS
SELECT * SELECT *
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L') WHERE item_code IN (
AND item_code NOT LIKE '%--'; SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+9 -2
View File
@@ -8,8 +8,15 @@
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real -- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+) -- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
-- while never dropping a row. -- while never dropping a row.
--
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
-- rather than the COALESCEd code: every IG harmonization target (M36) is
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
-- the column that exists on every row.
CREATE OR REPLACE VIEW ig_long_harmonized AS CREATE OR REPLACE VIEW ig_long_harmonized AS
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code) SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L') WHERE item_code IN (
AND item_code NOT LIKE '%--'; SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+22
View File
@@ -0,0 +1,22 @@
-- Cash and security holdings, classified by crosswalk MEMBERSHIP on
-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot
-- do this job -- the X and Y families each span revenue, expenditure AND
-- balance).
--
-- These rows are STOCKS: a balance at a point in time, not a flow over a
-- fiscal year. Summing a stock with a flow is meaningless, which is why they
-- live behind a third view rather than as a subtype of either money view, and
-- why neither spending_long nor revenue_long can reach them.
--
-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era
-- aggregate-only holdings codes (X40/X41) are deliberately outside this view;
-- they are reachable only through the recipe path, which bypasses this filter
-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2).
CREATE OR REPLACE VIEW balance_long AS
SELECT *
FROM long
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'balance'
)
AND NOT is_aggregate;
-3
View File
@@ -1,3 +0,0 @@
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW representation AS
SELECT *
FROM read_parquet('{url}data/representation.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW code_set AS
SELECT *
FROM read_parquet('{url}data/code_set.parquet');
+16
View File
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW balance_annotated AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.balance_subtype
FROM balance_long s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
+122
View File
@@ -0,0 +1,122 @@
% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/balances.R
\name{cog_balances}
\alias{cog_balances}
\title{Cash and security holdings for one or more governments}
\usage{
cog_balances(
govid = NULL,
years,
category = NULL,
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
state = NULL,
type = NULL,
limit = NULL,
offset = NULL
)
}
\arguments{
\item{govid}{Canonical govid(s): a character vector, or a data frame with a
`canonical_govid` column (e.g. from [cog_gov_search()]). `NULL` to name
the cohort by `state`/`type` instead.}
\item{years}{Integer vector of fiscal years.}
\item{category}{Optional character vector of categories to keep. One of
`"Fund Balances"`, `"Insurance Trust Balances"`,
`"Retirement System Holdings"`. There is deliberately no `subtype`
argument: for holdings, `category` is a strict coarsening of
`balance_subtype` (unlike the money verbs, where the two axes cross), so
every combination would be either redundant or empty.
`category = "Fund Balances"` is exactly the `general` family
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
one `dplyr::filter()` away. The reserved pseudo-category
`"All Categories"` (see [cog_spending()]) is **not** supported here and
errors with class `uscogdata_all_categories_unsupported`: it sums a
concept's subtype scope, and holdings are a stock with no concept
vocabulary to sum across. Omit `category` to get every category broken
out instead.}
\item{per_capita}{Divide holdings by population. Note this is a **stock per
resident** (reserves per person), which is *not* comparable to
[cog_spending()]'s per-capita figures -- those are a flow per person.}
\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).}
\item{basis}{Accepted for uniformity with the money verbs, but currently a
**no-op**: `harmonization_map` carries no balance-code rows, so harmonized
and raw space are identical for holdings. Reported in
`provenance$basis_note`.}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]).
`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
wide era to the modern one.}
\item{state, type}{Name the cohort by predicate instead of by id: `state` is
a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
`"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
the same vocabulary, and the same internal coercion, as
[cog_gov_search()]. Both default to `NULL`.
The cohort is then expressed as a subquery against `canonical_fips_xwalk`
inside each statement rather than round-tripped through R as a literal id
list. For a fleet-scale cohort that is the difference between a
301,591-character `IN` list re-parsed in 5--8 statements per call and a
constant-size predicate: measured at **94 ms versus 449 ms** for the same
FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
7% of the no-filter floor.
Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
governments in `govid` that also match the predicate -- rather than one
silently taking precedence. Naming no cohort at all (`govid`, `state` and
`type` all `NULL`) aborts with class `uscogdata_no_cohort`.
When the cohort is named by predicate, `provenance$scope$govids_found`
and `govids_missing` are empty -- there is no id list to report against --
and `provenance$scope$cohort` carries `state`, `type` and
`n_governments` instead. A `govid`-named cohort reports exactly as before.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
rather than applied after materializing every row. `NULL` (default)
returns everything. Cannot be combined with `recipe` -- see `offset` and
`total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`balance_subtype`, `category`, `amt_nominal`, `codes_included`,
`aggregate_fallback`, plus optional `amt_per_capita_nominal` and
`pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
`adjust_to_year` is set), and optional `amt_per_capita_real` (only when
**both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
nominal per-capita column to deflate otherwise). Amounts are full US
dollars.
Carries a `provenance` attribute matching
`inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
`not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
every balance subtype in the mounted corpus, not only the observed ones)
and `truncated` (the observed subtypes whose coverage falls short of the
requested years). `expenditure_concept`/`revenue_concept` are `NA` --
holdings are a stock, not a flow, so neither concept vocabulary applies.
When `limit` is set, also carries a `total_rows` attribute: the full
unpaginated row count, computed by the same query (`COUNT(*) OVER()`)
rather than a second scan.
}
\description{
Returns Census cash-and-security holdings (`category_type = "balance"`):
fund balances, retirement system holdings and insurance trust balances.
}
\section{Holdings are not GAAP fund balance}{
Census holdings are **gross** -- no liabilities are netted -- so a reserve
ratio built from them overstates what is actually available. They are not
comparable to a GAAP fund balance from an ACFR.
}
+15 -4
View File
@@ -7,8 +7,8 @@
cog_categories(type = NULL, pattern = NULL) cog_categories(type = NULL, pattern = NULL)
} }
\arguments{ \arguments{
\item{type}{Either `NULL` (default, return both spending and revenue \item{type}{Either `NULL` (default, every row: expenditure, revenue and
rows), `"spending"`, or `"revenue"`.} balance), `"spending"`, `"revenue"`, or `"balance"`.}
\item{pattern}{Optional regex matched case-insensitively against the \item{pattern}{Optional regex matched case-insensitively against the
`category` column (e.g. `"Police"` or `"Tax"`).} `category` column (e.g. `"Police"` or `"Tax"`).}
@@ -16,13 +16,24 @@ rows), `"spending"`, or `"revenue"`.}
\value{ \value{
Tibble with columns `category`, `category_type`, `subtype`, Tibble with columns `category`, `category_type`, `subtype`,
`n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
`category_type`, `category`, `subtype`. `category_type`, `category`, `subtype`. Includes one row per flow for the
reserved pseudo-category `"All Categories"`, which carries `NA` for
`subtype`, `n_codes` and `item_codes` because it is a query mode rather
than a crosswalk entry — see [cog_spending()]'s `category` argument.
} }
\description{ \description{
Returns the category taxonomy exposed by the corpus's Returns the category taxonomy exposed by the corpus's
`summary_categories` view, grouped to one row per `summary_categories` view, grouped to one row per
`(category, subtype)` pair. Use this to discover valid `category` `(category, subtype)` pair. Use this to discover valid `category`
values for [cog_spending()] / [cog_revenue()] / values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
[cog_geographic_rollup()] and to audit which Census item codes feed [cog_geographic_rollup()] and to audit which Census item codes feed
each category. each category.
} }
\details{
`subtype` COALESCEs the crosswalk's three subtype columns, so it carries
`spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
`balance_subtype` on balance rows. Note that [cog_balances()] itself takes
no `subtype` argument — for holdings, `category` is a strict coarsening of
`balance_subtype` — but the value is surfaced here because it is the
discovery surface downstream consumers build their vocabulary from.
}
+30
View File
@@ -21,3 +21,33 @@ Prints the structured provenance attached to a tibble returned by any
`cog_*` verb, or returns it as a list for downstream use (MCP tools, `cog_*` verb, or returns it as a list for downstream use (MCP tools,
dashboards, JSON export). dashboards, JSON export).
} }
\section{Two kinds of series break}{
Catalogued breaks reach you without being asked for, in two disjoint
fields, because a caveat about one series and a caveat about the whole
corpus are different claims:
* **`series_break_refs`** — breaks matched against the item codes actually
present in this result. A break in one code you queried.
* **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
which are statements about the corpus rather than about any one code:
dollar precision across the 1976/1977 boundary (`SB085`), imputation
exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
representation change (`SB194`), and the FY2017 government-identifier
change (`SB086`). These are selected on the break-year window alone.
`SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
crosses the boundary where an absent cell stops meaning "Census published
$0" and starts meaning "not reported".
}
\section{Other provenance blocks}{
`transformations$units_conversion` records the `$1,000s`-to-dollars
multiply that every amount column has already had applied.
`transformations$per_capita` records the population denominator and its
year range. `coverage` and `coverage_mode` appear on multi-government
results (see [cog_geographic_rollup()]). `completion` appears when
`complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
}
+10 -1
View File
@@ -11,7 +11,8 @@ cog_find_peers(
same_state = FALSE, same_state = FALSE,
pop_range = c(0.7, 1.3), pop_range = c(0.7, 1.3),
is_ratio = TRUE, is_ratio = TRUE,
max_peers = 10L max_peers = 10L,
coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
@@ -34,6 +35,14 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
`pop_range` is interpreted as absolute population counts.} `pop_range` is interpreted as absolute population counts.}
\item{max_peers}{Integer cap on the number of peers returned.} \item{max_peers}{Integer cap on the number of peers returned.}
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
most recent census year with an observed population, so a cohort is not
built from a sample year in which most of the candidate universe is
absent. `"consistent"` needs a year range, which cohort selection does not
have, so it selects like `"all"` and is carried on the result as
`attr(x, "coverage")` for [cog_peer_compare()].}
} }
\value{ \value{
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`, Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
+53 -8
View File
@@ -10,7 +10,8 @@ cog_geographic_rollup(
years, years,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL, adjust_to_year = NULL,
expenditure_concept = c("direct", "total") expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
@@ -19,7 +20,11 @@ cog_geographic_rollup(
`canonical_govid` values. At least one layer required.} `canonical_govid` values. At least one layer required.}
\item{category}{Single category name or character vector (passed through \item{category}{Single category name or character vector (passed through
to [cog_spending()]).} to [cog_spending()]), or the reserved `"All Categories"` for one summed
row per `(year, canonical_govid, subtype)` covering every category in the
concept's scope. `"All Categories"` is the efficient way to build a
geographic total: without it a caller must issue one rollup per category
and sum the results themselves.}
\item{years}{Integer vector of years.} \item{years}{Integer vector of years.}
@@ -29,12 +34,32 @@ are excluded from the result.}
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.} \item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only \item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
single-government queries but cannot be used here because combining Total refused here because combining Total across multiple layers of
across multiple layers of government double-counts intergovernmental government double-counts intergovernmental transfers (a state's payment
transfers (a state's payment to a school district is the same dollar the to a school district is the same dollar the district reports as its own
district reports as its own Direct spending).} Direct spending); `"primary"` and `"direct"` combine safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.}
} }
\value{ \value{
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
@@ -59,3 +84,23 @@ the result. The dropped govids are recorded in
(gov type 4) and school districts (gov type 5) from per-capita rollups (gov type 4) and school districts (gov type 5) from per-capita rollups
by design — see `vignette('population-denominators')`. by design — see `vignette('population-denominators')`.
} }
\section{Reading `coverage`}{
`provenance$coverage` reports `n_units_reporting` against
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
it counts governments with rows for the category you asked for, not
governments collected that year.** A government that was surveyed and
genuinely spends nothing in that category is indistinguishable here from one
that was never surveyed.
The ratio is therefore **not a response rate** and must not be used as one.
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
`category = "Police"`; the 174-city gap is overwhelmingly cities that
contract policing to the county sheriff, not non-response.
The comparison that *is* valid is the same category across a census year
(ending in 2 or 7) and a sample year, where the real-zero component is
roughly constant and the difference reflects the survey cycle. `is_census_year`
marks which is which.
}
+32 -8
View File
@@ -4,7 +4,13 @@
\alias{cog_gov_search} \alias{cog_gov_search}
\title{Search for governments by name, state, and/or type} \title{Search for governments by name, state, and/or type}
\usage{ \usage{
cog_gov_search(name = NULL, state = NULL, type = NULL) cog_gov_search(
name = NULL,
state = NULL,
type = NULL,
limit = NULL,
offset = NULL
)
} }
\arguments{ \arguments{
\item{name}{Character vector of place name(s). Length 1 = utility mode; \item{name}{Character vector of place name(s). Length 1 = utility mode;
@@ -19,12 +25,26 @@ all entries; otherwise must match `length(name)`.}
in basket mode (recycles from length 1). Excluded types `4`/`5` (or in basket mode (recycles from length 1). Excluded types `4`/`5` (or
`"special_district"` / `"school_district"`) trigger an explanatory `"special_district"` / `"school_district"`) trigger an explanatory
message and an empty result.} message and an empty result.}
\item{limit}{Maximum number of rows to return, applied in SQL. `NULL`
(default) returns every match -- which, with no other filter, is the
entire crosswalk. Utility mode only: pagination has no meaning in basket
mode, where the result is one resolved row per requested name in input
order, and is refused there with class
`uscogdata_basket_pagination_conflict`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
} }
\value{ \value{
A tibble of `canonical_fips_xwalk` rows. In utility mode, all A tibble of `canonical_fips_xwalk` rows. In utility mode, all
matches sorted by `population_acs` desc. In basket mode, resolved matches sorted by `population_acs` desc, ties broken by
rows in input order, with `attr(., "resolution")` set to the `canonical_govid`. In basket mode, resolved rows in input order, with
sidecar tibble. `attr(., "resolution")` set to the sidecar tibble.
When `limit` is set, carries a `total_rows` attribute: the full
unpaginated match count, computed by the same query (`COUNT(*) OVER()`)
rather than a second scan.
} }
\description{ \description{
Resolves human-readable place names into rows of `canonical_fips_xwalk`, Resolves human-readable place names into rows of `canonical_fips_xwalk`,
@@ -32,8 +52,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
} }
\details{ \details{
* **Utility mode** (single `name`, the original behavior): returns all * **Utility mode** (single `name`, the original behavior): returns all
rows whose `gov_name` matches the regex case-insensitively, sorted by rows whose `gov_name` contains `name` as a **literal, case-insensitive
`population_acs` descending. Useful for exploratory lookups. substring**, sorted by `population_acs` descending. Useful for
exploratory lookups. Regex metacharacters in `name` are escaped, so a
government is findable by its own complete name even when that name
contains parentheses or a period.
* **Basket mode** (`length(name) > 1`): resolves each input row to a * **Basket mode** (`length(name) > 1`): resolves each input row to a
single canonical govid and returns a tibble in input order, suitable single canonical govid and returns a tibble in input order, suitable
for piping straight into [cog_spending()] / [cog_revenue()] / for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -45,7 +68,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`. 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
2. **Exact pass:** case-insensitive equality against `gov_name`. 2. **Exact pass:** case-insensitive equality against `gov_name`.
Single hit -> resolved. Multiple -> step 4. Single hit -> resolved. Multiple -> step 4.
3. **Substring fallback:** case-insensitive regex against `gov_name`. 3. **Substring fallback:** case-insensitive literal substring against
`gov_name` (metacharacters escaped).
Single hit -> resolved (`match_method = "substring"`). Zero hits -> Single hit -> resolved (`match_method = "substring"`). Zero hits ->
`status = "no_match"`. Multiple hits -> step 4. `status = "no_match"`. Multiple hits -> step 4.
4. **Disambiguation:** if matches share one `govs_type`, pick the 4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -58,7 +82,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
} }
\examples{ \examples{
\dontrun{ \dontrun{
# Utility mode — exploratory regex lookup # Utility mode — exploratory substring lookup
cog_gov_search("broward", state = "FL") cog_gov_search("broward", state = "FL")
# Basket mode — resolve a known cohort # Basket mode — resolve a known cohort
+82 -6
View File
@@ -11,7 +11,8 @@ cog_peer_compare(
years, years,
per_capita = TRUE, per_capita = TRUE,
adjust_to_year = NULL, adjust_to_year = NULL,
expenditure_concept = c("direct", "total") expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
@@ -29,10 +30,37 @@ population.}
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.} \item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only \item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
single-government queries but cannot be used here because combining Total refused here because combining Total across peer sets counts
across peer sets counts intergovernmental transfers twice.} intergovernmental transfers twice; `"primary"` and `"direct"` combine
safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.
The comparison target is exempt from `"consistent"` balancing -- it is the
subject of the comparison, not a member of the cohort -- and the
`summary_*` quantiles are computed AFTER the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against
the cohort size: "3 of your 15 peers reported in FY2019".}
} }
\value{ \value{
Tibble matching [cog_spending()]'s columns, plus a `role` Tibble matching [cog_spending()]'s columns, plus a `role`
@@ -43,11 +71,59 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`, vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
`cohort_year`, and `cohort_govids`. `cohort_year`, and `cohort_govids`.
**The `summary_*` rows are per-category quantiles: they are not additive.**
Each one is computed **within each `(year, spend_subtype,
category)` cell** across the peer set, so a `summary_p50` row is *the
median peer's value in that one category*, not *the value of the median
peer's total*. The median peer for Police and the median peer for Fire
are usually different governments, so summing `summary_*` rows across
categories does not give any peer's total and misstates the band it
appears to describe — measured at −32.7% to +251.0% across 24 years on
one cohort, with a sign flip at FY2012.
Facet by `role` **and** `category` (the documented use, and what the
rows are built for). For a genuine "median peer's total spending" line,
sum each peer's own categories first and take the quantile of those
per-government totals:
```r
library(dplyr)
cmp |>
filter(role %in% c("target", "peer")) |>
group_by(year, role, canonical_govid) |>
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
filter(role == "peer") |>
group_by(year) |>
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
```
} }
\description{ \description{
Pulls spending for the target plus a peer set (either a Pulls spending for the target plus a peer set (either a
[cog_find_peers()] result or a character vector of `canonical_govid`) and [cog_find_peers()] result or a character vector of `canonical_govid`) and
appends peer-distribution summary rows (`summary_p25`, `summary_p50`, appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
`summary_p75`) so the result can be faceted by `role` in a single ggplot `summary_p75`) so the result can be faceted by `role` in a single ggplot
call. call. Those summary rows are quantiles **within each category**, not
quantiles of each peer's total — see the `@return` section before summing
them.
} }
\section{Reading `coverage`}{
`provenance$coverage` reports `n_units_reporting` against
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
it counts cohort members with rows for the category you asked for, not
cohort members collected that year.** A government that was surveyed and
genuinely spends nothing in that category is indistinguishable here from one
that was never surveyed.
The ratio is therefore **not a response rate** and must not be used as one.
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
`category = "Police"`; the 174-city gap is overwhelmingly cities that
contract policing to the county sheriff, not non-response.
The comparison that *is* valid is the same category across a census year
(ending in 2 or 7) and a sample year, where the real-zero component is
roughly constant and the difference reflects the survey cycle. `is_census_year`
marks which is which.
}
+105 -5
View File
@@ -5,22 +5,39 @@
\title{Summarized revenue by category} \title{Summarized revenue by category}
\usage{ \usage{
cog_revenue( cog_revenue(
govid, govid = NULL,
years, years,
category = NULL, category = NULL,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL, adjust_to_year = NULL,
basis = c("harmonized", "raw"), basis = c("harmonized", "raw"),
recipe = NULL recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE,
limit = NULL,
offset = NULL,
state = NULL,
type = NULL
) )
} }
\arguments{ \arguments{
\item{govid}{Character vector of `canonical_govid` values.} \item{govid}{Character vector of `canonical_govid` values, or `NULL` to name
the cohort by `state`/`type` instead. One of `govid`, `state`, or `type`
is required.}
\item{years}{Integer vector of years.} \item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from \item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.} `summary_categories.category`), or `NULL` for all categories broken out
one row each. The reserved value `"All Categories"` instead returns a
single summed row per `(year, canonical_govid, subtype)`, covering every
category inside the requested concept's subtype scope. It cannot be
combined with other category names, and it is not the same thing as
`revenue_concept = "total"`: the concept chooses which subtypes are in
scope, `"All Categories"` chooses whether rows inside that scope are
broken out or summed. Because the result keeps one row per
`revenue_subtype`, filtering the returned frame to
`revenue_subtype == "own_source"` gives an own-source revenue total.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and \item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
@@ -55,12 +72,95 @@ argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied = `basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.} possibly-misleading `"harmonized"`/`"raw"` value.}
\item{revenue_concept}{Which of Census's two published revenue concepts to
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
values -- never as item-code first letters, which cannot classify
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
prefix `X` does the same):
* `"general"` (default) -- Census General Revenue: `own_source` +
`federal` + `state` + `local_aid`. The manual defines this concept by
subtraction (section 4.3: *"General revenue comprises all revenue
except that classified as liquor store, utility, or insurance trust
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
insurance trust revenue are all excluded.
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
`general` plus utility, liquor store, and insurance trust revenue
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
`X01`/`X02`/`X05`/`X08`).
The two are related by Census's own identity, `Total Revenue = General +
Utility + Liquor Store + Insurance Trust`.
Note that the employee-retirement (`X`) codes stop at FY2016, when those
systems moved out of the annual finance file into the separate Annual
Survey of Public Pensions, so a `"total"` series steps down at the
FY2016/FY2017 seam for reasons that are about collection scope rather
than revenue (series breaks `SB197`-`SB202`).}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
result is materialized. `NULL` (the default) returns every matching row,
exactly as before this parameter existed. Mutually exclusive with
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
\item{state, type}{Name the cohort by predicate instead of by id: `state` is
a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
`"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
the same vocabulary, and the same internal coercion, as
[cog_gov_search()]. Both default to `NULL`.
The cohort is then expressed as a subquery against `canonical_fips_xwalk`
inside each statement rather than round-tripped through R as a literal id
list. For a fleet-scale cohort that is the difference between a
301,591-character `IN` list re-parsed in 5--8 statements per call and a
constant-size predicate: measured at **94 ms versus 449 ms** for the same
FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
7% of the no-filter floor.
Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
governments in `govid` that also match the predicate -- rather than one
silently taking precedence. Naming no cohort at all (`govid`, `state` and
`type` all `NULL`) aborts with class `uscogdata_no_cohort`.
When the cohort is named by predicate, `provenance$scope$govids_found`
and `govids_missing` are empty -- there is no id list to report against --
and `provenance$scope$cohort` carries `state`, `type` and
`n_governments` instead. A `govid`-named cohort reports exactly as before.}
} }
\value{ \value{
Tibble with columns `year`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `canonical_govid`, `gov_name`,
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`, `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`. optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
} }
\description{ \description{
Mirror of [cog_spending()] for revenue categories. One row per Mirror of [cog_spending()] for revenue categories. One row per
+129 -33
View File
@@ -5,23 +5,39 @@
\title{Summarized spending by category} \title{Summarized spending by category}
\usage{ \usage{
cog_spending( cog_spending(
govid, govid = NULL,
years, years,
category = NULL, category = NULL,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL, adjust_to_year = NULL,
basis = c("harmonized", "raw"), basis = c("harmonized", "raw"),
recipe = NULL, recipe = NULL,
expenditure_concept = c("direct", "total") expenditure_concept = c("primary", "direct", "total"),
complete = FALSE,
limit = NULL,
offset = NULL,
state = NULL,
type = NULL
) )
} }
\arguments{ \arguments{
\item{govid}{Character vector of `canonical_govid` values.} \item{govid}{Character vector of `canonical_govid` values, or `NULL` to name
the cohort by `state`/`type` instead. One of `govid`, `state`, or `type`
is required.}
\item{years}{Integer vector of years.} \item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from \item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.} `summary_categories.category`), or `NULL` for all categories broken out
one row each. The reserved value `"All Categories"` instead returns a
single summed row per `(year, canonical_govid, subtype)`, covering every
category inside the requested concept's subtype scope. It cannot be
combined with other category names, and it is not the same thing as
`expenditure_concept = "total"`: the concept chooses which subtypes are in
scope, `"All Categories"` chooses whether rows inside that scope are
broken out or summed. Because the result keeps one row per
`spend_subtype`, filtering the returned frame to
`spend_subtype == "operations"` gives an operating-expenditure total.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and \item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
@@ -57,41 +73,121 @@ argument is ignored and the result's provenance reports
FALSE`, pointing at the `recipe` block instead) rather than a FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.} possibly-misleading `"harmonized"`/`"raw"` value.}
\item{expenditure_concept}{`"direct"` (default) returns only the \item{expenditure_concept}{Which spending concept to return. Concepts are
government's own direct spending (item codes `E`/`F`/`G`), unchanged defined as sets of the crosswalk's `spend_subtype` values -- never as
from prior releases. `"total"` additionally UNIONs in the item-code first letters, which cannot classify correctly (prefix `Y`
intergovernmental leg -- payments to local governments (`M` codes) and alone spans revenue, expenditure, and balance codes):
to the state government (`L` codes, excluding the `L--` family-total
rollup) -- so results gain rows with `spend_subtype ==
"intergovernmental"`. Requires the active corpus's `summary_categories`
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
`uscogdata_ig_categories_unsupported` on an older corpus rather than
silently under-reporting. Mutually exclusive with `recipe` (a recipe
already defines its own component codes). **Do not sum `"total"`
results across levels of government** (e.g. state + county + city):
a state's `M12` payment to a school district is the same dollar the
district reports as its own direct `E12`, so summing both double-counts
it. This matters in particular with [cog_geographic_rollup()], which
sums across exactly that kind of multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY * `"primary"` (default) -- the government's own service provision:
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05` `operations` + `capital` + `assistance` subtypes.
split), which the Direct leg excludes by construction but the IG leg * `"direct"` -- Census's published Direct Expenditure: `primary` plus
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"` `interest` (interest on debt) and `insurance_benefits` (insurance
query, any (year, category) where this leaves intergovernmental rows trust benefit payments, e.g. pensions -- Census manual section
with NO Direct counterpart is flagged: the affected rows' `notes` 5.2.2.1 includes payments to retirees in Direct).
name the harmonization recipe that recovers the missing Direct * `"total"` -- `direct` plus the intergovernmental leg: payments to
component (when one exists), and local governments (`M` codes), to the state government (`L` codes,
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the excluding the `L--` family-total rollup), and state payments to
figure in those rows is the intergovernmental leg alone, not Direct + school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
IG.} `spend_subtype == "intergovernmental"`. Requires the active corpus's
`summary_categories` to carry M/L rows (added by cog_pipeline PR
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
older corpus rather than silently under-reporting. Mutually
exclusive with `recipe` (a recipe already defines its own component
codes).
**Do not sum `"total"` results across levels of government** (e.g.
state + county + city): a state's `M12` payment to a school district is
the same dollar the district reports as its own direct `E12`, so
summing both double-counts it. This matters in particular with
[cog_geographic_rollup()], which sums across exactly that kind of
multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
split), which the Direct leg excludes by construction but the IG leg
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
query, any (year, category) where this leaves intergovernmental rows
with NO Direct counterpart is flagged: the affected rows' `notes`
name the harmonization recipe that recovers the missing Direct
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG. When `category = "All Categories"` is combined with
`expenditure_concept = "total"`, this detection cannot run (it keys on
per-category rows, which all-categories mode collapses to one literal
value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
possibly-false `FALSE`; query an explicit `category` to get a real
answer.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
result is materialized. `NULL` (the default) returns every matching row,
exactly as before this parameter existed. Mutually exclusive with
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
\item{state, type}{Name the cohort by predicate instead of by id: `state` is
a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
`"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
the same vocabulary, and the same internal coercion, as
[cog_gov_search()]. Both default to `NULL`.
The cohort is then expressed as a subquery against `canonical_fips_xwalk`
inside each statement rather than round-tripped through R as a literal id
list. For a fleet-scale cohort that is the difference between a
301,591-character `IN` list re-parsed in 5--8 statements per call and a
constant-size predicate: measured at **94 ms versus 449 ms** for the same
FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
7% of the no-filter floor.
Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
governments in `govid` that also match the predicate -- rather than one
silently taking precedence. Naming no cohort at all (`govid`, `state` and
`type` all `NULL`) aborts with class `uscogdata_no_cohort`.
When the cohort is named by predicate, `provenance$scope$govids_found`
and `govids_missing` are empty -- there is no id list to report against --
and `provenance$scope$cohort` carries `state`, `type` and
`n_governments` instead. A `govid`-named cohort reports exactly as before.}
} }
\value{ \value{
Tibble with columns `year`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`, `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`. optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`. and `value_source` when `complete = TRUE`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
whose `completion` block reports `applied`, `rows_filled`, and the
per-year `absence_means` rule that was applied. When `limit` is set,
also carries a `total_rows` attribute: the full unpaginated row count,
computed by the same query (`COUNT(*) OVER()`) rather than a second
round trip -- so a caller walking pages never has to ask "how many are
there" separately.
} }
\description{ \description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
File diff suppressed because it is too large Load Diff
+14
View File
@@ -0,0 +1,14 @@
# Project journal
Append-only, newest first. **Entries are never edited** — the value of this file is
that it records what was believed at the time, including the parts that turned out
wrong. Where things stand *today* is in `STATUS.md`, which is generated.
Four lines per entry. The analysis belongs in the issue or the decision record; this
file carries the reasoning and the pointers.
- **Why** — the driver. The one line git cannot reconstruct later.
- **Obligates** — issues this change created elsewhere. Numbers, not prose.
- **Refs** — commits, issues, decision records.
---
+67
View File
@@ -0,0 +1,67 @@
# Project status
> Between the compass markers is generated. Edit the sources, not this.
<!-- compass:begin -->
<!-- compass:board -->
## Where this stands
uscogdata is at 0.4.0 and its public surface is settled: the query verbs, the cohort
predicates added in this release, and the provenance contract every verb returns.
The six open issues split cleanly. Two are API work carried out of the #9 review pass
and deliberately deferred there rather than fixed in that branch. Three concern the
corpus layer, and the largest of them, partition-level caching, was named the single
highest-leverage change on the remote path before being deferred. One, the
data-correction intake (#52), is a decision rather than a task: it was parked during
the 0.3.0 design, and the API announcement waits on it, because without it the corpus
cannot make the "traceable and correctable" claim that most distinguishes it from
Census's own files.
Nothing here is blocked on anything else, so the ordering is a judgement about value
rather than a dependency graph.
Compass's own files moved out of `docs/` this session. They were sitting inside
pkgdown's output directory, and `pkgdown::clean_site()` deletes every top-level entry
there except `CNAME` and `dev` — asked directly, it listed `docs/pm` and
`docs/decisions` among the 28 it would remove, with the guard that would have stopped
it satisfied by `docs/pkgdown.yml`. They are in `pm/` now. Nothing was lost: the
journal had no entries and there were no decision records yet, which made this the
cheapest moment to move. The `.gitignore` workaround that re-included two children of
an excluded `docs/` is gone with it.
## Ready to work on next
- **#34** cog_revenue() offers expenditure recipes as suggestions: scope the candidate query by category_type · `ws/api` — nothing is blocking it; something is currently wrong
- **#36** n_units_reporting is category-conditional and cannot be read as a response rate · `ws/corpus` — nothing is blocking it; owed work from an earlier change
- **#2** Extend population data to be households as an alternate spending denominator · `ws/corpus` — nothing is blocking it
- **#33** Decompose .build_suggestions() (106 lines) into named helpers · `ws/api` — nothing is blocking it
- **#52** Release 11/11: design the data-correction intake (deferred; gates the API announcement) · `ws/corpus` — nothing is blocking it
- **#64** Partition-level caching: R/cache.R is still a stub, and the remote path pays for it every session · `ws/corpus` — nothing is blocking it
## Workstreams
| Stream | Commits since | Open | Debt | Owes docs |
|---|---|---|---|---|
| Query verbs and results | 77 | 2 | 0 | no |
| Corpus, mirror, provenance | 39 | 4 | 1 | no |
| Vignettes and guides | 34 | 0 | 0 | **yes** |
## CI
![R-CMD-check](https://gitea.civilytics.org/Civilytics/uscogdata/actions/workflows/ci.yml/badge.svg?branch=main)
![Mirror to GitHub](https://gitea.civilytics.org/Civilytics/uscogdata/actions/workflows/mirror-github.yml/badge.svg?branch=main)
<details>
<summary>Dependency graph and detail</summary>
_Nothing blocks anything else, so there is no graph to draw._
- Marker: `none` (no journal entry yet)
- Commits since: 165
- Open issues: 6
</details>
<!-- compass:end -->
+44
View File
@@ -0,0 +1,44 @@
[project]
name = "uscogdata"
forge = "Civilytics/uscogdata"
# Three strands that go stale independently: what the verbs return, what the
# corpus is and how it is mounted, and how both are explained to a reader.
[[workstream]]
id = "api"
title = "Query verbs and results"
paths = [
"R/revenue.R", "R/spending.R", "R/balances.R", "R/peers.R", "R/search.R",
"R/categories.R", "R/recipes.R", "R/rollup.R", "R/explain.R", "R/basket.R",
"R/suggestions.R", "R/suppression.R", "R/complete.R", "R/cohort.R",
"R/basis.R", "R/adjust.R", "R/pagination.R",
]
docs = ["vignettes/*.Rmd", "README.md"]
[[workstream]]
id = "corpus"
title = "Corpus, mirror, provenance"
paths = [
"R/manifest.R", "R/mirror.R", "R/cache.R", "R/session.R", "R/provenance.R",
"R/coverage.R", "R/config.R", "R/views.R", "R/series_breaks.R",
"R/balance_caveats.R", "R/zzz.R", "data-raw/**", "inst/sql/**",
]
docs = ["vignettes/*.Rmd", "NEWS.md"]
[[workstream]]
id = "docs"
title = "Vignettes and guides"
paths = ["vignettes/**", "README.md", "_pkgdown.yml", "NEWS.md"]
docs = []
[roborev]
project_guidelines = [
"Every verb calls .ensure_session() first, then queries via DBI::dbGetQuery().",
"A verb's return value is always a tbl_df carrying a provenance attribute.",
"govid inputs always go through .coerce_govid_input(); it accepts a character vector or a data frame.",
"SQL has two layers: view definitions are numbered .sql files in inst/sql/ registered by .register_views(); query construction is inline sprintf() in R. Add a view as a file; build a query in R.",
"No arrow dependency -- DuckDB reads parquet natively.",
"withr is Suggests-only and must appear in tests alone.",
"Tests must pass offline against the bundled fixture; tests/testthat/setup.R sets USCOGDATA_URL for that.",
]
+12
View File
@@ -0,0 +1,12 @@
# Decisions
One file per decision, numbered and immutable. A decision that changes is superseded
by a new record, never edited in place — the old reasoning is the point.
The table below is **generated** by `compass:decide`. Do not hand-edit it.
<!-- compass:begin decisions -->
| # | Date | Decision | Status |
|---|---|---|---|
| — | — | *No decisions recorded yet.* | — |
<!-- compass:end decisions -->
+281
View File
@@ -0,0 +1,281 @@
# `cog_balances()` — a reader surface for cash and security holdings
**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26`
**Date:** 2026-08-03 · **Status:** design, awaiting approval
Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped
with `#11`/`#12` and is asserted at both view and verb level. This spec covers
requirement 2 only: a way to query holdings.
## Decision: a verb, not an argument
`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`.
Holdings are a **stock** — a balance at a point in time — while the money verbs
return **flows** over a fiscal year. The flow verbs' whole argument vocabulary
is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe
which flows Census aggregates into a published total, and `complete=` fills a
grid of fiscal-year cells. Overloading a money verb would put a stock behind
arguments that all assume a flow.
## The 14 codes
Measured against the published corpus 2026-08-03, not transcribed from the
issue. `year_min`/`year_max` are observed row extents.
| `balance_subtype` | `category` | codes | observed years |
|---|---|---|---|
| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 |
| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 |
| | | `X47` | 1988–2016 |
| | | `X30`, `Z77`, `Z78` | 2012–2016 |
| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 |
| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 |
| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 |
## Architecture
### Two new views
Mirroring the `revenue_long` / `revenue_annotated` pair exactly:
- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate`
- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and
`summary_categories`, exposing `category`, `category_type`, `balance_subtype`
`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register
with no new registration code.
### A third gate list in `R/views.R`
`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails
at registration time, not at query time. `46-balance_annotated.sql` selects
`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`.
That arrived without a `schema_version` bump, so neither existing gate applies:
`.harmonization_view_files` keys on `schema_version`, `.representation_view_files`
on the presence of a *file*. The discriminator here is a **column on an existing
table**.
```r
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
```
gated by probing `summary_categories` for `balance_subtype`, with
`cog_balances()` erroring cleanly via `.require_balance_support()` on an older
corpus — mirroring how `.require_schema_v5()` gates the harmonized views.
### `R/balances.R` — a dedicated path, not `.verb_spendrev()`
`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and
`complete=` grid are all flow-specific, and four verbs depend on it. Threading a
third mode through it adds branching to shared code for no reuse benefit.
Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`,
`.build_corpus_break_refs()`, the population join, `.inflate()`, and
`.coerce_govid_input()`.
Following the package's real two-layer convention: **view definitions** live in
`inst/sql/`; **query construction** is inline `sprintf()` in R, as in
`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R
files", which the verb layer has never obeyed. Corrected in a separate commit —
see Out of scope.)
## Signature
```r
cog_balances(govid, years,
category = NULL, # Fund Balances | Insurance Trust Balances |
# Retirement System Holdings
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL)
```
Returns a `tbl_df` with a `provenance` attribute, like every other verb.
**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`,
and `subtype` — see below.
**`per_capita` is offered.** Holdings per resident is a real measure (pension
assets per capita, fund balance per resident). The roxygen `@param` states
plainly that this is a *stock per resident* and is **not** comparable to
`cog_spending()`'s per-capita figures.
**`basis` is currently a no-op** — `harmonization_map` has zero balance-code
rows, so harmonized and raw are identical for holdings. Kept for uniformity
with the money verbs (the API would otherwise special-case), and
`provenance$basis_note` says so outright rather than letting it look meaningful.
**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
to the modern one:
```
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
```
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
other basis view. That is by design, not a defect:
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
era exposes these split families *only* as aggregates, and that the recipe join
must therefore **not** filter `is_aggregate` — safe by construction, because
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
every component is year-scoped, so no double-count is possible. § 1 records the
matching decision that the planned `X40→Z77` harmonization *map* rows were
dropped and the continuity ships as recipes instead, which is why
`harmonization_map` has no balance-code rows.
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
as intended.
### No `subtype` argument: `category` is a strict coarsening
`balance` is the only `category_type` in which `category` and the subtype column
are **not** orthogonal. Measured against the published crosswalk:
| `category_type` | subtypes spanning more than one category |
|---|---|
| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) |
| revenue | 1 of 7 (`own_source`) |
| **balance** | **0 of 5** |
For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire)
× *economic character* (operations, capital) — so both earn their place. For
balance the relation is a strict tree:
```
Fund Balances = {general} W01 W31 W61
Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78
Insurance Trust Balances = {unemployment_trust,
workers_comp_trust,
other_insurance_trust} Y07 Y08 Y21 Y61
```
Exposing both would therefore admit no useful combination. Of the 15 possible
pairs, 3 are redundant (the subtype already implies its category) and **12 are
guaranteed empty for every government in every year** — and an impossible query
would fail by returning an empty tibble, which reads as "this government holds
none" rather than "you asked a contradiction."
Dropping `subtype` also keeps the verb aligned with the rest of the package: no
uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in
`.verb_spendrev()`, and the API layers its own `subtype` row filter on top
(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for
`/balances`.
`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the
`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is
isolating one of the three insurance funds in a single argument;
`balance_subtype` remains a returned column, so that is one `dplyr::filter()`
away.
## Caveat surfacing
`provenance$balance_caveats`, always present, plus one `cli_inform()` per
session per caveat class when a query actually touches an affected family or
year. Structured so `cog-api#26` can forward the fields verbatim.
Verified against `series_breaks.csv`, not assumed:
| # | Caveat | Covered by existing machinery? |
|---|---|---|
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed |
On caveat 4's second condition: `.build_series_break_refs()` matches
`break_year BETWEEN min(years) AND max(years)`, so a request spanning only
2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a
series sits entirely after the change, on one consistent basis, and flagging a
break it never crosses would be noise. The same rule is applied deliberately in
`.build_corpus_break_refs()`. An earlier draft of this row omitted the span
condition and overclaimed.
`coverage_window` is derived per observed subtype family from the corpus, never
hardcoded, so it stays correct as the corpus grows.
`series_break_refs` and `corpus_break_refs` are otherwise populated by the
existing code-driven builders and need no change.
## Testing
New `tests/testthat/test-balances.R`. The bundled fixture covers all four
fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y`
throughout — so every test below runs offline.
- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing
the already-asserted forward guard. Absence is verified against the raw corpus
via `read_parquet` on `data/long`, never through the verb that creates it.
- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019;
`coverage_window` reports the termination and the console message fires once.
- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the
measured table above; the FY2002 valuation caveat fires only when the year
range crosses 2002 *and* touches `employee_retirement`.
- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`.
- **`category = "Fund Balances"` is the `general` family.** Returns exactly
`W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted
rather than assumed.
- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to
exactly one `category`. Asserted against the crosswalk so that an upstream
change breaking the tree — which would silently make `category` lossy —
fails here rather than in a user's analysis.
- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
`X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
regression here would silently truncate a 45-year series to five.
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
FY2002 carries both in `provenance$series_break_refs`, so the book → market
basis change is disclosed wherever `X40`/`X41` are actually observed.
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
`summary_categories` lacks `balance_subtype`.
## Out of scope, tracked separately
1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
the seven holdings codes in `series_breaks.csv`. There is currently **no**
entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
`docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
exactly this — *"candidate new `series_breaks.csv` entries (recommend
`with_caution` documentation rows, no map action)"*. The follow-through never
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
what keeps this out of the `joinable = "no"` identity-change rule, which
would otherwise oblige a harmonization-map row.
Verify the break corpus-wide and census-to-census before writing the rows.
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
signpost when it lands.
**Superseded:** an earlier draft of this spec proposed adding
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
are the documented § 1 decision, and the recipe path reaches them by design.
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
`param_contract`, and the `plumber.R` route signature. Lands after this.
**Two contract facts the API must carry forward**, both settled during
implementation and easy to get wrong from the outside:
- `provenance$balance_caveats$coverage_window` is **corpus-scoped, not
result-scoped**. It reports the observed year extent of *every* balance
subtype in the corpus, not only the subtypes a given query returned — so a
`category = "Fund Balances"` query still returns all five windows. That is
deliberate: the windows describe what the corpus holds, which is what a
consumer needs in order to know what it did *not* ask for. The sibling
field `truncated` is the result-scoped one. Documented in
`inst/schemas/provenance-v1.json` and mutation-guarded against silent
inversion.
- `balance_caveats` appears **only** on `cog_balances()` results. It is
absent from `cog_spending()`/`cog_revenue()` provenance, and the schema
says so — an API layer that assumes it is universal will read `NULL`.
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7
SQL views (there are 21), 181 tests (716), a two-year fixture (four years),
and a "never inline SQL" rule the verb layer does not follow.
+296
View File
@@ -0,0 +1,296 @@
# `uscogdata` 0.3.0 — public release
**Date:** 2026-08-08 · **Status:** design, awaiting approval
**Scope:** release-readiness, README, NEWS. Distribution mechanics recorded here as
decided, sequenced after the package is clean.
`uscogdata` is feature-complete and the corpus it reads has been public on
HuggingFace since 2026-08-07 (294 downloads as of this writing). The API built on
it is live. What does not exist is a public *package*: the repo is private, there
is no install path, and — measured, not assumed — **a stranger who installed it
today could not read the corpus at all.**
This spec covers making that untrue.
## Decisions locked
| Decision | Choice |
|---|---|
| Canonical source | `gitea.civilytics.org/Civilytics/uscogdata`, flipped public |
| Public mirror | `github.com/civilytics/uscogdata` — issues, PRs, multi-OS check, CDN |
| Mirror mechanism | Gitea Actions non-force `git push` (not a push mirror) |
| Binaries | `civilytics.r-universe.dev`, registry pinned to a release tag |
| Author of record | Jared E. Knowles `<jared@civilytics.com>`, ORCID `0000-0003-0005-9478` |
| Copyright | Civilytics Consulting LLC (`cph`, `fnd`) |
| License | MIT (package) · CC-BY-4.0 (corpus) |
| Corrections intake | Deferred — see *Out of scope* |
| Other packages | Parked until this one walks the path end to end |
## P0 — the corpus is unreachable
Two independent faults, either of which alone is fatal.
**No corpus URL exists.** `R/config.R` defaults to the literal
`REPLACE_WITH_SHARE_TOKEN` sentinel, and no file in the repo supplies a working
one. A new user calling any verb gets `uscogdata_url_not_configured` with no path
to resolution.
**Remote reads are broken regardless.** Every partitioned view globs:
```sql
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true)
```
DuckDB 1.5.5 refuses globs over generic HTTP. Its suggested
`allow_asterisks_in_http_paths` escape hatch does not help — it forwards the
literal `**/*` as a filename and 404s, because plain HTTP exposes no directory
listing to expand against.
The package therefore works only against a **local path**. That is how the API
runs it (`CORPUS_HOST_PATH` is a host mount on maxwell) and how the tests run
(bundled fixture), which is why the fault went unnoticed. The README's headline
claim — *"Reads the published corpus directly from Nextcloud via DuckDB httpfs —
no local bulk downloads required"* — is currently false.
### Fix: enumerate from the manifest, do not glob
`manifest.json` already lists every partition under `files.long_partitions[]`
with `path`, `year`, `sha256`, `row_count` and `size_bytes` — 56 of them.
Substituting an explicit file list for the glob was measured against the
published corpus on 2026-08-08:
| Path | Result |
|---|---|
| `https://…/data/long/**/*.parquet` (default) | error — globs unsupported over HTTP |
| same, `allow_asterisks_in_http_paths = true` | error — literal `**/*` 404s |
| `hf://datasets/civilytics/us-cog-finance/…` glob | 46,148,034 rows |
| **explicit list over plain https** | **46,148,034 rows** |
`hive_partitioning = true` still recovers `year` from the paths under
enumeration, so no downstream view or verb changes.
Enumeration is preferred over `hf://` deliberately. It is **host-agnostic** —
Nextcloud, HuggingFace, or any static server take the same code path — where
`hf://` would tie the default to one vendor's protocol and still need
special-casing, since manifest fetching goes through `httr2`, which cannot speak
`hf://`. Enumeration also *removes* a dependency (globbing) rather than adding
one, and the manifest's per-file `sha256` becomes available for integrity
checking later.
Views are registered from `inst/sql/` with `{url}` substitution in
`R/views.R:.register_views()`. The list must be built once per session from the
already-fetched manifest and substituted the same way, so the change is confined
to view registration and does not touch verb code.
### Fix: ship a working default
`R/config.R`'s default becomes the public HuggingFace `resolve/main/` URL:
CC-BY-4.0, no token to publish, CDN-backed, and it keeps maxwell's uplink out of
the path — the same reasoning behind the GitHub mirror and r-universe.
This means `library(uscogdata)` followed by a verb works with **zero
configuration**, which is what makes the package demonstrable in a README and
later in a post. `USCOGDATA_URL` and `options(uscogdata.url=)` continue to
override, so the Nextcloud copy and local mirrors are unaffected.
The `uscogdata_url_not_configured` error class stays — it still fires for an
explicitly-set empty or placeholder URL — but ceases to be the default
experience.
### Consequence: `cog_mirror()` is promoted
Measured cost of the remote default, from efron on a good connection:
| | |
|---|---|
| Whole corpus | **190.6 MB**, 56 partitions, 46,148,034 rows, FY1967–FY2024 |
| One government, one year | 1.5 s |
| One government, all 56 years | 2.8 s |
| Disk written | **0.00 MB** — range requests only; `external_file_cache` is in-memory |
Nothing persists locally beyond the shared `httpfs` extension in `~/.duckdb` (a
few MB, once per machine, across all DuckDB use). Costs are RAM and per-query
bandwidth, since nothing caches between sessions.
Those timings are raw scans. Real verbs additionally join crosswalks, resolve
categories and assemble provenance, so end-to-end verb latency will be higher and
**must be re-measured once the fix lands** — it cannot be measured today.
The corpus being only 190.6 MB makes `cog_mirror()` a first-class option rather
than a developer footnote. The README presents **both paths**:
- **Remote (default, zero setup)** — trying it out, teaching, one-off questions.
- **Mirrored (`cog_mirror()`, 190 MB once)** — repeated or heavy analysis,
offline work, reproducibility, or preferring not to depend on HuggingFace.
The second is also the honest answer to the vendor-dependency question raised by
defaulting to HuggingFace: **the escape hatch is one function call and 190 MB**,
after which no analysis touches an external service. The README says so
explicitly. That is the difference between a convenience default and lock-in.
## Release-readiness fixes
| # | Issue | Fix |
|---|---|---|
| 1 | `MaxCorpusSchema: 5` in DESCRIPTION; `.validate_schema()` accepts `4,5,6,7`; published corpus is **7** | `MaxCorpusSchema: 7` |
| 2 | `^vignettes$` in `.Rbuildignore` — both vignettes absent from the installed package, while README tells users to run `vignette("total-spending")` | Remove `^vignettes$`, `^doc$`, `^Meta$`. Both vignettes build offline (`total-spending` reads the bundled fixture; `population-denominators` is `eval = FALSE`) |
| 3 | `_pkgdown.yml` reference index covers 6 of 14 exports — pkgdown errors on missing topics | Add `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`; set `url:` |
| 4 | No `URL:` / `BugReports:` in DESCRIPTION | Add both, pointing at the GitHub mirror |
| 5 | No `LICENSE.md`; `LICENSE` holder reads `Civilytics` | `usethis::use_mit_license("Civilytics Consulting LLC")` |
| 6 | README instructs stripping the fixture at release | Delete that section — see below |
| 7 | `Authors@R` is an org with no human | Jared E. Knowles `aut`/`cre` + ORCID; Civilytics Consulting LLC `cph`/`fnd` |
**On #6.** The advice to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`
is CRAN-sized thinking (5 MB limit) and this package is not going to CRAN.
Stripping the 15 MB fixture would break `total-spending.Rmd`, which reads from
it, and would leave r-universe and GitHub Actions unable to run the 28 test files
without a corpus credential. **The fixture is what lets `R CMD check` pass
anywhere with zero secrets** — precisely what public CI needs. It ships.
## README
The current README addresses someone standing inside the repo tree: status reads
"Under active development (Phase 2 of the cog_pipeline project)", it points at
`../cog_pipeline/docs/reader-specification.md`, the install line is commented
out, and developer, testing and release sections sit above anything a user needs.
Restructured around a stranger, in this order:
1. **What this is** — one paragraph, and what the corpus covers (types 0–3,
FY1967–FY2024, 46M rows, 190.6 MB).
2. **Install** — r-universe first (binaries), git second.
3. **Quickstart that actually runs** — resolve a government, get its history,
print provenance. No configuration step.
4. **Two ways to read the corpus** — remote default vs `cog_mirror()`, with the
measured numbers and the independence note.
5. **Amounts are in full US dollars** — kept near the top. This is the errata
most likely to produce a wrong answer that looks plausible.
6. **Concepts** — primary/direct/total spending, general/total revenue,
coverage. Condensed, linking to the vignettes for the full treatment.
7. **How to cite** — `citation("uscogdata")`, corpus CC-BY-4.0 attribution.
8. **Contributing** — canonical-on-Gitea PR flow.
Developer notes, testing instructions and release procedure move to
`CONTRIBUTING.md`. Every path reference to a sibling repo is removed or replaced
with a URL that resolves for someone who has only this repo.
## NEWS.md
`NEWS.md` currently holds two sections. `0.2.0` is a legitimate changelog — the
`"All Categories"` reserved value, the coverage-signposting fix, the
`n_units_reporting` documentation — and it stays. Beneath it,
`0.1.0 (development)` is a pre-release churn log: changes described relative to
states no user has ever seen ("Breaking: corpus schema_version 4", "the package
now requires…"), spanning the package's entire pre-release development. To a
newcomer deciding whether to depend on this, that section reads as instability.
**A new `0.3.0` section is added at the top, framed as the first public
release**: what the package does, what the corpus covers, and the caveats that
are genuinely load-bearing. **`0.2.0` is kept verbatim.** **`0.1.0 (development)`
is dropped** — that history stays in git, where it belongs.
The version is `0.3.0` rather than `0.2.0` because this release changes
user-visible behaviour: remote corpus reads go from broken to working, and the
default URL from a dead placeholder to a live corpus. It is also not `1.0.0` —
the corpus still excludes government types 4 and 5 pending validation, so a
stability promise would overclaim. No git tag exists for any prior version;
`chore: release 0.2.0` bumped `DESCRIPTION` and `NEWS` only.
The substantive content is migrated, not deleted. These are hard-won and belong
in documentation rather than buried in a changelog:
| Content | Destination |
|---|---|
| Coverage disclosure on multi-government aggregates (census vs sample years) | README concepts + `cog_geographic_rollup()` docs |
| `complete = TRUE` three-way absence semantics (`reported` / `census_zero` / `not_reported`) | `cog_spending()` / `cog_revenue()` docs |
| Series-break and corpus-break surfacing | README + `cog_explain()` docs |
| $1,000s → full dollars conversion | README, already prominent |
| Per-year F-33 population denominators | `population-denominators` vignette, already there |
This also makes NEWS reusable as raw material for the release announcement,
which is the stated downstream purpose.
## Distribution mechanics
Recorded as decided; executed after the package is clean and checks are green.
**Sequence matters.** r-universe publishes check results the moment a package is
registered. Registering before the fixes above land means a red badge on day one,
which is a worse first impression than a week's delay.
1. `gitleaks` over full history. A coarse grep found nothing across 140 commits
and the default corpus URL is still the placeholder sentinel, but a proper
scan is the gate on an irreversible action.
2. Flip the Gitea repo public. Disable Gitea issues on it, so there is exactly
one inbox.
3. Create `github.com/civilytics/uscogdata`. Add `.github/workflows/` for the
Windows/macOS/Linux `R CMD check` matrix — the platforms the Gitea runner
cannot provide, and which this package has never been tested on despite
depending on duckdb and httr2. Gitea reads `.gitea/workflows`, GitHub reads
`.github/workflows`; both live in one tree without colliding.
4. Gitea Actions workflow pushing to GitHub **without `--force`**, so divergence
fails loudly in CI rather than silently overwriting.
5. Add `jared@civilytics.com` as a verified secondary email on the GitHub
account — r-universe links maintainer identity by matching DESCRIPTION's email
against registered GitHub emails, and the association only takes effect on the
next build.
6. Tag `v0.3.0`. Create `github.com/civilytics/civilytics.r-universe.dev` with a
`packages.json` pinned to the tag, pointing at the GitHub mirror rather than
Gitea so clone traffic stays off maxwell. Install the r-universe app.
### PR flow
Never press Merge on GitHub. A merge there is overwritten by the next sync, the
PR still displays "Merged", and nothing says otherwise.
```sh
git remote add github https://github.com/civilytics/uscogdata.git
git config --add remote.github.fetch '+refs/pull/*/head:refs/remotes/github/pr/*'
git fetch github
git switch -c pr-42 github/pr/42 # test
git switch main && git merge --no-ff pr-42
git push origin main # Gitea -> mirror -> GitHub
```
GitHub auto-closes a PR as merged once its head commit becomes an ancestor of the
base branch, so `--no-ff` — which preserves the contributor's SHAs — makes the PR
close itself when the mirror pushes. **For external PRs, merge; do not squash or
rebase.** Squashing rewrites the SHAs, the auto-close never fires, and closing by
hand reads to a first-time contributor as rejection.
`CONTRIBUTING.md` states this, and a GitHub Action comments it on incoming PRs.
No CLA; no DCO.
## Verification
The release is not done until all of these pass:
1. `R CMD check --as-cran` clean on Linux, and on Windows and macOS via the
GitHub matrix. This package has never been checked on the latter two.
2. Full test suite (28 files) green against the **bundled fixture**, offline,
with no credentials — the property public CI depends on.
3. Full test suite green against the **live corpus**, which additionally
exercises the enumeration fix that the fixture's local path cannot.
4. `pkgdown::build_site()` completes.
5. Both vignettes present in the built tarball and
`vignette("total-spending", package = "uscogdata")` resolves from an
installed copy.
6. **Cold-start check on a machine that has never seen this package:** install
from r-universe, `library(uscogdata)`, run the README quickstart verbatim with
no environment variables set. This is the only test that catches the P0 class
of fault, and its absence is why the fault survived.
7. End-to-end verb latency re-measured against the live corpus and the README's
numbers updated if they moved.
## Out of scope
- **Corrections intake.** Deferred by decision. Consequence: the release cannot
invite data-error reports or make the "traceable and correctable" claim that
most distinguishes this corpus from Census's own files. `BugReports:` points at
package issues only. A verified correction should eventually terminate as a
`lineage_event` or `series_break` row so it propagates through provenance to
every consumer — that design is unstarted.
- **Announcement posts.** Deferred. The API announcement is gated on corrections
landing and merits a Civic Pulse edition.
- **The rest of the R package backlog.** Parked until this one completes the path.
- **`cog_pipeline` publication.** Stays private.
+96
View File
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
if (nzchar(p)) paste0(p, "/") else "" if (nzchar(p)) paste0(p, "/") else ""
} }
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
# or "" when it isn't there.
#
# Tests that assert on documentation content have to read the sources, and the
# sources only exist when the suite runs from a checkout. Under R CMD check the
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
# shipped and `../../README.md` does not resolve -- so those tests must skip
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
# rcmdcheck, so the assertions are still enforced on every push; this only
# stops them from failing a context that structurally cannot satisfy them.
source_tree_path <- function(...) {
p <- testthat::test_path("..", "..", ...)
if (file.exists(p)) p else ""
}
# Skip unless every named source file is present (see source_tree_path()).
skip_if_no_source_tree <- function(...) {
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
character(1))
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
testthat::skip_if(
any(missing),
"package source tree not available (running against the installed package)"
)
invisible(paths)
}
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL). # Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
skip_if_no_corpus <- function() { skip_if_no_corpus <- function() {
p <- fixture_corpus_path() p <- fixture_corpus_path()
@@ -58,6 +85,42 @@ with_doctored_schema_version <- function(version, code) {
force(code) force(code)
} }
# Copy the bundled fixture to a temp dir with representation.parquet and
# code_set.parquet removed (and dropped from the manifest's metadata list),
# then run `code` against it. Models a corpus published BEFORE sparsification:
# schema_version is left alone deliberately, because it was never bumped for
# that change -- the pre-sparsification fixture this package shipped until
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
# is therefore the only honest signal, and this helper is what proves the
# package keys off it rather than off the version number.
with_corpus_missing_representation <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
dropped <- c("representation.parquet", "code_set.parquet")
file.remove(file.path(tmp, "data", dropped))
manifest_path <- file.path(tmp, "manifest.json")
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
m$files$metadata <- Filter(
function(f) !basename(f$path) %in% dropped, m$files$metadata
)
writeLines(
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
manifest_path
)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
# Copy the bundled fixture to a temp dir with summary_categories.parquet # Copy the bundled fixture to a temp dir with summary_categories.parquet
# rewritten to drop every M/L (intergovernmental) row, then run `code` # rewritten to drop every M/L (intergovernmental) row, then run `code`
# against it with a clean session (mirrors with_fixture_corpus()/ # against it with a clean session (mirrors with_fixture_corpus()/
@@ -91,3 +154,36 @@ with_corpus_missing_ig_categories <- function(code) {
}, add = TRUE) }, add = TRUE)
force(code) force(code)
} }
# Copy the bundled fixture to a temp dir with summary_categories.parquet
# rewritten to DROP the balance_subtype column, then run `code` against it.
# Models a corpus published before cog_pipeline #76/#77. schema_version is
# left untouched deliberately: that change shipped without a version bump, so
# column presence is the only honest signal -- this helper is what proves the
# package keys off it. Mirrors with_corpus_missing_ig_categories().
with_corpus_missing_balance_subtype <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf(
"COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s))
TO %s (FORMAT PARQUET)",
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
))
file.remove(cats_path)
file.rename(filtered_path, cats_path)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
@@ -0,0 +1,58 @@
test_that('cog_geographic_rollup() accepts "All Categories" and agrees with per-category sums', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
expect_gt(nrow(govs), 1L)
ids <- list(city = utils::head(govs$canonical_govid, 25L))
by_cat <- cog_geographic_rollup(ids, category = NULL, years = 2019L)
total <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L)
expect_setequal(unique(total$category), "All Categories")
# one row per (govid, subtype) that appears in the per-category result
key_by_cat <- unique(paste(by_cat$canonical_govid, by_cat$spend_subtype))
key_total <- paste(total$canonical_govid, total$spend_subtype)
expect_setequal(key_total, key_by_cat)
lhs <- tapply(by_cat$amt_nominal, paste(by_cat$canonical_govid, by_cat$spend_subtype), sum)
rhs <- tapply(total$amt_nominal, key_total, sum)
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
})
test_that('"All Categories" survives per_capita and inflation adjustment through the rollup', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = utils::head(govs$canonical_govid, 10L))
r <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
per_capita = TRUE, adjust_to_year = 2020L)
expect_true(all(c("amt_per_capita_nominal", "amt_real", "amt_per_capita_real") %in% names(r)))
expect_setequal(unique(r$category), "All Categories")
expect_true(all(is.finite(r$amt_real)))
})
test_that('cog_geographic_rollup() still refuses expenditure_concept = "total" with "All Categories"', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = utils::head(govs$canonical_govid, 5L))
expect_error(
cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
expenditure_concept = "total")
)
})
test_that("n_units_reporting is category-conditional, not a response rate", {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = govs$canonical_govid)
police <- cog_geographic_rollup(ids, category = "Police", years = 2012L)
allcat <- cog_geographic_rollup(ids, category = "All Categories", years = 2012L)
cov_police <- cog_explain(police, format = "list")$coverage
cov_all <- cog_explain(allcat, format = "list")$coverage
# Same year, same requested govids, same collection -- yet a single category
# reports fewer units than the all-categories query. That gap is real zeros,
# not non-response, which is exactly why the ratio is not a response rate.
expect_lte(cov_police$n_units_reporting, cov_all$n_units_reporting)
expect_identical(cov_police$n_units_expected, cov_all$n_units_expected)
})
+267
View File
@@ -0,0 +1,267 @@
# Baseline at branch point: 843 PASS / 0 FAIL / 0 SKIP / 0 WARN (2026-08-05, origin/main 2fc9e75)
test_that(".build_verb_sql emits a literal category and no category filter in all-categories mode", {
sql <- uscogdata:::.build_verb_sql(
view = "spending_annotated",
subtype_col = "spend_subtype",
cohort = uscogdata:::.make_cohort("552025209777"),
years = 2019L,
category = NULL,
subtype_scope = c("operations", "capital"),
all_categories = TRUE
)
expect_match(sql, "'All Categories' AS category", fixed = TRUE)
# no category filter of any kind
expect_false(grepl("AND category IN", sql, fixed = TRUE))
# category is not a grouping key
expect_false(grepl("GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
sql, fixed = TRUE))
# the subtype allowlist still applies -- this is what makes the sum a concept
expect_match(sql, "AND spend_subtype IN ('operations','capital')", fixed = TRUE)
})
test_that(".build_verb_sql is unchanged when all_categories is FALSE", {
args <- list(
view = "spending_annotated", subtype_col = "spend_subtype",
cohort = uscogdata:::.make_cohort("552025209777"), years = 2019L, category = NULL,
subtype_scope = c("operations", "capital")
)
old <- do.call(uscogdata:::.build_verb_sql, args)
new <- do.call(uscogdata:::.build_verb_sql, c(args, list(all_categories = FALSE)))
expect_identical(old, new)
expect_match(new, "GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
fixed = TRUE)
})
test_that(".ALL_CATEGORIES is the exact reserved string", {
expect_identical(uscogdata:::.ALL_CATEGORIES, "All Categories")
})
test_that('cog_spending(category = "All Categories") sums to the per-category total', {
gov <- "552025209777"
by_cat <- cog_spending(gov, 2019L)
total <- cog_spending(gov, 2019L, category = "All Categories")
expect_true(nrow(total) > 0L)
expect_setequal(unique(total$category), "All Categories")
# one row per subtype present in the by-category result
expect_setequal(unique(total$spend_subtype), unique(by_cat$spend_subtype))
expect_equal(nrow(total), length(unique(by_cat$spend_subtype)))
# the dollars agree, per subtype
lhs <- tapply(by_cat$amt_nominal, by_cat$spend_subtype, sum)
rhs <- tapply(total$amt_nominal, total$spend_subtype, sum)
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
})
test_that('"All Categories" respects expenditure_concept', {
gov <- "552025209777"
prim <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "primary")
dir <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "direct")
# direct = primary plus interest and insurance benefits, so it is never smaller
expect_gte(sum(dir$amt_nominal), sum(prim$amt_nominal))
})
test_that('"All Categories" works on revenue and respects revenue_concept', {
gov <- "552025209777"
gen <- cog_revenue(gov, 2019L, category = "All Categories",
revenue_concept = "general")
tot <- cog_revenue(gov, 2019L, category = "All Categories",
revenue_concept = "total")
expect_setequal(unique(gen$category), "All Categories")
expect_gte(sum(tot$amt_nominal), sum(gen$amt_nominal))
})
test_that('"All Categories" cannot be combined with another category', {
expect_error(
cog_spending("552025209777", 2019L, category = c("All Categories", "Police")),
class = "uscogdata_all_categories_not_combinable"
)
})
test_that('"All Categories" is recorded in provenance', {
r <- cog_spending("552025209777", 2019L, category = "All Categories")
expect_identical(cog_explain(r, format = "list")$category, "All Categories")
})
test_that('"All Categories" combines with subtype to give operating totals', {
gov <- "552025209777"
ops_by_cat <- cog_spending(gov, 2019L)
ops_by_cat <- ops_by_cat[ops_by_cat$spend_subtype == "operations", ]
ops_total <- cog_spending(gov, 2019L, category = "All Categories")
ops_total <- ops_total[ops_total$spend_subtype == "operations", ]
expect_equal(sum(ops_total$amt_nominal), sum(ops_by_cat$amt_nominal),
tolerance = 1e-8)
})
test_that('cog_categories() advertises "All Categories" for both flows', {
all <- cog_categories()
rows <- all[all$category == "All Categories", ]
expect_setequal(rows$category_type, c("expenditure", "revenue"))
expect_true(all(is.na(rows$subtype)))
expect_true(all(is.na(rows$n_codes)))
})
test_that('cog_categories(type=) still scopes, including the pseudo-category', {
sp <- cog_categories(type = "spending")
expect_setequal(unique(sp$category_type), "expenditure")
expect_true("All Categories" %in% sp$category)
rev <- cog_categories(type = "revenue")
expect_setequal(unique(rev$category_type), "revenue")
expect_true("All Categories" %in% rev$category)
# balances have no concept vocabulary, so no pseudo-category
bal <- cog_categories(type = "balance")
expect_false("All Categories" %in% bal$category)
})
test_that('cog_categories(pattern=) matches the pseudo-category', {
hit <- cog_categories(pattern = "^All Categories$")
expect_equal(nrow(hit), 2L)
})
# --- final whole-branch review fixes ---------------------------------------
test_that('complete = TRUE is refused when combined with "All Categories"', {
# .completion_grid_sql() would emit `AND c.category IN ('All Categories')`,
# match zero crosswalk rows, and the early return in .complete_result()
# would stamp completion$applied = TRUE, rows_filled = 0 -- reading as "the
# grid was checked and nothing was missing" when nothing was actually
# checked. Filling a summed row has no defined semantics, so the verb must
# refuse the combination outright (finding 2).
expect_error(
cog_spending("552025209777", 2019L, category = "All Categories",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
expect_error(
cog_revenue("552025209777", 2019L, category = "All Categories",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
})
test_that('cog_balances() rejects "All Categories" instead of silently returning zero rows', {
# cog_balances() reuses .validate_verb_inputs() but did not pass
# allow_all_categories = TRUE, so "All Categories" used to become
# `AND category IN ('All Categories')` against balance_annotated -- 0
# matching crosswalk rows, 0 rows back, no error (finding 3). Holdings are
# a stock with no concept vocabulary to sum across, so the honest answer is
# to refuse, the same way cog_spending()/cog_revenue() refuse other
# nonsensical combinations.
expect_error(
cog_balances("552025209777", 2019L, category = "All Categories"),
class = "uscogdata_all_categories_unsupported"
)
# An ordinary category still works -- this is not a blanket regression.
r <- suppressMessages(
cog_balances("552025209777", 2019L, category = "Fund Balances")
)
expect_gt(nrow(r), 0L)
})
test_that('expenditure_concept_direct_suppressed is NA, not FALSE, when categories are collapsed', {
# .detect_direct_suppressed() keys on
# paste(year, canonical_govid, category, sep = "\r"). In all-categories
# mode every row carries the literal "All Categories" value, so an IG-only
# row's key collides with any ordinary Direct row for the same
# (year, govid) -- has_direct reads TRUE whenever the government has ANY
# direct spending at all, candidate is always empty, and the detector can
# never fire. Before the fix this silently reported FALSE, an affirmative
# claim the code did not actually compute (finding 1). NA is the honest
# answer: cog_explain(x, format = "list") is required here, since without
# format = "list" it returns the result tibble, not the provenance list.
gov <- "552025209777"
t <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "total")
prov <- cog_explain(t, format = "list")
expect_true(is.na(prov$expenditure_concept_direct_suppressed))
expect_false(isTRUE(prov$expenditure_concept_direct_suppressed))
expect_match(prov$expenditure_concept_note, "unavailable", fixed = TRUE)
# A per-category "total" query on the same government/year is unaffected --
# the detector can still key correctly and reports a strict logical.
t_by_cat <- cog_spending(gov, 2019L, expenditure_concept = "total")
prov_by_cat <- cog_explain(t_by_cat, format = "list")
expect_false(is.na(prov_by_cat$expenditure_concept_direct_suppressed))
})
test_that('"All Categories" still signposts coverage gaps (finding 6, final whole-branch review)', {
# .build_suggestions()'s candidate sub-select used to be keyed on
# `category`, e.g. `WHERE category IN ('All Categories')`. Since
# .ALL_CATEGORIES is never itself a row in summary_categories.category,
# that sub-select always came back empty in all-categories mode, so
# `candidates` was empty and .build_suggestions() short-circuited to
# list() -- coverage signposting was structurally impossible for the one
# mode whose whole selling point is "you cannot sum the wrong scope"
# (uscogdata#9's entire point, silently defeated).
#
# AL state government, FY2011, category = "Corrections": this category has
# no legacy leaf rows in FY2011 (aggregate-flagged E04/E05 family), so the
# per-category query returns 0 rows and 3 recipe-hint suggestions fire
# (empty_year path). All-categories mode does not have an empty year --
# the government has other primary spending in FY2011 -- but the same
# suppressed Corrections dollars are still excluded from the summed total,
# so the fix (scoping the candidate sub-select by subtype_col/subtype_scope
# instead of by category, symmetric with .build_verb_sql()) must still
# surface them via the suppressed_component path.
gov <- "010000226085"
by_cat <- suppressMessages(cog_spending(gov, 2011L, category = "Corrections"))
sugg_by_cat <- cog_explain(by_cat, format = "list")$suggestions
expect_gt(length(sugg_by_cat), 0L)
all_cat <- suppressMessages(cog_spending(gov, 2011L, category = "All Categories"))
sugg_all_cat <- cog_explain(all_cat, format = "list")$suggestions
expect_gt(length(sugg_all_cat), 0L)
# The same Corrections recipe that fired per-category must also fire in
# all-categories mode -- not just some unrelated recipe.
ids_by_cat <- vapply(sugg_by_cat, function(s) s$recipe_id %||% "", character(1))
ids_all_cat <- vapply(sugg_all_cat, function(s) s$recipe_id %||% "", character(1))
expect_true("corrections_combined" %in% ids_by_cat)
expect_true("corrections_combined" %in% ids_all_cat)
# In all-categories mode the government DOES have other primary spending
# in FY2011 (the year itself is not a gap), so the suggestion can only have
# fired via the suppressed_component path, not empty_year.
corr_all <- sugg_all_cat[[which(ids_all_cat == "corrections_combined")]]
expect_identical(corr_all$trigger, "suppressed_component")
expect_gt(corr_all$suppressed_amount, 0)
})
test_that('"All Categories" candidate scoping is symmetric with .build_verb_sql() -- subtype, not category', {
# Direct assertion on the mechanism itself (finding 6): in all-categories
# mode .build_suggestions() must scope its candidate recipe sub-select by
# subtype_col/subtype_scope, not by the literal "All Categories" value.
# Passing all_categories = FALSE with the identical category value proves
# the branch -- not merely the subtype_col/subtype_scope arguments' mere
# presence -- is what changes the query.
con <- uscogdata:::.ensure_session()
none <- uscogdata:::.build_suggestions(
con, cohort = uscogdata:::.make_cohort("010000226085"), years = 2011L,
category = "All Categories", result = NULL, basis = "harmonized",
flow_prefixes = c("E", "F", "G"),
long_view = "spending_long_harmonized",
all_categories = FALSE,
subtype_col = "spend_subtype",
subtype_scope = c("operations", "capital", "assistance")
)
expect_length(none, 0L)
scoped <- uscogdata:::.build_suggestions(
con, cohort = uscogdata:::.make_cohort("010000226085"), years = 2011L,
category = "All Categories", result = NULL, basis = "harmonized",
flow_prefixes = c("E", "F", "G"),
long_view = "spending_long_harmonized",
all_categories = TRUE,
subtype_col = "spend_subtype",
subtype_scope = c("operations", "capital", "assistance")
)
expect_gt(length(scoped), 0L)
})
+11 -5
View File
@@ -17,7 +17,16 @@
# cog-api's llms.txt, which is silent on units). # cog-api's llms.txt, which is silent on units).
test_that("returned amounts are documented as full US dollars where readers meet the package", { test_that("returned amounts are documented as full US dollars where readers meet the package", {
testthat::skip("Blocked on uscogdata#15 (finding F-004)")
# README and vignettes ship only in the source tree, not in the installed
# package, so these assertions cannot run under R CMD check -- CI's earlier
# testthat::test_local() step is what enforces them. See
# skip_if_no_source_tree() in helper-fixture.R.
docs <- skip_if_no_source_tree(
"README.md",
c("vignettes", "total-spending.Rmd"),
c("vignettes", "population-denominators.Rmd")
)
says_units <- function(path) { says_units <- function(path) {
txt <- paste(readLines(path, warn = FALSE), collapse = " ") txt <- paste(readLines(path, warn = FALSE), collapse = " ")
@@ -25,10 +34,7 @@ test_that("returned amounts are documented as full US dollars where readers meet
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE) grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
} }
expect_true(says_units(testthat::test_path("..", "..", "README.md"))) for (path in docs) expect_true(says_units(path))
expect_true(says_units(testthat::test_path("..", "..", "vignettes", "total-spending.Rmd")))
expect_true(says_units(testthat::test_path("..", "..", "vignettes",
"population-denominators.Rmd")))
# Pin the documented claim to the actual behaviour, so the two cannot drift. # Pin the documented claim to the actual behaviour, so the two cannot drift.
# The expected raw amount is read straight from the corpus's parquet # The expected raw amount is read straight from the corpus's parquet
+539
View File
@@ -0,0 +1,539 @@
test_that("balance views register and carry only balance codes", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)$table_name
expect_true(all(c("balance_long", "balance_annotated") %in% views))
# Every item_code in balance_long is a category_type = 'balance' member.
leak <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM balance_long
WHERE item_code NOT IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance')"
)$n
expect_identical(as.integer(leak), 0L)
# balance_annotated exposes the subtype column the verb groups on.
cols <- DBI::dbGetQuery(con,
"SELECT column_name FROM information_schema.columns
WHERE table_name = 'balance_annotated'"
)$column_name
expect_true(all(c("category", "category_type", "balance_subtype") %in% cols))
})
test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", {
# Every category_type = 'balance' item_code in the bundled fixture has
# is_aggregate = FALSE for every row of every year -- there is no real row
# that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An
# assertion against the live fixture (`WHERE is_aggregate` returns 0) is
# therefore vacuous: it passes identically whether or not the view's
# predicate is present. As with the 22-/23- and 24-/25- tests above, this
# reads the real inst/sql/26-balance_long.sql text off disk and executes it
# -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against
# a synthetic hive-partitioned parquet tree that DOES contain an aggregate
# row under a real balance item_code (W01), so a regression that drops the
# predicate changes which rows survive.
skip_if_no_corpus()
tmp <- withr::local_tempdir()
part_dir <- file.path(tmp, "data", "long", "year=2004")
dir.create(part_dir, recursive = TRUE)
part_path <- file.path(part_dir, "part-0.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives
('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate`
) AS t(canonical_govid, item_code, amt, is_aggregate)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('W01', 'Fund Balances', 'balance', NULL, NULL, 'general')
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
uscogdata:::.render_view_sql(txt, paste0(tmp, "/"))
}
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("26-balance_long.sql"))
rows <- DBI::dbGetQuery(con,
"SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid"
)
expect_equal(nrow(rows), 1L)
expect_equal(rows$canonical_govid, "bal-A")
expect_equal(rows$amt, 100)
})
test_that("balance views are skipped on a corpus without balance_subtype", {
skip_if_no_corpus()
with_corpus_missing_balance_subtype({
con <- cog_open()
on.exit(cog_close())
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)$table_name
# Registration must SKIP them, not error -- an older corpus stays usable.
expect_false(any(c("balance_long", "balance_annotated") %in% views))
expect_true("revenue_long" %in% views)
# ...and calling the verb on such a corpus must hit
# .require_balance_support()'s curated abort (spec § Testing: "Gating"),
# not a DuckDB binder error naming a view that was never registered.
# Asserted on the CLASS: removing the guard still errors, so a bare
# expect_error() would pass on the regression.
expect_error(
cog_balances("550000227544", 2019),
class = "uscogdata_no_balance_support"
)
})
})
test_that("cog_balances returns holdings for a government that has them", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019)
expect_s3_class(r, "tbl_df")
expect_true(nrow(r) > 0L)
expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype",
"category", "amt_nominal") %in% names(r)))
expect_identical(sort(unique(r$category)),
c("Fund Balances", "Insurance Trust Balances"))
expect_false(is.null(attr(r, "provenance")))
expect_identical(attr(r, "provenance")$verb, "cog_balances")
})
})
test_that('category = "Fund Balances" is exactly the general family', {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
expect_identical(unique(r$balance_subtype), "general")
codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
expect_identical(codes, c("W01", "W31", "W61"))
})
})
test_that("no flow code can reach cog_balances", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
# The expected set is read from the RAW corpus, never from the verb --
# verifying an absence through the filter that creates it proves nothing.
# A fresh, direct DuckDB connection against the raw parquet files (never
# cog_open()'s session, never balance_long/balance_annotated) reads
# parquet natively -- no arrow dependency needed (see CLAUDE.md).
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
balance_codes <- DBI::dbGetQuery(con2, sprintf(
"SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'",
uscogdata:::.sql_lit_chr(cats_path)
))$item_code
expect_true(length(got) > 0L)
expect_true(all(got %in% balance_codes))
})
})
test_that("every balance_subtype maps to exactly one category", {
skip_if_no_corpus()
# Dropping the `subtype` argument is only safe while this tree holds. If the
# pipeline ever gives a balance subtype a second category, `category` becomes
# a lossy filter -- fail HERE rather than in a user's analysis. Read via a
# fresh direct DuckDB connection against the raw parquet file, not through
# any registered view.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
b <- DBI::dbGetQuery(con2, sprintf(
"SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'",
uscogdata:::.sql_lit_chr(cats_path)
))
per_subtype <- tapply(b$category, b$balance_subtype,
function(x) length(unique(x)))
expect_true(all(per_subtype == 1L))
})
test_that("cog_balances records found + missing govids in provenance", {
skip_if_no_corpus()
with_fixture_corpus({
suppressMessages(
r <- cog_balances(c("550000227544", "XXXINVALID"), 2019)
)
prov <- attr(r, "provenance")
expect_equal(sort(prov$scope$govids_found), "550000227544")
expect_equal(sort(prov$scope$govids_missing), "XXXINVALID")
})
})
test_that("per_capita divides holdings by population", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_balances("550000227544", 2019, category = "Fund Balances")
pc <- cog_balances("550000227544", 2019, category = "Fund Balances",
per_capita = TRUE)
expect_true("amt_per_capita_nominal" %in% names(pc))
expect_true("pop_source" %in% names(pc))
expect_identical(pc$amt_nominal, plain$amt_nominal)
# Assert against the denominator read from the corpus, NOT against a
# quantity derived from amt_per_capita_nominal itself -- dividing the
# column back out would be tautological and would pass on any value.
pop <- DBI::dbGetQuery(cog_open(), sprintf(
"SELECT population FROM gov_population_yearly
WHERE canonical_govid = %s AND year = 2019",
uscogdata:::.sql_lit_chr("550000227544")
))$population
expect_length(pop, 1L)
expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop,
tolerance = 1e-8)
prov <- attr(pc, "provenance")
expect_true(prov$transformations$per_capita$applied)
})
})
test_that("adjust_to_year adds real dollars", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
adjust_to_year = 2020)
expect_true("amt_real" %in% names(r))
# 2012 dollars inflated to 2020 must exceed nominal.
expect_true(all(r$amt_real > r$amt_nominal))
prov <- attr(r, "provenance")
expect_true(prov$transformations$inflation$applied)
expect_identical(prov$transformations$inflation$base_year, 2020L)
})
})
test_that("per_capita and adjust_to_year compose", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
per_capita = TRUE, adjust_to_year = 2020)
expect_true("amt_per_capita_real" %in% names(r))
# The per-capita column must be deflated by the SAME factor as the level
# column -- this is what the ordering at R/balances.R:101-103 guarantees.
# .attach_real_dollars() silently no-ops on the per-capita leg when
# amt_per_capita_nominal does not exist yet (R/spending.R:664), so
# reversing those two calls drops this column with no error at all.
expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal,
r$amt_real / r$amt_nominal, tolerance = 1e-8)
# And the documented condition is a conjunction: adjust_to_year ALONE
# must not produce amt_per_capita_real (pins the @return wording).
r2 <- cog_balances("550000227544", 2012, category = "Fund Balances",
adjust_to_year = 2020)
expect_true("amt_real" %in% names(r2))
expect_false("amt_per_capita_real" %in% names(r2))
})
})
# --- input validation ------------------------------------------------------
test_that("cog_balances validates its inputs like the money verbs", {
skip_if_no_corpus()
with_fixture_corpus({
G <- "550000227544"
# Pinned to the message, not bare expect_error(): every one of these
# already produces *some* error or *some* quiet wrong answer today --
# years = integer(0) leaks `Parser Error ... AND year IN ()` with the
# generated SQL, recipe = c("a","b") throws "the condition has length > 1",
# and the govid/category cases return 0 rows with no error at all.
expect_error(cog_balances(G, integer(0)), "non-empty integer vector")
expect_error(cog_balances(character(0), 2019), "non-empty character vector")
expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL")
expect_error(cog_balances(G, 2019, recipe = c("a", "b")),
"length-1 character string")
})
})
test_that("validation runs after govid coercion, so a data frame still works", {
skip_if_no_corpus()
with_fixture_corpus({
# .validate_verb_inputs() asserts is.character(govid); it must therefore
# run AFTER .coerce_govid_input(), never before, or the documented
# data-frame input (cog_gov_search() output) would abort.
df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE)
r <- suppressMessages(cog_balances(df, 2019))
expect_true(nrow(r) > 0L)
expect_identical(unique(r$canonical_govid), "550000227544")
})
})
test_that("recipe and category are mutually exclusive", {
skip_if_no_corpus()
with_fixture_corpus({
expect_error(
cog_balances("550000227544", c(2011, 2012),
category = "Fund Balances",
recipe = "cash_securities_z77_wide"),
class = "uscogdata_recipe_category_conflict"
)
})
})
# --- recipe = : the wide-era holdings bridge -------------------------------
test_that("recipe bridges the wide era into the modern one", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012),
recipe = "cash_securities_z77_wide")
# .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait,
# not specific to this recipe -- see the money-verb recipe tests, which
# only ever assert on it with expect_equal), so compare numerically rather
# than with expect_identical()'s type-strict comparison.
expect_equal(sort(r$year), c(2011, 2012))
# The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE
# and therefore invisible to balance_long. If the recipe path ever starts
# filtering aggregates, a 45-year series silently truncates to five --
# this is the regression guard for phase_r_harmonization_review.md § 0.2.
codes <- attr(r, "provenance")$codes_summed$observed
expect_true("X40" %in% codes)
expect_true("Z77" %in% codes)
expect_true(all(r$amt_nominal > 0))
prov <- attr(r, "provenance")
expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide")
})
})
test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", {
skip_if_no_corpus()
with_fixture_corpus({
# 2002 is in the year vector deliberately, and must stay -- do not
# "simplify" this back to c(2011, 2012).
#
# .build_series_break_refs() (R/series_breaks.R, shared with every verb)
# matches breaks with `break_year BETWEEN min(years) AND max(years)`, and
# SB195's break_year is 2002. A c(2011, 2012) span never crosses the
# FY2002 book -> market change -- that whole span sits after it, on one
# consistent basis -- so NOT disclosing SB195 there is correct behaviour,
# not a gap (same reasoning as the "a request that never crosses the
# boundary is not affected by it" comment on .build_corpus_break_refs()).
#
# The property actually worth testing is: a recipe query that observes
# X40 AND spans FY2002 discloses SB195. This fixture has no 2002
# partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020
# partitions exist), so including 2002 in `years` widens the
# break-matching window without changing which rows the recipe join
# returns -- verified empirically: r$year below is exactly {2011, 2012}
# whether or not 2002 is in the request (see task-4-report.md).
# Removing 2002 would silently turn this back into the non-crossing case
# above and destroy the test's purpose.
r <- cog_balances("550000227544", c(2002, 2011, 2012),
recipe = "cash_securities_z77_wide")
expect_equal(sort(r$year), c(2011, 2012))
refs <- attr(r, "provenance")$series_break_refs
# SB195 sits on fin_code X40; it can only fire where X40 is observed,
# which is exactly the recipe path.
expect_true("SB195" %in% refs)
})
})
test_that("the second holdings bridge works too", {
skip_if_no_corpus()
with_fixture_corpus({
# X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011
# and Z78 in 2012, so both legs are exercised.
r <- cog_balances("550000227544", c(2011, 2012),
recipe = "cash_securities_z78_wide")
codes <- attr(r, "provenance")$codes_summed$observed
expect_true(all(c("X41", "Z78") %in% codes))
expect_equal(sort(r$year), c(2011, 2012))
})
})
test_that("an unknown recipe id is rejected", {
skip_if_no_corpus()
with_fixture_corpus({
# Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84).
# Without it the test is non-discriminating: deleting the validation call
# leaves .recipe_components() returning 0 rows and comps$label[[1]]
# throwing "subscript out of bounds", which a bare expect_error() accepts
# while the user loses the curated "valid recipe ids are ..." message.
expect_error(
cog_balances("550000227544", 2019, recipe = "no_such_recipe"),
class = "uscogdata_unknown_recipe"
)
})
})
# --- balance_caveats: GAAP disclosure + measured coverage windows ----------
test_that("balance_caveats is always present and flags the GAAP distinction", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019)
cav <- attr(r, "provenance")$balance_caveats
expect_false(is.null(cav))
expect_true(cav$not_gaap)
})
})
test_that("coverage_window is computed from the corpus, not hardcoded", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
cav <- attr(r, "provenance")$balance_caveats
# Read the "general" family's true year extent independently, via a
# fresh DuckDB connection against the raw parquet files (never through
# balance_long/.balance_caveats() itself, and never via arrow -- this
# package reads parquet through DuckDB only, see CLAUDE.md). Replicates
# the same predicates 26-balance_long.sql applies (category_type =
# 'balance', NOT is_aggregate) so this is a faithful, independent
# measurement rather than a re-statement of the view under test.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet")
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
obs <- DBI::dbGetQuery(con2, sprintf(
"SELECT MIN(l.year) AS y0, MAX(l.year) AS y1
FROM read_parquet(%s, hive_partitioning = true) l
JOIN read_parquet(%s) c USING (item_code)
WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate",
uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path)
))
expect_identical(as.integer(cav$coverage_window$general),
c(as.integer(obs$y0), as.integer(obs$y1)))
})
})
test_that("coverage_window covers every corpus subtype, not just observed ones", {
skip_if_no_corpus()
with_fixture_corpus({
# Deliberate contract (provenance-v1.json): the window block is corpus-
# scoped so a caller can ask "is there a family I missed?", while
# `truncated` is the observed-scoped field. A single-category query must
# therefore still report every balance family in the mounted corpus.
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
expect_identical(unique(r$balance_subtype), "general")
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
all_subtypes <- DBI::dbGetQuery(con2, sprintf(
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
WHERE balance_subtype IS NOT NULL",
uscogdata:::.sql_lit_chr(cats_path)
))$balance_subtype
cav <- attr(r, "provenance")$balance_caveats
expect_setequal(names(cav$coverage_window), all_subtypes)
expect_true(length(all_subtypes) > 1L)
# ...while `truncated` stays scoped to what this query actually observed.
expect_true(all(cav$truncated %in% unique(r$balance_subtype)))
})
})
test_that("the corpus-constant coverage windows are memoised per session", {
skip_if_no_corpus()
with_fixture_corpus({
# The windows query has no govid/year predicate: its answer depends only
# on which corpus is mounted, so re-running the full balance_long scan on
# every call is pure waste (35% of verb runtime on the fixture). Same
# memoise-and-invalidate pattern as .uscogdata_env$manifest.
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
suppressMessages(cog_balances("550000227544", 2019))
memo <- uscogdata:::.uscogdata_env$balance_coverage_windows
expect_false(is.null(memo))
expect_true("general" %in% names(memo))
uscogdata:::cog_close()
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
})
})
test_that("a request past a family's coverage window is flagged", {
skip_if_no_corpus()
with_fixture_corpus({
# employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in
# the LIVE corpus -- Census moved employee retirement reporting to the
# Annual Survey of Public Pensions after that year. This bundled FIXTURE
# doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present),
# so the family's *observed* max here is 2012, not 2016. Either way the
# requested span (2012, 2019) reaches past what the family covers in
# THIS corpus, which is what makes .balance_caveats() flag it -- the
# assertion below is about the fixture's measured window, not the FY2016
# live-corpus cutoff.
r <- cog_balances("550000227544", c(2012, 2019))
cav <- attr(r, "provenance")$balance_caveats
expect_true("employee_retirement" %in% cav$truncated)
})
})
test_that("the provenance schema documents balance_caveats", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE
)
expect_true("balance_caveats" %in% names(sch$properties))
})
test_that("cog_explain surfaces the balance caveats", {
skip_if_no_corpus()
with_fixture_corpus({
# Asserted on the RENDERED text, not on prov$balance_caveats: the field
# is already covered above, and the once-per-session cli_inform() means
# cog_explain() is the only surface a caller who missed (or suppressed)
# the first message can still audit.
r <- suppressMessages(cog_balances("550000227544", c(2012, 2019)))
# Both streams: cli routes most of its output through conditions that
# land on stderr, so a stdout-only capture would be empty (the pattern
# used throughout test-explain.R).
out <- paste(c(capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")),
collapse = "\n")
expect_match(out, "GAAP")
expect_match(out, "employee_retirement")
})
})
test_that("cog_explain on a money-verb result has no balance caveat section", {
skip_if_no_corpus()
with_fixture_corpus({
r <- suppressMessages(cog_spending("550000227544", 2019))
out <- paste(c(capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")),
collapse = "\n")
# Guard against the capture itself being vacuous: the section must be
# absent from output that demonstrably contains the rest of the report.
expect_match(out, "Data vintage")
expect_false(grepl("GAAP", out))
})
})
test_that("the caveat message fires once per session", {
skip_if_no_corpus()
with_fixture_corpus({
expect_message(cog_balances("550000227544", 2019), "not.*GAAP")
expect_no_message(cog_balances("550000227544", 2020))
})
})
+71 -4
View File
@@ -8,14 +8,32 @@ test_that("cog_categories returns all categories grouped by subtype", {
expect_gt(nrow(r), 10L) expect_gt(nrow(r), 10L)
# corpus preserves Census-native "expenditure" vocabulary; the API takes # corpus preserves Census-native "expenditure" vocabulary; the API takes
# "spending" as a friendlier alias. # "spending" as a friendlier alias.
expect_setequal(unique(r$category_type), c("expenditure", "revenue")) #
# `balance` joined as a third category_type with the cash-and-security
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
# money verb, so it surfaces every category_type the corpus carries -- the
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
# return a balance row.
expect_setequal(unique(r$category_type),
c("expenditure", "revenue", "balance"))
}) })
test_that("cog_categories(type = 'spending') returns only expenditure rows", { test_that("cog_categories(type = 'spending') returns only expenditure rows", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories(type = "spending") r <- cog_categories(type = "spending")
expect_true(all(r$category_type == "expenditure")) expect_true(all(r$category_type == "expenditure"))
expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental"))) # "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
# carrying dollars now maps to a category.
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
# joined with the I/Q/Y flow batch -- the last two characters of Census's
# expenditure taxonomy. `interest` is what makes the three-concept model
# computable: primary = direct minus debt service.
# Exclude pseudo-category which has NA for subtype
r_crosswalk <- r[r$category != "All Categories", ]
expect_true(all(r_crosswalk$subtype %in%
c("operations", "capital", "intergovernmental", "assistance",
"interest", "insurance_benefits")))
}) })
test_that("cog_categories surfaces the intergovernmental spending subtype", { test_that("cog_categories surfaces the intergovernmental spending subtype", {
@@ -33,8 +51,16 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories(type = "revenue") r <- cog_categories(type = "revenue")
expect_true(all(r$category_type == "revenue")) expect_true(all(r$category_type == "revenue"))
expect_true(all(r$subtype %in% # The four non-general subtypes are deliberately NOT own_source: Census's
c("own_source", "federal", "state", "local_aid"))) # General Revenue excludes insurance trust (Y01 alone is $1.31T corpus-wide,
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
# (A90) revenue by definition, which is what makes both of its published
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
# Exclude pseudo-category which has NA for subtype
r_crosswalk <- r[r$category != "All Categories", ]
expect_true(all(r_crosswalk$subtype %in%
c("own_source", "federal", "state", "local_aid",
"insurance_trust", "utility", "liquor_store")))
}) })
test_that("cog_categories(pattern = ...) filters case-insensitively", { test_that("cog_categories(pattern = ...) filters case-insensitively", {
@@ -47,6 +73,8 @@ test_that("cog_categories(pattern = ...) filters case-insensitively", {
test_that("cog_categories has one row per (category, subtype)", { test_that("cog_categories has one row per (category, subtype)", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories() r <- cog_categories()
# Exclude pseudo-category which is not a crosswalk entry
r <- r[r$category != "All Categories", ]
key <- paste(r$category, r$subtype, sep = "|") key <- paste(r$category, r$subtype, sep = "|")
expect_equal(length(key), length(unique(key))) expect_equal(length(key), length(unique(key)))
}) })
@@ -54,6 +82,8 @@ test_that("cog_categories has one row per (category, subtype)", {
test_that("cog_categories item_codes is non-empty comma-separated string", { test_that("cog_categories item_codes is non-empty comma-separated string", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories() r <- cog_categories()
# Exclude pseudo-category which has NA for n_codes and item_codes
r <- r[r$category != "All Categories", ]
expect_true(all(nzchar(r$item_codes))) expect_true(all(nzchar(r$item_codes)))
expect_true(all(r$n_codes >= 1L)) expect_true(all(r$n_codes >= 1L))
# n_codes should equal count of commas + 1 # n_codes should equal count of commas + 1
@@ -71,3 +101,40 @@ test_that("cog_categories sorted by category_type, category, subtype", {
test_that("cog_categories rejects invalid type", { test_that("cog_categories rejects invalid type", {
expect_error(cog_categories(type = "both"), "type") expect_error(cog_categories(type = "both"), "type")
}) })
test_that("cog_categories() surfaces balance subtypes", {
skip_if_no_corpus()
with_fixture_corpus({
cc <- cog_categories()
b <- cc[cc$category_type == "balance", ]
expect_true(nrow(b) > 0L)
# Every balance row must carry its subtype. Before the COALESCE included
# balance_subtype these were all NA, which silently made the balance
# taxonomy undiscoverable -- cog-api derives its subtype vocabulary from
# this function, so an NA here becomes an unusable API parameter.
expect_false(any(is.na(b$subtype)))
# The exact set, read independently from the crosswalk rather than from
# the function under test.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
p <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
want <- DBI::dbGetQuery(con2, sprintf(
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
WHERE category_type = 'balance' AND balance_subtype IS NOT NULL
ORDER BY 1", uscogdata:::.sql_lit_chr(p)))$balance_subtype
expect_true(length(want) > 1L)
expect_identical(sort(unique(b$subtype)), sort(want))
})
})
test_that('cog_categories(type = "balance") filters to holdings', {
skip_if_no_corpus()
with_fixture_corpus({
b <- cog_categories(type = "balance")
expect_true(nrow(b) > 0L)
expect_identical(unique(b$category_type), "balance")
expect_false(any(is.na(b$subtype)))
})
})
+218
View File
@@ -0,0 +1,218 @@
# The money verbs accepting a cohort by predicate (uscogdata#58).
#
# The load-bearing property is EQUIVALENCE: naming the same set of governments
# by id and by state/type must return the same rows. Everything else here is
# about the ways that equivalence could silently break -- the postal/FIPS
# translation, the intersection rule, and provenance no longer having an id
# list to describe.
#
# Fixture cohorts used: RI (fips 44) cities = 8 governments, DE (fips 10)
# counties = 3. Small on purpose; the size of the win is measured against the
# production corpus, not here.
# --- equivalence -----------------------------------------------------------
test_that("a predicate cohort returns exactly what the same ids return", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "RI", type = "city")$canonical_govid
expect_gt(length(ids), 1L)
by_id <- cog_spending(govid = ids, years = 2019)
by_pred <- cog_spending(years = 2019, state = "RI", type = "city")
# Compare the data itself, ignoring the provenance attribute -- which is
# SUPPOSED to differ (see the scope tests below).
expect_equal(
as.data.frame(by_id[order(by_id$canonical_govid, by_id$category), ]),
as.data.frame(by_pred[order(by_pred$canonical_govid, by_pred$category), ]),
ignore_attr = TRUE
)
expect_gt(nrow(by_pred), 0L)
})
})
test_that("equivalence holds for cog_revenue()", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "DE", type = "county")$canonical_govid
by_id <- cog_revenue(govid = ids, years = 2019)
by_pred <- cog_revenue(years = 2019, state = "DE", type = "county")
expect_equal(nrow(by_id), nrow(by_pred))
expect_equal(sum(by_id$amt_nominal), sum(by_pred$amt_nominal))
})
})
test_that("equivalence holds for cog_balances()", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "DE", type = "county")$canonical_govid
by_id <- cog_balances(govid = ids, years = 2019)
by_pred <- cog_balances(years = 2019, state = "DE", type = "county")
expect_equal(nrow(by_id), nrow(by_pred))
expect_equal(sum(by_id$amt_nominal), sum(by_pred$amt_nominal))
})
})
test_that("equivalence survives per_capita, adjust_to_year and pagination", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "RI", type = "city")$canonical_govid
# per_capita now keys its population lookup on the rows in the result
# rather than on the requested cohort; these must stay identical.
by_id <- cog_spending(govid = ids, years = 2019, per_capita = TRUE,
adjust_to_year = 2020)
by_pred <- cog_spending(years = 2019, state = "RI", type = "city",
per_capita = TRUE, adjust_to_year = 2020)
expect_equal(by_id$amt_per_capita_nominal, by_pred$amt_per_capita_nominal)
expect_equal(by_id$amt_per_capita_real, by_pred$amt_per_capita_real)
expect_equal(by_id$pop_source, by_pred$pop_source)
paged_id <- cog_spending(govid = ids, years = 2019, limit = 5, offset = 5)
paged_pred <- cog_spending(years = 2019, state = "RI", type = "city",
limit = 5, offset = 5)
expect_equal(as.data.frame(paged_id), as.data.frame(paged_pred),
ignore_attr = TRUE)
expect_identical(attr(paged_id, "total_rows"), attr(paged_pred, "total_rows"))
})
})
test_that("a state-only predicate spans every type in that state", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "DE", type = NULL)$canonical_govid
by_id <- cog_spending(govid = ids, years = 2019)
by_pred <- cog_spending(years = 2019, state = "DE")
expect_equal(nrow(by_id), nrow(by_pred))
})
})
# --- the postal/FIPS trap --------------------------------------------------
test_that("a postal abbreviation resolves to rows, not to silence", {
skip_if_no_corpus()
with_fixture_corpus({
# canonical_fips_xwalk.fips_state holds "44", not "RI". A predicate built
# from the raw parameter matches nothing and returns an empty result that
# reads as "these governments reported nothing" -- the exact trap cog-api
# hit. A zero-row result here is the regression.
r <- cog_spending(years = 2019, state = "RI", type = "city")
expect_gt(nrow(r), 0L)
# And the FIPS form is accepted as the same cohort.
expect_equal(nrow(cog_spending(years = 2019, state = "44", type = "city")),
nrow(r))
})
})
test_that("an unknown state abbreviation aborts with a message that names the problem", {
skip_if_no_corpus()
with_fixture_corpus({
# Regression: `.state_abbrev_to_fips` is a named character vector, so `[[`
# on an absent name threw base R's "subscript out of bounds" and the
# curated message was unreachable.
expect_error(cog_spending(years = 2019, state = "ZZ"),
class = "uscogdata_unknown_state")
expect_error(cog_spending(years = 2019, state = "ZZ"),
"Unknown state abbreviation")
})
})
# --- naming the cohort -----------------------------------------------------
test_that("naming no cohort at all is refused", {
skip_if_no_corpus()
with_fixture_corpus({
expect_error(cog_spending(years = 2019), class = "uscogdata_no_cohort")
expect_error(cog_revenue(years = 2019), class = "uscogdata_no_cohort")
expect_error(cog_balances(years = 2019), class = "uscogdata_no_cohort")
})
})
test_that("an empty govid vector still fails as it always did", {
skip_if_no_corpus()
with_fixture_corpus({
# Supplied-but-empty is a caller error, not "cohort named some other way".
expect_error(cog_spending(character(0), 2019),
"must be a non-empty character vector")
})
})
test_that("govid and state/type together intersect", {
skip_if_no_corpus()
with_fixture_corpus({
cities <- cog_gov_search(NULL, state = "RI", type = "city")$canonical_govid
counties <- cog_gov_search(NULL, state = "DE", type = "county")$canonical_govid
# The documented rule: the governments in `govid` that ALSO match the
# predicate -- never one silently taking precedence over the other.
both <- cog_spending(govid = c(cities, counties), years = 2019,
state = "RI", type = "city")
only <- cog_spending(govid = cities, years = 2019)
expect_equal(as.data.frame(both), as.data.frame(only), ignore_attr = TRUE)
# A disjoint intersection is empty, not "whichever one won".
none <- cog_spending(govid = counties, years = 2019,
state = "RI", type = "city")
expect_equal(nrow(none), 0L)
})
})
# --- provenance ------------------------------------------------------------
test_that("a govid cohort's provenance scope is untouched", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "DE", type = "county")$canonical_govid
prov <- attr(cog_spending(govid = ids, years = 2019), "provenance")
expect_setequal(prov$scope$govids_found, ids)
expect_length(prov$scope$govids_missing, 0L)
# No cohort block: govids_found already describes this cohort exactly.
expect_null(prov$scope$cohort)
})
})
test_that("a predicate cohort describes itself instead of listing ids", {
skip_if_no_corpus()
with_fixture_corpus({
prov <- attr(cog_spending(years = 2019, state = "DE", type = "county"),
"provenance")
# Deliberately NOT the resolved id list: a fleet-scale cohort would put
# 20,000 govids into every response body.
expect_length(prov$scope$govids_found, 0L)
expect_length(prov$scope$govids_missing, 0L)
expect_identical(prov$scope$cohort$state, "DE")
expect_identical(prov$scope$cohort$type, "county")
expect_identical(prov$scope$cohort$n_governments, 3L)
})
})
test_that("the cohort block counts the intersection, not the predicate alone", {
skip_if_no_corpus()
with_fixture_corpus({
counties <- cog_gov_search(NULL, state = "DE", type = "county")$canonical_govid
prov <- attr(
cog_spending(govid = counties[1], years = 2019, state = "DE", type = "county"),
"provenance"
)
expect_identical(prov$scope$cohort$n_governments, 1L)
})
})
# --- the SQL actually changed ----------------------------------------------
test_that("a predicate cohort never renders the ids into the query", {
skip_if_no_corpus()
with_fixture_corpus({
ids <- cog_gov_search(NULL, state = "RI", type = "city")$canonical_govid
prov <- attr(cog_spending(years = 2019, state = "RI", type = "city"),
"provenance")
sql <- prov$sql_query %||% prov$sql
skip_if(is.null(sql), "provenance carries no SQL for this verb")
# The whole point: cohort size does not enter the SQL string.
for (id in ids) expect_false(grepl(id, sql, fixed = TRUE))
expect_match(sql, "SELECT canonical_govid FROM canonical_fips_xwalk",
fixed = TRUE)
})
})
+83
View File
@@ -0,0 +1,83 @@
# A cohort is how every verb names the set of governments it queries. It can be
# named by explicit id, by a predicate over canonical_fips_xwalk, or by both
# (intersection). These tests cover the SQL construction itself -- pure string
# building, no corpus needed -- because that is where the postal/FIPS trap and
# the 40k-literal blowup both live.
test_that(".make_cohort() keeps an explicit id vector as ids", {
ch <- .make_cohort(govid = c("550000227544", "060000000001"))
expect_identical(ch$ids, c("550000227544", "060000000001"))
expect_null(ch$state_fips)
expect_null(ch$type_int)
expect_false(.cohort_by_predicate(ch))
})
test_that(".make_cohort() translates a postal abbreviation to FIPS", {
# The trap this whole issue exists to avoid: canonical_fips_xwalk.fips_state
# holds "55", not "WI". A predicate written against the raw parameter matches
# nothing and returns an empty result that reads as "reported nothing".
ch <- .make_cohort(state = "WI")
expect_identical(ch$state_fips, "55")
expect_true(.cohort_by_predicate(ch))
})
test_that(".make_cohort() translates a type label to its integer code", {
expect_identical(.make_cohort(type = "city")$type_int, 2L)
expect_identical(.make_cohort(type = "state")$type_int, 0L)
expect_identical(.make_cohort(type = 1)$type_int, 1L)
})
test_that(".make_cohort() reuses the search verb's coercers for invalid input", {
expect_error(.make_cohort(state = "ZZ"), "Unknown state abbreviation")
expect_error(.make_cohort(type = "special_district"), "Unknown type")
# Out-of-scope types (4 = special district, 5 = school district) are refused
# by the numeric branch, with the v0.1-scope message.
expect_error(.make_cohort(type = 4), "type must be 0, 1, 2, or 3")
})
test_that(".make_cohort() rejects naming no cohort at all", {
expect_error(.make_cohort(), class = "uscogdata_no_cohort")
})
test_that(".cohort_sql() renders an id cohort as a literal IN list", {
sql <- .cohort_sql(.make_cohort(govid = c("a", "b")))
expect_identical(sql, "canonical_govid IN ('a','b')")
})
test_that(".cohort_sql() renders a predicate cohort as an xwalk subquery", {
# The point of the issue: the cohort never becomes a literal list, so its
# size does not enter the SQL string at all.
sql <- .cohort_sql(.make_cohort(state = "WI", type = "city"))
expect_match(sql, "SELECT canonical_govid FROM canonical_fips_xwalk", fixed = TRUE)
expect_match(sql, "fips_state = '55'", fixed = TRUE)
expect_match(sql, "govs_type = 2", fixed = TRUE)
expect_false(grepl("'WI'", sql, fixed = TRUE))
})
test_that(".cohort_sql() renders ids and a predicate as an intersection", {
sql <- .cohort_sql(.make_cohort(govid = c("a", "b"), type = "city"))
expect_match(sql, "canonical_govid IN ('a','b')", fixed = TRUE)
expect_match(sql, "AND canonical_govid IN (SELECT", fixed = TRUE)
})
test_that(".cohort_sql() honours a column alias", {
# Several call sites join the xwalk under an alias (`l.`, `x.`, `v.`), so the
# predicate has to be able to name the qualified column.
sql <- .cohort_sql(.make_cohort(state = "WI"), col = "l.canonical_govid")
expect_match(sql, "l.canonical_govid IN (SELECT", fixed = TRUE)
# The subquery's own column stays unqualified -- it selects from the xwalk,
# not from the outer relation.
expect_match(sql, "SELECT canonical_govid FROM", fixed = TRUE)
})
test_that(".cohort_sql() escapes quotes in ids", {
sql <- .cohort_sql(.make_cohort(govid = "o'brien"))
expect_match(sql, "'o''brien'", fixed = TRUE)
})
test_that("a predicate cohort's SQL does not grow with cohort size", {
# The regression this guards: 20,106 ids rendered to a 301,591-character
# IN list, embedded in 5-8 statements per call.
wide <- .cohort_sql(.make_cohort(state = "CA", type = "city"))
expect_lt(nchar(wide), 200L)
})
+193
View File
@@ -0,0 +1,193 @@
# tests/testthat/test-complete.R
#
# uscogdata#18. The published corpus no longer stores the wide era's explicit
# zeros (cog_pipeline#64, series break SB194), so absence means two different
# things:
#
# <= FY2011 (dense_source) : cell absent => Census published $0
# >= FY2012 (sparse_source): cell absent => not reported, unknown
#
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
# row's `value_source` so the two are distinguishable. Expected row sets here
# are built from the corpus parquet directly, never from the verb under test --
# verifying what a filter does through that same filter proves nothing.
# The (subtype, category) cells that SHOULD exist for one government-year:
# every code in force for that government's type, mapped through
# summary_categories, matching the verb's crosswalk subtype scope (the
# default concept, `primary`, is operations/capital/assistance -- see
# uscogdata#11) and excluding aggregate-flagged codes (which
# spending_long/revenue_long drop).
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
wt_raw_query(sprintf(
"SELECT DISTINCT c.%s AS subtype, c.category
FROM %s cs
JOIN %s x ON x.govs_type = cs.type
JOIN %s c ON c.item_code = cs.item_code
WHERE x.canonical_govid = '%s'
AND cs.year = %d
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%s IN (%s)",
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
q("summary_categories.parquet"), govid, year,
subtype_col, paste0("'", subtypes, "'", collapse = ",")
))
}
# The default expenditure concept's subtype scope, mirrored from
# R/spending.R's .spend_subtypes_primary.
primary_subtypes <- c("operations", "capital", "assistance")
test_that("complete = FALSE is the default and changes nothing", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
expect_equal(nrow(plain), nrow(explicit))
expect_false("value_source" %in% names(plain))
})
})
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2011 is dense_source: before sparsification this government carried a
# row for every code in force, most of them $0. complete = TRUE must
# reproduce that cell set exactly.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$spend_subtype, r$category),
key(expected$subtype, expected$category))
expect_gt(nrow(expected), 0L)
# Every filled cell in a dense-source year is a Census-published $0 --
# never "unknown", which is what the modern era's absences mean.
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
})
})
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
full <- cog_spending("121011212191", 2011L, complete = TRUE)
# Filling adds rows; it must never alter or drop one.
expect_gt(nrow(full), nrow(plain))
reported <- full[full$value_source == "reported", ]
expect_equal(nrow(reported), nrow(plain))
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
# ... and the total is unchanged, because every added cell is $0.
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
})
})
test_that("a sparse-source year's absences are unknown, not zero", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2019 is sparse_source: an absent cell means the government did not
# report, which is NOT a zero. Filling those with 0 would invent data --
# the exact error the representation contract exists to prevent.
r <- cog_spending("121011212191", 2019L, complete = TRUE)
filled <- r[r$value_source != "reported", ]
expect_gt(nrow(filled), 0L)
expect_true(all(filled$value_source == "not_reported"))
expect_true(all(is.na(filled$amt_nominal)))
expect_false(any(r$value_source == "census_zero"))
})
})
test_that("the fill is scoped to each government's own type", {
skip_if_no_corpus()
with_fixture_corpus({
# Filling against the union of all types would invent cells for codes a
# county can never report. Every filled category must be one that
# code_set puts in force for type 1 (county) specifically.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
county_cells <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
expect_true(all(r$category %in% county_cells$category))
})
})
test_that("complete = TRUE respects the category filter", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011L, category = "Police",
complete = TRUE)
expect_true(all(r$category == "Police"))
expect_true("value_source" %in% names(r))
})
})
test_that("cog_revenue() completes on its own flow", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
c("own_source", "federal", "state", "local_aid"),
"revenue_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$revenue_subtype, r$category),
key(expected$subtype, expected$category))
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
})
})
test_that("provenance records the completion and its absence rule", {
skip_if_no_corpus()
with_fixture_corpus({
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
"provenance")
expect_true(prov$completion$applied)
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
expect_gt(prov$completion$rows_filled, 0L)
off <- attr(cog_spending("121011212191", 2011L), "provenance")
expect_false(off$completion$applied)
expect_equal(off$completion$rows_filled, 0L)
})
})
test_that("complete = TRUE is refused where the fill would be guesswork", {
skip_if_no_corpus()
with_fixture_corpus({
# A recipe defines its own component codes and does not go through
# summary_categories at all, so there is no grid to fill from.
expect_error(
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
# The intergovernmental leg keeps aggregate rows by design
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
expect_error(
cog_spending("121011212191", 2011L, expenditure_concept = "total",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
})
})
test_that("complete = TRUE aborts on a corpus with no representation contract", {
skip_if_no_corpus()
# A corpus published before sparsification carries neither table, so there
# is nothing to fill from and no rule saying what an absence means. That
# must abort rather than guess.
with_corpus_missing_representation({
expect_error(
cog_spending("121011212191", 2011L, complete = TRUE),
class = "uscogdata_representation_unavailable"
)
# ... while an ordinary query on the same corpus still works.
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
})
})
+106
View File
@@ -64,3 +64,109 @@ test_that(".resolve_url does not invent a slash for an empty setting", {
withr::local_options(uscogdata.url = "") withr::local_options(uscogdata.url = "")
expect_equal(.resolve_url(), "") expect_equal(.resolve_url(), "")
}) })
test_that("the default corpus URL is real, not a placeholder", {
# setup.R points USCOGDATA_URL at the bundled fixture for the whole suite,
# so both the env var and the option have to be cleared to see the default.
withr::local_envvar(USCOGDATA_URL = NA)
withr::local_options(uscogdata.url = NULL)
url <- .resolve_url()
expect_false(grepl("REPLACE_WITH", url, fixed = TRUE))
expect_match(url, "^https://")
expect_match(url, "/$")
})
test_that("an explicitly-set sentinel URL still aborts", {
# The guard must survive the default change: a user who half-edited a
# copied config still gets the actionable error.
withr::local_envvar(
USCOGDATA_URL = "https://other.example/s/REPLACE_WITH_SHARE_TOKEN/x/"
)
expect_error(
.check_url_configured(.resolve_url()),
class = "uscogdata_url_not_configured"
)
})
test_that("DESCRIPTION carries release metadata", {
skip_if_no_source_tree("DESCRIPTION")
d <- read.dcf(source_tree_path("DESCRIPTION"))
fields <- colnames(d)
expect_true(all(c("URL", "BugReports") %in% fields))
expect_match(d[1, "Authors@R"], "Knowles", fixed = TRUE)
expect_match(d[1, "Authors@R"], "0000-0003-0005-9478", fixed = TRUE)
expect_match(d[1, "Authors@R"], "Civilytics Consulting LLC", fixed = TRUE)
# The gate in .validate_schema() accepts up to 7 and the published corpus
# IS 7; DESCRIPTION must not claim otherwise.
expect_equal(as.integer(d[1, "MaxCorpusSchema"]), 7L)
# Authors@R must actually parse -- a malformed person() call is only
# caught at citation()/build time otherwise.
people <- eval(parse(text = d[1, "Authors@R"]))
expect_s3_class(people, "person")
expect_true("cre" %in% unlist(lapply(people, function(p) p$role)))
})
test_that("LICENSE and LICENSE.md name the same copyright holder", {
skip_if_no_source_tree("LICENSE", "LICENSE.md")
holder <- sub("^COPYRIGHT HOLDER:\\s*", "",
grep("^COPYRIGHT HOLDER:", readLines(source_tree_path("LICENSE"),
warn = FALSE), value = TRUE))
full <- paste(readLines(source_tree_path("LICENSE.md"), warn = FALSE), collapse = "\n")
expect_equal(holder, "Civilytics Consulting LLC")
expect_match(full, holder, fixed = TRUE)
# usethis::use_mit_license() writes LICENSE.md but leaves an existing
# LICENSE alone, which is how the two came to disagree in the first place.
expect_match(full, "MIT License", fixed = TRUE)
})
test_that("vignettes are not excluded from the build", {
skip_if_no_source_tree(".Rbuildignore")
ignore <- readLines(source_tree_path(".Rbuildignore"), warn = FALSE)
expect_false(any(grepl("^\\^vignettes\\$$", ignore)))
# The fixture is what lets R CMD check run offline with no credentials on
# r-universe and GitHub Actions. It must never be excluded.
expect_false(any(grepl("fixture_corpus", ignore, fixed = TRUE)))
# doc/ and Meta/ ARE build artefacts of devtools::build_vignettes() and must
# stay excluded -- R CMD build regenerates inst/doc/ from vignettes/ on its
# own, and leaving them in earns a "non-standard file at top level" NOTE.
expect_true(any(grepl("^\\^doc\\$$", ignore)))
expect_true(any(grepl("^\\^Meta\\$$", ignore)))
})
test_that("_pkgdown.yml indexes every exported topic", {
skip_if_no_source_tree("_pkgdown.yml", "NAMESPACE")
exports <- grep("^export\\(", readLines(source_tree_path("NAMESPACE"), warn = FALSE),
value = TRUE)
exports <- sub("^export\\((.*)\\)$", "\\1", exports)
yml <- paste(readLines(source_tree_path("_pkgdown.yml"), warn = FALSE), collapse = "\n")
missing <- exports[!vapply(exports,
function(e) grepl(paste0("\\b", e, "\\b"), yml),
logical(1))]
# pkgdown errors on topics missing from the index, so an unlisted export
# means the docs site does not build at all.
expect_equal(missing, character(0))
})
test_that("README is written for a stranger, not a repo insider", {
skip_if_no_source_tree("README.md")
r <- paste(readLines(source_tree_path("README.md"), warn = FALSE), collapse = "\n")
# No paths that only resolve inside a maintainer's checkout.
expect_false(grepl("../cog_pipeline", r, fixed = TRUE))
# A real, uncommented install line.
expect_match(r, "install.packages", fixed = TRUE)
expect_false(grepl("# pak::pkg_install", r, fixed = TRUE))
# The errata most likely to produce a plausible-looking wrong answer.
expect_match(r, "full US dollars", fixed = TRUE)
# The release advice that conflicts with public CI is gone.
expect_false(grepl("Rbuildignore", r, fixed = TRUE))
# Both read paths documented.
expect_match(r, "cog_mirror", fixed = TRUE)
# cog_spending() has no default for `years`; a quickstart that omits it
# errors on the reader's first call.
expect_match(r, "years\\s*=", perl = TRUE)
})
+94
View File
@@ -0,0 +1,94 @@
# tests/testthat/test-corpus-breaks.R
#
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
# are caveats about the corpus itself rather than about one item code:
#
# SB085 1977 dollar precision across the 1976/1977 boundary
# SB087 2002 imputation exclusion FY2002-2006
# SB194 2012 dense -> sparse representation change
# SB086 2017 government ID scheme change
#
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
# and no row's item_code is ever the literal "ALL", so none of them could
# ever reach a user. They now travel in their own provenance field,
# `corpus_break_refs`, which keeps them distinguishable from the
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
# result, not one series).
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
skip_if_no_corpus()
with_fixture_corpus({
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
# DoD 4 intended to reach users.
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_true("SB194" %in% prov$corpus_break_refs)
})
})
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
skip_if_no_corpus()
with_fixture_corpus({
# 2019-2020 spans no catalogued corpus-wide break.
r <- cog_spending("121011212191", 2019:2020, "Police")
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
})
})
test_that("corpus_break_refs and series_break_refs stay disjoint", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_type(prov$series_break_refs, "character")
expect_type(prov$corpus_break_refs, "character")
# An ALL caveat must never masquerade as a break in a specific series.
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
expect_false("SB194" %in% prov$series_break_refs)
})
})
test_that(".build_corpus_break_refs matches on the break_year window alone", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
# series_breaks table is a full cross-vintage registry, so the matching
# logic is testable there even though no long partition covers it.
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1975:1980, schema_version = 6L
))
# ... and does not fire for a range that misses it, unlike a filter keyed
# on the era rather than the boundary.
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1978:1980, schema_version = 6L
))
# Unlike code-specific refs, these do not depend on which codes a result
# happens to contain -- that dependency is the whole defect.
expect_setequal(
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
"SB087"
)
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
expect_equal(
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
character(0)
)
})
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
out <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
expect_match(out, "SB194", fixed = TRUE)
})
})
+14 -3
View File
@@ -30,7 +30,6 @@ wt_coverage <- function(x) {
} }
test_that("multi-government aggregates disclose reporting coverage on every result", { test_that("multi-government aggregates disclose reporting coverage on every result", {
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
# -- F-020: geographic rollups ------------------------------------------- # -- F-020: geographic rollups -------------------------------------------
# Wisconsin's city/village universe is 608 governments. On the bundled # Wisconsin's city/village universe is 608 governments. On the bundled
@@ -49,10 +48,22 @@ test_that("multi-government aggregates disclose reporting coverage on every resu
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L)) expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE)) expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
# Cross-check against the raw partitions, scoped to the SAME universe the
# rollup was given -- the 608 govids above. Scoping instead on the long
# table's own `type`/`fips_state` asks a different question and answers 595:
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
# identity, when they were townships) while the xwalk lists them as
# govs_type = 2 (their present identity, as villages). Schema v6 made the
# long table's geography present-harmonized and moved as-of-year to the
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
# in R/manifest.R. n_units_reporting counts against the requested universe,
# so 597 is the number that answers "how many of the governments I asked
# about reported".
raw_2012 <- wt_raw_query(paste0( raw_2012 <- wt_raw_query(paste0(
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ", "SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ", "WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate")) "AND canonical_govid IN (",
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]])) expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
# -- F-023: peer cohorts -------------------------------------------------- # -- F-023: peer cohorts --------------------------------------------------
+134
View File
@@ -0,0 +1,134 @@
# tests/testthat/test-duckdb-limits.R
#
# uscogdata#60. cog_open() used to connect with a bare dbConnect() and set no
# resource pragmas, so DuckDB claimed every visible core. That is right for one
# interactive session on a dedicated machine and wrong for a server: cog-api
# runs two replicas on an 8-core host budgeted 4, and without a cap each
# replica independently claims all 8 and they fight.
#
# The consumer-side workaround this replaces reached into the namespace at
# boot -- getFromNamespace(".ensure_session", "uscogdata")() followed by a
# manual SET threads -- which depends on a private name AND on the session
# already being open.
#
# The load-bearing property is the NEGATIVE one: unset must emit no pragma at
# all, so an unconfigured session is byte-identical to pre-#60 behaviour.
# Open a session under a given configuration and read a DuckDB setting back.
# Each call closes first, because both settings are session-scoped: an
# already-open connection would be reused by .ensure_session() and report the
# PREVIOUS test's value, which is exactly the false pass to avoid here.
setting_under <- function(setting, envvars = character(0), opts = list()) {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
withr::with_envvar(envvars, {
withr::with_options(opts, {
con <- uscogdata:::cog_open()
DBI::dbGetQuery(
con, sprintf("SELECT current_setting('%s') AS v", setting)
)$v[[1]]
})
})
}
test_that("USCOGDATA_DUCKDB_THREADS caps the connection's thread count", {
skip_if_no_corpus()
expect_equal(
as.integer(setting_under("threads", c(USCOGDATA_DUCKDB_THREADS = "2"))),
2L
)
})
test_that("the option spelling works, and the env var beats it", {
skip_if_no_corpus()
expect_equal(
as.integer(setting_under("threads",
c(USCOGDATA_DUCKDB_THREADS = NA),
list(uscogdata.duckdb_threads = 3L))),
3L
)
# Same precedence .cfg() gives every other setting: env var > option.
expect_equal(
as.integer(setting_under("threads",
c(USCOGDATA_DUCKDB_THREADS = "1"),
list(uscogdata.duckdb_threads = 3L))),
1L
)
})
test_that("unset leaves DuckDB's own default in place", {
skip_if_no_corpus()
# Not asserting a specific number -- the default is core-count-dependent and
# a literal would fail on a different machine. The claim is that NO pragma
# was issued, so the session sees whatever DuckDB would have chosen on its
# own. Compared against a plain connection opened the pre-#60 way.
unset <- setting_under("threads",
c(USCOGDATA_DUCKDB_THREADS = NA),
list(uscogdata.duckdb_threads = NULL))
bare <- local({
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbGetQuery(con, "SELECT current_setting('threads') AS v")$v[[1]]
})
expect_equal(as.integer(unset), as.integer(bare))
})
test_that("USCOGDATA_DUCKDB_MEMORY_LIMIT is applied", {
skip_if_no_corpus()
v <- setting_under("memory_limit", c(USCOGDATA_DUCKDB_MEMORY_LIMIT = "2GB"))
# DuckDB does not echo back the string it was given: it stores bytes and
# reports BINARY units, so "2GB" (2e9 bytes) comes back as "1.8 GiB". Assert
# the magnitude it actually means rather than the spelling this package sent
# -- matching on "2" passes for the wrong reason and fails on the right one.
expect_match(as.character(v), "GiB", fixed = TRUE)
# DuckDB also truncates the display to one decimal ("1.8 GiB" for 1.863), so
# the tolerance covers rounding, not slack in the setting itself.
gib <- as.numeric(sub("\\s*GiB$", "", as.character(v)))
expect_equal(gib, 2e9 / 1024^3, tolerance = 0.05)
})
# --- Validation -------------------------------------------------------------
# .cfg() returns an env var as CHARACTER. Without coercion here,
# sprintf("SET threads TO %d", "4") aborts inside the connection path with an
# error about the pragma rather than about the setting the operator got wrong.
test_that(".resolve_duckdb_threads coerces a character env var to integer", {
withr::local_envvar(USCOGDATA_DUCKDB_THREADS = "4")
expect_identical(uscogdata:::.resolve_duckdb_threads(), 4L)
})
test_that(".resolve_duckdb_threads returns NULL when unset or empty", {
withr::local_options(uscogdata.duckdb_threads = NULL)
withr::local_envvar(USCOGDATA_DUCKDB_THREADS = NA)
expect_null(uscogdata:::.resolve_duckdb_threads())
withr::local_envvar(USCOGDATA_DUCKDB_THREADS = "")
expect_null(uscogdata:::.resolve_duckdb_threads())
})
test_that(".resolve_duckdb_threads rejects values that are not positive integers", {
for (bad in c("0", "-1", "two", "1.5.2")) {
withr::local_envvar(USCOGDATA_DUCKDB_THREADS = bad)
expect_error(uscogdata:::.resolve_duckdb_threads(),
class = "uscogdata_invalid_duckdb_threads")
}
})
test_that(".resolve_duckdb_memory_limit accepts size strings and rejects junk", {
withr::local_envvar(USCOGDATA_DUCKDB_MEMORY_LIMIT = "4GB")
expect_identical(uscogdata:::.resolve_duckdb_memory_limit(), "4GB")
withr::local_envvar(USCOGDATA_DUCKDB_MEMORY_LIMIT = "1.5GB")
expect_identical(uscogdata:::.resolve_duckdb_memory_limit(), "1.5GB")
# A SQL fragment must not reach the connection as one.
withr::local_envvar(USCOGDATA_DUCKDB_MEMORY_LIMIT = "4GB'; DROP TABLE x; --")
expect_error(uscogdata:::.resolve_duckdb_memory_limit(),
class = "uscogdata_invalid_duckdb_memory_limit")
})
test_that(".apply_duckdb_limits issues no statement when both are NULL", {
# The negative property, asserted directly rather than inferred: a connection
# that would ERROR on any statement proves none was sent.
expect_silent(uscogdata:::.apply_duckdb_limits(NULL, NULL, NULL))
})
+28 -9
View File
@@ -13,9 +13,13 @@ test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
} }
}) })
test_that("expenditure_concept defaults to direct and preserves today's numbers", { test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
gov <- "010000226085" # Alabama state government gov <- "010000226085" # Alabama state government
base <- cog_spending(gov, years = 2019, category = "Police") base <- cog_spending(gov, years = 2019, category = "Police")
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
# Police maps only to operations/capital codes (E62/F62/G62), so the
# direct concept's extra subtypes (interest, insurance_benefits) cannot
# contribute and the two concepts must agree exactly here.
expl <- cog_spending(gov, years = 2019, category = "Police", expl <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "direct") expenditure_concept = "direct")
expect_equal(base$amt_nominal, expl$amt_nominal) expect_equal(base$amt_nominal, expl$amt_nominal)
@@ -59,7 +63,9 @@ test_that("the IG leg never includes the L-- family total", {
codes <- DBI::dbGetQuery(con, codes <- DBI::dbGetQuery(con,
"SELECT DISTINCT item_code FROM ig_long")$item_code "SELECT DISTINCT item_code FROM ig_long")$item_code
expect_false(any(grepl("--$", codes))) expect_false(any(grepl("--$", codes)))
expect_true(all(substr(codes, 1, 1) %in% c("M", "L"))) # Q joined the IG family with the crosswalk-membership rewrite
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
}) })
test_that("expenditure_concept rejects unknown values", { test_that("expenditure_concept rejects unknown values", {
@@ -246,9 +252,12 @@ test_that("both cross-government verbs still accept the direct default", {
}) })
test_that("provenance always records the expenditure concept", { test_that("provenance always records the expenditure concept", {
d <- cog_spending("010000226085", years = 2019, category = "Police") p <- cog_spending("010000226085", years = 2019, category = "Police")
d <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending("010000226085", years = 2019, category = "Police", t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total") expenditure_concept = "total")
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
expect_equal(attr(d, "provenance")$expenditure_concept, "direct") expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
expect_equal(attr(t, "provenance")$expenditure_concept, "total") expect_equal(attr(t, "provenance")$expenditure_concept, "total")
# The note explains the non-obvious part: how legacy IG was assembled. # The note explains the non-obvious part: how legacy IG was assembled.
@@ -298,15 +307,25 @@ test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real # (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
# Direct/Total pairing. The flow-family gate in # Direct/Total pairing. The flow-family gate in
# .attach_ig_counterparts() must keep ig_recipe_id NULL here. # .attach_ig_counterparts() must keep ig_recipe_id NULL here.
#
# Anchored on FL state government, not AL. Coverage is presence-based: a
# recipe is only suggested when its component codes have rows for the
# requested government-year. AL state's only FY2011 B47 cell was an
# explicit zero, which the corpus no longer stores after sparsification
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
# FL state carries a real FY2011 B47 amount, so this exercises the guard
# against a suggestion that genuinely fires.
#
# Issue #34: "IG Federal" maps to B-prefixed codes in summary_categories
# with category_type = 'revenue'. A spending verb (flow_prefixes E/F/G)
# now scopes its candidate query by category_type = 'expenditure', so it
# correctly finds NO candidates for this revenue-only category -- the
# suggestion machinery cannot fire, and no M/L counterpart is attached.
r <- suppressMessages( r <- suppressMessages(
cog_spending("010000226085", years = c(2005, 2011), category = "IG Federal") cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
) )
sugg <- attr(r, "provenance")$suggestions sugg <- attr(r, "provenance")$suggestions
expect_gt(length(sugg), 0L) expect_length(sugg, 0L)
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
expect_true("ig_federal_b47_wide" %in% ids)
ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id))
expect_length(ig, 0L)
}) })
test_that("C1: 'total' on a legacy aggregate-only family reports the IG-only figure honestly, not as Direct + IG", { test_that("C1: 'total' on a legacy aggregate-only family reports the IG-only figure honestly, not as Direct + IG", {
+43 -6
View File
@@ -19,8 +19,6 @@
# also check the FY2022 numbers above. # also check the FY2022 numbers above.
test_that("expenditure concepts classify on spend_type, not item-code prefix", { test_that("expenditure concepts classify on spend_type, not item-code prefix", {
testthat::skip("Blocked on uscogdata#11 (findings F-012, F-017, F-018)")
mad <- "552025209777" # MADISON CITY, WI mad <- "552025209777" # MADISON CITY, WI
wi_state <- "550000227544" # WISCONSIN (state government) wi_state <- "550000227544" # WISCONSIN (state government)
@@ -61,15 +59,54 @@ test_that("expenditure concepts classify on spend_type, not item-code prefix", {
# -- F-018: prefix Y splits revenue from expenditure, by spend_type --------- # -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit # Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
# payments. All four share the first letter `Y` and the spend_type # payments. All four share the first letter `Y`, so no first-letter allowlist
# "Insurance Trust", so this pair of assertions is the concrete proof that # can route them. The proof that classification is crosswalk-keyed:
# classification is no longer keyed on the first letter. # Y05 lands in `total` spending (insurance_benefits is inside `direct`),
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
# therefore can never appear in a spending result.
#
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
# DEFAULT stays Census General Revenue and so excludes insurance-trust
# revenue; Y01's revenue-side classification is asserted against the
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
# revenue concept argument is uscogdata#12.
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L) wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
spend_codes <- wt_codes_included(wi_total) spend_codes <- wt_codes_included(wi_total)
rev_codes <- wt_codes_included(wi_revenue) rev_codes <- wt_codes_included(wi_revenue)
expect_true("Y05" %in% spend_codes) expect_true("Y05" %in% spend_codes)
expect_false("Y05" %in% rev_codes) expect_false("Y05" %in% rev_codes)
expect_true("Y01" %in% rev_codes)
expect_false("Y01" %in% spend_codes) expect_false("Y01" %in% spend_codes)
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
con <- uscogdata:::.ensure_session()
y_class <- DBI::dbGetQuery(con,
"SELECT item_code, category_type, spend_subtype, revenue_subtype
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
})
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
# balances) share first letters with flow codes, so this could never be
# guaranteed under prefix classification; under crosswalk membership it
# falls out structurally -- asserted here at the verb level, on a
# government-year the fixture gives real balance rows (Wisconsin carries
# Y07/Y08/Y21/Y61-type balances in FY2019).
wi_state <- "550000227544"
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
expect_gt(nrow(balance), 0L)
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
rev <- cog_revenue(wi_state, 2019L)
expect_false(any(spend$category %in% balance$category))
expect_false(any(rev$category %in% balance$category))
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
}) })
+1 -1
View File
@@ -83,7 +83,7 @@ test_that("cog_explain prints the expenditure concept (I1)", {
capture.output(cog_explain(t)), capture.output(cog_explain(t)),
capture.output(cog_explain(t), type = "message") capture.output(cog_explain(t), type = "message")
), collapse = "\n") ), collapse = "\n")
expect_true(grepl("Concept: direct", txt_d)) expect_true(grepl("Concept: primary", txt_d))
expect_true(grepl("Concept: total", txt_t)) expect_true(grepl("Concept: total", txt_t))
}) })
+112
View File
@@ -0,0 +1,112 @@
# tests/testthat/test-fixture-vintage.R
#
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
# every test in this package -- plus the whole cog-api suite -- runs against
# it. When the published corpus changes shape and the fixture does not, both
# suites stay green against a corpus that no longer exists (uscogdata#18).
#
# These tests pin the structural facts that distinguish the current published
# vintage from its predecessor, so a stale fixture fails loudly instead of
# passing quietly. They assert shape, never dollar values: re-running
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
# keep them green.
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
# not the package session: these assertions are about what the fixture
# CONTAINS, and routing them through the reader's own views would let a
# filter hide the very absence being checked.
fixture_query <- function(sql, ...) {
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
path <- function(rel) {
sprintf("read_parquet(%s)",
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
}
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
}
test_that("fixture ships every metadata table the publish tree does", {
skip_if_no_corpus()
# representation/code_set are what make a sparse corpus interpretable; a
# fixture without them predates sparsification (cog_pipeline#64).
expected <- c(
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet", "code_set.parquet",
"harmonization_map.parquet", "harmonization_recipes.parquet",
"lineage_events.parquet", "representation.parquet",
"series_breaks.parquet", "summary_categories.parquet"
)
on_disk <- basename(list.files(
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
))
expect_true(all(expected %in% on_disk))
# The manifest must list them too -- consumers read the manifest, not ls().
in_manifest <- with_fixture_corpus(
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
)
expect_true(all(expected %in% in_manifest))
})
test_that("fixture carries the dense/sparse representation contract", {
skip_if_no_corpus()
rep <- fixture_query(
"SELECT year, representation, absence_means FROM %s
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
"data/representation.parquet"
)
expect_equal(nrow(rep), 4L)
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
})
test_that("the fixture's wide era is sparse, not zero-padded", {
skip_if_no_corpus()
# FY2011 is a dense_source year: the corpus publishes only the cells Census
# reported non-zero, and an absent cell means Census published $0. Before
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
# zeros. A single explicit zero here means the fixture predates the change.
zeros_2011 <- fixture_query(
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
"data/long/year=2011/part-0.parquet"
)$n
expect_equal(zeros_2011, 0L)
# The modern era is a different regime: a reported zero there is real data
# (the government filed $0), so zeros legitimately survive and must not be
# asserted away.
expect_gt(
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
0L
)
})
test_that("code_set covers every fixture year with the reader-spec columns", {
skip_if_no_corpus()
cs <- fixture_query(
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
"data/code_set.parquet"
)
expect_true(all(
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
%in% names(cs)
))
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
})
test_that("every flow code carrying dollars has a category, J-prefix included", {
skip_if_no_corpus()
# The J (assistance/benefit) codes were uncategorised until the crosswalk
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
# duplication fix landed). Their absence is how a pre-crosswalk fixture
# gives itself away.
j <- fixture_query(
"SELECT item_code, category, category_type, spend_subtype FROM %s
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
"data/summary_categories.parquet"
)
expect_true("J19" %in% j$item_code)
expect_true(all(j$category_type == "expenditure"))
expect_true(all(j$spend_subtype == "assistance"))
expect_false(any(is.na(j$category)))
})
@@ -18,7 +18,6 @@
# semantics, not a row the fix makes findable. # semantics, not a row the fix makes findable.
test_that("cog_gov_search() matches name literally, not as an unescaped regex", { test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
testthat::skip("Blocked on uscogdata#16 (finding F-025)")
# -- correctness (1): a government must be findable by its own exact name --- # -- correctness (1): a government must be findable by its own exact name ---
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex # FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
+62
View File
@@ -0,0 +1,62 @@
# Network-gated. Set USCOGDATA_LIVE_TEST=true to run.
#
# This file exists because the defect fixed for 0.3.0 -- no remote corpus was
# readable at all, because DuckDB cannot expand a glob over generic HTTP --
# survived precisely because every other test path used a LOCAL corpus (the
# bundled fixture), and so did the API in production (a host mount). Nothing
# ever exercised the package the way a new user does.
skip_live <- function() {
testthat::skip_if_not(
identical(tolower(Sys.getenv("USCOGDATA_LIVE_TEST", "")), "true"),
"live-corpus test: set USCOGDATA_LIVE_TEST=true to run"
)
}
# The suite's setup.R pins USCOGDATA_URL to the bundled fixture, so reaching
# the default requires clearing both the env var and the option.
with_default_corpus <- function(code) {
withr::local_envvar(
USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA,
.local_envir = parent.frame()
)
withr::local_options(uscogdata.url = NULL, .local_envir = parent.frame())
cog_close()
withr::defer(cog_close(), envir = parent.frame())
force(code)
}
test_that("the package reads the public corpus with no configuration at all", {
skip_live()
with_default_corpus({
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
expect_gt(nrow(g), 0)
s <- cog_spending(g$canonical_govid[1], years = 2022)
expect_gt(nrow(s), 0)
expect_true(all(c("amt_nominal", "year", "category") %in% names(s)))
# Amounts are full dollars, already x1000. A city's annual spending is
# millions, not thousands -- this catches a regression that dropped or
# doubled the conversion.
expect_gt(sum(s$amt_nominal, na.rm = TRUE), 1e6)
p <- attr(s, "provenance")
expect_true(isTRUE(p$transformations$units_conversion$applied))
expect_equal(p$transformations$units_conversion$multiplier, 1000)
})
})
test_that("a multi-decade query reads across many partitions", {
skip_live()
with_default_corpus({
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
# `years` is required on cog_spending() -- there is no full-history
# default at the reader level (the API's /profile route supplies one).
s <- cog_spending(g$canonical_govid[1], years = 2000:2022)
# Enumeration builds one read_parquet() path per requested partition. If
# the list were truncated, or silently collapsed to a single file, the
# returned span is what catches it.
expect_gt(diff(range(s$year)), 10)
expect_gt(length(unique(s$year)), 5)
})
})
+97
View File
@@ -0,0 +1,97 @@
test_that(".long_files_sql enumerates every partition the manifest lists", {
manifest <- list(files = list(long_partitions = list(
list(year = 2011L, path = "data/long/year=2011/part-0.parquet"),
list(year = 2012L, path = "data/long/year=2012/part-0.parquet")
)))
expect_equal(
uscogdata:::.long_files_sql("https://example.org/corpus/", manifest),
paste0(
"['https://example.org/corpus/data/long/year=2011/part-0.parquet',",
"'https://example.org/corpus/data/long/year=2012/part-0.parquet']"
)
)
})
test_that(".long_files_sql falls back to the glob when no partition list is present", {
# test-views.R registers views with a hand-built manifest that has no
# `files` element. That must keep working: the glob is valid for the
# local paths such a manifest is used with.
expect_equal(
uscogdata:::.long_files_sql("/tmp/corpus/", list(schema_version = 4L)),
"'/tmp/corpus/data/long/**/*.parquet'"
)
expect_equal(
uscogdata:::.long_files_sql("/tmp/corpus/", list(files = list(long_partitions = list()))),
"'/tmp/corpus/data/long/**/*.parquet'"
)
})
test_that("the enumerated list matches the bundled fixture's partition count", {
skip_if_no_corpus()
m <- jsonlite::fromJSON(
file.path(fixture_corpus_path(), "manifest.json"), simplifyVector = FALSE
)
out <- uscogdata:::.long_files_sql(fixture_corpus_path(), m)
expect_equal(
lengths(regmatches(out, gregexpr("part-0\\.parquet", out))),
length(m$files$long_partitions)
)
})
test_that("no view SQL survives rendering with an unsubstituted token", {
# Introducing {long_files} broke four test sites that had hand-rolled the
# {url} substitution -- each failed with a DuckDB parser error on the
# surviving brace. This asserts the whole SQL directory renders clean, so
# a future token cannot reintroduce that silently.
sql_dir <- system.file("sql", package = "uscogdata")
for (f in list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE)) {
rendered <- uscogdata:::.render_view_sql(
paste(readLines(f, warn = FALSE), collapse = "\n"), "/tmp/corpus/"
)
expect_false(grepl("\\{[a-z_]+\\}", rendered), label = basename(f))
}
})
test_that("registered `long` view reads through the enumerated list", {
skip_if_no_corpus()
with_fixture_corpus({
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con, "SELECT count(*) AS n FROM long")$n
expect_gt(n, 0)
yrs <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
expect_true(all(c(2011, 2012, 2019, 2020) %in% yrs))
})
})
test_that("a Windows-style corpus path survives token substitution", {
# gsub() in regex mode treats backslashes in the REPLACEMENT as escape
# sequences and silently drops them, so a Windows path went in as
# C:\Users\RUNNER\... and came out as C:UsersRUNNER..., after which every
# DuckDB read failed with "No files found that match the pattern".
#
# That made a LOCAL corpus unreadable on Windows -- the bundled fixture
# included, so the whole suite failed there -- while remote https URLs
# worked fine, having no backslashes. It went unnoticed for the life of the
# package because nothing ever ran on Windows.
#
# Reproducible on any platform: this is string handling, not a filesystem
# behaviour, so it does not need a Windows runner to catch.
win <- "C:\\Users\\RUNNER~1\\AppData\\Local\\Temp\\Rtmp123/"
out <- uscogdata:::.render_view_sql(
"FROM read_parquet('{url}data/summary_categories.parquet')", win
)
expect_true(grepl("C:\\Users\\RUNNER~1\\AppData", out, fixed = TRUE))
expect_false(grepl("C:Users", out, fixed = TRUE))
# The same must hold through the {long_files} path, which embeds the url
# once per enumerated partition.
manifest <- list(files = list(long_partitions = list(
list(year = 2011L, path = "data/long/year=2011/part-0.parquet")
)))
out2 <- uscogdata:::.render_view_sql(
"FROM read_parquet({long_files}, hive_partitioning = true)", win, manifest
)
expect_true(grepl("C:\\Users\\RUNNER~1\\AppData", out2, fixed = TRUE))
expect_false(grepl("C:Users", out2, fixed = TRUE))
})
+21 -7
View File
@@ -4,12 +4,14 @@
# protect users from silent failures when USCOGDATA_URL is misconfigured # protect users from silent failures when USCOGDATA_URL is misconfigured
# or returns non-JSON content. # or returns non-JSON content.
test_that("cog_open aborts with actionable error when URL is the placeholder default", { test_that("cog_open aborts with actionable error when URL contains the sentinel", {
uscogdata:::cog_close() uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE) on.exit(uscogdata:::cog_close(), add = TRUE)
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/" # No longer the package default (that is the public HF corpus). This is a
withr::with_envvar(c(USCOGDATA_URL = placeholder), { # user who copied a config template and did not finish editing it.
sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
withr::with_envvar(c(USCOGDATA_URL = sentinel_url), {
expect_error( expect_error(
uscogdata:::cog_open(), uscogdata:::cog_open(),
class = "uscogdata_url_not_configured" class = "uscogdata_url_not_configured"
@@ -35,8 +37,10 @@ test_that("placeholder guard error names both env var and option as remediation"
uscogdata:::cog_close() uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE) on.exit(uscogdata:::cog_close(), add = TRUE)
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/" # No longer the package default (that is the public HF corpus). This is a
withr::with_envvar(c(USCOGDATA_URL = placeholder), { # user who copied a config template and did not finish editing it.
sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
withr::with_envvar(c(USCOGDATA_URL = sentinel_url), {
msg <- tryCatch(uscogdata:::cog_open(), error = conditionMessage) msg <- tryCatch(uscogdata:::cog_open(), error = conditionMessage)
expect_match(msg, "USCOGDATA_URL", fixed = TRUE) expect_match(msg, "USCOGDATA_URL", fixed = TRUE)
expect_match(msg, "uscogdata.url", fixed = TRUE) expect_match(msg, "uscogdata.url", fixed = TRUE)
@@ -111,7 +115,7 @@ test_that("cog_manifest returns the active session's parsed manifest", {
}) })
}) })
test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others", { test_that(".validate_schema accepts schema_version 4 through 7, rejects others", {
expect_silent(uscogdata:::.validate_schema(list(schema_version = 4L))) expect_silent(uscogdata:::.validate_schema(list(schema_version = 4L)))
expect_silent(uscogdata:::.validate_schema(list(schema_version = 5L))) expect_silent(uscogdata:::.validate_schema(list(schema_version = 5L)))
# v6 = FIPS geography harmonization (2026-07-22): _code -> _asof rename + # v6 = FIPS geography harmonization (2026-07-22): _code -> _asof rename +
@@ -119,12 +123,22 @@ test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others",
# renamed columns and its geography comes from the xwalk, so v6 is accepted # renamed columns and its geography comes from the xwalk, so v6 is accepted
# without behavioural change -- see .validate_schema()'s note. # without behavioural change -- see .validate_schema()'s note.
expect_silent(uscogdata:::.validate_schema(list(schema_version = 6L))) expect_silent(uscogdata:::.validate_schema(list(schema_version = 6L)))
# v7 = `data_year` APPENDED as column 29 (cog_pipeline #80, 2026-08-03), the
# most recent fiscal year contributing to a collapsed key. Appended, never
# inserted: canonical_govid stays at position 26, so nothing this package
# reads shifts. Verified against the real v7 corpus before widening the
# allow-list -- cog_spending()/cog_balances() return correctly for FY2024 AND
# for FY2012, so the new column is inert here.
expect_silent(uscogdata:::.validate_schema(list(schema_version = 7L)))
expect_error( expect_error(
uscogdata:::.validate_schema(list(schema_version = 3L)), uscogdata:::.validate_schema(list(schema_version = 3L)),
"schema_version" "schema_version"
) )
# The upper bound still has to be ENFORCED, not just moved. Without this the
# test would no longer prove that an unknown future schema is refused, and a
# v8 corpus with a genuinely breaking change would sail through.
expect_error( expect_error(
uscogdata:::.validate_schema(list(schema_version = 7L)), uscogdata:::.validate_schema(list(schema_version = 8L)),
"schema_version" "schema_version"
) )
}) })

Some files were not shown because too many files have changed in this diff Show More