Compare commits

...
Author SHA1 Message Date
jared 8bf9c4ccc1 refactor: split .suppressed_components() into R/suppression.R (#9)
R-CMD-check / check (pull_request) Successful in 3m40s
R-CMD-check / check (push) Successful in 3m46s
R/suggestions.R crossed the project's 400-line limit (424 lines).
Pure move of .suppressed_components() and its roxygen block per the
plan's Task 5 Step 2 remedy; no logic, SQL, or wording changed.
2026-08-05 08:11:15 -04:00
jaredandClaude Opus 5 77074621d8 revert: drop the I3(b) suppression pre-check gate (#9)
Scoped re-review measured .needs_suppression_query() against the fixture
and found it doesn't pay for itself: it skips the round trip on ~3% of
healthy candidate-bearing calls, ~0% of the multi-govid batch shape
(cog_geographic_rollup()/cog_peer_compare()) it was meant to help, and
reaching the gate costs an unconditional metadata query that on its own
roughly cancels the expected saving -- net slower on the fixture. The
gate was also correct (0 unsound skips) but left an untested exactness
invariant (result$codes_included and the anti-join sharing the harmonized
item_code space) whose silent violation would kill signposting, which is
the exact failure class uscogdata#9 exists to prevent.

Owner's call: revert it and keep the code simple. A batch-aware
optimization, if warranted, is a separate issue.

Removes .needs_suppression_query() entirely (function, roxygen, call
site, comp_rows/flow_components), restoring .build_suggestions() to call
.suppressed_components() directly -- unchanged from f77adb6 except that
it still threads flow_prefixes through (I1, kept). Also removes the two
tests that existed solely to exercise the gate (the five-branch synthetic
test and the local_mocked_bindings call-counter test); no test asserting
real signposting behavior was touched.

I1 (flow_prefixes filter), I2 (schema wording + ig_recipe_id required),
I3(a) (restated NOT EXISTS literals for partition pruning), and M4
(reworded scope claims) are all untouched by this revert.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 07:59:19 -04:00
jaredandClaude Opus 5 4b749205a5 fix: scope suppressed-dollar measurement to the calling verb's own flow family (#9)
Final whole-branch review fix wave for the partial-coverage signposting
feature:

- I1: .suppressed_components() now filters measured recipe components to
  the calling verb's own flow_prefixes. Without this, a candidate recipe
  from the OTHER flow family was always absent from the verb's own view by
  construction and so was always reported as "suppressed" -- fabricating a
  dollar claim across flow families (cog_revenue(category = "Corrections")
  claimed $3.63B excluded that cog_spending() actually reports in full).
- I2: reworded provenance-v1.json's trigger/suppressed_amount descriptions
  to describe what the code actually measures (the verb's underlying long
  view, not "the result"), and to note suppressed_amount can be negative.
  Added ig_recipe_id to the suggestions items' required list, matching the
  key's always-set/nullable runtime behavior.
- I3(a): restated the year/govid literals inside .suppressed_components()'s
  NOT EXISTS subquery so DuckDB can partition-prune that side too (verified
  via EXPLAIN: Scanning Files 1/4 instead of an unfiltered full scan;
  all.equal(old, new) results confirmed unchanged).
- I3(b): added .needs_suppression_query(), a free, exact pre-check reusing
  the verb's own already-computed result$codes_included to skip the anti-
  join round trip on the common fully-covered path, without weakening the
  "suppression can fire with zero gap years" guarantee.
- M4: corrected the overbroad "confines every fire to 2011" scope claim in
  R/suggestions.R and NEWS.md -- the suppressed-dollar measurement is now
  flow-scoped (post-I1), but the empty_year trigger itself is not, and can
  still fire in modern years for a mis-scoped cross-flow-family category.

Added a regression test for I1 plus direct unit-test coverage for the new
flow-family filter and the I3(b) pre-check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 23:22:52 -04:00
jared f77adb6c83 fix: revert DESCRIPTION to original roxygen2 7.3.3 config
R-CMD-check / check (push) Successful in 3m37s
R-CMD-check / check (pull_request) Successful in 3m34s
devtools::document() should not change DESCRIPTION when only JSON/Markdown/test files are edited. Restore the original RoxygenNote: 7.3.3 and remove the Config/roxygen2/version auto-generated line that resulted from running devtools::document() locally.
2026-08-04 22:16:56 -04:00
jared 230f3401c4 docs: document the suggestion trigger and suppressed-dollar fields (#9) 2026-08-04 22:11:11 -04:00
jared 7522b48a08 feat: report suppressed component dollars in the signpost message (#9)
.inform_suggestions() and cog_explain() now render suppressed_amount /
suppressed_years / suppressed_codes as a continuation line on each
suggestion bullet whenever suppressed_amount > 0 (an empty_year fire can
carry them too, so this keys off the amount, not trigger). Also renames
the cli header from "Coverage gap detected" to "Incomplete coverage" --
a partial-coverage fire is not a gap, the year has rows, they're just
short.
2026-08-04 22:02:01 -04:00
jared 693f8d81a6 fix: signpost aggregate-suppressed components in a category that still has rows (#9) 2026-08-04 21:49:23 -04:00
jared db35fa9058 feat: measure structurally-suppressed recipe component dollars (#9) 2026-08-04 21:38:22 -04:00
jared cabe2e2799 docs: plan for partial-coverage signposting (#9) 2026-08-04 21:33:09 -04:00
jared 6cd219a291 Merge pull request 'ci: fetch apt indexes over HTTPS so the install step stops hanging' (#31) from ci/apt-https into main
R-CMD-check / check (push) Successful in 5m2s
Reviewed-on: #31
2026-08-04 12:28:49 -04:00
jared e067a5930f test: accept schema_version 7, and keep the upper bound enforced
R-CMD-check / check (pull_request) Successful in 6m3s
R-CMD-check / check (push) Has been cancelled
The merged schema-v7 fix (b59b79b) widened .validate_schema()'s allow-list but
left this test asserting that 7 is REJECTED, so main went red. CI had been
hanging on the apt step before ever reaching the suite, which is why the
failure only surfaced once the HTTPS fix let the job get that far.

Flips 7L from expect_error to expect_silent, and ADDS an 8L rejection case.
That second part is the point: simply deleting the 7L expectation would have
left the test unable to prove any upper bound is enforced at all, so a future
v8 corpus with a genuinely breaking change would pass validation silently. The
test should assert the boundary moved, not that it disappeared.

Suite: 796 PASS, 0 FAIL, 0 WARN, 0 SKIP.
2026-08-04 12:07:59 -04:00
jared 342debaefa ci: fetch apt indexes over HTTPS so the install step stops hanging
R-CMD-check / check (push) Has been cancelled
R-CMD-check / check (pull_request) Has been cancelled
The "Install system libraries" step was stalling indefinitely. It was not
deadlocked on a config prompt and not slow-but-progressing: measured inside the
live runner container, /var/cache/apt/archives stayed at 0 .deb files after 3+
minutes, with apt's http workers parked in S state waiting on the network.

Root cause is the http:// mirror path being pathologically slow from this
runner, not broken. Measured 2026-08-04 from inside the CI container, same
index file, back to back:

  http://archive.ubuntu.com/ubuntu/dists/noble/Release   20.1s
  https://archive.ubuntu.com/ubuntu/dists/noble/Release    3.1s

apt fetches many indexes serially, so ~20s apiece compounds into what looks
like a hang. Rewriting the deb822 sources to https makes the step complete.

Verified before committing, in the running CI container (rocker/r-ver:4.4):
- ca-certificates present and apt 2.8.3 ships the https method built in, so
  nothing has to be installed over http first to bootstrap TLS
- the sed rewrites both URIs (archive + security); the only remaining http://
  is an inert comment line
- '#' is used as the sed delimiter deliberately: '|' collides with the
  alternation and fails with "unknown option to `s'"
- the regex survives YAML block-scalar parsing with backslashes intact

`|| true` guards each sed because the step runs under `sh -e`, so a
missing-sources-file on some other base image must not kill the job.
2026-08-04 11:57:49 -04:00
jared b59b79b2d5 Merge pull request 'fix: accept corpus schema_version 7 (#80)' (#30) from fix/schema-v7 into main
R-CMD-check / check (push) Has been cancelled
Reviewed-on: #30
2026-08-04 11:38:45 -04:00
jared 5668d6b102 fix: accept corpus schema_version 7 (#80)
R-CMD-check / check (pull_request) Has been cancelled
R-CMD-check / check (push) Has been cancelled
2026-08-04 11:22:28 -04:00
jared 0a6d878a36 Merge pull request 'fix: cog_categories() surfaces balance subtypes and accepts type = "balance"' (#29) from fix/cog-categories-balance-subtype into main
R-CMD-check / check (push) Successful in 3m28s
Reviewed-on: #29
2026-08-03 12:20:10 -04:00
jared da726a61f6 fix: cog_categories() surfaces balance subtypes and accepts type = "balance"
R-CMD-check / check (pull_request) Successful in 3m39s
R-CMD-check / check (push) Successful in 3m39s
The balance work added category_type = "balance" rows to the corpus and
cog_balances() to read them, but left cog_categories() -- the discovery
surface -- unable to describe them:

- subtype COALESCEd only spend_subtype and revenue_subtype, so every balance
  row came back with subtype = NA
- type rejected "balance", so there was no way to ask for the holdings
  taxonomy at all

Both matter downstream: cog-api derives its subtype vocabulary from
cog_categories(), so an NA subtype becomes an unusable API parameter. Found
while implementing cog-api#26.

Note cog_balances() itself still takes no subtype argument -- for holdings
category is a strict coarsening of balance_subtype -- but the value belongs
in the discovery surface regardless.

Tests read the expected subtype set independently from the crosswalk parquet
rather than from the function under test.
2026-08-03 12:02:26 -04:00
jared 03c313b46d Merge pull request 'feat: cog_balances(), a reader surface for cash and security holdings (#25)' (#28) from feat/cog-balances-25 into main
R-CMD-check / check (push) Successful in 3m18s
Reviewed-on: #28
2026-08-03 11:52:13 -04:00
jared 2c532bde19 docs: record the two balance_caveats contract facts cog-api#26 must carry
R-CMD-check / check (push) Successful in 3m38s
R-CMD-check / check (pull_request) Successful in 3m35s
Both were settled during implementation and are easy to get wrong from
outside the package:

- coverage_window is corpus-scoped, not result-scoped. It reports the observed
  year extent of every balance subtype, not only those a query returned. The
  sibling field `truncated` is the result-scoped one.
- balance_caveats is present only on cog_balances() results; an API layer that
  assumes it is universal will read NULL from the money verbs.
2026-08-03 11:45:43 -04:00
jared a9e80858d4 docs: correct coverage_window scope and the stale CLAUDE.md Current State block (#25)
F-7: inst/schemas/provenance-v1.json described coverage_window as mapping
each *observed* balance_subtype, but the query at R/balance_caveats.R has
no predicate tied to the query's codes and always returns every subtype in
the mounted corpus. Took option (b) of the two the review offered -- change
the doc, not the code. Reporting all windows is the better product
behaviour (it answers 'is there a family I missed?'), it is what cog-api#26
already forwards verbatim, and option (a) would make the block empty for a
0-row result. Reworded to say the windows are corpus-wide and that
'truncated' is the query-scoped field. Pinned by a new test either way.

F-10: the 'Current State' block was self-contradictory after a partial
update -- headed 2026-04-27, claiming branch main @ d65e9fe, with a
2026-08-03 test count measured on feat/cog-balances-25 underneath it, and
listing README.md / _pkgdown.yml as outstanding when both exist and
_pkgdown.yml was edited by this branch. All numbers below re-measured on
the final tree after every other fix in this wave, not before:
788 tests (testthat::test_local()), 14 exports (NAMESPACE), 14 man/*.Rd,
2 vignettes, no docs/ (pkgdown::build_site() genuinely still outstanding,
as is the .Rbuildignore fixture entry -- both kept in the list).

The related deferred README.md item is closed with no change, per the
review's ruling: README.md enumerates no verbs at all, so naming
cog_balances would make it the only non-cog_spending verb mentioned.
2026-08-03 11:37:15 -04:00
jared fde62eb6cc test(balances): pin the behaviours the final review found untested or weakly asserted (#25)
Findings F-1..F-8. Every assertion below was verified to FAIL before its
fix (or under mutation, where the behaviour already worked) and pass after.

F-3: 'an unknown recipe id is rejected' used a bare expect_error(). Deleting
.validate_recipe_id() leaves .recipe_components() returning 0 rows and
comps$label[[1]] throwing 'subscript out of bounds' -- still an error, so
the test passed on the regression while the user lost the curated message.
Now asserts class = 'uscogdata_unknown_recipe'. Mutation-checked.

F-4: no test ever set per_capita and adjust_to_year together, so the
load-bearing ordering comment at R/balances.R was unverified. Reversing
those two calls silently drops amt_per_capita_real (.attach_real_dollars()
no-ops when amt_per_capita_nominal does not exist yet). New test asserts
presence AND that the per-capita column is deflated by the same factor as
the level column; mutation-checked by reversing the order (2 failures).

F-5: the spec's 'Gating' requirement had no test -- nothing ever called
cog_balances() on a corpus without balance_subtype. Extended the existing
with_corpus_missing_balance_subtype() block to assert class =
'uscogdata_no_balance_support'; mutation-checked by dropping the guard.

F-1/F-2: added mutual-exclusivity and four-argument validation tests, each
pinned to the message or class (all four inputs already produced *some*
error or *some* quiet wrong answer, so bare expect_error() was useless
here). Plus an ordering guard: a data-frame govid must still work, which
is what fails if validation is put before .coerce_govid_input().

F-6: asserts on the RENDERED cog_explain() text (both streams -- cli
routes through conditions that land on stderr), with a negative case
proving money-verb output is unaffected and that the capture is not vacuous.

F-7: pins that coverage_window is corpus-scoped while truncated is
query-scoped; mutation-checked by scoping the windows to observed subtypes.

F-8: pins the memo slot is populated on first call and cleared by
cog_close().
2026-08-03 11:36:32 -04:00
jared 22c2478634 fix(balances): validate the full signature, surface caveats in cog_explain, memoise coverage windows (#25)
Final-review findings F-1, F-2, F-6, F-8 (plus the F-9 @return reword,
which shares R/balances.R).

F-2: .validate_balance_inputs() checked 2 of cog_balances()' 7 arguments.
years = integer(0) leaked a raw DuckDB 'Parser Error ... AND year IN ()'
with the generated SQL echoed back; govid = character(0) and a non-character
category returned 0 rows with no error at all; recipe = c("a","b") threw
'the condition has length > 1' from inside .validate_recipe_id(). Replaced
with a call to the money verbs' own .validate_verb_inputs() (R/spending.R),
which validates the exact superset needed. Deleted the local copy rather
than extending it -- two validators is how they drift. Placed AFTER
.coerce_govid_input(), because .validate_verb_inputs() asserts
is.character(govid) and a data-frame govid is not unwrapped before that.
This is helper reuse of the same kind as .build_verb_sql()/.attach_per_capita();
the verb still does NOT route through .verb_spendrev().

F-1: falls out of F-2 for free -- the recipe/category mutual-exclusivity
guard lives inside .validate_verb_inputs(). Previously recipe silently
discarded category AND overwrote provenance$category with the recipe label,
so a caller asking for Fund Balances got X40/Z77 insurance-trust holdings
with no trace of the dropped filter.

F-6: cog_explain() rendered every provenance caveat block except
balance_caveats. Since .emit_balance_caveats() fires at most once per
session -- and is routinely consumed by a suppressMessages() call or an
unread knitr chunk -- cog_explain() is the only surface left for a caller
who deliberately audits the result. Added a 'Holdings caveats' section
guarded on !is.null(prov$balance_caveats). Also relabels the cosmetic
'Concept: NA' line on balance results as 'not applicable (holdings are a
stock, not a flow)'.

F-8: the coverage-window query has no govid and no year predicate -- its
answer depends only on the mounted corpus -- yet it scanned all of
balance_long on every call (35% of verb runtime on the fixture, and a
per-request throughput ceiling for cog-api#26). Memoised in
.uscogdata_env$balance_coverage_windows, invalidated by cog_close(), the
same pattern as .uscogdata_env$manifest.
2026-08-03 11:36:18 -04:00
jared 225cd60968 docs: fix stale test count and phantom notes column in cog_balances docs (#25)
Re-measured CLAUDE.md's test count on the final tree (764, not 763 --
the earlier number predated the balance_caveats schema test). Removed
notes from cog_balances()'s @return block: it was copied from
cog_spending()'s @return style without checking cog_balances() never
calls .verb_spendrev(), the only place that sets notes. Verified the
remaining documented columns against colnames() observed across every
argument combination (bare, per_capita, adjust_to_year, both, recipe,
category filter).
2026-08-03 11:10:46 -04:00
jared b03f095e49 docs: document cog_balances() and correct stale CLAUDE.md claims (#25)
Adds the NEWS entry, a Financial data pkgdown reference section (none
existed for cog_spending/cog_revenue), and corrects CLAUDE.md's SQL-layer
claim, view count, test count and fixture-year description against
measured values. Also documents balance_caveats in
inst/schemas/provenance-v1.json (test-first: added a schema-documentation
test to test-balances.R, confirmed it failed, then fixed the schema) and
fleshes out cog_balances()'s @return roxygen to enumerate its conditional
columns, regenerating man/cog_balances.Rd.
2026-08-03 11:02:16 -04:00
jared 724b6bd58b docs: disambiguate live-corpus vs fixture year claim in balances test comment (#25) 2026-08-03 10:50:46 -04:00
jared 82e4face4e feat: balance_caveats provenance + once-per-session disclosure (#25) 2026-08-03 10:48:11 -04:00
jared 6c5bdb3048 test: clarify why 2002 must stay in the SB195 recipe test's year vector 2026-08-03 10:40:23 -04:00
jared b8189aeb7f docs: caveat 4 needs a year span crossing FY2002, not just a recipe query
Task 4's implementer found that SB195 does not surface for a
recipe query spanning only 2011-2012. .build_series_break_refs() matches
break_year BETWEEN min(years) AND max(years), and SB195's break_year is 2002.

That is correct behaviour rather than a gap: a series lying entirely after the
book -> market change sits on one consistent basis, so disclosing a break it
never crosses would be noise. .build_corpus_break_refs() applies the same rule
deliberately.

The spec's caveat table overclaimed by omitting the span condition. Corrected.
2026-08-03 10:39:09 -04:00
jared 90d2e6019e feat: recipe= bridges the wide-era holdings series (#25) 2026-08-03 10:37:54 -04:00
jared 769164c824 feat: per_capita and adjust_to_year for cog_balances() (#25) 2026-08-03 10:24:59 -04:00
jared de3a58d105 fix: attach govids_found/govids_missing to cog_balances() provenance
Mirrors R/spending.R:465-466 -- .check_govids_in_scope()'s return was
previously captured only for its message side effect. Also drops a
redundant duplicate assertion in the flow-code guard test.
2026-08-03 10:19:32 -04:00
jared cdb574d3d0 test: drop arrow dependency from cog_balances tests, use direct DuckDB reads
Also add explicit non-empty assertion to the flow-code guard test so it
cannot pass vacuously on a zero-row result.
2026-08-03 10:12:57 -04:00
jared a281a9621f feat: cog_balances() core verb (#25) 2026-08-03 10:01:40 -04:00
jared 825ac394f2 test: replace vacuous is_aggregate assertion with synthetic-parquet coverage
The bundled fixture has no balance item_code with is_aggregate = TRUE, so
asserting COUNT(*) FROM balance_long WHERE is_aggregate = 0 passed whether
or not the view's AND NOT is_aggregate predicate existed. Follows the
synthetic hive-partitioned parquet pattern already used for the 22-/23-
and 24-/25- view predicates in test-views.R: reads the real
inst/sql/26-balance_long.sql text off disk and executes it against a
synthetic corpus containing both an aggregate and non-aggregate row under
a real balance item_code (W01).
2026-08-03 09:54:53 -04:00
jared d09bfd6aef feat: register balance_long / balance_annotated behind a column gate (#25) 2026-08-03 09:46:36 -04:00
jared a11e29a0e0 docs: use Wisconsin state govt (550000227544) as the cog_balances test government
Standardises on the identifier other agents use for state governments, which
is stable across corpus vintages and is the same id used against the live API.

Verified in the bundled fixture, and it is strictly better coverage than the
previous pick: Wisconsin reaches four of the five balance subtypes (adds
workers_comp_trust via Y21) and carries BOTH wide->modern recipe bridges
(X40->Z77 and X41->Z78), so a second recipe test is added. Y61
(other_insurance_trust) is absent for Wisconsin; no test depends on it.

Also notes not to assert on gov_name -- the fixture carries both "WISCONSIN"
and "WISCONSIN STATE GOVT" and the verb COALESCEs them.
2026-08-03 09:42:47 -04:00
jared 7ac4dc6882 docs: implementation plan for cog_balances() (#25)
Six TDD tasks: the two views + registration gate, the core verb, per_capita
and adjust_to_year, recipe=, balance_caveats provenance, docs.

Every internal the plan calls was verified to exist with the signature used
(.build_verb_sql, .shape_recipe_result, .attach_per_capita, .run_recipe,
.require_schema_v5, ...), so the tasks reuse the shared machinery rather than
reimplementing it. The verb deliberately does not route through
.verb_spendrev(), whose concept scoping, IG leg and complete= grid are all
flow-specific.

Test government is ALABAMA STATE GOVT (010000226085), which covers every case
in the bundled fixture: W01/W31/W61 in 2012/2019/2020, X21+Z77 in 2012,
Y07/Y08 throughout, and X40 in 2011 -- so the wide-era recipe bridge is
testable offline.
2026-08-03 09:36:42 -04:00
jared 57212e3399 docs: restore recipe= to cog_balances(); the pipeline was right
Corrects this spec. The earlier draft deferred recipe= and proposed adding
summary_categories rows for X40/X41. Both were wrong, and the pipeline state
they were meant to fix is correct and documented.

cog_pipeline/docs/phase_r_harmonization_review.md records the decisions:

- Sec 0.2: the wide era exposes these split families ONLY as aggregates, so
  the recipe join deliberately does NOT filter is_aggregate. Safe by
  construction -- wide rows are aggregate-only, modern rows leaf-only, every
  component year-scoped.
- Sec 1: the planned X40->Z77 harmonization MAP rows were dropped on purpose;
  continuity ships as recipes instead. That is why harmonization_map carries
  no balance-code rows.

The reader already implements this (R/recipes.R, R/spending.R). Verified
against the live corpus rather than trusting the comment: corrections_combined
FY2007, whose wide leg E05 is likewise aggregate-only, returns $906,743,000.

Also withdraws the claim that SB155/156 and SB195/196 contradict each other.
X40 rows after FY2002 are the wide-era SAS column persisting through the era
boundary; they say nothing about a classification-level rename. The two sets
describe different layers.

What survives is one narrow, non-blocking gap: no series_breaks row exists at
2016/2017 for Z77/Z78/X30, though review doc Sec 2 recommended exactly that.
Recorded as out-of-scope item 1 with the SB197-SB202 precedent.
2026-08-03 09:28:52 -04:00
jared 9f9d40e1c3 docs: drop the subtype argument from cog_balances()
balance is the only category_type whose subtype column is not orthogonal to
category. Measured against the crosswalk: 5 of 6 expenditure subtypes and 1 of
7 revenue subtypes span more than one category, but 0 of 5 balance subtypes do.
Balance is a strict tree -- Fund Balances = {general}, Retirement System
Holdings = {employee_retirement}, Insurance Trust Balances = the three trust
subtypes.

Exposing both arguments would admit no useful combination: of the 15 pairs, 3
are redundant and 12 are guaranteed empty for every government in every year,
failing as an empty tibble that reads as "holds none" rather than as a
contradiction.

Dropping it also keeps the verb aligned -- no uscogdata verb exposes a subtype
argument; the API layers its own subtype row filter on top, which cog-api#26
can do for /balances. #25's one-filter requirement is still met, since
category = "Fund Balances" is exactly W01/W31/W61.

Adds two tests: that one-filter equivalence, and an assertion that the
subtype -> category tree holds, so an upstream change making category lossy
fails here rather than in a user's analysis.
2026-08-03 09:15:50 -04:00
jared d7e14156ff docs: design spec for cog_balances(), the uscogdata#25 holdings surface
Requirement 1 of #25 shipped with #11/#12. This specs requirement 2 only.

Records three upstream gaps found while measuring the corpus, which change
the shipping scope:

- X40/X41 carry ~42.7K rows (1967-2011) but have no summary_categories row,
  so they cannot appear in a category_type='balance' view. Both holdings
  recipes span X40/X41 + Z77/Z78, so recipe= would silently return only the
  2012-2016 leg. recipe= is therefore deferred to v2.
- SB195/SB196 attach to fin_code X40/X41, outside the balance view.
- SB197-SB202 attach to flow codes, not the holdings codes, so the FY2016
  termination of X21/X30/X42/X44/X47/Z77/Z78 has no catalogued break.

Caveats 2-4 are therefore surfaced reader-side via a computed coverage_window
rather than through the existing series-break builders.
2026-08-03 08:44:18 -04:00
jared de7ccbebc7 Merge pull request 'feat: revenue_concept = c("general", "total") off the crosswalk (#12)' (#27) from feat/revenue-concepts-12 into main
R-CMD-check / check (push) Successful in 3m9s
Reviewed-on: #27
2026-07-30 22:09:44 -04:00
jared 4b23dbd9f4 feat: revenue_concept = c("general", "total") off the crosswalk (#12)
R-CMD-check / check (push) Successful in 3m47s
R-CMD-check / check (pull_request) Successful in 3m27s
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.

`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:

  general = own_source + federal + state + local_aid   (the default)
  total   = general + utility + liquor_store + insurance_trust

The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.

- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
  subtype; the concept narrows in R via the existing subtype_scope
  machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
  printed `expenditure_concept` unconditionally, so a cog_revenue()
  caller was told "Concept: primary" -- a spending concept their result
  has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).

Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).

Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.

Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.

Closes #12
2026-07-30 20:49:49 -04:00
jared 93300ae0c1 feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00
jared 5d77d39711 Merge pull request 'chore: regenerate fixture corpus at pipeline_commit e64a046 (#11 groundwork)' (#26) from feat/expenditure-concepts-11 into main
R-CMD-check / check (push) Successful in 3m3s
Reviewed-on: #26
2026-07-30 16:22:19 -04:00
jared 7d798b9937 chore: regenerate fixture corpus at pipeline_commit e64a046
R-CMD-check / check (pull_request) Successful in 3m18s
R-CMD-check / check (push) Successful in 4m6s
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.

Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.

test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.

Suite: 0 failures, 2 skips (the #11 and #12 blocks).
2026-07-30 16:04:35 -04:00
jared 915a4d0678 Merge pull request 'feat: coverage argument + always-on reporting-coverage metadata (#13)' (#24) from feat/coverage-disclosure-13 into main
R-CMD-check / check (push) Successful in 3m30s
2026-07-30 12:06:53 -04:00
jared 6f98d061a9 Merge pull request 'feat: complete = TRUE fills absent cells with their meaning (#18)' (#23) from feat/complete-argument-18 into main
R-CMD-check / check (push) Successful in 4m14s
Reviewed-on: #23
2026-07-30 12:04:27 -04:00
jared d95c9032c5 feat: coverage argument + always-on reporting-coverage metadata (#13)
R-CMD-check / check (pull_request) Successful in 3m13s
R-CMD-check / check (push) Successful in 3m18s
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".

On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.

Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.

Decisions worth stating:

  - n_units_expected is the universe the CALLER named, not the national one.
    That is what makes the ratio mean something: "597 of the 608 Wisconsin
    cities you asked about". For peers it is the cohort size, counted over
    peer rows only -- including the target would inflate every count by one
    and make a cohort that has entirely stopped reporting look non-empty.

  - The coverage table is built from the REQUESTED years, not the years
    present in the result, so a year in which nothing reported still appears
    with n_units_reporting = 0. A year that vanishes silently is precisely
    the disclosure failure at issue.

  - "census" filters years BEFORE the query, and aborts when the range holds
    no census year rather than returning an empty result for a query the
    caller believes they made.

  - "consistent" exempts the peer-comparison target: it is the subject of the
    comparison, not a member of the cohort being balanced, and dropping it
    would leave nothing to compare. The summary_* quantiles are computed
    AFTER the filter so they describe the cohort actually returned.

  - is_census_year is documented as a statement about the survey CALENDAR,
    never a claim of completeness -- FY1967 is a census year in which only 97
    of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
    that tells the truth.

On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().

One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.

Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
2026-07-30 11:57:11 -04:00
60 changed files with 5303 additions and 336 deletions
+15
View File
@@ -11,6 +11,21 @@ jobs:
steps:
- name: Install system libraries and Node.js (required by actions/checkout)
run: |
# Switch apt to HTTPS mirrors. Measured from this runner on
# 2026-08-04: the SAME index file takes 20.1s over http:// and 3.1s
# over https://. apt fetches many indexes serially, so http:// does
# not read as "slow" -- it reads as a hang (zero bytes in
# /var/cache/apt/archives after 3+ minutes, apt's http workers parked
# in S state). rocker/r-ver:4.4 already ships ca-certificates and
# apt 2.8.3 has the https method built in, so nothing needs to be
# installed over http first to bootstrap this.
# `|| true` because the step runs under `sh -e`: on an image whose
# sources live in the other location, the missing-file sed must not
# kill the job.
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
/etc/apt/sources.list.d/ubuntu.sources 2>/dev/null || true
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
/etc/apt/sources.list 2>/dev/null || true
apt-get update -qq
apt-get install -y --no-install-recommends \
nodejs git \
File diff suppressed because it is too large Load Diff
+36 -14
View File
@@ -28,8 +28,23 @@ USCOGDATA_URL (local path or https://)
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated`
- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
`revenue_annotated`/`ig_annotated`), and metadata views
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
`representation`, `code_set`)
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
**stock** (a balance at a point in time) rather than a **flow** (activity
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
disclosure + measured per-subtype coverage windows).
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
@@ -45,28 +60,31 @@ USCOGDATA_URL (local path or https://)
Any value without `://` is treated as a local path by `.is_local_path()` and reads
`manifest.json` directly from disk (no HTTP, no TTL cache).
## Current State (2026-04-27)
## Current State (2026-08-03)
**Version:** 0.1.0 (pre-release)
**Branch:** `main`, commit `d65e9fe`
**Tests:** 181 PASS / 0 FAIL / 0 SKIP
**Branch:** `feat/cog-balances-25`, commit `fde62eb`
**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
### Completed (Tasks 2.1–2.7)
All 8 exported verbs implemented and tested:
`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`,
`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`,
plus `cog_categories`.
All **14** exports implemented and tested (measured from `NAMESPACE`):
`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years
2019+2020, all 50 states). Tests run fully offline — no credentials needed.
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
### Remaining to v0.1 release
1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports;
full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`.
Vignettes can be stubbed for v0.1.
1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
`README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
`total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
`pkgdown::build_site()` has never been run (no `docs/`).
2. **Phase 3 — cog_explorer bridge:** create
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
@@ -99,6 +117,10 @@ devtools::test()
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
- Return value is always a `tbl_df` with a `provenance` attribute
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
- SQL lives in `inst/sql/` — never inline SQL strings in R files
- SQL has two layers. **View definitions** live in `inst/sql/` and are
registered by `.register_views()`, which globs the directory in sorted order
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
a numbered `.sql` file; build a query in R.
- No arrow dependency — DuckDB reads parquet natively
- `withr` is a Suggests-only dep; only used in tests
+1
View File
@@ -1,5 +1,6 @@
# Generated by roxygen2: do not edit by hand
export(cog_balances)
export(cog_basket_resolution)
export(cog_basket_unresolved)
export(cog_categories)
+80
View File
@@ -1,5 +1,85 @@
# uscogdata 0.1.0 (development)
## Signposting now catches partially-suppressed categories
* A coverage suggestion used to fire only when a category returned **no rows
at all** in a requested year. That missed the more dangerous case: a
category that still returns rows while silently dropping component codes
the wide era publishes only as aggregates (#9). `cog_spending(category =
"Public Welfare")` for FY2011 returned a plausible figure that omitted
`E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true
$5,261,404,000, a 39% understatement, with `provenance$suggestions` empty.
* Suggestions now also fire on **partial** coverage, and every suggestion
carries `trigger` (`"empty_year"` or `"suppressed_component"`),
`suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller
can see how much is missing and decide whether to re-run with the recipe.
* `cog_revenue()` gets the same fix through the shared verb path. Alaska's
FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping
$1,899,995,000 of aggregate-published `U4-` rents and royalties.
* The trigger stays recipe-driven, so it only fires where a harmonization
recipe actually exists to name the fix. `higher_ed_e18_wide` and
`general_gov_e89_wide` stay silent in every year measured on the bundled
fixture, because their components are ordinary classified leaves even
pre-2012.
* The `suppressed_component` trigger (and any `suppressed_amount`/
`suppressed_codes` an `empty_year` fire also carries) is scoped to the
calling verb's own flow family: `cog_spending()` only ever measures E/F/G
component dollars, `cog_revenue()` only T/A/U/B/C/D. A component from the
OTHER flow family reports `suppressed_amount = 0` rather than a fabricated
claim. The `empty_year` trigger itself is not flow-scoped -- a category
belonging to the other flow (e.g. `cog_spending(category = "IG Local")`)
still returns zero rows and can still fire, in any year including modern
ones, naming the recipe whose own generic join finds real data for this
government. That is a mis-scoped query, not a corpus-format gap, so its
`suppressed_amount` is correctly 0.
## New: `cog_balances()` for cash-and-security holdings
* New `cog_balances()` exposes the 14 cash-and-security holding codes
(`category_type = "balance"`): fund balances, retirement system holdings and
insurance trust balances (#25). Holdings are a stock, not a flow, so the verb
has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and
no `subtype` argument either -- for holdings, `category` is a strict
coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly
the `general` family (`W01`/`W31`/`W61`).
* `cog_balances()` results carry `provenance$balance_caveats`, recording that
Census holdings are gross rather than GAAP fund balance, and the measured
coverage window of each subtype family.
## Multi-government aggregates now disclose their reporting coverage
* The Census of Governments is a **complete census only in years ending in 2
and 7**; every other year is a sample, and the sample varies enormously. On
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
return value said nothing about, so a statewide total resting on a fifth of
the universe looked exactly like one resting on all of it.
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
`coverage`:
| value | effect |
|---|---|
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
* **Regardless of mode**, every result now carries `provenance$coverage` with
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
section. So the default mode can no longer mislead silently.
* `is_census_year` is a statement about the **survey calendar**, never a claim
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
cities report. `n_units_reporting` is the number that tells the truth.
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
it is the subject of the comparison, not a member of the cohort — and the
`summary_*` quantiles are computed after the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against the
cohort size.
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
is `NULL`: `"census"` snaps to the most recent census year with an observed
population, so a cohort is not built from a sample year in which most of the
candidate universe is absent.
## `complete = TRUE`: absent cells, labelled with why they are absent
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
+122
View File
@@ -0,0 +1,122 @@
# R/balance_caveats.R
#
# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and
# security holdings. Each one silently invalidates an obvious analysis, so
# they travel in provenance (machine-readable, for cog-api#26) rather than
# living only in prose.
#
# Two of the four are already carried by the code-driven series-break
# builders and are deliberately NOT duplicated here:
# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via
# series_break_refs on the recipe path, the only path that observes those
# codes.
# What remains is the GAAP distinction (a constant) and the coverage windows
# (measured, never hardcoded, so they stay correct as the corpus grows).
#' Per-subtype observed year extents, plus which requested families are
#' truncated relative to the requested span.
#' @noRd
.balance_caveats <- function(con, codes_observed, years) {
cw <- .balance_coverage_windows(con)
observed_subtypes <- if (length(codes_observed) == 0L) {
character(0)
} else {
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT balance_subtype FROM summary_categories
WHERE item_code IN (%s) AND balance_subtype IS NOT NULL",
.sql_lit_chr(codes_observed)
))$balance_subtype
}
# A family is "truncated" when the caller asked for years outside the span
# that family actually covers -- the FY2016 employee-retirement termination
# and the FY2021 end of the W family are both this shape.
truncated <- character(0)
if (length(years) > 0L) {
for (s in observed_subtypes) {
w <- cw[[s]]
if (is.null(w)) next
if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s)
}
}
list(
not_gaap = TRUE,
not_gaap_note = paste0(
"Census holdings are gross -- no liabilities are netted -- and are NOT ",
"GAAP fund balance. A reserve ratio built from them overstates what is ",
"actually available."
),
coverage_window = cw,
truncated = sort(unique(truncated))
)
}
#' Per-subtype [min year, max year] extents for EVERY balance subtype in the
#' mounted corpus, memoised for the session.
#'
#' The query carries no govid and no year predicate -- its answer is a property
#' of the mounted corpus alone and cannot change between calls -- but it scans
#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime
#' on the bundled fixture and would be a per-request throughput ceiling once
#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and
#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`.
#'
#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking
#' "is there a family I missed?" needs every window. The observed-scoped field
#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json.
#' @noRd
.balance_coverage_windows <- function(con) {
cached <- .uscogdata_env$balance_coverage_windows
if (!is.null(cached)) return(cached)
windows <- DBI::dbGetQuery(con,
"SELECT c.balance_subtype AS subtype,
MIN(l.year) AS year_min,
MAX(l.year) AS year_max
FROM balance_long l
JOIN summary_categories c USING (item_code)
WHERE c.balance_subtype IS NOT NULL
GROUP BY 1
ORDER BY 1"
)
cw <- stats::setNames(
lapply(seq_len(nrow(windows)),
function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))),
windows$subtype
)
.uscogdata_env$balance_coverage_windows <- cw
cw
}
#' TRUE the first time `key` is seen this session, FALSE thereafter.
#' Reset by cog_close().
#' @noRd
.balance_caveat_once <- function(key) {
seen <- .uscogdata_env$balance_caveats_shown
if (is.null(seen)) seen <- character(0)
if (key %in% seen) return(FALSE)
.uscogdata_env$balance_caveats_shown <- c(seen, key)
TRUE
}
#' Emit at most one message per caveat class per session.
#' @noRd
.emit_balance_caveats <- function(caveats) {
if (.balance_caveat_once("not_gaap")) {
cli::cli_inform(c(
"!" = "Census holdings are gross and are {.strong not} GAAP fund balance.",
"i" = "No liabilities are netted; a reserve ratio built from them overstates available funds."
))
}
if (length(caveats$truncated) > 0L &&
.balance_caveat_once("coverage_window")) {
cli::cli_inform(c(
"!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.",
"i" = "See {.code provenance$balance_caveats$coverage_window}."
))
}
invisible(NULL)
}
+159
View File
@@ -0,0 +1,159 @@
# R/balances.R
#
# Cash and security holdings. A third verb rather than an argument on a money
# verb because holdings are a STOCK -- a balance at a point in time -- while
# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money
# verbs' whole argument vocabulary (expenditure_concept, revenue_concept,
# complete=) describes flows and is meaningless here, so this deliberately
# does NOT route through .verb_spendrev().
#' Cash and security holdings for one or more governments
#'
#' Returns Census cash-and-security holdings (`category_type = "balance"`):
#' fund balances, retirement system holdings and insurance trust balances.
#'
#' @section Holdings are not GAAP fund balance:
#' Census holdings are **gross** -- no liabilities are netted -- so a reserve
#' ratio built from them overstates what is actually available. They are not
#' comparable to a GAAP fund balance from an ACFR.
#'
#' @param govid Canonical govid(s): a character vector, or a data frame with a
#' `canonical_govid` column (e.g. from [cog_gov_search()]).
#' @param years Integer vector of fiscal years.
#' @param category Optional character vector of categories to keep. One of
#' `"Fund Balances"`, `"Insurance Trust Balances"`,
#' `"Retirement System Holdings"`. There is deliberately no `subtype`
#' argument: for holdings, `category` is a strict coarsening of
#' `balance_subtype` (unlike the money verbs, where the two axes cross), so
#' every combination would be either redundant or empty.
#' `category = "Fund Balances"` is exactly the `general` family
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
#' one `dplyr::filter()` away.
#' @param per_capita Divide holdings by population. Note this is a **stock per
#' resident** (reserves per person), which is *not* comparable to
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
#' @param adjust_to_year Deflate to this year's dollars (CPI-U).
#' @param basis Accepted for uniformity with the money verbs, but currently a
#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized
#' and raw space are identical for holdings. Reported in
#' `provenance$basis_note`.
#' @param recipe Optional harmonization recipe id (see [cog_recipes()]).
#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
#' wide era to the modern one.
#'
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`,
#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and
#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when
#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
#' nominal per-capita column to deflate otherwise). Amounts are full US
#' dollars.
#'
#' Carries a `provenance` attribute matching
#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
#' every balance subtype in the mounted corpus, not only the observed ones)
#' and `truncated` (the observed subtypes whose coverage falls short of the
#' requested years). `expenditure_concept`/`revenue_concept` are `NA` --
#' holdings are a stock, not a flow, so neither concept vocabulary applies.
#' @export
cog_balances <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL) {
call <- match.call()
basis <- match.arg(basis, c("harmonized", "raw"))
# Coerce FIRST, validate second: .validate_verb_inputs() asserts
# is.character(govid), and a data-frame govid (cog_gov_search() output) has
# not been unwrapped yet at this point.
govid <- .coerce_govid_input(govid)
# The money verbs' validator, reused rather than re-implemented (R/spending.R).
# It covers the exact superset cog_balances() needs -- including the
# recipe/category mutual-exclusivity guard -- so a second local copy would
# only be a place for the two to drift apart. This is the same kind of
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
# route the verb through .verb_spendrev(), which stays deliberately unused
# here because its flow vocabulary is meaningless for a stock.
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
con <- .ensure_session()
.require_balance_support(con)
scope <- .check_govids_in_scope(govid)
basis_note <- paste0(
"`basis` has no effect on holdings: harmonization_map carries no ",
"balance-code rows, so harmonized and raw space are identical here."
)
manifest <- .uscogdata_env$manifest
recipe_block <- NULL
category_for_prov <- category
if (!is.null(recipe)) {
.require_schema_v5(con, manifest, "recipe =")
.validate_recipe_id(con, recipe)
comps <- .recipe_components(con, recipe)
recipe_label <- comps$label[[1]]
result <- .run_recipe(con, recipe, govid, years)
sql <- attr(result, "sql_query")
result <- .shape_recipe_result(result, "balance_subtype", recipe_label)
recipe_block <- list(
recipe_id = recipe, label = recipe_label,
components = .df_to_row_list(comps)
)
category_for_prov <- recipe_label
} else {
sql <- .build_verb_sql("balance_annotated", "balance_subtype",
govid, years, category,
ig_view = NULL, subtype_scope = NULL)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
# Order matters (matches .verb_spendrev()): per-capita first, so
# .attach_real_dollars() deflates the nominal per-capita column into
# amt_per_capita_real rather than needing amt_per_capita_nominal recomputed.
if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid)
if (!is.null(adjust_to_year)) {
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
}
prov <- .build_provenance(
verb = "cog_balances", call = call, govid = govid, years = years,
category = category_for_prov, per_capita = per_capita,
adjust_to_year = adjust_to_year, result = result, sql = sql,
subtype_col = "balance_subtype",
basis = basis, basis_note = basis_note,
# Neither concept vocabulary applies to a stock.
expenditure_concept = NA_character_,
revenue_concept = NA_character_,
recipe = recipe_block
)
prov$scope$govids_found <- scope$found
prov$scope$govids_missing <- scope$missing
prov$balance_caveats <- .balance_caveats(
con, prov$codes_summed$observed, years
)
.emit_balance_caveats(prov$balance_caveats)
attr(result, "provenance") <- prov
result
}
#' Abort unless the mounted corpus classifies balance codes.
#'
#' `balance_subtype` arrived with cog_pipeline #76/#77 without a
#' schema_version bump, so the check is on the column, not the version.
#' @noRd
.require_balance_support <- function(con) {
if (.corpus_has_balance_subtype(con)) return(invisible(TRUE))
cli::cli_abort(
c("This corpus does not classify cash and security holdings.",
i = "`summary_categories` has no {.field balance_subtype} column.",
i = "Republish from cog_pipeline at #76/#77 or later."),
class = "uscogdata_no_balance_support"
)
}
+15 -6
View File
@@ -44,11 +44,18 @@
#' Count + sum item-level rows that basis="harmonized" excludes because they
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
#' requested flow type (spending or revenue), govids, and years. Only
#' meaningful when the resolved basis is "harmonized"; returns an
#' applied = FALSE stub otherwise (raw basis never excludes rows this way).
#' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
#' the same subtype-membership classification the verb SQL uses, never
#' item-code prefixes), govids, and years. Only meaningful when the resolved
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
#' basis never excludes rows this way).
#'
#' The intergovernmental leg is deliberately outside this count even for
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
#' @noRd
.build_harmonization_block <- function(con, govid, years, resolved, flow_prefixes) {
.build_harmonization_block <- function(con, govid, years, resolved,
subtype_col, subtype_scope) {
if (!identical(resolved$basis, "harmonized")) {
return(list(
applied = FALSE,
@@ -63,9 +70,11 @@
FROM long
WHERE canonical_govid IN (%s) AND year IN (%s)
AND NOT is_aggregate AND harmonized_code IS NULL
AND LEFT(item_code, 1) IN (%s)",
AND item_code IN (
SELECT item_code FROM summary_categories WHERE %s IN (%s)
)",
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
.sql_lit_chr(flow_prefixes)
subtype_col, .sql_lit_chr(subtype_scope)
)
na <- DBI::dbGetQuery(con, sql)
+13 -6
View File
@@ -5,12 +5,19 @@
#' Returns the category taxonomy exposed by the corpus's
#' `summary_categories` view, grouped to one row per
#' `(category, subtype)` pair. Use this to discover valid `category`
#' values for [cog_spending()] / [cog_revenue()] /
#' values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
#' [cog_geographic_rollup()] and to audit which Census item codes feed
#' each category.
#'
#' @param type Either `NULL` (default, return both spending and revenue
#' rows), `"spending"`, or `"revenue"`.
#' `subtype` COALESCEs the crosswalk's three subtype columns, so it carries
#' `spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
#' `balance_subtype` on balance rows. Note that [cog_balances()] itself takes
#' no `subtype` argument — for holdings, `category` is a strict coarsening of
#' `balance_subtype` — but the value is surfaced here because it is the
#' discovery surface downstream consumers build their vocabulary from.
#'
#' @param type Either `NULL` (default, every row: expenditure, revenue and
#' balance), `"spending"`, `"revenue"`, or `"balance"`.
#' @param pattern Optional regex matched case-insensitively against the
#' `category` column (e.g. `"Police"` or `"Tax"`).
#' @return Tibble with columns `category`, `category_type`, `subtype`,
@@ -20,8 +27,8 @@
cog_categories <- function(type = NULL, pattern = NULL) {
if (!is.null(type)) {
if (!is.character(type) || length(type) != 1L ||
!type %in% c("spending", "revenue")) {
cli::cli_abort('`type` must be NULL, "spending", or "revenue".')
!type %in% c("spending", "revenue", "balance")) {
cli::cli_abort('`type` must be NULL, "spending", "revenue", or "balance".')
}
}
if (!is.null(pattern) &&
@@ -48,7 +55,7 @@ cog_categories <- function(type = NULL, pattern = NULL) {
sql <- paste(
"SELECT category, category_type,
COALESCE(spend_subtype, revenue_subtype) AS subtype,
COALESCE(spend_subtype, revenue_subtype, balance_subtype) AS subtype,
COUNT(DISTINCT item_code) AS n_codes,
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS item_codes
FROM summary_categories",
+8 -7
View File
@@ -44,7 +44,9 @@
#' The cells a government-year COULD carry: every code in force for that
#' government's own type, mapped through `summary_categories`, restricted to
#' the calling verb's flow prefixes and (when given) its category filter.
#' the calling verb's crosswalk subtype scope (the same subtype-membership
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
#' operations/capital/assistance) and (when given) its category filter.
#'
#' Scoped by `govs_type` deliberately. Filling against the union of all types
#' would invent cells that the government can never report -- a county row for
@@ -56,7 +58,7 @@
#' never returns, so every one of them would fill as a phantom $0.
#' @noRd
.completion_grid_sql <- function(subtype_col, govid, years, category,
flow_prefixes) {
subtype_scope) {
category_pred <- if (is.null(category)) {
""
} else {
@@ -77,13 +79,12 @@
WHERE x.canonical_govid IN (%2$s)
AND cs.year IN (%3$s)
AND NOT cs.is_aggregate
AND LEFT(cs.item_code, 1) IN (%4$s)
AND c.category IS NOT NULL
AND c.%1$s IS NOT NULL
AND c.%1$s IN (%4$s)
%5$s",
subtype_col, .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","),
.sql_lit_chr(flow_prefixes), category_pred
.sql_lit_chr(subtype_scope), category_pred
)
}
@@ -94,9 +95,9 @@
#' must never alter or drop what the corpus actually published.
#' @noRd
.complete_result <- function(result, con, subtype_col, govid, years, category,
flow_prefixes) {
subtype_scope) {
grid <- tibble::as_tibble(DBI::dbGetQuery(
con, .completion_grid_sql(subtype_col, govid, years, category, flow_prefixes)
con, .completion_grid_sql(subtype_col, govid, years, category, subtype_scope)
))
result$value_source <- rep("reported", nrow(result))
+107
View File
@@ -0,0 +1,107 @@
# R/coverage.R
#
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
# findings F-020 and F-023).
#
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
# 7. Every other year is a sample, and the sample varies enormously: on the
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
# what the verbs have always done -- correctly -- but the return value said
# nothing about it, so a statewide total resting on 18% of the universe looked
# exactly like one resting on 98%.
#
# Owner's settled design: a `coverage` argument selecting WHICH units to
# include, plus always-on metadata saying how many there were either way. The
# principle behind it: using these verbs correctly must not require the caller
# to know the survey calendar.
# Years ending in 2 or 7 are full censuses of every government; all others are
# samples.
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
#' @noRd
.is_census_year <- function(years) {
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
}
#' @noRd
.validate_coverage <- function(coverage) {
tryCatch(
match.arg(coverage, c("all", "census", "consistent")),
error = function(e) {
cli::cli_abort(
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
class = "uscogdata_invalid_coverage", parent = e
)
}
)
}
#' Restrict `years` to census years for `coverage = "census"`.
#'
#' Aborts rather than returning an empty result when the requested range holds
#' no census year: silently handing back zero rows for a query the caller
#' believes they made is the failure mode this whole issue is about.
#' @noRd
.apply_census_years <- function(years, coverage, verb) {
if (!identical(coverage, "census")) return(as.integer(years))
keep <- as.integer(years)[.is_census_year(years)]
if (length(keep) == 0L) {
cli::cli_abort(c(
"{.code coverage = \"census\"} leaves no years to query.",
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
), class = "uscogdata_no_census_years")
}
sort(keep)
}
#' Keep only units that report in EVERY requested year (a balanced panel).
#'
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
#' filter (the peer-comparison target, which is the subject of the comparison
#' rather than a member of the cohort being balanced).
#' @noRd
.filter_consistent <- function(result, years, id_col = "canonical_govid",
keep_ids = character(0)) {
years <- unique(as.integer(years))
if (nrow(result) == 0L || length(years) <= 1L) return(result)
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
present <- vapply(ids, function(g) {
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
}, logical(1))
consistent <- c(ids[present], keep_ids)
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
drop = FALSE]
}
#' Per-year coverage metadata, always attached regardless of mode.
#'
#' Built from the REQUESTED years rather than the years present in the result,
#' so a year in which nothing reported still appears -- with
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
#' missing year fails to make.
#'
#' `n_units_reporting` describes the result the caller actually received, so
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
#' `n_units_reporting` is the number that tells the truth.
#' @noRd
.coverage_table <- function(result, years, n_expected,
id_col = "canonical_govid", rows = NULL) {
years <- sort(unique(as.integer(years)))
src <- if (is.null(rows)) result else rows
reporting <- vapply(years, function(y) {
ids <- src[[id_col]][as.integer(src$year) == y]
length(unique(ids[!is.na(ids)]))
}, integer(1))
tibble::tibble(
year = years,
n_units_reporting = as.integer(reporting),
n_units_expected = rep(as.integer(n_expected), length(years)),
is_census_year = .is_census_year(years)
)
}
+64 -2
View File
@@ -60,7 +60,20 @@ cog_explain <- function(result, format = c("print", "list")) {
cli::cli_text("Basis: {prov$basis}{note}")
}
if (!is.null(prov$expenditure_concept)) {
# Each verb reports its OWN concept. Both fields are always present (each
# defaults to its concept's default), so printing `expenditure_concept`
# unconditionally would tell a cog_revenue() caller "Concept: primary",
# which names a spending concept their result has nothing to do with.
if (identical(prov$verb, "cog_revenue")) {
if (!is.null(prov$revenue_concept)) {
cli::cli_text("Concept: {prov$revenue_concept} revenue")
}
} else if (identical(prov$verb, "cog_balances")) {
# Both concept fields are deliberately NA here (a stock has no flow
# concept). Printing the raw NA reads as a missing value rather than an
# intentional one, so say what it means instead.
cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)")
} else if (!is.null(prov$expenditure_concept)) {
concept_note <- if (!is.null(prov$expenditure_concept_note) &&
!is.na(prov$expenditure_concept_note)) {
sprintf(" (%s)", prov$expenditure_concept_note)
@@ -112,12 +125,36 @@ cog_explain <- function(result, format = c("print", "list")) {
if (length(prov$suggestions) > 0L) {
cli::cli_h2("Suggestions")
sugg_lines <- vapply(prov$suggestions, function(s) {
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
s$available_years[1], s$available_years[2], s$hint)
if (isTRUE(s$suppressed_amount > 0)) {
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
line
}, character(1))
cli::cli_ul(sugg_lines)
}
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
cli::cli_h2("Reporting coverage")
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
cov <- prov$coverage
cli::cli_ul(sprintf(
"%d: %d of %d units reporting (%.0f%%) -- %s year",
cov$year, cov$n_units_reporting, cov$n_units_expected,
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
ifelse(cov$is_census_year, "census", "sample")
))
if (any(!cov$is_census_year)) {
cli::cli_text(
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
)
}
}
if (isTRUE(prov$completion$applied)) {
cli::cli_h2("Completion")
cli::cli_text(
@@ -149,6 +186,31 @@ cog_explain <- function(result, format = c("print", "list")) {
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
}
# Balance results only (NULL on money-verb provenance, so they are
# unaffected). This is the ONLY on-demand surface for the GAAP disclosure:
# .emit_balance_caveats() fires at most once per session, and is routinely
# consumed by a suppressMessages() call or by a knitted chunk nobody reads,
# so a caller who deliberately audits a result with cog_explain() must still
# be told.
bc <- prov$balance_caveats
if (!is.null(bc)) {
cli::cli_h2("Holdings caveats")
if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note)
if (length(bc$truncated) > 0L) {
cli::cli_text(
"Requested years extend beyond what these families actually cover:"
)
cli::cli_ul(vapply(bc$truncated, function(s) {
w <- bc$coverage_window[[s]]
if (length(w) == 2L) {
sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2])
} else {
s
}
}, character(1)))
}
}
cli::cli_h2("Transformations")
uc <- prov$transformations$units_conversion
if (isTRUE(uc$applied)) {
+1 -1
View File
@@ -138,7 +138,7 @@
#' year, matching canonical_fips_xwalk) rather than as-of-year; as-of-year
#' moved to the *_asof columns. This package's own geography always came from
#' the xwalk (already present-based), so behaviour is unchanged.
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L)) {
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L, 7L)) {
if (!manifest$schema_version %in% supported) {
cli::cli_abort(c(
"Corpus schema version mismatch.",
+88 -9
View File
@@ -19,6 +19,13 @@
#' target's population at `year` to produce absolute bounds. If `FALSE`,
#' `pop_range` is interpreted as absolute population counts.
#' @param max_peers Integer cap on the number of peers returned.
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
#' most recent census year with an observed population, so a cohort is not
#' built from a sample year in which most of the candidate universe is
#' absent. `"consistent"` needs a year range, which cohort selection does not
#' have, so it selects like `"all"` and is carried on the result as
#' `attr(x, "coverage")` for [cog_peer_compare()].
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
#' `attr(x, "cohort_year")`.
@@ -29,7 +36,9 @@ cog_find_peers <- function(target_govid,
same_state = FALSE,
pop_range = c(0.7, 1.3),
is_ratio = TRUE,
max_peers = 10L) {
max_peers = 10L,
coverage = c("all", "census", "consistent")) {
coverage <- .validate_coverage(coverage)
if (!is.character(target_govid) || length(target_govid) != 1L) {
cli::cli_abort("`target_govid` must be a length-1 character string.")
}
@@ -59,7 +68,7 @@ cog_find_peers <- function(target_govid,
))
}
cohort_year <- .resolve_cohort_year(con, target_govid, year)
cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
pop_sql <- sprintf(
"SELECT population FROM gov_population_yearly
@@ -107,12 +116,34 @@ cog_find_peers <- function(target_govid,
attr(peers, "cohort_year") <- as.integer(cohort_year)
attr(peers, "pop_range") <- as.numeric(pop_range)
attr(peers, "is_ratio") <- isTRUE(is_ratio)
attr(peers, "coverage") <- coverage
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
peers
}
# `coverage` picks the cohort vintage when the caller did not name one.
# "census" snaps to the most recent CENSUS year with an observed population,
# so a cohort is not silently built from a sample year in which most of the
# candidate universe is absent. "consistent" is a comparison-time concept --
# it needs a year RANGE, which cohort selection does not have -- so it selects
# like "all" here and is carried on the result for cog_peer_compare().
#' @noRd
.resolve_cohort_year <- function(con, target_govid, year) {
.resolve_cohort_year <- function(con, target_govid, year,
coverage = "all") {
if (!is.null(year)) return(as.integer(year))
if (identical(coverage, "census")) {
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
.sql_lit_chr(target_govid)
)
y <- DBI::dbGetQuery(con, sql)$y
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
cli::cli_abort(c(
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
), class = "uscogdata_no_census_years")
}
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s",
@@ -145,10 +176,36 @@ cog_find_peers <- function(target_govid,
#' @param per_capita Default `TRUE` — peer compare usually normalizes by
#' population.
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
#' single-government queries but cannot be used here because combining Total
#' across peer sets counts intergovernmental transfers twice.
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across peer sets counts
#' intergovernmental transfers twice; `"primary"` and `"direct"` combine
#' safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#'
#' The comparison target is exempt from `"consistent"` balancing -- it is the
#' subject of the comparison, not a member of the cohort -- and the
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
#' cohort actually returned. `n_units_reporting` counts peers only, against
#' the cohort size: "3 of your 15 peers reported in FY2019".
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
@@ -186,9 +243,11 @@ cog_find_peers <- function(target_govid,
#' @export
cog_peer_compare <- function(target_govid, peers, category, years,
per_capita = TRUE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_peer_compare")
}
@@ -211,9 +270,21 @@ cog_peer_compare <- function(target_govid, peers, category, years,
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
all_govids <- unique(c(target_govid, peer_govids))
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
years <- .apply_census_years(years, coverage, "cog_peer_compare")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
# The target is exempt from balancing: it is the subject of the comparison,
# not a member of the cohort being balanced, and dropping it would leave a
# peer comparison with nothing to compare. Filtering happens BEFORE the
# quantiles below, so a "consistent" cohort's summary rows describe that
# cohort rather than the unbalanced one.
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years, keep_ids = target_govid)
}
value_col <- .peer_value_col(per_capita, adjust_to_year)
summary_rows <- .peer_summary_rows(r, value_col)
@@ -234,6 +305,14 @@ cog_peer_compare <- function(target_govid, peers, category, years,
canonical_govid = target_govid,
gov_name = unique(r$gov_name[r$role == "target"])
)
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
# reported in FY2019". Including the target would inflate every count by one
# and make a cohort that has entirely stopped reporting look non-empty.
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(
out, years, length(peer_govids),
rows = r[r$role == "peer", , drop = FALSE]
)
attr(out, "provenance") <- prov
out
}
+3 -1
View File
@@ -6,9 +6,10 @@
per_capita, adjust_to_year, result, sql,
subtype_col, basis = NA_character_,
basis_note = NA_character_,
expenditure_concept = "direct",
expenditure_concept = "primary",
expenditure_concept_note = NA_character_,
expenditure_concept_direct_suppressed = FALSE,
revenue_concept = "general",
harmonization = NULL, recipe = NULL,
suggestions = list(),
completion = NULL) {
@@ -67,6 +68,7 @@
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note,
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
revenue_concept = revenue_concept,
harmonization = harmonization %||% list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
note = NA_character_
+31
View File
@@ -8,6 +8,31 @@
#' multiplies by 1000 and records the conversion in `provenance`).
#'
#' @inheritParams cog_spending
#' @param revenue_concept Which of Census's two published revenue concepts to
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
#' values -- never as item-code first letters, which cannot classify
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
#' prefix `X` does the same):
#'
#' * `"general"` (default) -- Census General Revenue: `own_source` +
#' `federal` + `state` + `local_aid`. The manual defines this concept by
#' subtraction (section 4.3: *"General revenue comprises all revenue
#' except that classified as liquor store, utility, or insurance trust
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
#' insurance trust revenue are all excluded.
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
#' `general` plus utility, liquor store, and insurance trust revenue
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
#' `X01`/`X02`/`X05`/`X08`).
#'
#' The two are related by Census's own identity, `Total Revenue = General +
#' Utility + Liquor Store + Insurance Trust`.
#'
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
#' systems moved out of the annual finance file into the separate Annual
#' Survey of Public Pensions, so a `"total"` series steps down at the
#' FY2016/FY2017 seam for reasons that are about collection scope rather
#' than revenue (series breaks `SB197`-`SB202`).
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
@@ -17,7 +42,12 @@
cog_revenue <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
# this verb's recipe families (see R/suggestions.R).
.verb_spendrev(
verb = "cog_revenue",
view_base = "revenue_annotated",
@@ -31,6 +61,7 @@ cog_revenue <- function(govid, years, category = NULL,
adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe,
revenue_concept = revenue_concept,
complete = complete
)
}
+44 -8
View File
@@ -25,12 +25,31 @@
#' population from `gov_population_yearly`. Govs with missing population
#' are excluded from the result.
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
#' single-government queries but cannot be used here because combining Total
#' across multiple layers of government double-counts intergovernmental
#' transfers (a state's payment to a school district is the same dollar the
#' district reports as its own Direct spending).
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across multiple layers of
#' government double-counts intergovernmental transfers (a state's payment
#' to a school district is the same dollar the district reports as its own
#' Direct spending); `"primary"` and `"direct"` combine safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
@@ -40,9 +59,11 @@
#' @export
cog_geographic_rollup <- function(govids, category, years,
per_capita = FALSE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_geographic_rollup")
}
@@ -59,11 +80,21 @@ cog_geographic_rollup <- function(govids, category, years,
layer = rep(layer_names, lengths(govids))
)
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
# coverage = "census" drops non-census years BEFORE the query rather than
# after: a sample year's rows are not wanted at all, and fetching them only
# to discard them would also let them into the coverage table.
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
relationship = "many-to-many")
r$scope_note <- .rollup_scope_note(r$layer)
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years)
}
excluded <- character(0)
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
drop <- r$pop_source == "unavailable"
@@ -82,6 +113,11 @@ cog_geographic_rollup <- function(govids, category, years,
included_govids = included,
excluded_govids = excluded
)
# n_units_expected is the universe the CALLER named -- the govids passed in
# -- not the national universe. That is what makes the ratio meaningful:
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
attr(r, "provenance") <- prov
r
+4 -1
View File
@@ -12,7 +12,7 @@ cog_open <- function(url = .resolve_url(),
DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;")
manifest <- .fetch_or_cache_manifest(url, cache_dir)
.validate_schema(manifest, supported = c(4L, 5L, 6L))
.validate_schema(manifest, supported = c(4L, 5L, 6L, 7L))
.validate_scope(manifest)
.register_views(con, url, manifest)
@@ -95,4 +95,7 @@ cog_close <- function() {
}
.uscogdata_env$con <- NULL
.uscogdata_env$manifest <- NULL
.uscogdata_env$balance_caveats_shown <- NULL
# Memoised corpus-constant; a different corpus may be mounted next.
.uscogdata_env$balance_coverage_windows <- NULL
}
+155 -33
View File
@@ -1,5 +1,57 @@
# R/spending.R
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
# `spend_subtype` values. Classification is crosswalk membership, never
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
# route it (finding F-018).
#
# primary = operations + capital + assistance (the default)
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
# total = direct + intergovernmental (via the ig_* views)
#
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
# than intergovernmental -- including payments to retirees, i.e. insurance
# trust benefits. Verified against Census's own published FY2020 state
# aggregates (20statetypepu.txt): `total` reproduces the published
# expenditure sum to the dollar; omitting insurance benefits understates
# California's Direct by 10.9%.
.spend_subtypes_primary <- c("operations", "capital", "assistance")
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
#' @noRd
.expenditure_concept_subtypes <- function(concept) {
switch(concept,
primary = .spend_subtypes_primary,
# "total" = the direct subtypes here PLUS the intergovernmental leg,
# which travels through the ig_* views rather than this scope (see
# .build_verb_sql()).
direct = ,
total = .spend_subtypes_direct
)
}
# The two revenue concepts (uscogdata#12), again as crosswalk subtype sets.
# Census's manual section 4.3 defines the first by SUBTRACTING from the second
# -- "General revenue comprises all revenue except that classified as liquor
# store, utility, or insurance trust revenue" -- giving the identity
#
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
#
# Verified against Census's own computed concept fields (IndFin FY2012,
# Wisconsin state): 31,410,686 + 0 + 0 + 4,469,906 = 35,880,592, exact.
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
.revenue_subtypes_total <- c(.revenue_subtypes_general, "utility",
"liquor_store", "insurance_trust")
#' @noRd
.revenue_concept_subtypes <- function(concept) {
switch(concept,
general = .revenue_subtypes_general,
total = .revenue_subtypes_total
)
}
#' Summarized spending by category
#'
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
@@ -42,22 +94,34 @@
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
#' FALSE`, pointing at the `recipe` block instead) rather than a
#' possibly-misleading `"harmonized"`/`"raw"` value.
#' @param expenditure_concept `"direct"` (default) returns only the
#' government's own direct spending (item codes `E`/`F`/`G`), unchanged
#' from prior releases. `"total"` additionally UNIONs in the
#' intergovernmental leg -- payments to local governments (`M` codes) and
#' to the state government (`L` codes, excluding the `L--` family-total
#' rollup) -- so results gain rows with `spend_subtype ==
#' "intergovernmental"`. Requires the active corpus's `summary_categories`
#' to carry M/L rows (added by cog_pipeline PR #59); aborts with class
#' `uscogdata_ig_categories_unsupported` on an older corpus rather than
#' silently under-reporting. Mutually exclusive with `recipe` (a recipe
#' already defines its own component codes). **Do not sum `"total"`
#' results across levels of government** (e.g. state + county + city):
#' a state's `M12` payment to a school district is the same dollar the
#' district reports as its own direct `E12`, so summing both double-counts
#' it. This matters in particular with [cog_geographic_rollup()], which
#' sums across exactly that kind of multi-layer government set.
#' @param expenditure_concept Which spending concept to return. Concepts are
#' defined as sets of the crosswalk's `spend_subtype` values -- never as
#' item-code first letters, which cannot classify correctly (prefix `Y`
#' alone spans revenue, expenditure, and balance codes):
#'
#' * `"primary"` (default) -- the government's own service provision:
#' `operations` + `capital` + `assistance` subtypes.
#' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
#' `interest` (interest on debt) and `insurance_benefits` (insurance
#' trust benefit payments, e.g. pensions -- Census manual section
#' 5.2.2.1 includes payments to retirees in Direct).
#' * `"total"` -- `direct` plus the intergovernmental leg: payments to
#' local governments (`M` codes), to the state government (`L` codes,
#' excluding the `L--` family-total rollup), and state payments to
#' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
#' `spend_subtype == "intergovernmental"`. Requires the active corpus's
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
#' older corpus rather than silently under-reporting. Mutually
#' exclusive with `recipe` (a recipe already defines its own component
#' codes).
#'
#' **Do not sum `"total"` results across levels of government** (e.g.
#' state + county + city): a state's `M12` payment to a school district is
#' the same dollar the district reports as its own direct `E12`, so
#' summing both double-counts it. This matters in particular with
#' [cog_geographic_rollup()], which sums across exactly that kind of
#' multi-layer government set.
#'
#' In the legacy wide era (<= FY2011), some functions are published ONLY
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
@@ -104,8 +168,12 @@
cog_spending <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total"),
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk subtype membership
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
# machinery to this verb's recipe families (see R/suggestions.R; the
# catalog only has E/F/G-component direct-expenditure recipes).
.verb_spendrev(
verb = "cog_spending",
view_base = "spending_annotated",
@@ -128,8 +196,9 @@ cog_spending <- function(govid, years, category = NULL,
.abort_concept_not_aggregatable <- function(verb) {
cli::cli_abort(c(
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
"*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\
comparison or sum that spans more than one government.",
"*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
{.code \"direct\"} for any comparison or sum that spans more than \\
one government.",
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
money it hands to other governments. The receiving government reports \\
that same dollar again as its own Direct when it actually spends it, \\
@@ -145,7 +214,8 @@ cog_spending <- function(govid, years, category = NULL,
govid, years, category,
per_capita, adjust_to_year,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total"),
expenditure_concept = c("primary", "direct", "total"),
revenue_concept = c("general", "total"),
complete = FALSE) {
basis_explicit <- length(basis) == 1L
basis <- match.arg(basis, c("harmonized", "raw"))
@@ -153,16 +223,39 @@ cog_spending <- function(govid, years, category = NULL,
# condition; wrap it so an invalid expenditure_concept aborts consistently
# with the rest of this package's validation (cli::cli_abort -> rlang_error).
expenditure_concept <- tryCatch(
match.arg(expenditure_concept, c("direct", "total")),
match.arg(expenditure_concept, c("primary", "direct", "total")),
error = function(e) {
cli::cli_abort(
"`expenditure_concept` must be one of {.val direct} or {.val total}.",
"`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
class = "uscogdata_invalid_expenditure_concept",
parent = e
)
}
)
revenue_concept <- tryCatch(
match.arg(revenue_concept, c("general", "total")),
error = function(e) {
cli::cli_abort(
"`revenue_concept` must be one of {.val general} or {.val total}.",
class = "uscogdata_invalid_revenue_concept",
parent = e
)
}
)
# The concept's subtype scope. Every code path below -- the verb SQL, the
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
# by crosswalk subtype membership, never by item-code prefix. The
# expenditure "total" concept's extra intergovernmental leg is the one
# exception: it travels through the ig_* views rather than this scope,
# because its legacy rows are aggregate-flagged.
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
.expenditure_concept_subtypes(expenditure_concept)
} else {
.revenue_concept_subtypes(revenue_concept)
}
govid <- .coerce_govid_input(govid, arg = "govid")
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
@@ -176,10 +269,10 @@ cog_spending <- function(govid, years, category = NULL,
}
# .verb_spendrev() is shared with cog_revenue(), which never exposes
# expenditure_concept and always resolves it to "direct" -- so nothing on
# the public API can reach this today. But it's a cheap guard against a
# expenditure_concept and always resolves it to the default -- so nothing
# on the public API can reach this today. But it's a cheap guard against a
# future call (direct or via a modified cog_revenue()) that would UNION
# the IG leg's expenditure M/L rows into a revenue result, which has no
# the IG leg's expenditure M/L/Q rows into a revenue result, which has no
# matching IG view and no sensible meaning.
if (identical(expenditure_concept, "total") &&
!identical(view_base, "spending_annotated")) {
@@ -240,7 +333,8 @@ cog_spending <- function(govid, years, category = NULL,
} else {
NULL
}
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view)
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
subtype_scope)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
@@ -251,7 +345,7 @@ cog_spending <- function(govid, years, category = NULL,
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
if (complete) {
result <- .complete_result(result, con, subtype_col, govid, years,
category, flow_prefixes)
category, subtype_scope)
completion <- attr(result, ".completion")
attr(result, ".completion") <- NULL
}
@@ -284,7 +378,7 @@ cog_spending <- function(govid, years, category = NULL,
basis_for_prov <- resolved$basis
basis_note_for_prov <- resolved$note
harmonization <- .build_harmonization_block(
con, govid, years, resolved, flow_prefixes
con, govid, years, resolved, subtype_col, subtype_scope
)
# C1(a): gap detection must run against the Direct leg alone. `result`
# can also carry UNION'd intergovernmental rows (expenditure_concept =
@@ -301,7 +395,8 @@ cog_spending <- function(govid, years, category = NULL,
}
suggestions <- .build_suggestions(con, govid, years, category,
direct_leg_result,
resolved$basis, flow_prefixes)
resolved$basis, flow_prefixes,
.select_long_view(view_base, resolved$basis))
}
# C1(b): when expenditure_concept = "total", flag any row where the IG
@@ -362,6 +457,7 @@ cog_spending <- function(govid, years, category = NULL,
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note_for_prov,
expenditure_concept_direct_suppressed = direct_suppressed_flag,
revenue_concept = revenue_concept,
harmonization = harmonization,
recipe = recipe_block,
suggestions = suggestions,
@@ -417,6 +513,18 @@ cog_spending <- function(govid, years, category = NULL,
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
}
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
#' anti-joins the LONG view rather than the annotated one: they have identical
#' row membership (the annotated views are the long views plus LEFT JOINs, see
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
#' the suppression test is asking about.
#' @noRd
.select_long_view <- function(view_base, basis) {
.select_view(sub("_annotated$", "_long", view_base), basis)
}
#' @noRd
.select_ig_view <- function(basis) {
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
@@ -462,7 +570,7 @@ cog_spending <- function(govid, years, category = NULL,
#' @noRd
.build_verb_sql <- function(view, subtype_col, govid, years, category,
ig_view = NULL) {
ig_view = NULL, subtype_scope = NULL) {
govid_lit <- .sql_lit_chr(govid)
years_lit <- paste(as.integer(years), collapse = ",")
category_pred <- if (is.null(category)) {
@@ -471,9 +579,22 @@ cog_spending <- function(govid, years, category = NULL,
sprintf("AND category IN (%s)", .sql_lit_chr(category))
}
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
# The base views carry every subtype of their flow (spending_annotated has
# all five non-IG expenditure subtypes); the concept narrows here. For
# "total", the IG leg's rows are 'intergovernmental', so that value joins
# the allowlist exactly when ig_view is present.
subtype_pred <- if (is.null(subtype_scope)) {
""
} else {
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
}
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
# never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L),
# so de-duplication would be pure cost, and a silent row-drop if two
# never UNION: the two legs are disjoint by crosswalk subtype (the direct
# view excludes 'intergovernmental'; the IG view is only that), so
# de-duplication would be pure cost, and a silent row-drop if two
# governments ever reported identical values.
source_expr <- if (is.null(ig_view)) {
view
@@ -505,9 +626,10 @@ cog_spending <- function(govid, years, category = NULL,
WHERE canonical_govid IN (%3$s)
AND year IN (%4$s)
%5$s
%6$s
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
ORDER BY year, canonical_govid, %1$s, category",
subtype_col, source_expr, govid_lit, years_lit, category_pred
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
)
}
+82 -16
View File
@@ -1,8 +1,16 @@
# R/suggestions.R
# Recipe-component-driven signposting: when a basis = "harmonized" query for
# a category comes back with a coverage gap in some requested years (the
# result has no rows at all in that year) that a harmonization recipe would
# actually fill for this government, surface that recipe as a suggestion.
# Recipe-component-driven signposting. When a basis = "harmonized" query for
# a category comes back incomplete in some requested year -- and a
# harmonization recipe would actually fill it for this government -- surface
# that recipe as a suggestion. "Incomplete" has two forms, and a recipe
# qualifies on either:
# 1. empty_year -- the result has no rows at all in that year.
# 2. suppressed_component -- the result HAS rows, but a component code
# carries dollars the verb's own long view structurally excludes
# (aggregate-published, or absent from summary_categories). This is
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
# aggregate-only E67/E68, so form 1 never fired and the caller got a
# number a third too low with no signpost at all.
#
# This is deliberately keyed off the recipe catalog's component codes, not
# off harmonization_map rows: no live map row carries a non-blank
@@ -48,11 +56,15 @@
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
#' lookup scoped to the calling verb's own flow family.
#' @param long_view Name of the verb's own long view (from
#' `.select_long_view()`), passed through to `.suppressed_components()` to
#' measure the second qualifying path (uscogdata#9).
#' @return List of `list(recipe_id, label, available_years, hint,
#' ig_recipe_id)`, possibly empty.
#' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
#' suppressed_codes)`, possibly empty.
#' @noRd
.build_suggestions <- function(con, govid, years, category, result, basis,
flow_prefixes) {
flow_prefixes, long_view) {
if (!identical(basis, "harmonized") || is.null(category)) return(list())
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
@@ -86,7 +98,28 @@
unique(as.integer(result$year))
}
gap_years <- setdiff(as.integer(years), result_years)
if (length(gap_years) == 0L) return(list())
# Path 2 (uscogdata#9): component dollars this government holds that the
# verb's own view structurally excludes. Measured across ALL requested
# years, not just gap years -- the whole point is that a year with rows can
# still be missing dollars. Scoped to the calling verb's own flow_prefixes
# (I1) -- see `.suppressed_components()`'s own roxygen for why.
#
# This runs unconditionally whenever there are candidates -- an earlier
# revision of this fix wave tried a free, in-memory pre-check
# (`.needs_suppression_query()`) to skip the round trip on an already-
# covered path, but a scoped re-review measured it against the fixture and
# found it didn't pay for itself (it skipped ~3% of healthy calls, ~0% of
# the multi-govid batch shape it was meant to help, at a net cost increase
# once its own always-run metadata query was counted) while adding an
# untested exactness invariant -- that `result$codes_included` and this
# anti-join share the harmonized `item_code` space -- whose silent
# violation would kill signposting, the exact failure class uscogdata#9
# exists to prevent. Owner's call: keep this simple; a batch-aware
# optimization, if one is worth building, is a separate issue.
supp <- .suppressed_components(con, candidates, govid, years, long_view, flow_prefixes)
if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
"SELECT recipe_id, any_value(label) AS label,
@@ -97,11 +130,12 @@
.sql_lit_chr(candidates)
)))
# Which (recipe_id, year) pairs the recipe's own generic join actually
# covers for this government, restricted to the gap years -- the same
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
# no is_aggregate filter), just checking existence instead of summing.
covered <- DBI::dbGetQuery(con, sprintf(
# Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
# covers for this government, restricted to the gap years.
covered <- if (length(gap_years) == 0L) {
data.frame(recipe_id = character(0), year = integer(0))
} else {
DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.year
FROM long l
JOIN harmonization_recipes r
@@ -116,16 +150,35 @@
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(gap_years, collapse = ",")
))
}
suggestions <- list()
for (rid in candidates) {
if (!rid %in% covered$recipe_id) next
empty_hit <- rid %in% covered$recipe_id
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
supp_hit <- nrow(s_rows) > 0L
if (!empty_hit && !supp_hit) next
m <- meta[meta$recipe_id == rid, ]
suggestions[[length(suggestions) + 1L]] <- list(
recipe_id = rid,
label = m$label[[1]],
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
hint = sprintf("re-run with recipe = '%s'", rid)
hint = sprintf("re-run with recipe = '%s'", rid),
# An empty year is the stronger claim -- the category returned nothing
# at all -- so it wins when both paths qualify. The suppressed_* fields
# are still populated, so an empty_year fire also reports its dollars.
trigger = if (empty_hit) "empty_year" else "suppressed_component",
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
suppressed_years = if (supp_hit) {
sort(unique(as.integer(s_rows$year)))
} else {
integer(0)
},
suppressed_codes = if (supp_hit) {
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
} else {
character(0)
}
)
}
.attach_ig_counterparts(con, suggestions, flow_prefixes)
@@ -232,12 +285,25 @@
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
#' continuation line is appended naming the intergovernmental counterpart
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
#' same bullet under cli, not a new bullet).
#' same bullet under cli, not a new bullet). Same treatment for
#' `suppressed_amount` (uscogdata#9): only present when dollars were
#' actually measured as excluded (an `empty_year` fire can carry them too --
#' see `.build_suggestions()` -- so this keys off the amount, not `trigger`).
#' @noRd
.inform_suggestions <- function(suggestions) {
bullets <- vapply(suggestions, function(s) {
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
s$available_years[1], s$available_years[2], s$hint)
# Only present when dollars were actually measured as excluded. An
# empty_year fire can carry them too -- the year had no rows AND the
# component was suppressed -- which is strictly more informative.
if (isTRUE(s$suppressed_amount > 0)) {
bullet <- paste0(bullet, sprintf(
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
if (!is.null(s$ig_recipe_id)) {
bullet <- paste0(bullet, sprintf(
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
@@ -245,7 +311,7 @@
bullet
}, character(1))
cli::cli_inform(c(
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
stats::setNames(bullets, rep("*", length(bullets)))
))
}
+115
View File
@@ -0,0 +1,115 @@
# R/suppression.R
# Split out of R/suggestions.R (2026-08-05) to keep files under the project's
# 400-line limit. Owns the second qualifying path for coverage signposting
# (uscogdata#9): measuring, per government, the component dollars the
# calling verb's own long view structurally excludes (aggregate-published,
# or absent from summary_categories). See R/suggestions.R for the
# orchestrator (`.build_suggestions()`) that calls this and the full
# uscogdata#9 background.
#' Measure, per (recipe, year), the component dollars this government holds
#' that the calling verb's own long view structurally excludes.
#'
#' This is the second qualifying path for a suggestion (uscogdata#9). The
#' first -- row absence -- only fires when a category returns NOTHING in a
#' requested year, which is how Corrections behaves in the wide era. Public
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
#' 2011, and absent from `summary_categories` entirely) are dropped. The
#' caller gets a plausible number a third too low, silently.
#'
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
#' view rather than restating its WHERE clause, so this stays correct if
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
#' anti-join is keyed on `item_code`, which is sound only because
#' harmonization never renames a recipe component -- asserted by the "no
#' recipe component is ever renamed by harmonization" test in
#' tests/testthat/test-recipes.R.
#'
#' Note what this deliberately does NOT count as suppressed: a component
#' excluded from the RESULT for scoping reasons -- because it belongs to a
#' different `category`, or because `expenditure_concept` narrowed the
#' subtypes -- is still present in the view, so it never fires. Suggesting a
#' recipe is a coverage fix, not a category redefinition.
#'
#' `flow_prefixes` (uscogdata#9 review, finding I1) restricts the measured
#' components to the CALLING VERB's own flow family (`c("E","F","G")` for
#' spending, `c("T","A","U","B","C","D")` for revenue). Without this, a
#' candidate recipe belonging to the OTHER flow family is always absent from
#' this verb's view (by construction -- `cog_revenue()`'s view never carries
#' an E-coded row) and so was always reported as "suppressed", fabricating a
#' dollar claim across flow families (`cog_revenue(category = "Corrections")`
#' claimed $3.63B excluded that `cog_spending()` reports and fully accounts
#' for). Filtering on `LEFT(r.component_code, 1)` also drops M/L-prefixed
#' components from measurement under `cog_spending()` (`flow_prefixes` never
#' includes "M"/"L") -- harmless today, because a recipe's own M/L components
#' (e.g. `corrections_ig_local_combined`'s M04/M05) are present in the view
#' in every year they exist and so never fired as suppressed anyway, but
#' worth recording since this filter is now the thing relied on to prevent
#' it.
#'
#' @param con Active DuckDB connection.
#' @param candidates Character vector of recipe ids to measure.
#' @param govid Character vector of canonical_govid values.
#' @param years Integer vector of requested years.
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
#' `.build_suggestions()`). Only recipe components whose first character is
#' in this set are measured.
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
#' nothing is suppressed.
#' @noRd
.suppressed_components <- function(con, candidates, govid, years, long_view,
flow_prefixes) {
empty <- tibble::tibble(
recipe_id = character(0), year = numeric(0),
suppressed_amount = numeric(0), suppressed_codes = character(0)
)
if (length(candidates) == 0L) return(empty)
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
# never be quoted safely. It is always internally derived from a fixed
# view_base, so an off-allowlist value is a programming error, not input.
if (!long_view %in% c("spending_long", "spending_long_harmonized",
"revenue_long", "revenue_long_harmonized")) {
cli::cli_abort(
"Internal error: unexpected `long_view` {.val {long_view}}.",
class = "uscogdata_internal_error"
)
}
sql <- sprintf(
"SELECT r.recipe_id,
l.year,
SUM(l.amt) * 1000.0 AS suppressed_amount,
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
AS suppressed_codes
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%1$s)
AND l.canonical_govid IN (%2$s)
AND l.year IN (%3$s)
AND l.amt <> 0
AND LEFT(r.component_code, 1) IN (%5$s)
AND NOT EXISTS (
SELECT 1 FROM %4$s v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year
AND v.item_code = l.item_code
AND v.year IN (%3$s) -- restated: enables partition pruning (I3a)
AND v.canonical_govid IN (%2$s) -- restated: pushes the govid filter (I3a)
)
GROUP BY 1, 2
ORDER BY 1, 2",
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","), long_view,
.sql_lit_chr(flow_prefixes)
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
+24
View File
@@ -45,6 +45,29 @@
"37-code_set.sql" = "code_set.parquet"
)
# Cash and security holdings (uscogdata#25). 46- selects
# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and
# WITHOUT a schema_version bump -- so neither existing gate applies:
# .harmonization_view_files keys on schema_version, .representation_view_files
# on the presence of a FILE. Here the discriminator is a COLUMN on a table
# that exists either way. CREATE VIEW resolves its source schema eagerly, so
# on an older corpus 46- would fail at registration with "Binder Error:
# Referenced column balance_subtype not found" rather than at query time.
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
#' Does the mounted corpus's `summary_categories` carry `balance_subtype`?
#' Probed against the live connection rather than the manifest, because the
#' manifest describes files, not columns.
#' @noRd
.corpus_has_balance_subtype <- function(con) {
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM information_schema.columns
WHERE table_name = 'summary_categories'
AND column_name = 'balance_subtype'"
)$n
isTRUE(as.integer(n) > 0L)
}
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
#' works identically for a local fixture and a remote share.
@@ -66,6 +89,7 @@
if (base %in% .harmonization_view_files && schema_version < 5L) next
if (base %in% names(.representation_view_files) &&
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
DBI::dbExecute(con, sql)
+48 -8
View File
@@ -46,23 +46,63 @@ than obviously wrong.
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
## Direct vs Total spending
## Primary vs Direct vs Total spending
`cog_spending(..., expenditure_concept = c("direct", "total"))` controls
whose spending a result counts. `"direct"` (the default) is a government's
own current operations, capital outlay, and other direct spending. `"total"`
additionally adds in the intergovernmental legs — money it hands to other
governments to spend on its behalf — which is meaningful for describing one
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
controls whose spending a result counts. Concepts are defined as sets of the
crosswalk's `spend_subtype` values — never item-code first letters, which
cannot classify correctly (the letter `Y` alone spans revenue, expenditure,
and balance codes):
- `"primary"` (the default) is the government's own service provision:
current operations, capital outlay, and assistance payments.
- `"direct"` is Census's published Direct Expenditure: `primary` plus
interest on debt and insurance trust benefit payments (e.g. pensions).
- `"total"` additionally adds the intergovernmental leg — money handed to
other governments to spend (`M`/`L` codes plus `Q11`/`Q12`/`Q18` state
payments to school systems) — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
**Rule of thumb: any figure that spans more than one government uses
`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this
by refusing `expenditure_concept = "total"`. See
`primary` or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()`
enforce this by refusing `expenditure_concept = "total"`. See
`vignette("total-spending", package = "uscogdata")` for the full
explanation with worked examples.
## General vs Total revenue
`cog_revenue(..., revenue_concept = c("general", "total"))` selects between
Census's two published revenue concepts, again defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:
- `"general"` (the default) is Census **General Revenue**: own-source
(taxes, charges, miscellaneous) plus federal, state and local
intergovernmental aid.
- `"total"` is Census **Total Revenue**: `general` plus utility revenue
(`A91`–`A94`), liquor store revenue (`A90`), and insurance trust revenue
(unemployment and workers' compensation `Y` codes plus the
employee-retirement `X` codes).
The manual defines the first by subtracting the other three from the second,
so the two are related by Census's own identity:
```
Total Revenue = General + Utility + Liquor Store + Insurance Trust
```
Two things worth knowing before switching to `"total"`:
- **Utility revenue is large for cities.** Measured on the bundled fixture,
utility plus liquor store revenue is 15.9% of city (type 2) revenue, versus
1.2% for states and 1.7% for counties. `general` excludes it by definition.
- **The employee-retirement (`X`) codes stop at FY2016**, when those systems
moved out of the annual finance file into the separate Annual Survey of
Public Pensions. A `"total"` series therefore steps down at the
FY2016/FY2017 seam for reasons of collection scope, not revenue (series
breaks `SB197`–`SB202`, in the corpus's `series_breaks` table).
## Developer notes
### Testing
+6
View File
@@ -3,6 +3,12 @@ template:
bootstrap: 5
reference:
- title: Financial data
desc: Spending, revenue and balance-sheet holdings for one or more governments.
contents:
- cog_spending
- cog_revenue
- cog_balances
- title: Search & basket
desc: Resolve place names into canonical govids.
contents:
Binary file not shown.
Binary file not shown.
+4 -4
View File
@@ -1,7 +1,7 @@
{
"schema_version": 6,
"built_at": "2026-07-30T14:07:36Z",
"pipeline_commit": "83f9715",
"built_at": "2026-07-31T00:47:27Z",
"pipeline_commit": "aadb46b",
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
"data_vintage": {
"source_vintages": {
@@ -105,12 +105,12 @@
},
{
"path": "data/series_breaks.parquet",
"sha256": "5ae050dd7a76c4d25e5f99e7c2e81c1896482e3504e0443b47ab5d78ba148953",
"sha256": "06dcc995ff533e57cc65fa25086cc9bf83ba592c58bf7cc99269dc2576f69944",
"description": "series_breaks.parquet"
},
{
"path": "data/summary_categories.parquet",
"sha256": "e71d6d70d767c26c983fe56213baf204355f879582aa94841e62d9aea1877f83",
"sha256": "e3b0efa00ce713b8f45829b89cfde24b55333f26101f0495df82d85997d18d8e",
"description": "summary_categories.parquet"
}
]
+62 -5
View File
@@ -14,20 +14,67 @@
"basis_note": { "type": ["string", "null"] },
"expenditure_concept": {
"type": "string",
"enum": ["direct", "total"],
"description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments."
"enum": ["primary", "direct", "total"],
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
},
"expenditure_concept_note": {
"type": ["string", "null"],
"description": "How the intergovernmental leg was assembled; null for 'direct'."
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
},
"expenditure_concept_direct_suppressed": {
"type": "boolean",
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
},
"revenue_concept": {
"type": "string",
"enum": ["general", "total"],
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
},
"harmonization": { "type": "object" },
"recipe": { "type": ["object", "null"] },
"suggestions": { "type": "array" },
"suggestions": {
"type": "array",
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
"items": {
"type": "object",
"required": ["recipe_id", "label", "available_years", "hint", "ig_recipe_id",
"trigger", "suppressed_amount", "suppressed_years", "suppressed_codes"],
"properties": {
"recipe_id": { "type": "string" },
"label": { "type": "string" },
"available_years": {
"type": "array",
"items": { "type": "integer" },
"description": "[year_min, year_max] of the recipe's component coverage."
},
"hint": { "type": "string" },
"ig_recipe_id": {
"type": ["string", "null"],
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
},
"trigger": {
"type": "string",
"enum": ["empty_year", "suppressed_component"],
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars this government reports in the requested years that the verb's underlying long view structurally excludes -- aggregate-published, carrying no harmonized code, or absent from summary_categories. This is NOT the same thing as 'excluded from the result': a component present in the view under a different category (a scoping choice, e.g. a different `category` or a narrower `expenditure_concept`) contributes 0 and never fires. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way, using the same underlying-view measurement, and can be 0 even on an 'empty_year' fire."
},
"suppressed_amount": {
"type": "number",
"description": "Full US dollars this government reports, in the recipe's component codes, in the requested years, that the verb's underlying long view structurally excludes (aggregate-published, carrying no harmonized code, or absent from summary_categories) -- summed across those years. This is NOT the same quantity as 'what the result excludes': a component present in the view under a different category or a narrower `expenditure_concept` is scoped out on purpose, counts as 0 here, and is not suppression. 0 does not always mean full coverage -- see 'trigger' and 'empty_year'. May be negative where Census publishes a negative `amt` for the excluded rows."
},
"suppressed_years": {
"type": "array",
"items": { "type": "integer" },
"description": "The requested years contributing to suppressed_amount."
},
"suppressed_codes": {
"type": "array",
"items": { "type": "string" },
"description": "The excluded component item codes, sorted."
}
}
}
},
"scope": { "type": "object" },
"codes_summed": { "type": "object" },
"aggregate_fallback": { "type": ["object", "null"] },
@@ -47,6 +94,16 @@
"items": { "type": "string" },
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
},
"balance_caveats": {
"type": ["object", "null"],
"description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.",
"properties": {
"not_gaap": { "type": "boolean" },
"not_gaap_note": { "type": "string" },
"coverage_window": { "type": "object" },
"truncated": { "type": "array", "items": { "type": "string" } }
}
},
"manifest": { "type": "object" },
"sql_query": { "type": "string" }
}
+7
View File
@@ -0,0 +1,7 @@
-- Category crosswalk. Numbered 11 (not with the other reference tables at
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
-- already exist when they register.
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+18 -1
View File
@@ -1,5 +1,22 @@
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
-- (summary_categories.category_type = 'expenditure'), never by item-code
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
-- every non-intergovernmental expenditure subtype: operations, capital,
-- assistance, interest, insurance_benefits.
--
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
-- 24-ig_long.sql).
CREATE OR REPLACE VIEW spending_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G')
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
)
AND NOT is_aggregate;
+14 -1
View File
@@ -1,5 +1,18 @@
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
--
-- Carries EVERY revenue subtype. Which of Census's two published concepts a
-- query actually returns is decided per revenue_concept in R
-- (.verb_spendrev), exactly as expenditure_concept narrows spending_long:
-- general = own_source + federal + state + local_aid (the default)
-- total = general + utility + liquor_store + insurance_trust
-- Census defines the first by subtracting the other three from the second
-- (manual section 4.3), so both concepts need all four families present here.
CREATE OR REPLACE VIEW revenue_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D')
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
)
AND NOT is_aggregate;
+11 -1
View File
@@ -1,6 +1,16 @@
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
-- classification, applied to harmonized_code (the code the row is folded
-- onto) rather than the published item_code. Safe because the harmonized
-- space is leaf-only and every harmonized_code in the corpus is a
-- summary_categories member (verified at fixture regen; a code the
-- crosswalk cannot classify would be silently dropped here).
CREATE OR REPLACE VIEW spending_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G');
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
);
+7 -1
View File
@@ -1,6 +1,12 @@
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
-- classification (every revenue subtype; the concept narrows in R), applied
-- to harmonized_code rather than the published item_code.
CREATE OR REPLACE VIEW revenue_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D');
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
);
+12 -5
View File
@@ -1,4 +1,6 @@
-- Intergovernmental expenditure rows (M = to local govts, L = to state govts).
-- Intergovernmental expenditure rows: crosswalk spend_subtype =
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
-- = state payments to school systems -- uscogdata#11, finding F-017).
--
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
@@ -9,10 +11,15 @@
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
-- the pipeline's recipe joins use.
--
-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up
-- the L-NN codes, so including it would double-count.
-- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
-- rolls up the L-NN codes, so including it would double-count. The crosswalk
-- deliberately carries no `--` family-total codes, so membership excludes it
-- (guarded by "the IG leg never includes the L-- family total" in
-- tests/testthat/test-expenditure-concept.R).
CREATE OR REPLACE VIEW ig_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L')
AND item_code NOT LIKE '%--';
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+9 -2
View File
@@ -8,8 +8,15 @@
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
-- while never dropping a row.
--
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
-- rather than the COALESCEd code: every IG harmonization target (M36) is
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
-- the column that exists on every row.
CREATE OR REPLACE VIEW ig_long_harmonized AS
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L')
AND item_code NOT LIKE '%--';
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+22
View File
@@ -0,0 +1,22 @@
-- Cash and security holdings, classified by crosswalk MEMBERSHIP on
-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot
-- do this job -- the X and Y families each span revenue, expenditure AND
-- balance).
--
-- These rows are STOCKS: a balance at a point in time, not a flow over a
-- fiscal year. Summing a stock with a flow is meaningless, which is why they
-- live behind a third view rather than as a subtype of either money view, and
-- why neither spending_long nor revenue_long can reach them.
--
-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era
-- aggregate-only holdings codes (X40/X41) are deliberately outside this view;
-- they are reachable only through the recipe path, which bypasses this filter
-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2).
CREATE OR REPLACE VIEW balance_long AS
SELECT *
FROM long
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'balance'
)
AND NOT is_aggregate;
-3
View File
@@ -1,3 +0,0 @@
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+16
View File
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW balance_annotated AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.balance_subtype
FROM balance_long s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
+76
View File
@@ -0,0 +1,76 @@
% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/balances.R
\name{cog_balances}
\alias{cog_balances}
\title{Cash and security holdings for one or more governments}
\usage{
cog_balances(
govid,
years,
category = NULL,
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL
)
}
\arguments{
\item{govid}{Canonical govid(s): a character vector, or a data frame with a
`canonical_govid` column (e.g. from [cog_gov_search()]).}
\item{years}{Integer vector of fiscal years.}
\item{category}{Optional character vector of categories to keep. One of
`"Fund Balances"`, `"Insurance Trust Balances"`,
`"Retirement System Holdings"`. There is deliberately no `subtype`
argument: for holdings, `category` is a strict coarsening of
`balance_subtype` (unlike the money verbs, where the two axes cross), so
every combination would be either redundant or empty.
`category = "Fund Balances"` is exactly the `general` family
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
one `dplyr::filter()` away.}
\item{per_capita}{Divide holdings by population. Note this is a **stock per
resident** (reserves per person), which is *not* comparable to
[cog_spending()]'s per-capita figures -- those are a flow per person.}
\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).}
\item{basis}{Accepted for uniformity with the money verbs, but currently a
**no-op**: `harmonization_map` carries no balance-code rows, so harmonized
and raw space are identical for holdings. Reported in
`provenance$basis_note`.}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]).
`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
wide era to the modern one.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`balance_subtype`, `category`, `amt_nominal`, `codes_included`,
`aggregate_fallback`, plus optional `amt_per_capita_nominal` and
`pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
`adjust_to_year` is set), and optional `amt_per_capita_real` (only when
**both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
nominal per-capita column to deflate otherwise). Amounts are full US
dollars.
Carries a `provenance` attribute matching
`inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
`not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
every balance subtype in the mounted corpus, not only the observed ones)
and `truncated` (the observed subtypes whose coverage falls short of the
requested years). `expenditure_concept`/`revenue_concept` are `NA` --
holdings are a stock, not a flow, so neither concept vocabulary applies.
}
\description{
Returns Census cash-and-security holdings (`category_type = "balance"`):
fund balances, retirement system holdings and insurance trust balances.
}
\section{Holdings are not GAAP fund balance}{
Census holdings are **gross** -- no liabilities are netted -- so a reserve
ratio built from them overstates what is actually available. They are not
comparable to a GAAP fund balance from an ACFR.
}
+11 -3
View File
@@ -7,8 +7,8 @@
cog_categories(type = NULL, pattern = NULL)
}
\arguments{
\item{type}{Either `NULL` (default, return both spending and revenue
rows), `"spending"`, or `"revenue"`.}
\item{type}{Either `NULL` (default, every row: expenditure, revenue and
balance), `"spending"`, `"revenue"`, or `"balance"`.}
\item{pattern}{Optional regex matched case-insensitively against the
`category` column (e.g. `"Police"` or `"Tax"`).}
@@ -22,7 +22,15 @@ Tibble with columns `category`, `category_type`, `subtype`,
Returns the category taxonomy exposed by the corpus's
`summary_categories` view, grouped to one row per
`(category, subtype)` pair. Use this to discover valid `category`
values for [cog_spending()] / [cog_revenue()] /
values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
[cog_geographic_rollup()] and to audit which Census item codes feed
each category.
}
\details{
`subtype` COALESCEs the crosswalk's three subtype columns, so it carries
`spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
`balance_subtype` on balance rows. Note that [cog_balances()] itself takes
no `subtype` argument — for holdings, `category` is a strict coarsening of
`balance_subtype` — but the value is surfaced here because it is the
discovery surface downstream consumers build their vocabulary from.
}
+10 -1
View File
@@ -11,7 +11,8 @@ cog_find_peers(
same_state = FALSE,
pop_range = c(0.7, 1.3),
is_ratio = TRUE,
max_peers = 10L
max_peers = 10L,
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -34,6 +35,14 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
`pop_range` is interpreted as absolute population counts.}
\item{max_peers}{Integer cap on the number of peers returned.}
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
most recent census year with an observed population, so a cohort is not
built from a sample year in which most of the candidate universe is
absent. `"consistent"` needs a year range, which cohort selection does not
have, so it selects like `"all"` and is carried on the result as
`attr(x, "coverage")` for [cog_peer_compare()].}
}
\value{
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
+28 -7
View File
@@ -10,7 +10,8 @@ cog_geographic_rollup(
years,
per_capita = FALSE,
adjust_to_year = NULL,
expenditure_concept = c("direct", "total")
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -29,12 +30,32 @@ are excluded from the result.}
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
single-government queries but cannot be used here because combining Total
across multiple layers of government double-counts intergovernmental
transfers (a state's payment to a school district is the same dollar the
district reports as its own Direct spending).}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across multiple layers of
government double-counts intergovernmental transfers (a state's payment
to a school district is the same dollar the district reports as its own
Direct spending); `"primary"` and `"direct"` combine safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.}
}
\value{
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
+33 -5
View File
@@ -11,7 +11,8 @@ cog_peer_compare(
years,
per_capita = TRUE,
adjust_to_year = NULL,
expenditure_concept = c("direct", "total")
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -29,10 +30,37 @@ population.}
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
single-government queries but cannot be used here because combining Total
across peer sets counts intergovernmental transfers twice.}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across peer sets counts
intergovernmental transfers twice; `"primary"` and `"direct"` combine
safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.
The comparison target is exempt from `"consistent"` balancing -- it is the
subject of the comparison, not a member of the cohort -- and the
`summary_*` quantiles are computed AFTER the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against
the cohort size: "3 of your 15 peers reported in FY2019".}
}
\value{
Tibble matching [cog_spending()]'s columns, plus a `role`
+27
View File
@@ -12,6 +12,7 @@ cog_revenue(
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE
)
}
@@ -57,6 +58,32 @@ argument is ignored and the result's provenance reports
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{revenue_concept}{Which of Census's two published revenue concepts to
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
values -- never as item-code first letters, which cannot classify
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
prefix `X` does the same):
* `"general"` (default) -- Census General Revenue: `own_source` +
`federal` + `state` + `local_aid`. The manual defines this concept by
subtraction (section 4.3: *"General revenue comprises all revenue
except that classified as liquor store, utility, or insurance trust
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
insurance trust revenue are all excluded.
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
`general` plus utility, liquor store, and insurance trust revenue
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
`X01`/`X02`/`X05`/`X08`).
The two are related by Census's own identity, `Total Revenue = General +
Utility + Liquor Store + Insurance Trust`.
Note that the employee-retirement (`X`) codes stop at FY2016, when those
systems moved out of the annual finance file into the separate Annual
Survey of Public Pensions, so a `"total"` series steps down at the
FY2016/FY2017 seam for reasons that are about collection scope rather
than revenue (series breaks `SB197`-`SB202`).}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
+29 -17
View File
@@ -12,7 +12,7 @@ cog_spending(
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
expenditure_concept = c("direct", "total"),
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE
)
}
@@ -58,22 +58,34 @@ argument is ignored and the result's provenance reports
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{expenditure_concept}{`"direct"` (default) returns only the
government's own direct spending (item codes `E`/`F`/`G`), unchanged
from prior releases. `"total"` additionally UNIONs in the
intergovernmental leg -- payments to local governments (`M` codes) and
to the state government (`L` codes, excluding the `L--` family-total
rollup) -- so results gain rows with `spend_subtype ==
"intergovernmental"`. Requires the active corpus's `summary_categories`
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
`uscogdata_ig_categories_unsupported` on an older corpus rather than
silently under-reporting. Mutually exclusive with `recipe` (a recipe
already defines its own component codes). **Do not sum `"total"`
results across levels of government** (e.g. state + county + city):
a state's `M12` payment to a school district is the same dollar the
district reports as its own direct `E12`, so summing both double-counts
it. This matters in particular with [cog_geographic_rollup()], which
sums across exactly that kind of multi-layer government set.
\item{expenditure_concept}{Which spending concept to return. Concepts are
defined as sets of the crosswalk's `spend_subtype` values -- never as
item-code first letters, which cannot classify correctly (prefix `Y`
alone spans revenue, expenditure, and balance codes):
* `"primary"` (default) -- the government's own service provision:
`operations` + `capital` + `assistance` subtypes.
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
`interest` (interest on debt) and `insurance_benefits` (insurance
trust benefit payments, e.g. pensions -- Census manual section
5.2.2.1 includes payments to retirees in Direct).
* `"total"` -- `direct` plus the intergovernmental leg: payments to
local governments (`M` codes), to the state government (`L` codes,
excluding the `L--` family-total rollup), and state payments to
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
`spend_subtype == "intergovernmental"`. Requires the active corpus's
`summary_categories` to carry M/L rows (added by cog_pipeline PR
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
older corpus rather than silently under-reporting. Mutually
exclusive with `recipe` (a recipe already defines its own component
codes).
**Do not sum `"total"` results across levels of government** (e.g.
state + county + city): a state's `M12` payment to a school district is
the same dollar the district reports as its own direct `E12`, so
summing both double-counts it. This matters in particular with
[cog_geographic_rollup()], which sums across exactly that kind of
multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
File diff suppressed because it is too large Load Diff
+281
View File
@@ -0,0 +1,281 @@
# `cog_balances()` — a reader surface for cash and security holdings
**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26`
**Date:** 2026-08-03 · **Status:** design, awaiting approval
Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped
with `#11`/`#12` and is asserted at both view and verb level. This spec covers
requirement 2 only: a way to query holdings.
## Decision: a verb, not an argument
`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`.
Holdings are a **stock** — a balance at a point in time — while the money verbs
return **flows** over a fiscal year. The flow verbs' whole argument vocabulary
is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe
which flows Census aggregates into a published total, and `complete=` fills a
grid of fiscal-year cells. Overloading a money verb would put a stock behind
arguments that all assume a flow.
## The 14 codes
Measured against the published corpus 2026-08-03, not transcribed from the
issue. `year_min`/`year_max` are observed row extents.
| `balance_subtype` | `category` | codes | observed years |
|---|---|---|---|
| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 |
| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 |
| | | `X47` | 1988–2016 |
| | | `X30`, `Z77`, `Z78` | 2012–2016 |
| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 |
| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 |
| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 |
## Architecture
### Two new views
Mirroring the `revenue_long` / `revenue_annotated` pair exactly:
- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate`
- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and
`summary_categories`, exposing `category`, `category_type`, `balance_subtype`
`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register
with no new registration code.
### A third gate list in `R/views.R`
`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails
at registration time, not at query time. `46-balance_annotated.sql` selects
`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`.
That arrived without a `schema_version` bump, so neither existing gate applies:
`.harmonization_view_files` keys on `schema_version`, `.representation_view_files`
on the presence of a *file*. The discriminator here is a **column on an existing
table**.
```r
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
```
gated by probing `summary_categories` for `balance_subtype`, with
`cog_balances()` erroring cleanly via `.require_balance_support()` on an older
corpus — mirroring how `.require_schema_v5()` gates the harmonized views.
### `R/balances.R` — a dedicated path, not `.verb_spendrev()`
`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and
`complete=` grid are all flow-specific, and four verbs depend on it. Threading a
third mode through it adds branching to shared code for no reuse benefit.
Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`,
`.build_corpus_break_refs()`, the population join, `.inflate()`, and
`.coerce_govid_input()`.
Following the package's real two-layer convention: **view definitions** live in
`inst/sql/`; **query construction** is inline `sprintf()` in R, as in
`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R
files", which the verb layer has never obeyed. Corrected in a separate commit —
see Out of scope.)
## Signature
```r
cog_balances(govid, years,
category = NULL, # Fund Balances | Insurance Trust Balances |
# Retirement System Holdings
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL)
```
Returns a `tbl_df` with a `provenance` attribute, like every other verb.
**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`,
and `subtype` — see below.
**`per_capita` is offered.** Holdings per resident is a real measure (pension
assets per capita, fund balance per resident). The roxygen `@param` states
plainly that this is a *stock per resident* and is **not** comparable to
`cog_spending()`'s per-capita figures.
**`basis` is currently a no-op** — `harmonization_map` has zero balance-code
rows, so harmonized and raw are identical for holdings. Kept for uniformity
with the money verbs (the API would otherwise special-case), and
`provenance$basis_note` says so outright rather than letting it look meaningful.
**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
to the modern one:
```
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
```
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
other basis view. That is by design, not a defect:
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
era exposes these split families *only* as aggregates, and that the recipe join
must therefore **not** filter `is_aggregate` — safe by construction, because
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
every component is year-scoped, so no double-count is possible. § 1 records the
matching decision that the planned `X40→Z77` harmonization *map* rows were
dropped and the continuity ships as recipes instead, which is why
`harmonization_map` has no balance-code rows.
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
as intended.
### No `subtype` argument: `category` is a strict coarsening
`balance` is the only `category_type` in which `category` and the subtype column
are **not** orthogonal. Measured against the published crosswalk:
| `category_type` | subtypes spanning more than one category |
|---|---|
| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) |
| revenue | 1 of 7 (`own_source`) |
| **balance** | **0 of 5** |
For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire)
× *economic character* (operations, capital) — so both earn their place. For
balance the relation is a strict tree:
```
Fund Balances = {general} W01 W31 W61
Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78
Insurance Trust Balances = {unemployment_trust,
workers_comp_trust,
other_insurance_trust} Y07 Y08 Y21 Y61
```
Exposing both would therefore admit no useful combination. Of the 15 possible
pairs, 3 are redundant (the subtype already implies its category) and **12 are
guaranteed empty for every government in every year** — and an impossible query
would fail by returning an empty tibble, which reads as "this government holds
none" rather than "you asked a contradiction."
Dropping `subtype` also keeps the verb aligned with the rest of the package: no
uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in
`.verb_spendrev()`, and the API layers its own `subtype` row filter on top
(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for
`/balances`.
`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the
`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is
isolating one of the three insurance funds in a single argument;
`balance_subtype` remains a returned column, so that is one `dplyr::filter()`
away.
## Caveat surfacing
`provenance$balance_caveats`, always present, plus one `cli_inform()` per
session per caveat class when a query actually touches an affected family or
year. Structured so `cog-api#26` can forward the fields verbatim.
Verified against `series_breaks.csv`, not assumed:
| # | Caveat | Covered by existing machinery? |
|---|---|---|
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed |
On caveat 4's second condition: `.build_series_break_refs()` matches
`break_year BETWEEN min(years) AND max(years)`, so a request spanning only
2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a
series sits entirely after the change, on one consistent basis, and flagging a
break it never crosses would be noise. The same rule is applied deliberately in
`.build_corpus_break_refs()`. An earlier draft of this row omitted the span
condition and overclaimed.
`coverage_window` is derived per observed subtype family from the corpus, never
hardcoded, so it stays correct as the corpus grows.
`series_break_refs` and `corpus_break_refs` are otherwise populated by the
existing code-driven builders and need no change.
## Testing
New `tests/testthat/test-balances.R`. The bundled fixture covers all four
fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y`
throughout — so every test below runs offline.
- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing
the already-asserted forward guard. Absence is verified against the raw corpus
via `read_parquet` on `data/long`, never through the verb that creates it.
- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019;
`coverage_window` reports the termination and the console message fires once.
- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the
measured table above; the FY2002 valuation caveat fires only when the year
range crosses 2002 *and* touches `employee_retirement`.
- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`.
- **`category = "Fund Balances"` is the `general` family.** Returns exactly
`W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted
rather than assumed.
- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to
exactly one `category`. Asserted against the crosswalk so that an upstream
change breaking the tree — which would silently make `category` lossy —
fails here rather than in a user's analysis.
- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
`X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
regression here would silently truncate a 45-year series to five.
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
FY2002 carries both in `provenance$series_break_refs`, so the book → market
basis change is disclosed wherever `X40`/`X41` are actually observed.
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
`summary_categories` lacks `balance_subtype`.
## Out of scope, tracked separately
1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
the seven holdings codes in `series_breaks.csv`. There is currently **no**
entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
`docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
exactly this — *"candidate new `series_breaks.csv` entries (recommend
`with_caution` documentation rows, no map action)"*. The follow-through never
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
what keeps this out of the `joinable = "no"` identity-change rule, which
would otherwise oblige a harmonization-map row.
Verify the break corpus-wide and census-to-census before writing the rows.
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
signpost when it lands.
**Superseded:** an earlier draft of this spec proposed adding
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
are the documented § 1 decision, and the recipe path reaches them by design.
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
`param_contract`, and the `plumber.R` route signature. Lands after this.
**Two contract facts the API must carry forward**, both settled during
implementation and easy to get wrong from the outside:
- `provenance$balance_caveats$coverage_window` is **corpus-scoped, not
result-scoped**. It reports the observed year extent of *every* balance
subtype in the corpus, not only the subtypes a given query returned — so a
`category = "Fund Balances"` query still returns all five windows. That is
deliberate: the windows describe what the corpus holds, which is what a
consumer needs in order to know what it did *not* ask for. The sibling
field `truncated` is the result-scoped one. Documented in
`inst/schemas/provenance-v1.json` and mutation-guarded against silent
inversion.
- `balance_caveats` appears **only** on `cog_balances()` results. It is
absent from `cog_spending()`/`cog_revenue()` provenance, and the schema
says so — an API layer that assumes it is universal will read `NULL`.
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7
SQL views (there are 21), 181 tests (716), a two-year fixture (four years),
and a "never inline SQL" rule the verb layer does not follow.
+33
View File
@@ -154,3 +154,36 @@ with_corpus_missing_ig_categories <- function(code) {
}, add = TRUE)
force(code)
}
# Copy the bundled fixture to a temp dir with summary_categories.parquet
# rewritten to DROP the balance_subtype column, then run `code` against it.
# Models a corpus published before cog_pipeline #76/#77. schema_version is
# left untouched deliberately: that change shipped without a version bump, so
# column presence is the only honest signal -- this helper is what proves the
# package keys off it. Mirrors with_corpus_missing_ig_categories().
with_corpus_missing_balance_subtype <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf(
"COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s))
TO %s (FORMAT PARQUET)",
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
))
file.remove(cats_path)
file.rename(filtered_path, cats_path)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
+539
View File
@@ -0,0 +1,539 @@
test_that("balance views register and carry only balance codes", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)$table_name
expect_true(all(c("balance_long", "balance_annotated") %in% views))
# Every item_code in balance_long is a category_type = 'balance' member.
leak <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM balance_long
WHERE item_code NOT IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance')"
)$n
expect_identical(as.integer(leak), 0L)
# balance_annotated exposes the subtype column the verb groups on.
cols <- DBI::dbGetQuery(con,
"SELECT column_name FROM information_schema.columns
WHERE table_name = 'balance_annotated'"
)$column_name
expect_true(all(c("category", "category_type", "balance_subtype") %in% cols))
})
test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", {
# Every category_type = 'balance' item_code in the bundled fixture has
# is_aggregate = FALSE for every row of every year -- there is no real row
# that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An
# assertion against the live fixture (`WHERE is_aggregate` returns 0) is
# therefore vacuous: it passes identically whether or not the view's
# predicate is present. As with the 22-/23- and 24-/25- tests above, this
# reads the real inst/sql/26-balance_long.sql text off disk and executes it
# -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against
# a synthetic hive-partitioned parquet tree that DOES contain an aggregate
# row under a real balance item_code (W01), so a regression that drops the
# predicate changes which rows survive.
skip_if_no_corpus()
tmp <- withr::local_tempdir()
part_dir <- file.path(tmp, "data", "long", "year=2004")
dir.create(part_dir, recursive = TRUE)
part_path <- file.path(part_dir, "part-0.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives
('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate`
) AS t(canonical_govid, item_code, amt, is_aggregate)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('W01', 'Fund Balances', 'balance', NULL, NULL, 'general')
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
}
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("26-balance_long.sql"))
rows <- DBI::dbGetQuery(con,
"SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid"
)
expect_equal(nrow(rows), 1L)
expect_equal(rows$canonical_govid, "bal-A")
expect_equal(rows$amt, 100)
})
test_that("balance views are skipped on a corpus without balance_subtype", {
skip_if_no_corpus()
with_corpus_missing_balance_subtype({
con <- cog_open()
on.exit(cog_close())
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)$table_name
# Registration must SKIP them, not error -- an older corpus stays usable.
expect_false(any(c("balance_long", "balance_annotated") %in% views))
expect_true("revenue_long" %in% views)
# ...and calling the verb on such a corpus must hit
# .require_balance_support()'s curated abort (spec § Testing: "Gating"),
# not a DuckDB binder error naming a view that was never registered.
# Asserted on the CLASS: removing the guard still errors, so a bare
# expect_error() would pass on the regression.
expect_error(
cog_balances("550000227544", 2019),
class = "uscogdata_no_balance_support"
)
})
})
test_that("cog_balances returns holdings for a government that has them", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019)
expect_s3_class(r, "tbl_df")
expect_true(nrow(r) > 0L)
expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype",
"category", "amt_nominal") %in% names(r)))
expect_identical(sort(unique(r$category)),
c("Fund Balances", "Insurance Trust Balances"))
expect_false(is.null(attr(r, "provenance")))
expect_identical(attr(r, "provenance")$verb, "cog_balances")
})
})
test_that('category = "Fund Balances" is exactly the general family', {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
expect_identical(unique(r$balance_subtype), "general")
codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
expect_identical(codes, c("W01", "W31", "W61"))
})
})
test_that("no flow code can reach cog_balances", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
# The expected set is read from the RAW corpus, never from the verb --
# verifying an absence through the filter that creates it proves nothing.
# A fresh, direct DuckDB connection against the raw parquet files (never
# cog_open()'s session, never balance_long/balance_annotated) reads
# parquet natively -- no arrow dependency needed (see CLAUDE.md).
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
balance_codes <- DBI::dbGetQuery(con2, sprintf(
"SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'",
uscogdata:::.sql_lit_chr(cats_path)
))$item_code
expect_true(length(got) > 0L)
expect_true(all(got %in% balance_codes))
})
})
test_that("every balance_subtype maps to exactly one category", {
skip_if_no_corpus()
# Dropping the `subtype` argument is only safe while this tree holds. If the
# pipeline ever gives a balance subtype a second category, `category` becomes
# a lossy filter -- fail HERE rather than in a user's analysis. Read via a
# fresh direct DuckDB connection against the raw parquet file, not through
# any registered view.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
b <- DBI::dbGetQuery(con2, sprintf(
"SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'",
uscogdata:::.sql_lit_chr(cats_path)
))
per_subtype <- tapply(b$category, b$balance_subtype,
function(x) length(unique(x)))
expect_true(all(per_subtype == 1L))
})
test_that("cog_balances records found + missing govids in provenance", {
skip_if_no_corpus()
with_fixture_corpus({
suppressMessages(
r <- cog_balances(c("550000227544", "XXXINVALID"), 2019)
)
prov <- attr(r, "provenance")
expect_equal(sort(prov$scope$govids_found), "550000227544")
expect_equal(sort(prov$scope$govids_missing), "XXXINVALID")
})
})
test_that("per_capita divides holdings by population", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_balances("550000227544", 2019, category = "Fund Balances")
pc <- cog_balances("550000227544", 2019, category = "Fund Balances",
per_capita = TRUE)
expect_true("amt_per_capita_nominal" %in% names(pc))
expect_true("pop_source" %in% names(pc))
expect_identical(pc$amt_nominal, plain$amt_nominal)
# Assert against the denominator read from the corpus, NOT against a
# quantity derived from amt_per_capita_nominal itself -- dividing the
# column back out would be tautological and would pass on any value.
pop <- DBI::dbGetQuery(cog_open(), sprintf(
"SELECT population FROM gov_population_yearly
WHERE canonical_govid = %s AND year = 2019",
uscogdata:::.sql_lit_chr("550000227544")
))$population
expect_length(pop, 1L)
expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop,
tolerance = 1e-8)
prov <- attr(pc, "provenance")
expect_true(prov$transformations$per_capita$applied)
})
})
test_that("adjust_to_year adds real dollars", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
adjust_to_year = 2020)
expect_true("amt_real" %in% names(r))
# 2012 dollars inflated to 2020 must exceed nominal.
expect_true(all(r$amt_real > r$amt_nominal))
prov <- attr(r, "provenance")
expect_true(prov$transformations$inflation$applied)
expect_identical(prov$transformations$inflation$base_year, 2020L)
})
})
test_that("per_capita and adjust_to_year compose", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
per_capita = TRUE, adjust_to_year = 2020)
expect_true("amt_per_capita_real" %in% names(r))
# The per-capita column must be deflated by the SAME factor as the level
# column -- this is what the ordering at R/balances.R:101-103 guarantees.
# .attach_real_dollars() silently no-ops on the per-capita leg when
# amt_per_capita_nominal does not exist yet (R/spending.R:664), so
# reversing those two calls drops this column with no error at all.
expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal,
r$amt_real / r$amt_nominal, tolerance = 1e-8)
# And the documented condition is a conjunction: adjust_to_year ALONE
# must not produce amt_per_capita_real (pins the @return wording).
r2 <- cog_balances("550000227544", 2012, category = "Fund Balances",
adjust_to_year = 2020)
expect_true("amt_real" %in% names(r2))
expect_false("amt_per_capita_real" %in% names(r2))
})
})
# --- input validation ------------------------------------------------------
test_that("cog_balances validates its inputs like the money verbs", {
skip_if_no_corpus()
with_fixture_corpus({
G <- "550000227544"
# Pinned to the message, not bare expect_error(): every one of these
# already produces *some* error or *some* quiet wrong answer today --
# years = integer(0) leaks `Parser Error ... AND year IN ()` with the
# generated SQL, recipe = c("a","b") throws "the condition has length > 1",
# and the govid/category cases return 0 rows with no error at all.
expect_error(cog_balances(G, integer(0)), "non-empty integer vector")
expect_error(cog_balances(character(0), 2019), "non-empty character vector")
expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL")
expect_error(cog_balances(G, 2019, recipe = c("a", "b")),
"length-1 character string")
})
})
test_that("validation runs after govid coercion, so a data frame still works", {
skip_if_no_corpus()
with_fixture_corpus({
# .validate_verb_inputs() asserts is.character(govid); it must therefore
# run AFTER .coerce_govid_input(), never before, or the documented
# data-frame input (cog_gov_search() output) would abort.
df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE)
r <- suppressMessages(cog_balances(df, 2019))
expect_true(nrow(r) > 0L)
expect_identical(unique(r$canonical_govid), "550000227544")
})
})
test_that("recipe and category are mutually exclusive", {
skip_if_no_corpus()
with_fixture_corpus({
expect_error(
cog_balances("550000227544", c(2011, 2012),
category = "Fund Balances",
recipe = "cash_securities_z77_wide"),
class = "uscogdata_recipe_category_conflict"
)
})
})
# --- recipe = : the wide-era holdings bridge -------------------------------
test_that("recipe bridges the wide era into the modern one", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012),
recipe = "cash_securities_z77_wide")
# .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait,
# not specific to this recipe -- see the money-verb recipe tests, which
# only ever assert on it with expect_equal), so compare numerically rather
# than with expect_identical()'s type-strict comparison.
expect_equal(sort(r$year), c(2011, 2012))
# The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE
# and therefore invisible to balance_long. If the recipe path ever starts
# filtering aggregates, a 45-year series silently truncates to five --
# this is the regression guard for phase_r_harmonization_review.md § 0.2.
codes <- attr(r, "provenance")$codes_summed$observed
expect_true("X40" %in% codes)
expect_true("Z77" %in% codes)
expect_true(all(r$amt_nominal > 0))
prov <- attr(r, "provenance")
expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide")
})
})
test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", {
skip_if_no_corpus()
with_fixture_corpus({
# 2002 is in the year vector deliberately, and must stay -- do not
# "simplify" this back to c(2011, 2012).
#
# .build_series_break_refs() (R/series_breaks.R, shared with every verb)
# matches breaks with `break_year BETWEEN min(years) AND max(years)`, and
# SB195's break_year is 2002. A c(2011, 2012) span never crosses the
# FY2002 book -> market change -- that whole span sits after it, on one
# consistent basis -- so NOT disclosing SB195 there is correct behaviour,
# not a gap (same reasoning as the "a request that never crosses the
# boundary is not affected by it" comment on .build_corpus_break_refs()).
#
# The property actually worth testing is: a recipe query that observes
# X40 AND spans FY2002 discloses SB195. This fixture has no 2002
# partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020
# partitions exist), so including 2002 in `years` widens the
# break-matching window without changing which rows the recipe join
# returns -- verified empirically: r$year below is exactly {2011, 2012}
# whether or not 2002 is in the request (see task-4-report.md).
# Removing 2002 would silently turn this back into the non-crossing case
# above and destroy the test's purpose.
r <- cog_balances("550000227544", c(2002, 2011, 2012),
recipe = "cash_securities_z77_wide")
expect_equal(sort(r$year), c(2011, 2012))
refs <- attr(r, "provenance")$series_break_refs
# SB195 sits on fin_code X40; it can only fire where X40 is observed,
# which is exactly the recipe path.
expect_true("SB195" %in% refs)
})
})
test_that("the second holdings bridge works too", {
skip_if_no_corpus()
with_fixture_corpus({
# X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011
# and Z78 in 2012, so both legs are exercised.
r <- cog_balances("550000227544", c(2011, 2012),
recipe = "cash_securities_z78_wide")
codes <- attr(r, "provenance")$codes_summed$observed
expect_true(all(c("X41", "Z78") %in% codes))
expect_equal(sort(r$year), c(2011, 2012))
})
})
test_that("an unknown recipe id is rejected", {
skip_if_no_corpus()
with_fixture_corpus({
# Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84).
# Without it the test is non-discriminating: deleting the validation call
# leaves .recipe_components() returning 0 rows and comps$label[[1]]
# throwing "subscript out of bounds", which a bare expect_error() accepts
# while the user loses the curated "valid recipe ids are ..." message.
expect_error(
cog_balances("550000227544", 2019, recipe = "no_such_recipe"),
class = "uscogdata_unknown_recipe"
)
})
})
# --- balance_caveats: GAAP disclosure + measured coverage windows ----------
test_that("balance_caveats is always present and flags the GAAP distinction", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", 2019)
cav <- attr(r, "provenance")$balance_caveats
expect_false(is.null(cav))
expect_true(cav$not_gaap)
})
})
test_that("coverage_window is computed from the corpus, not hardcoded", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
cav <- attr(r, "provenance")$balance_caveats
# Read the "general" family's true year extent independently, via a
# fresh DuckDB connection against the raw parquet files (never through
# balance_long/.balance_caveats() itself, and never via arrow -- this
# package reads parquet through DuckDB only, see CLAUDE.md). Replicates
# the same predicates 26-balance_long.sql applies (category_type =
# 'balance', NOT is_aggregate) so this is a faithful, independent
# measurement rather than a re-statement of the view under test.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet")
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
obs <- DBI::dbGetQuery(con2, sprintf(
"SELECT MIN(l.year) AS y0, MAX(l.year) AS y1
FROM read_parquet(%s, hive_partitioning = true) l
JOIN read_parquet(%s) c USING (item_code)
WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate",
uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path)
))
expect_identical(as.integer(cav$coverage_window$general),
c(as.integer(obs$y0), as.integer(obs$y1)))
})
})
test_that("coverage_window covers every corpus subtype, not just observed ones", {
skip_if_no_corpus()
with_fixture_corpus({
# Deliberate contract (provenance-v1.json): the window block is corpus-
# scoped so a caller can ask "is there a family I missed?", while
# `truncated` is the observed-scoped field. A single-category query must
# therefore still report every balance family in the mounted corpus.
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
expect_identical(unique(r$balance_subtype), "general")
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
all_subtypes <- DBI::dbGetQuery(con2, sprintf(
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
WHERE balance_subtype IS NOT NULL",
uscogdata:::.sql_lit_chr(cats_path)
))$balance_subtype
cav <- attr(r, "provenance")$balance_caveats
expect_setequal(names(cav$coverage_window), all_subtypes)
expect_true(length(all_subtypes) > 1L)
# ...while `truncated` stays scoped to what this query actually observed.
expect_true(all(cav$truncated %in% unique(r$balance_subtype)))
})
})
test_that("the corpus-constant coverage windows are memoised per session", {
skip_if_no_corpus()
with_fixture_corpus({
# The windows query has no govid/year predicate: its answer depends only
# on which corpus is mounted, so re-running the full balance_long scan on
# every call is pure waste (35% of verb runtime on the fixture). Same
# memoise-and-invalidate pattern as .uscogdata_env$manifest.
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
suppressMessages(cog_balances("550000227544", 2019))
memo <- uscogdata:::.uscogdata_env$balance_coverage_windows
expect_false(is.null(memo))
expect_true("general" %in% names(memo))
uscogdata:::cog_close()
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
})
})
test_that("a request past a family's coverage window is flagged", {
skip_if_no_corpus()
with_fixture_corpus({
# employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in
# the LIVE corpus -- Census moved employee retirement reporting to the
# Annual Survey of Public Pensions after that year. This bundled FIXTURE
# doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present),
# so the family's *observed* max here is 2012, not 2016. Either way the
# requested span (2012, 2019) reaches past what the family covers in
# THIS corpus, which is what makes .balance_caveats() flag it -- the
# assertion below is about the fixture's measured window, not the FY2016
# live-corpus cutoff.
r <- cog_balances("550000227544", c(2012, 2019))
cav <- attr(r, "provenance")$balance_caveats
expect_true("employee_retirement" %in% cav$truncated)
})
})
test_that("the provenance schema documents balance_caveats", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE
)
expect_true("balance_caveats" %in% names(sch$properties))
})
test_that("cog_explain surfaces the balance caveats", {
skip_if_no_corpus()
with_fixture_corpus({
# Asserted on the RENDERED text, not on prov$balance_caveats: the field
# is already covered above, and the once-per-session cli_inform() means
# cog_explain() is the only surface a caller who missed (or suppressed)
# the first message can still audit.
r <- suppressMessages(cog_balances("550000227544", c(2012, 2019)))
# Both streams: cli routes most of its output through conditions that
# land on stderr, so a stdout-only capture would be empty (the pattern
# used throughout test-explain.R).
out <- paste(c(capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")),
collapse = "\n")
expect_match(out, "GAAP")
expect_match(out, "employee_retirement")
})
})
test_that("cog_explain on a money-verb result has no balance caveat section", {
skip_if_no_corpus()
with_fixture_corpus({
r <- suppressMessages(cog_spending("550000227544", 2019))
out <- paste(c(capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")),
collapse = "\n")
# Guard against the capture itself being vacuous: the section must be
# absent from output that demonstrably contains the rest of the report.
expect_match(out, "Data vintage")
expect_false(grepl("GAAP", out))
})
})
test_that("the caveat message fires once per session", {
skip_if_no_corpus()
with_fixture_corpus({
expect_message(cog_balances("550000227544", 2019), "not.*GAAP")
expect_no_message(cog_balances("550000227544", 2020))
})
})
+58 -3
View File
@@ -8,7 +8,14 @@ test_that("cog_categories returns all categories grouped by subtype", {
expect_gt(nrow(r), 10L)
# corpus preserves Census-native "expenditure" vocabulary; the API takes
# "spending" as a friendlier alias.
expect_setequal(unique(r$category_type), c("expenditure", "revenue"))
#
# `balance` joined as a third category_type with the cash-and-security
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
# money verb, so it surfaces every category_type the corpus carries -- the
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
# return a balance row.
expect_setequal(unique(r$category_type),
c("expenditure", "revenue", "balance"))
})
test_that("cog_categories(type = 'spending') returns only expenditure rows", {
@@ -18,8 +25,13 @@ test_that("cog_categories(type = 'spending') returns only expenditure rows", {
# "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
# carrying dollars now maps to a category.
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
# joined with the I/Q/Y flow batch -- the last two characters of Census's
# expenditure taxonomy. `interest` is what makes the three-concept model
# computable: primary = direct minus debt service.
expect_true(all(r$subtype %in%
c("operations", "capital", "intergovernmental", "assistance")))
c("operations", "capital", "intergovernmental", "assistance",
"interest", "insurance_benefits")))
})
test_that("cog_categories surfaces the intergovernmental spending subtype", {
@@ -37,8 +49,14 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
skip_if_no_corpus()
r <- cog_categories(type = "revenue")
expect_true(all(r$category_type == "revenue"))
# The four non-general subtypes are deliberately NOT own_source: Census's
# General Revenue excludes insurance trust (Y01 alone is $1.31T corpus-wide,
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
# (A90) revenue by definition, which is what makes both of its published
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
expect_true(all(r$subtype %in%
c("own_source", "federal", "state", "local_aid")))
c("own_source", "federal", "state", "local_aid",
"insurance_trust", "utility", "liquor_store")))
})
test_that("cog_categories(pattern = ...) filters case-insensitively", {
@@ -75,3 +93,40 @@ test_that("cog_categories sorted by category_type, category, subtype", {
test_that("cog_categories rejects invalid type", {
expect_error(cog_categories(type = "both"), "type")
})
test_that("cog_categories() surfaces balance subtypes", {
skip_if_no_corpus()
with_fixture_corpus({
cc <- cog_categories()
b <- cc[cc$category_type == "balance", ]
expect_true(nrow(b) > 0L)
# Every balance row must carry its subtype. Before the COALESCE included
# balance_subtype these were all NA, which silently made the balance
# taxonomy undiscoverable -- cog-api derives its subtype vocabulary from
# this function, so an NA here becomes an unusable API parameter.
expect_false(any(is.na(b$subtype)))
# The exact set, read independently from the crosswalk rather than from
# the function under test.
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
p <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
want <- DBI::dbGetQuery(con2, sprintf(
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
WHERE category_type = 'balance' AND balance_subtype IS NOT NULL
ORDER BY 1", uscogdata:::.sql_lit_chr(p)))$balance_subtype
expect_true(length(want) > 1L)
expect_identical(sort(unique(b$subtype)), sort(want))
})
})
test_that('cog_categories(type = "balance") filters to holdings', {
skip_if_no_corpus()
with_fixture_corpus({
b <- cog_categories(type = "balance")
expect_true(nrow(b) > 0L)
expect_identical(unique(b$category_type), "balance")
expect_false(any(is.na(b$subtype)))
})
})
+14 -9
View File
@@ -14,9 +14,11 @@
# The (subtype, category) cells that SHOULD exist for one government-year:
# every code in force for that government's type, mapped through
# summary_categories, matching the verb's flow prefixes and excluding
# aggregate-flagged codes (which spending_long/revenue_long drop).
raw_expected_cells <- function(govid, year, prefixes, subtype_col) {
# summary_categories, matching the verb's crosswalk subtype scope (the
# default concept, `primary`, is operations/capital/assistance -- see
# uscogdata#11) and excluding aggregate-flagged codes (which
# spending_long/revenue_long drop).
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
wt_raw_query(sprintf(
@@ -27,15 +29,18 @@ raw_expected_cells <- function(govid, year, prefixes, subtype_col) {
WHERE x.canonical_govid = '%s'
AND cs.year = %d
AND NOT cs.is_aggregate
AND LEFT(cs.item_code, 1) IN (%s)
AND c.category IS NOT NULL
AND c.%s IS NOT NULL",
AND c.%s IN (%s)",
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
q("summary_categories.parquet"), govid, year,
paste0("'", prefixes, "'", collapse = ","), subtype_col
subtype_col, paste0("'", subtypes, "'", collapse = ",")
))
}
# The default expenditure concept's subtype scope, mirrored from
# R/spending.R's .spend_subtypes_primary.
primary_subtypes <- c("operations", "capital", "assistance")
test_that("complete = FALSE is the default and changes nothing", {
skip_if_no_corpus()
with_fixture_corpus({
@@ -54,7 +59,7 @@ test_that("complete = TRUE round-trips a dense-source year to the pre-sparsifica
# reproduce that cell set exactly.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
c("E", "F", "G"), "spend_subtype")
primary_subtypes, "spend_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$spend_subtype, r$category),
@@ -108,7 +113,7 @@ test_that("the fill is scoped to each government's own type", {
# code_set puts in force for type 1 (county) specifically.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
county_cells <- raw_expected_cells("121011212191", 2011L,
c("E", "F", "G"), "spend_subtype")
primary_subtypes, "spend_subtype")
expect_true(all(r$category %in% county_cells$category))
})
})
@@ -128,7 +133,7 @@ test_that("cog_revenue() completes on its own flow", {
with_fixture_corpus({
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
c("T", "A", "U", "B", "C", "D"),
c("own_source", "federal", "state", "local_aid"),
"revenue_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$revenue_subtype, r$category),
+14 -3
View File
@@ -30,7 +30,6 @@ wt_coverage <- function(x) {
}
test_that("multi-government aggregates disclose reporting coverage on every result", {
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
# -- F-020: geographic rollups -------------------------------------------
# Wisconsin's city/village universe is 608 governments. On the bundled
@@ -49,10 +48,22 @@ test_that("multi-government aggregates disclose reporting coverage on every resu
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
# Cross-check against the raw partitions, scoped to the SAME universe the
# rollup was given -- the 608 govids above. Scoping instead on the long
# table's own `type`/`fips_state` asks a different question and answers 595:
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
# identity, when they were townships) while the xwalk lists them as
# govs_type = 2 (their present identity, as villages). Schema v6 made the
# long table's geography present-harmonized and moved as-of-year to the
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
# in R/manifest.R. n_units_reporting counts against the requested universe,
# so 597 is the number that answers "how many of the governments I asked
# about reported".
raw_2012 <- wt_raw_query(paste0(
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ",
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate"))
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
"AND canonical_govid IN (",
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
# -- F-023: peer cohorts --------------------------------------------------
+12 -3
View File
@@ -13,9 +13,13 @@ test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
}
})
test_that("expenditure_concept defaults to direct and preserves today's numbers", {
test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
gov <- "010000226085" # Alabama state government
base <- cog_spending(gov, years = 2019, category = "Police")
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
# Police maps only to operations/capital codes (E62/F62/G62), so the
# direct concept's extra subtypes (interest, insurance_benefits) cannot
# contribute and the two concepts must agree exactly here.
expl <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "direct")
expect_equal(base$amt_nominal, expl$amt_nominal)
@@ -59,7 +63,9 @@ test_that("the IG leg never includes the L-- family total", {
codes <- DBI::dbGetQuery(con,
"SELECT DISTINCT item_code FROM ig_long")$item_code
expect_false(any(grepl("--$", codes)))
expect_true(all(substr(codes, 1, 1) %in% c("M", "L")))
# Q joined the IG family with the crosswalk-membership rewrite
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
})
test_that("expenditure_concept rejects unknown values", {
@@ -246,9 +252,12 @@ test_that("both cross-government verbs still accept the direct default", {
})
test_that("provenance always records the expenditure concept", {
d <- cog_spending("010000226085", years = 2019, category = "Police")
p <- cog_spending("010000226085", years = 2019, category = "Police")
d <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
expect_equal(attr(t, "provenance")$expenditure_concept, "total")
# The note explains the non-obvious part: how legacy IG was assembled.
+43 -6
View File
@@ -19,8 +19,6 @@
# also check the FY2022 numbers above.
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
testthat::skip("Blocked on uscogdata#11 (findings F-012, F-017, F-018)")
mad <- "552025209777" # MADISON CITY, WI
wi_state <- "550000227544" # WISCONSIN (state government)
@@ -61,15 +59,54 @@ test_that("expenditure concepts classify on spend_type, not item-code prefix", {
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
# payments. All four share the first letter `Y` and the spend_type
# "Insurance Trust", so this pair of assertions is the concrete proof that
# classification is no longer keyed on the first letter.
# payments. All four share the first letter `Y`, so no first-letter allowlist
# can route them. The proof that classification is crosswalk-keyed:
# Y05 lands in `total` spending (insurance_benefits is inside `direct`),
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
# therefore can never appear in a spending result.
#
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
# DEFAULT stays Census General Revenue and so excludes insurance-trust
# revenue; Y01's revenue-side classification is asserted against the
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
# revenue concept argument is uscogdata#12.
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
spend_codes <- wt_codes_included(wi_total)
rev_codes <- wt_codes_included(wi_revenue)
expect_true("Y05" %in% spend_codes)
expect_false("Y05" %in% rev_codes)
expect_true("Y01" %in% rev_codes)
expect_false("Y01" %in% spend_codes)
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
con <- uscogdata:::.ensure_session()
y_class <- DBI::dbGetQuery(con,
"SELECT item_code, category_type, spend_subtype, revenue_subtype
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
})
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
# balances) share first letters with flow codes, so this could never be
# guaranteed under prefix classification; under crosswalk membership it
# falls out structurally -- asserted here at the verb level, on a
# government-year the fixture gives real balance rows (Wisconsin carries
# Y07/Y08/Y21/Y61-type balances in FY2019).
wi_state <- "550000227544"
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
expect_gt(nrow(balance), 0L)
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
rev <- cog_revenue(wi_state, 2019L)
expect_false(any(spend$category %in% balance$category))
expect_false(any(rev$category %in% balance$category))
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
})
+1 -1
View File
@@ -83,7 +83,7 @@ test_that("cog_explain prints the expenditure concept (I1)", {
capture.output(cog_explain(t)),
capture.output(cog_explain(t), type = "message")
), collapse = "\n")
expect_true(grepl("Concept: direct", txt_d))
expect_true(grepl("Concept: primary", txt_d))
expect_true(grepl("Concept: total", txt_t))
})
+12 -2
View File
@@ -111,7 +111,7 @@ test_that("cog_manifest returns the active session's parsed manifest", {
})
})
test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others", {
test_that(".validate_schema accepts schema_version 4 through 7, rejects others", {
expect_silent(uscogdata:::.validate_schema(list(schema_version = 4L)))
expect_silent(uscogdata:::.validate_schema(list(schema_version = 5L)))
# v6 = FIPS geography harmonization (2026-07-22): _code -> _asof rename +
@@ -119,12 +119,22 @@ test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others",
# renamed columns and its geography comes from the xwalk, so v6 is accepted
# without behavioural change -- see .validate_schema()'s note.
expect_silent(uscogdata:::.validate_schema(list(schema_version = 6L)))
# v7 = `data_year` APPENDED as column 29 (cog_pipeline #80, 2026-08-03), the
# most recent fiscal year contributing to a collapsed key. Appended, never
# inserted: canonical_govid stays at position 26, so nothing this package
# reads shifts. Verified against the real v7 corpus before widening the
# allow-list -- cog_spending()/cog_balances() return correctly for FY2024 AND
# for FY2012, so the new column is inert here.
expect_silent(uscogdata:::.validate_schema(list(schema_version = 7L)))
expect_error(
uscogdata:::.validate_schema(list(schema_version = 3L)),
"schema_version"
)
# The upper bound still has to be ENFORCED, not just moved. Without this the
# test would no longer prove that an unknown future schema is refused, and a
# v8 corpus with a genuinely breaking change would sail through.
expect_error(
uscogdata:::.validate_schema(list(schema_version = 7L)),
uscogdata:::.validate_schema(list(schema_version = 8L)),
"schema_version"
)
})
+252
View File
@@ -214,3 +214,255 @@ test_that("no signposting under basis = 'raw'", {
prov <- attr(r, "provenance")
expect_length(prov$suggestions, 0L)
})
# --- uscogdata#9: partial-coverage signposting ------------------------------
test_that("no recipe component is ever renamed by harmonization", {
# The suppression trigger anti-joins the verb's long view on item_code.
# That is only sound because harmonization never rewrites a recipe
# component's code -- every component whose harmonized_code differs has
# harmonized_code IS NULL (and is aggregate-flagged). If this ever fails,
# .suppressed_components() would report reachable dollars as suppressed.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS renamed FROM long
WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes)
AND harmonized_code IS NOT NULL
AND harmonized_code <> item_code")$renamed
expect_equal(as.integer(n), 0L)
})
test_that(".select_long_view maps annotated view bases to their long views", {
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "harmonized"),
"spending_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("revenue_annotated", "harmonized"),
"revenue_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "raw"),
"spending_long")
})
test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2011L,
long_view = "spending_long_harmonized",
flow_prefixes = c("E", "F", "G"))
expect_s3_class(s, "tbl_df")
expect_equal(nrow(s), 2L)
s <- s[order(s$recipe_id), ]
expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
expect_equal(s$suppressed_amount, c(1803872000, 271589000))
expect_equal(s$suppressed_codes, c("E67", "E68"))
})
test_that(".suppressed_components finds nothing in a modern year", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2019L,
long_view = "spending_long_harmonized",
flow_prefixes = c("E", "F", "G"))
expect_equal(nrow(s), 0L)
})
test_that(".suppressed_components rejects a long_view outside the allowlist", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
expect_error(
uscogdata:::.suppressed_components(
con, candidates = "welfare_cash_e67_wide", govid = "061037123085",
years = 2011L, long_view = "long; DROP TABLE x",
flow_prefixes = c("E", "F", "G")),
class = "uscogdata_internal_error")
})
test_that(".suppressed_components never measures a component from the other flow family (I1)", {
# uscogdata#9 review, finding I1: without the flow_prefixes filter, a
# candidate recipe entirely outside the calling verb's own flow family is
# ALWAYS absent from that verb's view (by construction), so it was always
# reported as "suppressed" -- fabricating a dollar claim. E67/E68 are
# Public Welfare EXPENDITURE codes; scoping the measurement to revenue's
# own flow_prefixes must find nothing for them.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2011L,
long_view = "revenue_long_harmonized",
flow_prefixes = c("T", "A", "U", "B", "C", "D"))
expect_equal(nrow(s), 0L)
})
test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", {
# The bug: E74/E79 return rows for FY2011, so there is no row-absence gap,
# so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were
# dropped for being aggregate-published. LA County reports $3,185,943,000
# and omits $2,075,461,000, a 39% understatement, silently.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 2L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]]
expect_equal(e67$trigger, "suppressed_component")
expect_equal(e67$suppressed_amount, 1803872000)
expect_equal(e67$suppressed_years, 2011L)
expect_equal(e67$suppressed_codes, "E67")
expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'")
e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]]
expect_equal(e68$trigger, "suppressed_component")
expect_equal(e68$suppressed_amount, 271589000)
expect_equal(e68$suppressed_codes, "E68")
})
test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", {
# Corrections is the case that already worked: zero rows in FY2011, so the
# row-absence path fires. It must keep firing, keep trigger = "empty_year",
# keep its IG counterpart -- and now also report what was suppressed.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Corrections"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 3L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("corrections_combined", "corrections_capital_combined",
"corrections_other_capital_combined"))
expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year"))
cc <- sugg[[which(ids == "corrections_combined")]]
expect_equal(cc$suppressed_amount, 1371460000)
expect_equal(cc$suppressed_codes, "E05")
expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined")
})
test_that("uscogdata#9: the revenue verb inherits the same trigger", {
# Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from
# U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-`
# rents and royalties -- the omission is LARGER than the reported figure.
skip_if_no_corpus()
r <- suppressMessages(
cog_revenue("020000227749", years = 2011L,
category = "Miscellaneous Revenue"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 1L)
expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide")
expect_equal(sugg[[1]]$trigger, "suppressed_component")
expect_equal(sugg[[1]]$suppressed_amount, 1899995000)
expect_equal(sugg[[1]]$suppressed_codes, "U4-")
# A revenue recipe must never be handed an M/L expenditure counterpart.
expect_null(sugg[[1]]$ig_recipe_id)
})
test_that("I1: cog_revenue never fabricates suppressed dollars for an expenditure-only recipe", {
# uscogdata#9 review, finding I1: Corrections is an expenditure-only
# category (E04/E05). cog_revenue() naturally returns zero rows for it, so
# corrections_combined still fires as an empty_year suggestion (its own
# generic join finds real E04/E05 data for this government) -- but before
# the flow_prefixes fix, .suppressed_components() measured E04/E05 against
# cog_revenue()'s OWN view (which can never contain an E-coded row by
# construction) and reported the full $3,631,945,000 as "suppressed",
# when cog_spending() for the same gov/years/category actually returns
# $3,691,029,000 -- nothing was suppressed at all.
skip_if_no_corpus()
r <- suppressMessages(
cog_revenue("061037123085", years = 2019:2020, category = "Corrections"))
sugg <- attr(r, "provenance")$suggestions
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_true("corrections_combined" %in% ids)
hit <- sugg[[which(ids == "corrections_combined")]]
expect_equal(hit$suppressed_amount, 0)
expect_equal(hit$suppressed_years, integer(0))
expect_equal(hit$suppressed_codes, character(0))
# And cog_spending() for the identical gov/years/category is unaffected --
# it actually finds the E04/E05 dollars the buggy measurement claimed were
# excluded.
sp <- suppressMessages(
cog_spending("061037123085", years = 2019:2020, category = "Corrections"))
expect_equal(sum(sp$amt_nominal), 3691029000)
})
test_that("uscogdata#9: no partial-coverage fire in a modern year", {
skip_if_no_corpus()
r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare")
expect_length(attr(r, "provenance")$suggestions, 0L)
})
test_that("uscogdata#9: leaf-and-classified wide-era families never fire", {
# higher_ed_e18_wide and general_gov_e89_wide are the control group: their
# components (E16/E18, E85/E89) are ordinary classified leaves even in the
# wide era, so widening the trigger must leave them silent. This is the
# measurement that refutes "it would fire on every category in every legacy
# year" -- corpus-wide on the fixture, these two produce zero suppressed rows.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide')
AND l.amt <> 0
AND NOT EXISTS (
SELECT 1 FROM spending_long_harmonized v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year AND v.item_code = l.item_code)")$n
expect_equal(as.integer(n), 0L)
})
test_that("uscogdata#9: the cli message reports the suppressed dollars", {
skip_if_no_corpus()
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"1,803,872,000", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"FY2011", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"E67", fixed = TRUE)
})
test_that("uscogdata#9: cog_explain() reports the suppressed dollars", {
# cog_explain()'s whole "print" output -- including the Suggestions
# section built from cli::cli_ul() -- is emitted on the message stream
# (verified empirically 2026-08-04: capture.output(..., type = "output")
# returns character(0) for this call; testthat::capture_messages() is what
# actually carries it), so that is the stream this test captures.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
out <- paste(testthat::capture_messages(cog_explain(r)), collapse = "")
expect_match(out, "271,589,000", fixed = TRUE)
})
test_that("the provenance schema documents the suggestion trigger fields", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE)
props <- sch$properties$suggestions$items$properties
expect_true(all(c("trigger", "suppressed_amount", "suppressed_years",
"suppressed_codes") %in% names(props)))
expect_setequal(unlist(props$trigger$enum),
c("empty_year", "suppressed_component"))
})
@@ -9,47 +9,131 @@
# a published Census revenue concept exactly the way I89 sits inside Census's
# Direct Expenditure concept (finding F-012).
#
# CAVEAT FOR WHOEVER PICKS THIS UP: the argument name below (`revenue_concept =
# "total"`) is this test's *proposal*, not a settled decision. The owner's
# 2026-07-28 resolution covers expenditure concepts only; no revenue-side
# naming has been ruled on. If the eventual argument is named differently,
# change the two calls here -- the asserted dollar invariants are what matter
# and are independent of the naming.
# RULED 2026-07-30. `revenue_concept = c("general", "total")` mirrors
# `expenditure_concept`, and the two values are Census's two published revenue
# concepts, related by the manual's own identity (section 4.3, which defines
# the first by SUBTRACTING from the second):
#
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
#
# so `general` is the four general subtypes (own_source/federal/state/
# local_aid) and `total` is every revenue subtype. Naming utility (A91-A94)
# and liquor store (A90) separately is what makes BOTH computable -- before
# cog_pipeline#79 they sat in own_source, so the default was really
# "General + Utility + Liquor", a concept Census does not publish.
#
# Fixture reproducibility: Madison's own X-prefix revenue (FY1970-FY1986,
# $15,098,000 nominal, $0 thereafter) is outside the bundled fixture's year
# window (2011/2012/2019/2020), so the same invariant is asserted on Wisconsin
# state government FY2012, where the fixture carries nonzero X01/X05/X08.
# state government FY2012, where the fixture carries nonzero X01/X02/X05/X08.
test_that("cog_revenue() can return Census Total Revenue including Insurance Trust (prefix X)", {
testthat::skip("Blocked on uscogdata#12 (finding F-014)")
wi_state <- "550000227544" # WISCONSIN (state government)
# Revenue-shaped Employee Retirement codes, read from the RAW corpus rather
# than through cog_revenue(), which is the filter under test:
# X01 local employee contribution, X04/X05 contributions and transfers from
# other governments, X08 earnings on investments.
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X04", "X05", "X08"))
expect_equal(x_revenue, 2038800) # 615,835 + 0 + 560,382 + 862,583 ($1,000s)
# X01/X02 employee contributions, X05 contributions from other governments,
# X08 total earnings on investments.
#
# X04 and X06 are deliberately NOT in this set, though an earlier draft of
# this test included X04. Both are exhibit codes for INTRAgovernmental
# transfers (the administering government paying into its own fund), which
# X05's own definition excludes by name. Census agrees: its computed "Total
# Emp Ret Rev" for this government-year is exactly the four codes below.
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X02", "X05", "X08"))
expect_equal(x_revenue, 2283883) # 615,835 + 245,083 + 560,382 + 862,583
# The Y-prefix insurance trust revenue (unemployment + workers comp), which
# is the other half of the same Census concept.
y_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("Y01", "Y11"))
expect_equal(y_revenue, 1259785)
general <- cog_revenue(govid = wi_state, years = 2012L)
expect_equal(attr(general, "provenance")$revenue_concept, "general")
expect_equal(sum(general$amt_nominal), 31338293000)
total <- cog_revenue(govid = wi_state, years = 2012L, revenue_concept = "total")
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal), x_revenue * 1000)
expect_equal(sum(total$amt_nominal), 33377093000)
expect_true(all(c("X01", "X05", "X08") %in% wt_codes_included(total)))
expect_equal(attr(total, "provenance")$revenue_concept, "total")
# total - general is the whole insurance trust leg, X and Y together.
# Asserted as a delta as well as a level so this stays correct however the
# utility/liquor families land (both are $0 for WI state in FY2012).
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
(x_revenue + y_revenue) * 1000)
expect_equal(sum(total$amt_nominal), 34881961000)
expect_true(all(c("X01", "X02", "X05", "X08") %in% wt_codes_included(total)))
# Sibling codes under the SAME first letter must stay out: X11/X12 are
# benefit payments (an expenditure) and X21/X30/X47 are cash and securities
# holdings (a balance-sheet stock). This is the F-018 point restated on the
# revenue side -- the split has to come from the crosswalk's spend_type, not
# from the letter X.
# revenue side -- the split comes from the crosswalk, not from the letter X.
expect_false(any(c("X11", "X12", "X21", "X30", "X47") %in% wt_codes_included(total)))
# Every returned row still resolves to a category. summary_categories has
# zero rows for prefix X today, so relaxing the prefix filter alone would
# produce category = NA rows -- see census_of_governments_finance_pipeline#60.
# Every returned row still resolves to a category (cog_pipeline#79 added the
# X crosswalk rows; relaxing a prefix filter alone would have produced
# category = NA rows).
expect_false(any(is.na(total$category)))
})
test_that("revenue_concept = 'general' is the default and is strict Census General Revenue", {
wi_state <- "550000227544"
default <- cog_revenue(govid = wi_state, years = 2012L)
explicit <- cog_revenue(govid = wi_state, years = 2012L,
revenue_concept = "general")
expect_equal(sum(default$amt_nominal), sum(explicit$amt_nominal))
# General Revenue excludes utility, liquor store AND insurance trust
# revenue. WI state carries $0 of utility/liquor in FY2012, so the level
# assertion above cannot see those two -- assert the subtype scope directly.
#
# A subset, not setequal: `state` means "intergovernmental revenue FROM the
# state government" (the C codes), which a STATE government does not receive
# from itself, so it is legitimately absent here.
expect_true(all(default$revenue_subtype %in%
c("own_source", "federal", "state", "local_aid")))
expect_false(any(c("utility", "liquor_store", "insurance_trust") %in%
default$revenue_subtype))
})
test_that("utility and liquor store revenue are inside `total` and outside `general`", {
# A city, where utility revenue is material: this is the case the WI state
# baseline structurally cannot exercise. Measured on the fixture, utility +
# liquor is 15.9% of what cog_revenue() returned for type-2 governments
# before the general/total split, so this is the largest behaviour change
# the concept split introduces.
con <- uscogdata:::.ensure_session()
gov <- DBI::dbGetQuery(con,
"SELECT canonical_govid, SUM(amt) amt FROM long
WHERE year = 2012 AND type = 2 AND NOT is_aggregate
AND item_code IN ('A91','A92','A93','A94')
GROUP BY 1 ORDER BY amt DESC LIMIT 1")$canonical_govid
util_raw <- wt_raw_amt(gov, 2012L, codes = c("A90", "A91", "A92", "A93", "A94"))
expect_gt(util_raw, 0)
general <- cog_revenue(govid = gov, years = 2012L)
total <- cog_revenue(govid = gov, years = 2012L, revenue_concept = "total")
expect_false(any(c("utility", "liquor_store") %in% general$revenue_subtype))
expect_true("utility" %in% total$revenue_subtype)
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
util_raw * 1000 +
wt_raw_amt(gov, 2012L, codes = c("Y01", "Y11", "X01", "X02",
"X05", "X08")) * 1000)
})
test_that("revenue_concept rejects unknown values and never returns a balance row", {
expect_error(
cog_revenue("550000227544", years = 2012L, revenue_concept = "gross"),
class = "uscogdata_invalid_revenue_concept"
)
# uscogdata#25 restated for the widest revenue concept: stocks are not
# flows, and `total` must not quietly admit the X/Y/W/Z balance families.
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
total <- cog_revenue("550000227544", years = 2012L, revenue_concept = "total")
expect_false(any(total$category %in% balance$category))
expect_length(intersect(wt_codes_included(total), balance$item_code), 0L)
})
+100 -29
View File
@@ -36,8 +36,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
# {url} exactly as .register_views() does, and executes them -- plus
# their 10-long.sql dependency -- against a synthetic hive-partitioned
# parquet tree written to a temp dir. A regression in any predicate (e.g.
# `NOT is_aggregate` dropped, the prefix list changed, the NULL guard
# removed) would change which of the rows below survive.
# `NOT is_aggregate` dropped, the crosswalk-membership subquery changed,
# the NULL guard removed) would change which of the rows below survive.
#
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
# PARQUET) rather than the arrow package: this package has no arrow
@@ -61,25 +61,45 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
-- 'S74' is outside BOTH flow families (E/F/G/K spending and
-- T/A/U/B/C/D revenue -- it mirrors the real corpus's own
-- non-flow-type codes like S74/Z61), so it can only leak into
-- EITHER view via the E/F/G/K or T/A/U/B/C/D prefix filter, never
-- both at once -- a prefix drawn from the other view's own family
-- (e.g. a real T-code for the spending row) would incorrectly
-- leak into the other view's assertion below and not discriminate
-- the predicate under test.
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by the E/F/G/K prefix filter
-- Revenue (T/A/U/B/C/D) rows, exercised against revenue_long_harmonized:
-- 'S74' and 'Z61' are classified `balance` in the synthetic
-- crosswalk below (mirroring the real corpus's own non-flow codes),
-- so each is excluded from its view ONLY by the crosswalk-membership
-- subquery -- the mechanism that replaced the prefix allowlists
-- (uscogdata#11) and keeps balance stocks out of both flows
-- (uscogdata#25).
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by crosswalk membership (balance)
-- Revenue rows, exercised against revenue_long_harmonized:
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by the T/A/U/B/C/D prefix filter
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by crosswalk membership (balance)
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The flow views classify by membership in summary_categories, so the
# synthetic corpus needs one too. Every flow code above is a member of its
# own flow (so is_aggregate / NULL-harmonized exclusions stay the SOLE
# excluder for those rows); S74/Z61 are members but classified balance, so
# membership itself is what excludes them.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('E36', 'Water Utilities', 'expenditure', 'operations', NULL),
('E38', 'Water Utilities', 'expenditure', 'operations', NULL),
('E05', 'Corrections', 'expenditure', 'operations', NULL),
('E99', 'Other', 'expenditure', 'operations', NULL),
('S74', 'Fund Balances', 'balance', NULL, NULL),
('U11', 'Interest Earnings','revenue', NULL, 'own_source'),
('U10', 'Interest Earnings','revenue', NULL, 'own_source'),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source'),
('T88', 'Other Taxes', 'revenue', NULL, 'own_source'),
('Z61', 'Fund Balances', 'balance', NULL, NULL)
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
@@ -89,6 +109,7 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
@@ -97,8 +118,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
GROUP BY item_code ORDER BY item_code"
)
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
# harmonized_code), and spend-E (wrong prefix family) must all be gone,
# and spend-A + spend-B must be folded together under E36.
# harmonized_code), and spend-E (balance, not an expenditure member) must
# all be gone, and spend-A + spend-B must be folded together under E36.
expect_equal(nrow(spend), 1L)
expect_equal(spend$item_code, "E36")
expect_equal(spend$amt, 150)
@@ -138,12 +159,27 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is
('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis
('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views
('ig-D', 'L--', 55555, false, 'L--'), -- family total: excluded from BOTH views
('ig-E', 'T29', 44444, false, 'T29') -- wrong prefix (revenue, not M/L): excluded from BOTH views
('ig-D', 'L--', 55555, false, 'L--'), -- family total: deliberately NOT a crosswalk member, excluded from BOTH views
('ig-E', 'T29', 44444, false, 'T29') -- revenue member, not intergovernmental: excluded from BOTH views
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The IG views classify by summary_categories membership
# (spend_subtype = 'intergovernmental'). L-- is deliberately absent --
# exactly as it is from the real crosswalk -- which is what excludes it.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('M04', 'Corrections', 'expenditure', 'intergovernmental', NULL),
('M38', 'Health', 'expenditure', 'intergovernmental', NULL),
('M36', 'Health', 'expenditure', 'intergovernmental', NULL),
('M47', 'IG Other', 'expenditure', 'intergovernmental', NULL),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source')
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
@@ -153,6 +189,7 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("24-ig_long.sql"))
DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql"))
@@ -160,8 +197,9 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
"SELECT item_code, SUM(amt) AS amt FROM ig_long
GROUP BY item_code ORDER BY item_code"
)
# L-- (family total) and T29 (wrong prefix) are gone; the aggregate row
# M47 survives -- proof `NOT is_aggregate` is absent from ig_long.
# L-- (family total, not a member) and T29 (revenue, not IG) are gone; the
# aggregate row M47 survives -- proof `NOT is_aggregate` is absent from
# ig_long.
expect_equal(raw$item_code, c("M04", "M38", "M47"))
expect_equal(raw$amt, c(100, 50, 99999))
@@ -322,14 +360,32 @@ test_that(".harmonization_view_files guard is necessary: registration against a
)
})
test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
test_that("spending_long carries exactly the non-IG expenditure crosswalk codes and excludes aggregates", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con,
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM spending_long"
)$pfx
expect_true(all(prefixes %in% c("E", "F", "G", "K")))
# Classification is crosswalk membership, not prefixes (uscogdata#11):
# every row's code must classify as expenditure and never as
# intergovernmental (which lives in ig_long).
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM spending_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'expenditure'
OR c.spend_subtype = 'intergovernmental'"
)$item_code
expect_length(stray, 0L)
# Balance codes are stocks, not flows -- they must never appear in a
# spending result (uscogdata#25). Prefix filtering could not guarantee
# this (W/X/Y/Z balance codes share letters with flow codes).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
@@ -337,14 +393,29 @@ test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
expect_equal(agg_count, 0)
})
test_that("revenue_long filters to T/A/U/B/C/D prefixes and excludes aggregates", {
test_that("revenue_long carries exactly the revenue crosswalk codes and excludes aggregates", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con,
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM revenue_long"
)$pfx
expect_true(all(prefixes %in% c("T", "A", "U", "B", "C", "D")))
# The view carries EVERY revenue subtype; which of Census's two published
# concepts a query returns is decided per `revenue_concept` in R
# (uscogdata#12), exactly as `expenditure_concept` narrows spending_long.
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM revenue_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'revenue'"
)$item_code
expect_length(stray, 0L)
# No balance stock ever appears in a revenue result (uscogdata#25).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
+94 -74
View File
@@ -1,8 +1,8 @@
---
title: "Total spending: Direct, Total, and when each is right"
title: "Total spending: Primary, Direct, Total, and when each is right"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}
%\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
@@ -17,17 +17,28 @@ knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
is about one government or several:
1. **"What did my county spend in total, a decade ago vs today?"** — one
government, tracked over time. Either `direct` or `total` spending answers
this correctly, as long as the same concept is used for both years.
government, tracked over time. Any concept answers this correctly, as
long as the same concept is used for both years.
2. **"How do all the counties in my state compare, a decade ago vs today,
against the neighboring state?"** — several governments, summed together.
Here only `direct` gives the right answer; summing `total` across
governments double-counts money that passes between them.
Here only a non-intergovernmental concept (`primary` or `direct`) gives
the right answer; summing `total` across governments double-counts money
that passes between them.
`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`)
controls which of these a query answers. This vignette walks through both
questions with code that actually runs against the package's bundled fixture
corpus, then explains why the second question refuses `"total"` outright.
`cog_spending()`'s `expenditure_concept` argument controls which of these a
query answers, via three nested concepts defined as sets of the crosswalk's
`spend_subtype` values (never item-code first letters — the letter `Y` alone
spans revenue, expenditure, and balance codes):
- `"primary"` (the default) — the government's own service provision:
`operations` + `capital` + `assistance`.
- `"direct"` — Census's published Direct Expenditure: `primary` plus
`interest` on debt and `insurance_benefits` (e.g. pension payments).
- `"total"` — `direct` plus the `intergovernmental` leg.
This vignette walks through both questions with code that actually runs
against the package's bundled fixture corpus, then explains why the second
question refuses `"total"` outright.
Before any of the numbers below: every amount column here — `amt_nominal`,
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
@@ -68,20 +79,23 @@ al_total <- cog_spending(
al_total
```
The `intergovernmental` rows are what `"total"` adds on top of `"direct"`
(`capital` + `operations`): Alabama's own payments out to counties and
cities for highway work. Because this query only ever concerns Alabama,
including that piece is safe -- there's no other government's number it
could be double-counted against.
The `intergovernmental` rows are what `"total"` adds on top of the
non-intergovernmental subtypes (here `capital` + `operations`): Alabama's
own payments out to counties and cities for highway work. Because this
query only ever concerns Alabama, including that piece is safe -- there's
no other government's number it could be double-counted against.
`"direct"` (the default) answers the same trend question just as validly:
`"primary"` (the default) answers the same trend question just as validly
(for Highways, which maps only to operations/capital codes, `"primary"` and
`"direct"` coincide -- there is no highway-specific interest or insurance
benefit to add):
```{r}
al_direct <- cog_spending(
al_primary <- cog_spending(
"010000226085", years = c(2012, 2020), category = "Highways"
# expenditure_concept = "direct" is the default; shown here for contrast
# expenditure_concept = "primary" is the default; shown here for contrast
)
al_direct
al_primary
```
Both are internally consistent series. What breaks the comparison is
@@ -93,8 +107,8 @@ every year in the series.
# Archetype 2: a cross-government rollup
`cog_geographic_rollup()` sums spending across state/county/city layers for
a place. Its default -- and, as shown below, its *only* accepted value for
`expenditure_concept` -- is `"direct"`:
a place. Its default is `"primary"`, and (as shown below) it accepts only
the non-intergovernmental concepts, `"primary"` and `"direct"`:
```{r}
fl_rollup <- cog_geographic_rollup(
@@ -145,15 +159,15 @@ shows up **twice** in the underlying corpus:
the county is the government that actually lets the contract and pays the
paving crew.
`direct` (item codes `E`/`F`/`G`) only ever counts the second of those --
the government that actually did the spending. `total` (Direct plus the
`M`/`L` intergovernmental legs) counts the first one *as well*, which is
exactly right for describing Alabama's own budget: Alabama's `total`
genuinely includes the $10M it committed to highways, whether it built the
road itself or paid the county to. But sum `total` across Alabama **and**
the county, and that $10M is counted twice -- once as Alabama's payment out,
once as the county's spending in -- reporting $20M of highway work for $10M
actually spent.
`primary` and `direct` (the crosswalk's non-intergovernmental expenditure
subtypes) only ever count the second of those -- the government that
actually did the spending. `total` (Direct plus the intergovernmental leg)
counts the first one *as well*, which is exactly right for describing
Alabama's own budget: Alabama's `total` genuinely includes the $10M it
committed to highways, whether it built the road itself or paid the county
to. But sum `total` across Alabama **and** the county, and that $10M is
counted twice -- once as Alabama's payment out, once as the county's
spending in -- reporting $20M of highway work for $10M actually spent.
This is exactly the shape of query `cog_geographic_rollup()` exists to run
(summing across layers of government), so it refuses `"total"` rather than
@@ -168,48 +182,51 @@ share of a government's own Direct spending is:
| Government type | Intergovernmental / Direct |
|---|---|
| State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) |
| County | 3.4%-5.1% (varies by year) |
| City | 2.6%-3.1% (varies by year) |
| State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) |
| County | 3.3%-4.8% (varies by year) |
| City | 2.4%-2.9% (varies by year) |
So the Direct/Total choice matters overwhelmingly for **state** governments
-- a state's Total genuinely differs from its Direct by a meaningful margin,
while for a county or city the two are close. The state range is also far
wider than a single flat figure would suggest: legacy wide-era years (2011:
48.4%) carry proportionally more intergovernmental spending than the modern
era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly
3x larger a decade earlier than it is today. That's also why the mistake
this vignette warns about is easy to make unnoticed at the county/city level
and costly at the state level: rolling up every government in a state using
`total` instead of `direct` overstates the true figure -- measured at 7.6%
for Alabama in FY2019, and 11.6% nationally.
So the Direct/Total choice matters overwhelmingly for **state**
governments -- a state's Total genuinely differs from its Direct by more
than a third, while for a county or city the two are close. (The state
share is much larger than pre-#11 measurements suggested, because the
intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state
payments to school systems -- for most states the single largest transfer
they make.) That's also why the mistake this vignette warns about is easy
to make unnoticed at the county/city level and costly at the state level:
rolling up every government using `total` instead of `primary`/`direct`
overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally.
# Why Total = Direct + M + L, not Direct + M
# Why Total = Direct + M + L + Q, not Direct + M
It's tempting to assume `total` only needs to add `M`. But `M` and `L` are
both money the queried government itself pays **out** -- they're not two
different accounts of a receiving government's revenue. `M` is what it
pays to other **local** governments (e.g. a county paying a city for a
shared paving contract); `L` is what it pays **up** to its **state**
government (e.g. a county's contribution to a state-administered program).
A local government's Total genuinely includes both legs, because both are
its own spending, just routed to a different kind of recipient. On the
bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for
state governments (a state has no "payments to the state government" leg of
its own) but is 43%-51% the size of `M` for counties (varies by year) and
144%-189% the size of `M` for cities (varies by year; 166% pooled across
all four) -- so a `total` that omitted `L` would silently undercount Total
specifically for local governments, and for cities `L` is often the
*larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes both legs (excluding
the `L--` family-total rollup row, which would double-count its own
components).
It's tempting to assume `total` only needs to add `M`. But the
intergovernmental leg has three families, all money the queried government
itself pays **out** -- they're not different accounts of a receiving
government's revenue. `M` is what it pays to other **local** governments
(e.g. a county paying a city for a shared paving contract); `L` is what it
pays **up** to its **state** government (e.g. a county's contribution to a
state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments
to **school systems** (K-12 and higher-ed aid -- for most states the
single largest transfer they make, and the piece the pre-#11 prefix
allowlist silently dropped, finding F-017). A government's Total genuinely
includes every leg it pays, because each is its own spending, just routed
to a different kind of recipient. On the bundled fixture corpus (all 50
states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has
no "payments to the state government" leg of its own) but is 43%-51% the
size of `M` for counties (varies by year) and 144%-189% the size of `M`
for cities (varies by year; 166% pooled across all four) -- so a `total`
that omitted `L` would silently undercount Total specifically for local
governments, and for cities `L` is often the *larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes every leg
(excluding the `L--` family-total rollup row, which would double-count its
own components).
# Composition rules
- `expenditure_concept` (whose spending counts -- Direct vs Direct plus
intergovernmental) is **orthogonal** to `basis` (which vintage of the
item-code space a query resolves against -- `"harmonized"` vs `"raw"`).
- `expenditure_concept` (whose spending counts -- Primary, Direct, or
Direct plus intergovernmental) is **orthogonal** to `basis` (which
vintage of the item-code space a query resolves against --
`"harmonized"` vs `"raw"`).
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
valid, meaningful query, and so is every other pairing.
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
@@ -225,12 +242,15 @@ components).
# Summary
- Comparing one government to itself over time: `"direct"` or `"total"`
both work -- pick one and hold it fixed across every year compared.
- Comparing one government to itself over time: any concept works -- pick
one and hold it fixed across every year compared.
- Comparing or summing across governments -- counties within a state, a
state against its neighbor, cities against counties: use `"direct"`.
`cog_geographic_rollup()` and `cog_peer_compare()` enforce this by
refusing `"total"`.
- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local
governments + `L` to the state government, excluding the `L--`
family-total row).
state against its neighbor, cities against counties: use `"primary"`
(the default) or `"direct"`. `cog_geographic_rollup()` and
`cog_peer_compare()` enforce this by refusing `"total"`.
- `"primary"` = operations + capital + assistance. `"direct"` = primary +
interest on debt + insurance trust benefits (Census's published Direct
Expenditure). `"total"` = direct + intergovernmental (`M` to local
governments, `L` to the state government excluding the `L--`
family-total row, and `Q11`/`Q12`/`Q18` state payments to school
systems).