Compare commits

..
Author SHA1 Message Date
jared 93300ae0c1 feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00
jared 7d798b9937 chore: regenerate fixture corpus at pipeline_commit e64a046
R-CMD-check / check (pull_request) Successful in 3m18s
R-CMD-check / check (push) Successful in 4m6s
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.

Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.

test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.

Suite: 0 failures, 2 skips (the #11 and #12 blocks).
2026-07-30 16:04:35 -04:00
jared 915a4d0678 Merge pull request 'feat: coverage argument + always-on reporting-coverage metadata (#13)' (#24) from feat/coverage-disclosure-13 into main
R-CMD-check / check (push) Successful in 3m30s
2026-07-30 12:06:53 -04:00
jared 6f98d061a9 Merge pull request 'feat: complete = TRUE fills absent cells with their meaning (#18)' (#23) from feat/complete-argument-18 into main
R-CMD-check / check (push) Successful in 4m14s
Reviewed-on: #23
2026-07-30 12:04:27 -04:00
jared d95c9032c5 feat: coverage argument + always-on reporting-coverage metadata (#13)
R-CMD-check / check (pull_request) Successful in 3m13s
R-CMD-check / check (push) Successful in 3m18s
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".

On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.

Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.

Decisions worth stating:

  - n_units_expected is the universe the CALLER named, not the national one.
    That is what makes the ratio mean something: "597 of the 608 Wisconsin
    cities you asked about". For peers it is the cohort size, counted over
    peer rows only -- including the target would inflate every count by one
    and make a cohort that has entirely stopped reporting look non-empty.

  - The coverage table is built from the REQUESTED years, not the years
    present in the result, so a year in which nothing reported still appears
    with n_units_reporting = 0. A year that vanishes silently is precisely
    the disclosure failure at issue.

  - "census" filters years BEFORE the query, and aborts when the range holds
    no census year rather than returning an empty result for a query the
    caller believes they made.

  - "consistent" exempts the peer-comparison target: it is the subject of the
    comparison, not a member of the cohort being balanced, and dropping it
    would leave nothing to compare. The summary_* quantiles are computed
    AFTER the filter so they describe the cohort actually returned.

  - is_census_year is documented as a statement about the survey CALENDAR,
    never a claim of completeness -- FY1967 is a census year in which only 97
    of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
    that tells the truth.

On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().

One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.

Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
2026-07-30 11:57:11 -04:00
jared af85a23ea7 feat: complete = TRUE fills absent cells with their meaning (#18)
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
2026-07-30 11:47:51 -04:00
jared 8db944e4a0 Merge pull request 'fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)' (#22) from fix/kodor-batch-14-15-16 into main
R-CMD-check / check (push) Successful in 3m11s
Reviewed-on: #22
2026-07-30 11:37:08 -04:00
jared 2e8383b098 fix: let the doc-content tests survive R CMD check
R-CMD-check / check (push) Successful in 3m5s
R-CMD-check / check (pull_request) Successful in 3m11s
CI failed on the previous commit. testthat::test_local() from a checkout was
green, but rcmdcheck was not: under R CMD check the suite runs against the
INSTALLED package, where README.md, vignettes/ and man/ do not exist. Both
newly-activated tests read them through test_path("..", "..", ...) and died
on `cannot open the connection`.

The defect was latent in the committed tests, not introduced here -- they
shipped skip()ped, so CI had never executed either one. Removing the skips
is what exposed it, which is the mechanism working as intended.

Guarded with skip_if_no_source_tree(), so they skip in the installed-package
context that structurally cannot satisfy them. They are NOT thereby unchecked
in CI: the workflow runs testthat::test_local() from the checkout as its own
step before rcmdcheck, and there the paths resolve and the assertions run.

Deliberately not split: test-peer-summary-scope.R's numeric pin needs only
the corpus and would survive check on its own, but it exists to protect the
sentence above it. Separating them would let the prose drift while the pin
kept passing.

Verified locally: test_local 629 pass / 0 fail / 3 skip; rcmdcheck
0 errors / 0 warnings / 0 notes.
2026-07-30 11:31:28 -04:00
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00
jared ebac39e6de Merge pull request 'fix: surface ALL-scoped series breaks in provenance (#19)' (#21) from fix/all-scoped-series-breaks-19 into main
R-CMD-check / check (push) Successful in 3m23s
2026-07-30 10:33:24 -04:00
jared 47dc08c4b0 Merge pull request 'fix: regenerate the bundled fixture against the sparsified corpus (#18)' (#20) from fix/regen-fixture-corpus-18 into main
R-CMD-check / check (push) Successful in 3m7s
2026-07-30 10:32:40 -04:00
jared 1d553a788f fix: surface ALL-scoped series breaks in provenance (#19)
R-CMD-check / check (push) Successful in 3m1s
R-CMD-check / check (pull_request) Successful in 3m1s
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:

  SB085  1977  dollar precision across the 1976/1977 boundary
  SB087  2002  imputation exclusion FY2002-2006
  SB194  2012  dense -> sparse representation change
  SB086  2017  government id scheme change

SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.

Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.

On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.

Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.

Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
2026-07-30 10:27:48 -04:00
jared c375c55da7 fix: regenerate the bundled fixture against the sparsified corpus (#18)
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m51s
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.

Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.

Three test repairs, each a real consequence of sparsification rather than
a number to bump:

  test-categories.R          "assistance" joined the spending subtype
                             vocabulary with the J-prefix codes.

  test-spending.R            The harmonization block counts rows that
                             exist. Broward's E21/F21/G21 were zero-pads
                             and are gone, so the anchor moves to FL state,
                             whose three NA-mapped rows carry $2.83B --
                             the amount accounting was previously asserted
                             only against 0 and could not have caught a
                             bug. Broward keeps a test of its own, now
                             asserting the zero-pads are absent.

  test-expenditure-concept.R Coverage-gap suggestions are presence-based.
                             AL state's only FY2011 B47 cell was an
                             explicit zero, so ig_federal_b47_wide stopped
                             being a candidate there; FL state carries a
                             real amount, so the counterpart guard is
                             exercised against a suggestion that fires.

test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.

Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
2026-07-30 10:20:09 -04:00
jared 82acda6f93 Merge pull request 'test: add failing tests for Madison walkthrough findings' (#17) from test/walkthrough-findings into main
R-CMD-check / check (push) Successful in 3m6s
2026-07-29 10:35:31 -04:00
57 changed files with 2119 additions and 350 deletions
+112
View File
@@ -1,5 +1,117 @@
# uscogdata 0.1.0 (development)
## Multi-government aggregates now disclose their reporting coverage
* The Census of Governments is a **complete census only in years ending in 2
and 7**; every other year is a sample, and the sample varies enormously. On
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
return value said nothing about, so a statewide total resting on a fifth of
the universe looked exactly like one resting on all of it.
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
`coverage`:
| value | effect |
|---|---|
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
* **Regardless of mode**, every result now carries `provenance$coverage` with
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
section. So the default mode can no longer mislead silently.
* `is_census_year` is a statement about the **survey calendar**, never a claim
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
cities report. `n_units_reporting` is the number that tells the truth.
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
it is the subject of the comparison, not a member of the cohort — and the
`summary_*` quantiles are computed after the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against the
cohort size.
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
is `NULL`: `"census"` snaps to the most recent census year with an observed
population, so a cohort is not built from a sample year in which most of the
candidate universe is absent.
## `complete = TRUE`: absent cells, labelled with why they are absent
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
(today's behaviour). With `complete = TRUE` the requested grid is filled
from the corpus's `code_set` table and every row carries a new
`value_source` column:
| `value_source` | meaning | `amt_nominal` |
|---|---|---|
| `reported` | the corpus carries this cell | as published |
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
The `NA` is deliberate and is the whole point: filling a modern absence
with `0` would invent data, which is precisely the error the corpus's
representation contract exists to prevent.
* This restores information the reader lost when the corpus was sparsified
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
had begun returning nothing at all — and improves on what came before it,
since the pre-sparsification corpus could not distinguish a published zero
from an unreported cell either.
* The grid is scoped to each government's **own type**, so a county is never
filled with cells only a state can report.
* Needs a corpus published from 2026-07-29 onward (when `representation` and
`code_set` began shipping); aborts with class
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
listing those tables rather than on `schema_version`, which was never
bumped for the change. Not available with `recipe` or
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
`absence_means` rule; `cog_explain()` prints a "Completion" section.
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
corpus as a whole rather than about one item code. `series_break_refs` is
built by matching `fin_code` against the item codes in the result, and no
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
representation change at FY2012) and `SB086` (the government id scheme
change at FY2017).
* Provenance gains `corpus_break_refs`, selected on the break-year window
alone and disjoint from `series_break_refs` by construction, so a consumer
can tell a whole-result caveat from a break in one series. `cog_explain()`
prints them under their own "Corpus-wide caveats" heading. cog-api passes
provenance through verbatim, so the field appears there without an API
change.
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
crosses the boundary where an absent cell stops meaning "Census published
`$0`" and starts meaning "not reported", and until now nothing said so.
## Bundled fixture regenerated against the sparsified corpus
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
2026-07-29 (`pipeline_commit 83f9715`, schema v6). The wide era no longer
stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of
which none are `$0`. **Absence now means two different things** — in a
`dense_source` year (≤ FY2011) an absent cell means Census published `$0`;
in a `sparse_source` year (≥ FY2012) it means not reported. The corpus
carries that rule in two new tables the fixture now ships,
`representation.parquet` and `code_set.parquet`, alongside
`census_collection_coverage.parquet` and `lineage_events.parquet`
(all ten publish-tree metadata tables, up from six). Catalogued upstream
as series break `SB194`.
* `cog_categories()` gains an `assistance` spending subtype: the J-prefix
aid/benefit codes (`J19`, `J67`, `J68`, `J85`) are categorised now that
the upstream crosswalk covers every flow code carrying dollars.
* Two consequences worth knowing about, both visible in provenance rather
than in returned dollars. The harmonization block's `na_rows_excluded`
counts only rows that exist, so wide-era codes that were zero-padded no
longer appear there. Coverage-gap `suggestions` are presence-based for the
same reason, so a recipe whose component codes were all `$0` for a given
government-year is no longer suggested for it.
* `tests/testthat/test-fixture-vintage.R` pins these structural facts, so a
fixture left behind by a future publish fails loudly instead of letting the
suite pass against a corpus that no longer exists.
## Breaking: corpus schema_version 4 (Phase P canonical ids)
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
+15 -6
View File
@@ -44,11 +44,18 @@
#' Count + sum item-level rows that basis="harmonized" excludes because they
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
#' requested flow type (spending or revenue), govids, and years. Only
#' meaningful when the resolved basis is "harmonized"; returns an
#' applied = FALSE stub otherwise (raw basis never excludes rows this way).
#' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
#' the same subtype-membership classification the verb SQL uses, never
#' item-code prefixes), govids, and years. Only meaningful when the resolved
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
#' basis never excludes rows this way).
#'
#' The intergovernmental leg is deliberately outside this count even for
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
#' @noRd
.build_harmonization_block <- function(con, govid, years, resolved, flow_prefixes) {
.build_harmonization_block <- function(con, govid, years, resolved,
subtype_col, subtype_scope) {
if (!identical(resolved$basis, "harmonized")) {
return(list(
applied = FALSE,
@@ -63,9 +70,11 @@
FROM long
WHERE canonical_govid IN (%s) AND year IN (%s)
AND NOT is_aggregate AND harmonized_code IS NULL
AND LEFT(item_code, 1) IN (%s)",
AND item_code IN (
SELECT item_code FROM summary_categories WHERE %s IN (%s)
)",
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
.sql_lit_chr(flow_prefixes)
subtype_col, .sql_lit_chr(subtype_scope)
)
na <- DBI::dbGetQuery(con, sql)
+149
View File
@@ -0,0 +1,149 @@
# R/complete.R
#
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
# cell the corpus does not carry still appears, labelled with WHY it is
# missing.
#
# The corpus stopped storing the wide era's explicit zeros
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
#
# <= FY2011 dense_source absent => Census published $0 (census_zero)
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
#
# Before sparsification a wide-era query whose cells were all $0 came back as
# explicit $0 rows; afterwards it came back empty, with nothing to say which
# of the two meanings applied. This restores that -- and improves on it,
# because the pre-sparsification corpus could not distinguish the two either.
#
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
# That difference is the entire point: writing 0 into a modern absence would
# invent data, which is the error the representation contract exists to stop.
#' @noRd
.abort_complete_unsupported <- function(reason, alternative) {
cli::cli_abort(c(
"{.code complete = TRUE} is not supported for this query.",
x = reason,
i = alternative
), class = "uscogdata_complete_unsupported")
}
#' @noRd
.require_representation <- function(con, manifest) {
needed <- c("representation.parquet", "code_set.parquet")
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
logical(1))]
if (length(missing) == 0L) return(invisible(TRUE))
cli::cli_abort(c(
"This corpus does not publish the representation contract.",
x = "Missing: {.file {missing}}.",
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
), class = "uscogdata_representation_unavailable")
}
#' The cells a government-year COULD carry: every code in force for that
#' government's own type, mapped through `summary_categories`, restricted to
#' the calling verb's crosswalk subtype scope (the same subtype-membership
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
#' operations/capital/assistance) and (when given) its category filter.
#'
#' Scoped by `govs_type` deliberately. Filling against the union of all types
#' would invent cells that the government can never report -- a county row for
#' "state IG transfer to school districts" -- and those inventions would then
#' be indistinguishable from real census zeros.
#'
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
#' aggregate rows. Without it the grid would offer cells the verb structurally
#' never returns, so every one of them would fill as a phantom $0.
#' @noRd
.completion_grid_sql <- function(subtype_col, govid, years, category,
subtype_scope) {
category_pred <- if (is.null(category)) {
""
} else {
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
}
sprintf(
"SELECT DISTINCT
cs.year,
x.canonical_govid,
x.gov_name,
c.%1$s AS subtype_value,
c.category,
r.absence_means
FROM code_set cs
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
JOIN summary_categories c ON c.item_code = cs.item_code
JOIN representation r ON r.year = cs.year
WHERE x.canonical_govid IN (%2$s)
AND cs.year IN (%3$s)
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%1$s IN (%4$s)
%5$s",
subtype_col, .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","),
.sql_lit_chr(subtype_scope), category_pred
)
}
#' Fill `result` out to the full grid, stamping `value_source` on every row.
#'
#' Returns the completed tibble with a `.completion` attribute carrying the
#' provenance block. Reported rows are passed through untouched -- filling
#' must never alter or drop what the corpus actually published.
#' @noRd
.complete_result <- function(result, con, subtype_col, govid, years, category,
subtype_scope) {
grid <- tibble::as_tibble(DBI::dbGetQuery(
con, .completion_grid_sql(subtype_col, govid, years, category, subtype_scope)
))
result$value_source <- rep("reported", nrow(result))
if (nrow(grid) == 0L) {
attr(result, ".completion") <- list(
applied = TRUE, rows_filled = 0L, absence_means = list()
)
return(result)
}
names(grid)[names(grid) == "subtype_value"] <- subtype_col
key <- function(d) {
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
}
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
if (nrow(missing) > 0L) {
filled <- tibble::tibble(
year = as.integer(missing$year),
canonical_govid = as.character(missing$canonical_govid),
gov_name = as.character(missing$gov_name),
category = as.character(missing$category),
# census_zero is a value Census published; not_reported is unknown and
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
amt_nominal = ifelse(missing$absence_means == "census_zero",
0, NA_real_),
codes_included = NA_character_,
aggregate_fallback = NA,
value_source = as.character(missing$absence_means)
)
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
if ("notes" %in% names(result)) filled$notes <- NA_character_
result <- dplyr::bind_rows(result, filled)
result <- result[order(result$year, result$canonical_govid,
result[[subtype_col]], result$category), ,
drop = FALSE]
}
rules <- unique(grid[, c("year", "absence_means")])
attr(result, ".completion") <- list(
applied = TRUE,
rows_filled = nrow(missing),
absence_means = stats::setNames(
as.list(as.character(rules$absence_means)), as.character(rules$year)
)
)
result
}
+107
View File
@@ -0,0 +1,107 @@
# R/coverage.R
#
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
# findings F-020 and F-023).
#
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
# 7. Every other year is a sample, and the sample varies enormously: on the
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
# what the verbs have always done -- correctly -- but the return value said
# nothing about it, so a statewide total resting on 18% of the universe looked
# exactly like one resting on 98%.
#
# Owner's settled design: a `coverage` argument selecting WHICH units to
# include, plus always-on metadata saying how many there were either way. The
# principle behind it: using these verbs correctly must not require the caller
# to know the survey calendar.
# Years ending in 2 or 7 are full censuses of every government; all others are
# samples.
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
#' @noRd
.is_census_year <- function(years) {
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
}
#' @noRd
.validate_coverage <- function(coverage) {
tryCatch(
match.arg(coverage, c("all", "census", "consistent")),
error = function(e) {
cli::cli_abort(
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
class = "uscogdata_invalid_coverage", parent = e
)
}
)
}
#' Restrict `years` to census years for `coverage = "census"`.
#'
#' Aborts rather than returning an empty result when the requested range holds
#' no census year: silently handing back zero rows for a query the caller
#' believes they made is the failure mode this whole issue is about.
#' @noRd
.apply_census_years <- function(years, coverage, verb) {
if (!identical(coverage, "census")) return(as.integer(years))
keep <- as.integer(years)[.is_census_year(years)]
if (length(keep) == 0L) {
cli::cli_abort(c(
"{.code coverage = \"census\"} leaves no years to query.",
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
), class = "uscogdata_no_census_years")
}
sort(keep)
}
#' Keep only units that report in EVERY requested year (a balanced panel).
#'
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
#' filter (the peer-comparison target, which is the subject of the comparison
#' rather than a member of the cohort being balanced).
#' @noRd
.filter_consistent <- function(result, years, id_col = "canonical_govid",
keep_ids = character(0)) {
years <- unique(as.integer(years))
if (nrow(result) == 0L || length(years) <= 1L) return(result)
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
present <- vapply(ids, function(g) {
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
}, logical(1))
consistent <- c(ids[present], keep_ids)
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
drop = FALSE]
}
#' Per-year coverage metadata, always attached regardless of mode.
#'
#' Built from the REQUESTED years rather than the years present in the result,
#' so a year in which nothing reported still appears -- with
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
#' missing year fails to make.
#'
#' `n_units_reporting` describes the result the caller actually received, so
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
#' `n_units_reporting` is the number that tells the truth.
#' @noRd
.coverage_table <- function(result, years, n_expected,
id_col = "canonical_govid", rows = NULL) {
years <- sort(unique(as.integer(years)))
src <- if (is.null(rows)) result else rows
reporting <- vapply(years, function(y) {
ids <- src[[id_col]][as.integer(src$year) == y]
length(unique(ids[!is.na(ids)]))
}, integer(1))
tibble::tibble(
year = years,
n_units_reporting = as.integer(reporting),
n_units_expected = rep(as.integer(n_expected), length(years)),
is_census_year = .is_census_year(years)
)
}
+43
View File
@@ -118,11 +118,54 @@ cog_explain <- function(result, format = c("print", "list")) {
cli::cli_ul(sugg_lines)
}
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
cli::cli_h2("Reporting coverage")
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
cov <- prov$coverage
cli::cli_ul(sprintf(
"%d: %d of %d units reporting (%.0f%%) -- %s year",
cov$year, cov$n_units_reporting, cov$n_units_expected,
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
ifelse(cov$is_census_year, "census", "sample")
))
if (any(!cov$is_census_year)) {
cli::cli_text(
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
)
}
}
if (isTRUE(prov$completion$applied)) {
cli::cli_h2("Completion")
cli::cli_text(
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
)
rules <- prov$completion$absence_means
if (length(rules) > 0L) {
cli::cli_ul(vapply(names(rules), function(y) {
sprintf("%s: an absent cell means %s", y,
if (identical(rules[[y]], "census_zero")) {
"Census published $0 (filled as 0)"
} else {
"the government did not report (filled as NA, not 0)"
})
}, character(1)))
}
}
if (length(prov$series_break_refs) > 0L) {
cli::cli_h2("Series breaks")
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
}
# Kept in a section of its own: these qualify the whole result, so folding
# them in with the per-code breaks above would invite reading them as a
# caveat about one series.
if (length(prov$corpus_break_refs) > 0L) {
cli::cli_h2("Corpus-wide caveats")
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
}
cli::cli_h2("Transformations")
uc <- prov$transformations$units_conversion
if (isTRUE(uc$applied)) {
+117 -10
View File
@@ -19,6 +19,13 @@
#' target's population at `year` to produce absolute bounds. If `FALSE`,
#' `pop_range` is interpreted as absolute population counts.
#' @param max_peers Integer cap on the number of peers returned.
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
#' most recent census year with an observed population, so a cohort is not
#' built from a sample year in which most of the candidate universe is
#' absent. `"consistent"` needs a year range, which cohort selection does not
#' have, so it selects like `"all"` and is carried on the result as
#' `attr(x, "coverage")` for [cog_peer_compare()].
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
#' `attr(x, "cohort_year")`.
@@ -29,7 +36,9 @@ cog_find_peers <- function(target_govid,
same_state = FALSE,
pop_range = c(0.7, 1.3),
is_ratio = TRUE,
max_peers = 10L) {
max_peers = 10L,
coverage = c("all", "census", "consistent")) {
coverage <- .validate_coverage(coverage)
if (!is.character(target_govid) || length(target_govid) != 1L) {
cli::cli_abort("`target_govid` must be a length-1 character string.")
}
@@ -59,7 +68,7 @@ cog_find_peers <- function(target_govid,
))
}
cohort_year <- .resolve_cohort_year(con, target_govid, year)
cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
pop_sql <- sprintf(
"SELECT population FROM gov_population_yearly
@@ -107,12 +116,34 @@ cog_find_peers <- function(target_govid,
attr(peers, "cohort_year") <- as.integer(cohort_year)
attr(peers, "pop_range") <- as.numeric(pop_range)
attr(peers, "is_ratio") <- isTRUE(is_ratio)
attr(peers, "coverage") <- coverage
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
peers
}
# `coverage` picks the cohort vintage when the caller did not name one.
# "census" snaps to the most recent CENSUS year with an observed population,
# so a cohort is not silently built from a sample year in which most of the
# candidate universe is absent. "consistent" is a comparison-time concept --
# it needs a year RANGE, which cohort selection does not have -- so it selects
# like "all" here and is carried on the result for cog_peer_compare().
#' @noRd
.resolve_cohort_year <- function(con, target_govid, year) {
.resolve_cohort_year <- function(con, target_govid, year,
coverage = "all") {
if (!is.null(year)) return(as.integer(year))
if (identical(coverage, "census")) {
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
.sql_lit_chr(target_govid)
)
y <- DBI::dbGetQuery(con, sql)$y
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
cli::cli_abort(c(
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
), class = "uscogdata_no_census_years")
}
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s",
@@ -133,7 +164,9 @@ cog_find_peers <- function(target_govid,
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
#' call.
#' call. Those summary rows are quantiles **within each category**, not
#' quantiles of each peer's total — see the `@return` section before summing
#' them.
#'
#' @param target_govid Character scalar.
#' @param peers A tibble from [cog_find_peers()] or a character vector of
@@ -143,10 +176,36 @@ cog_find_peers <- function(target_govid,
#' @param per_capita Default `TRUE` — peer compare usually normalizes by
#' population.
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
#' single-government queries but cannot be used here because combining Total
#' across peer sets counts intergovernmental transfers twice.
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across peer sets counts
#' intergovernmental transfers twice; `"primary"` and `"direct"` combine
#' safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#'
#' The comparison target is exempt from `"consistent"` balancing -- it is the
#' subject of the comparison, not a member of the cohort -- and the
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
#' cohort actually returned. `n_units_reporting` counts peers only, against
#' the cohort size: "3 of your 15 peers reported in FY2019".
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
@@ -155,12 +214,40 @@ cog_find_peers <- function(target_govid,
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
#' `cohort_year`, and `cohort_govids`.
#'
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
#' Each one is computed **within each `(year, spend_subtype,
#' category)` cell** across the peer set, so a `summary_p50` row is *the
#' median peer's value in that one category*, not *the value of the median
#' peer's total*. The median peer for Police and the median peer for Fire
#' are usually different governments, so summing `summary_*` rows across
#' categories does not give any peer's total and misstates the band it
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
#' one cohort, with a sign flip at FY2012.
#'
#' Facet by `role` **and** `category` (the documented use, and what the
#' rows are built for). For a genuine "median peer's total spending" line,
#' sum each peer's own categories first and take the quantile of those
#' per-government totals:
#'
#' ```r
#' library(dplyr)
#' cmp |>
#' filter(role %in% c("target", "peer")) |>
#' group_by(year, role, canonical_govid) |>
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
#' filter(role == "peer") |>
#' group_by(year) |>
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
#' ```
#' @export
cog_peer_compare <- function(target_govid, peers, category, years,
per_capita = TRUE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_peer_compare")
}
@@ -183,9 +270,21 @@ cog_peer_compare <- function(target_govid, peers, category, years,
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
all_govids <- unique(c(target_govid, peer_govids))
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
years <- .apply_census_years(years, coverage, "cog_peer_compare")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
# The target is exempt from balancing: it is the subject of the comparison,
# not a member of the cohort being balanced, and dropping it would leave a
# peer comparison with nothing to compare. Filtering happens BEFORE the
# quantiles below, so a "consistent" cohort's summary rows describe that
# cohort rather than the unbalanced one.
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years, keep_ids = target_govid)
}
value_col <- .peer_value_col(per_capita, adjust_to_year)
summary_rows <- .peer_summary_rows(r, value_col)
@@ -206,6 +305,14 @@ cog_peer_compare <- function(target_govid, peers, category, years,
canonical_govid = target_govid,
gov_name = unique(r$gov_name[r$role == "target"])
)
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
# reported in FY2019". Including the target would inflate every count by one
# and make a cohort that has entirely stopped reporting look non-empty.
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(
out, years, length(peer_govids),
rows = r[r$role == "peer", , drop = FALSE]
)
attr(out, "provenance") <- prov
out
}
+20 -2
View File
@@ -10,7 +10,8 @@
expenditure_concept_note = NA_character_,
expenditure_concept_direct_suppressed = FALSE,
harmonization = NULL, recipe = NULL,
suggestions = list()) {
suggestions = list(),
completion = NULL) {
manifest <- .uscogdata_env$manifest
codes <- result[["codes_included"]]
@@ -37,11 +38,20 @@
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
con <- .uscogdata_env$con
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) {
have_con <- !is.null(con) && DBI::dbIsValid(con)
break_refs <- if (have_con) {
.build_series_break_refs(con, codes_observed, years, schema_version)
} else {
character(0)
}
# Corpus-wide caveats travel separately: they qualify the whole result
# rather than one series, and they do not depend on codes_observed (see
# .build_corpus_break_refs()).
corpus_refs <- if (have_con) {
.build_corpus_break_refs(con, years, schema_version)
} else {
character(0)
}
list(
verb = verb,
@@ -116,6 +126,14 @@
)
),
series_break_refs = break_refs,
corpus_break_refs = corpus_refs,
# What `complete = TRUE` filled, and the rule it filled by. Always
# present so a consumer can read `completion$applied` without testing
# for the key -- an absent block and applied = FALSE would otherwise be
# indistinguishable from an older reader version.
completion = completion %||% list(
applied = FALSE, rows_filled = 0L, absence_means = list()
),
manifest = list(
schema_version = as.integer(manifest$schema_version),
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
+10 -3
View File
@@ -11,11 +11,17 @@
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' and `value_source` when `complete = TRUE`.
#' @export
cog_revenue <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL) {
basis = c("harmonized", "raw"), recipe = NULL,
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
# this verb's recipe families (see R/suggestions.R).
.verb_spendrev(
verb = "cog_revenue",
view_base = "revenue_annotated",
@@ -28,6 +34,7 @@ cog_revenue <- function(govid, years, category = NULL,
per_capita = per_capita,
adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe
recipe = recipe,
complete = complete
)
}
+44 -8
View File
@@ -25,12 +25,31 @@
#' population from `gov_population_yearly`. Govs with missing population
#' are excluded from the result.
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
#' single-government queries but cannot be used here because combining Total
#' across multiple layers of government double-counts intergovernmental
#' transfers (a state's payment to a school district is the same dollar the
#' district reports as its own Direct spending).
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across multiple layers of
#' government double-counts intergovernmental transfers (a state's payment
#' to a school district is the same dollar the district reports as its own
#' Direct spending); `"primary"` and `"direct"` combine safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
@@ -40,9 +59,11 @@
#' @export
cog_geographic_rollup <- function(govids, category, years,
per_capita = FALSE, adjust_to_year = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_geographic_rollup")
}
@@ -59,11 +80,21 @@ cog_geographic_rollup <- function(govids, category, years,
layer = rep(layer_names, lengths(govids))
)
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
# coverage = "census" drops non-census years BEFORE the query rather than
# after: a sample year's rows are not wanted at all, and fetching them only
# to discard them would also let them into the coverage table.
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
relationship = "many-to-many")
r$scope_note <- .rollup_scope_note(r$layer)
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years)
}
excluded <- character(0)
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
drop <- r$pop_source == "unavailable"
@@ -82,6 +113,11 @@ cog_geographic_rollup <- function(govids, category, years,
included_govids = included,
excluded_govids = excluded
)
# n_units_expected is the universe the CALLER named -- the govids passed in
# -- not the national universe. That is what makes the ratio meaningful:
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
attr(r, "provenance") <- prov
r
+19 -7
View File
@@ -6,8 +6,11 @@
#' the cross-vintage canonical-government registry. Operates in two modes:
#'
#' * **Utility mode** (single `name`, the original behavior): returns all
#' rows whose `gov_name` matches the regex case-insensitively, sorted by
#' `population_acs` descending. Useful for exploratory lookups.
#' rows whose `gov_name` contains `name` as a **literal, case-insensitive
#' substring**, sorted by `population_acs` descending. Useful for
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
#' government is findable by its own complete name even when that name
#' contains parentheses or a period.
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
#' single canonical govid and returns a tibble in input order, suitable
#' for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -19,7 +22,8 @@
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
#' Single hit -> resolved. Multiple -> step 4.
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`.
#' 3. **Substring fallback:** case-insensitive literal substring against
#' `gov_name` (metacharacters escaped).
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
#' `status = "no_match"`. Multiple hits -> step 4.
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -48,7 +52,7 @@
#' [cog_spending()], [cog_revenue()].
#' @examples
#' \dontrun{
#' # Utility mode — exploratory regex lookup
#' # Utility mode — exploratory substring lookup
#' cog_gov_search("broward", state = "FL")
#'
#' # Basket mode — resolve a known cohort
@@ -98,9 +102,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
if (!is.character(name) || length(name) != 1L) {
cli::cli_abort("`name` must be a length-1 character string.")
}
# Escaped, so `name` is a literal case-insensitive substring -- the same
# treatment basket mode has always given it. Interpolating it raw made a
# government unfindable by its own name whenever that name contains a
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
# match-everything wildcard, and let malformed pattern text reach the
# engine as an error -- which cog-api surfaced as a 500, reachable by
# typing a real name one character at a time (uscogdata#16, F-025).
preds <- c(preds,
sprintf("regexp_matches(gov_name, %s, 'i')",
.sql_lit_chr(name)))
.sql_lit_chr(.escape_regex(name))))
}
if (!is.null(state)) {
st_fips <- .coerce_state_to_fips(state)
@@ -136,8 +147,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
#' @noRd
.escape_regex <- function(x) {
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
# literal substring in the DuckDB regexp_matches call (substring fallback
# only; utility-mode intentionally preserves regex behavior).
# literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
# utility mode used to interpolate raw, which was a defect rather than a
# feature -- see the call site and uscogdata#16.
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
}
+33 -1
View File
@@ -14,9 +14,41 @@
sql <- sprintf(
"SELECT DISTINCT break_id
FROM series_breaks_pq
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
AND break_year BETWEEN %d AND %d
ORDER BY break_id",
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
)
DBI::dbGetQuery(con, sql)$break_id
}
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
#' `"ALL"` rather than an item code. They qualify the whole result, so they
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
#' before uscogdata#19. Selection is on the break_year window alone: which
#' codes a result happens to contain is irrelevant to a caveat about the
#' corpus.
#'
#' All four catalogued entries are *boundary* caveats (dollar precision
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
#' representation change at 2012, the id scheme change at 2017), so the same
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
#' path uses is the right one -- a request that never crosses the boundary
#' is not affected by it.
#'
#' Returned separately from `series_break_refs` so a consumer can tell a
#' whole-result caveat from a break in one series; the two are disjoint by
#' construction.
#' @noRd
.build_corpus_break_refs <- function(con, years, schema_version) {
if (schema_version < 5L || length(years) == 0L) return(character(0))
sql <- sprintf(
"SELECT DISTINCT break_id
FROM series_breaks_pq
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
ORDER BY break_id",
min(as.integer(years)), max(as.integer(years))
)
DBI::dbGetQuery(con, sql)$break_id
}
+170 -35
View File
@@ -1,5 +1,40 @@
# R/spending.R
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
# `spend_subtype` values. Classification is crosswalk membership, never
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
# route it (finding F-018).
#
# primary = operations + capital + assistance (the default)
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
# total = direct + intergovernmental (via the ig_* views)
#
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
# than intergovernmental -- including payments to retirees, i.e. insurance
# trust benefits. Verified against Census's own published FY2020 state
# aggregates (20statetypepu.txt): `total` reproduces the published
# expenditure sum to the dollar; omitting insurance benefits understates
# California's Direct by 10.9%.
.spend_subtypes_primary <- c("operations", "capital", "assistance")
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
#' @noRd
.expenditure_concept_subtypes <- function(concept) {
switch(concept,
primary = .spend_subtypes_primary,
# "total" = the direct subtypes here PLUS the intergovernmental leg,
# which travels through the ig_* views rather than this scope (see
# .build_verb_sql()).
direct = ,
total = .spend_subtypes_direct
)
}
# cog_revenue()'s single concept (until uscogdata#12 adds more): Census
# General Revenue -- every crosswalk revenue subtype except insurance_trust.
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
#' Summarized spending by category
#'
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
@@ -42,22 +77,34 @@
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
#' FALSE`, pointing at the `recipe` block instead) rather than a
#' possibly-misleading `"harmonized"`/`"raw"` value.
#' @param expenditure_concept `"direct"` (default) returns only the
#' government's own direct spending (item codes `E`/`F`/`G`), unchanged
#' from prior releases. `"total"` additionally UNIONs in the
#' intergovernmental leg -- payments to local governments (`M` codes) and
#' to the state government (`L` codes, excluding the `L--` family-total
#' rollup) -- so results gain rows with `spend_subtype ==
#' "intergovernmental"`. Requires the active corpus's `summary_categories`
#' to carry M/L rows (added by cog_pipeline PR #59); aborts with class
#' `uscogdata_ig_categories_unsupported` on an older corpus rather than
#' silently under-reporting. Mutually exclusive with `recipe` (a recipe
#' already defines its own component codes). **Do not sum `"total"`
#' results across levels of government** (e.g. state + county + city):
#' a state's `M12` payment to a school district is the same dollar the
#' district reports as its own direct `E12`, so summing both double-counts
#' it. This matters in particular with [cog_geographic_rollup()], which
#' sums across exactly that kind of multi-layer government set.
#' @param expenditure_concept Which spending concept to return. Concepts are
#' defined as sets of the crosswalk's `spend_subtype` values -- never as
#' item-code first letters, which cannot classify correctly (prefix `Y`
#' alone spans revenue, expenditure, and balance codes):
#'
#' * `"primary"` (default) -- the government's own service provision:
#' `operations` + `capital` + `assistance` subtypes.
#' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
#' `interest` (interest on debt) and `insurance_benefits` (insurance
#' trust benefit payments, e.g. pensions -- Census manual section
#' 5.2.2.1 includes payments to retirees in Direct).
#' * `"total"` -- `direct` plus the intergovernmental leg: payments to
#' local governments (`M` codes), to the state government (`L` codes,
#' excluding the `L--` family-total rollup), and state payments to
#' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
#' `spend_subtype == "intergovernmental"`. Requires the active corpus's
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
#' older corpus rather than silently under-reporting. Mutually
#' exclusive with `recipe` (a recipe already defines its own component
#' codes).
#'
#' **Do not sum `"total"` results across levels of government** (e.g.
#' state + county + city): a state's `M12` payment to a school district is
#' the same dollar the district reports as its own direct `E12`, so
#' summing both double-counts it. This matters in particular with
#' [cog_geographic_rollup()], which sums across exactly that kind of
#' multi-layer government set.
#'
#' In the legacy wide era (<= FY2011), some functions are published ONLY
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
@@ -70,16 +117,46 @@
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
#' figure in those rows is the intergovernmental leg alone, not Direct +
#' IG.
#' @param complete If `TRUE`, fill the requested grid so that a cell the
#' corpus does not carry still appears, labelled with **why** it is
#' missing, and add a `value_source` column to every row:
#'
#' * `"reported"` — the corpus carries this cell.
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
#' Census published `$0`. `amt_nominal` is `0`.
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
#' government did not report, and the value is unknown. `amt_nominal` is
#' `NA`, **not** `0` — writing a zero there would invent data.
#'
#' The grid comes from the corpus's `code_set` table, scoped to each
#' government's own type, so a county is never filled with cells only a
#' state can report. Reported rows are passed through untouched.
#'
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
#' when `representation`/`code_set` began shipping; aborts with class
#' `uscogdata_representation_unavailable` otherwise. Not available with
#' `recipe` or with `expenditure_concept = "total"` (class
#' `uscogdata_complete_unsupported`) — neither draws its cells from
#' `code_set`.
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' and `value_source` when `complete = TRUE`.
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
#' whose `completion` block reports `applied`, `rows_filled`, and the
#' per-year `absence_means` rule that was applied.
#' @export
cog_spending <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk subtype membership
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
# machinery to this verb's recipe families (see R/suggestions.R; the
# catalog only has E/F/G-component direct-expenditure recipes).
.verb_spendrev(
verb = "cog_spending",
view_base = "spending_annotated",
@@ -93,7 +170,8 @@ cog_spending <- function(govid, years, category = NULL,
adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe,
expenditure_concept = expenditure_concept
expenditure_concept = expenditure_concept,
complete = complete
)
}
@@ -101,8 +179,9 @@ cog_spending <- function(govid, years, category = NULL,
.abort_concept_not_aggregatable <- function(verb) {
cli::cli_abort(c(
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
"*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\
comparison or sum that spans more than one government.",
"*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
{.code \"direct\"} for any comparison or sum that spans more than \\
one government.",
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
money it hands to other governments. The receiving government reports \\
that same dollar again as its own Direct when it actually spends it, \\
@@ -118,23 +197,36 @@ cog_spending <- function(govid, years, category = NULL,
govid, years, category,
per_capita, adjust_to_year,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("direct", "total")) {
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE) {
basis_explicit <- length(basis) == 1L
basis <- match.arg(basis, c("harmonized", "raw"))
# match.arg() itself throws a base `simpleError`, not an rlang-classed
# condition; wrap it so an invalid expenditure_concept aborts consistently
# with the rest of this package's validation (cli::cli_abort -> rlang_error).
expenditure_concept <- tryCatch(
match.arg(expenditure_concept, c("direct", "total")),
match.arg(expenditure_concept, c("primary", "direct", "total")),
error = function(e) {
cli::cli_abort(
"`expenditure_concept` must be one of {.val direct} or {.val total}.",
"`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
class = "uscogdata_invalid_expenditure_concept",
parent = e
)
}
)
# The concept's subtype scope. Every code path below -- the verb SQL, the
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
# by crosswalk subtype membership, never by item-code prefix. For revenue
# there is a single concept today (General Revenue; uscogdata#12 will add
# more). "total"'s extra intergovernmental leg travels through the ig_*
# views, not through this scope.
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
.expenditure_concept_subtypes(expenditure_concept)
} else {
.revenue_subtypes_general
}
govid <- .coerce_govid_input(govid, arg = "govid")
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
@@ -148,10 +240,10 @@ cog_spending <- function(govid, years, category = NULL,
}
# .verb_spendrev() is shared with cog_revenue(), which never exposes
# expenditure_concept and always resolves it to "direct" -- so nothing on
# the public API can reach this today. But it's a cheap guard against a
# expenditure_concept and always resolves it to the default -- so nothing
# on the public API can reach this today. But it's a cheap guard against a
# future call (direct or via a modified cog_revenue()) that would UNION
# the IG leg's expenditure M/L rows into a revenue result, which has no
# the IG leg's expenditure M/L/Q rows into a revenue result, which has no
# matching IG view and no sensible meaning.
if (identical(expenditure_concept, "total") &&
!identical(view_base, "spending_annotated")) {
@@ -165,12 +257,27 @@ cog_spending <- function(govid, years, category = NULL,
)
}
complete <- isTRUE(complete)
if (complete && !is.null(recipe)) {
.abort_complete_unsupported(
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
)
}
if (complete && identical(expenditure_concept, "total")) {
.abort_complete_unsupported(
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
)
}
years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
con <- .ensure_session()
manifest <- .uscogdata_env$manifest
scope <- .check_govids_in_scope(govid)
if (complete) .require_representation(con, manifest)
resolved <- .resolve_basis(basis, basis_explicit, manifest)
@@ -197,10 +304,23 @@ cog_spending <- function(govid, years, category = NULL,
} else {
NULL
}
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view)
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
subtype_scope)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
# Fill BEFORE per_capita / inflation so the added cells get the same
# treatment as reported ones: a census_zero stays $0 per capita and in real
# dollars, and a not_reported stays NA through both rather than becoming a
# spurious 0.
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
if (complete) {
result <- .complete_result(result, con, subtype_col, govid, years,
category, subtype_scope)
completion <- attr(result, ".completion")
attr(result, ".completion") <- NULL
}
if (per_capita) result <- .attach_per_capita(result, con, govid)
if (!is.null(adjust_to_year)) {
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
@@ -229,7 +349,7 @@ cog_spending <- function(govid, years, category = NULL,
basis_for_prov <- resolved$basis
basis_note_for_prov <- resolved$note
harmonization <- .build_harmonization_block(
con, govid, years, resolved, flow_prefixes
con, govid, years, resolved, subtype_col, subtype_scope
)
# C1(a): gap detection must run against the Direct leg alone. `result`
# can also carry UNION'd intergovernmental rows (expenditure_concept =
@@ -309,7 +429,8 @@ cog_spending <- function(govid, years, category = NULL,
expenditure_concept_direct_suppressed = direct_suppressed_flag,
harmonization = harmonization,
recipe = recipe_block,
suggestions = suggestions
suggestions = suggestions,
completion = completion
)
prov$scope$govids_found <- scope$found
prov$scope$govids_missing <- scope$missing
@@ -406,7 +527,7 @@ cog_spending <- function(govid, years, category = NULL,
#' @noRd
.build_verb_sql <- function(view, subtype_col, govid, years, category,
ig_view = NULL) {
ig_view = NULL, subtype_scope = NULL) {
govid_lit <- .sql_lit_chr(govid)
years_lit <- paste(as.integer(years), collapse = ",")
category_pred <- if (is.null(category)) {
@@ -415,9 +536,22 @@ cog_spending <- function(govid, years, category = NULL,
sprintf("AND category IN (%s)", .sql_lit_chr(category))
}
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
# The base views carry every subtype of their flow (spending_annotated has
# all five non-IG expenditure subtypes); the concept narrows here. For
# "total", the IG leg's rows are 'intergovernmental', so that value joins
# the allowlist exactly when ig_view is present.
subtype_pred <- if (is.null(subtype_scope)) {
""
} else {
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
}
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
# never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L),
# so de-duplication would be pure cost, and a silent row-drop if two
# never UNION: the two legs are disjoint by crosswalk subtype (the direct
# view excludes 'intergovernmental'; the IG view is only that), so
# de-duplication would be pure cost, and a silent row-drop if two
# governments ever reported identical values.
source_expr <- if (is.null(ig_view)) {
view
@@ -449,9 +583,10 @@ cog_spending <- function(govid, years, category = NULL,
WHERE canonical_govid IN (%3$s)
AND year IN (%4$s)
%5$s
%6$s
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
ORDER BY year, canonical_govid, %1$s, category",
subtype_col, source_expr, govid_lit, years_lit, category_pred
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
)
}
+27 -1
View File
@@ -32,6 +32,29 @@
"45-ig_annotated_harmonized.sql"
)
# The representation contract (cog_pipeline#64): two parquet tables that say
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
# on schema_version, because the sparsification that introduced them did not
# bump the version -- the pre-sparsification corpus this package shipped
# against until 2026-07-30 was already schema v6 and carried neither table.
# Keying off the version number would therefore register a view over a file
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
# check exists to tolerate.
.representation_view_files <- c(
"36-representation.sql" = "representation.parquet",
"37-code_set.sql" = "code_set.parquet"
)
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
#' works identically for a local fixture and a remote share.
#' @noRd
.corpus_has_table <- function(manifest, file) {
paths <- vapply(manifest$files$metadata %||% list(),
function(f) as.character(f$path %||% ""), character(1))
file %in% basename(paths)
}
#' Register DuckDB views from inst/sql/ SQL files
#' @noRd
.register_views <- function(con, url, manifest) {
@@ -39,7 +62,10 @@
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
for (f in files) {
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next
base <- basename(f)
if (base %in% .harmonization_view_files && schema_version < 5L) next
if (base %in% names(.representation_view_files) &&
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
DBI::dbExecute(con, sql)
+40 -11
View File
@@ -19,26 +19,55 @@ package implements.
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
```
## Amounts are in full US dollars
Every amount column this package returns — `amt_nominal`, `amt_real`,
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
The raw Census source files report **thousands of dollars**, and the corpus's
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
you never have to. The conversion is recorded in every result:
```r
r <- cog_spending("552025209777", 2020L)
attr(r, "provenance")$transformations$units_conversion
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
```
**Do not multiply again.** If you have read elsewhere that COG amounts are in
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
twice overstates every figure by 1000x, and the result looks plausible rather
than obviously wrong.
## Configuration
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
## Direct vs Total spending
## Primary vs Direct vs Total spending
`cog_spending(..., expenditure_concept = c("direct", "total"))` controls
whose spending a result counts. `"direct"` (the default) is a government's
own current operations, capital outlay, and other direct spending. `"total"`
additionally adds in the intergovernmental legs — money it hands to other
governments to spend on its behalf — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
controls whose spending a result counts. Concepts are defined as sets of the
crosswalk's `spend_subtype` values — never item-code first letters, which
cannot classify correctly (the letter `Y` alone spans revenue, expenditure,
and balance codes):
- `"primary"` (the default) is the government's own service provision:
current operations, capital outlay, and assistance payments.
- `"direct"` is Census's published Direct Expenditure: `primary` plus
interest on debt and insurance trust benefit payments (e.g. pensions).
- `"total"` additionally adds the intergovernmental leg — money handed to
other governments to spend (`M`/`L` codes plus `Q11`/`Q12`/`Q18` state
payments to school systems) — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
**Rule of thumb: any figure that spans more than one government uses
`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this
by refusing `expenditure_concept = "total"`. See
`primary` or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()`
enforce this by refusing `expenditure_concept = "total"`. See
`vignette("total-spending", package = "uscogdata")` for the full
explanation with worked examples.
+38 -34
View File
@@ -11,12 +11,14 @@
# Each partition is a full year (all states/govs) as published, so
# Broward County FL and every other previously-pinned government stay
# covered without any per-gov slicing logic.
# 2. Copies the full canonical_fips_xwalk.parquet, canonical_alias.parquet,
# summary_categories.parquet, harmonization_map.parquet,
# harmonization_recipes.parquet, and series_breaks.parquet metadata
# tables as-is (these are small cross-vintage registries, not
# partitioned by year, so the fixture ships the complete tables rather
# than a year-scoped subset).
# 2. Copies every metadata parquet the publish tree ships (see
# .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
# registries, not partitioned by year, so the fixture ships the complete
# tables rather than a year-scoped subset. representation.parquet and
# code_set.parquet are what make the sparse wide era interpretable --
# absence means "$0" in a dense_source year and "not reported" in a
# sparse_source one -- so a fixture without them cannot represent the
# published corpus.
# 3. Resyncs the four reference docs (data_dictionary.md,
# reader-specification.md, README.md, series_breaks.md) from the
# publish tree's docs/.
@@ -38,6 +40,22 @@
# source("data-raw/regenerate_fixture_corpus.R")
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
# Every metadata parquet the publish tree ships, in the order they appear in
# the corpus manifest. Single source of truth for both the copy step and the
# fixture manifest, so the two can never drift apart.
.FIXTURE_METADATA_FILES <- c(
"canonical_alias.parquet",
"canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet",
"code_set.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"lineage_events.parquet",
"representation.parquet",
"series_breaks.parquet",
"summary_categories.parquet"
)
regenerate_fixture_corpus <- function(
publish_cache_dir = file.path(
"..", "cog_pipeline", "_targets", "publish_cache"
@@ -100,20 +118,11 @@ regenerate_fixture_corpus <- function(
invisible(NULL)
}
# Copy the full (not year-scoped) canonical_fips_xwalk, canonical_alias,
# summary_categories, and (schema v5+) harmonization_map/
# harmonization_recipes/series_breaks parquet tables.
# Copy the full (not year-scoped) metadata tables listed in
# .FIXTURE_METADATA_FILES.
#' @noRd
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
files <- c(
"canonical_fips_xwalk.parquet",
"canonical_alias.parquet",
"summary_categories.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"series_breaks.parquet"
)
for (f in files) {
for (f in .FIXTURE_METADATA_FILES) {
src <- file.path(publish_cache_dir, "data", f)
dst <- file.path(fixture_dir, "data", f)
if (!file.exists(src)) {
@@ -179,15 +188,7 @@ regenerate_fixture_corpus <- function(
)
})
metadata_files <- c(
"canonical_alias.parquet",
"canonical_fips_xwalk.parquet",
"summary_categories.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"series_breaks.parquet"
)
metadata <- lapply(metadata_files, function(f) {
metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
rel <- file.path("data", f)
path <- file.path(fixture_dir, rel)
list(
@@ -203,13 +204,16 @@ regenerate_fixture_corpus <- function(
pipeline_commit = source_manifest$pipeline_commit,
fixture_note = paste(
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
"corpus available via USCOGDATA_URL. Regenerated for Phase R2",
"(schema_version 5, harmonization_map/harmonization_recipes/",
"series_breaks parquet tables added). 2011/2012 straddle the",
"wide-aggregate -> modern-leaf format boundary exercised by basis=",
"\"harmonized\" and recipe= queries; 2019/2020 retain the prior",
"per-capita/CPI regression anchors. Full canonical_fips_xwalk master",
"and canonical_alias lookup table included via",
"corpus available via USCOGDATA_URL. Regenerated from the sparsified",
"schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
"zeros, so FY2011 absence means Census published $0 while FY2012+",
"absence means not reported. representation.parquet and",
"code_set.parquet carry that rule and ship in full, as do every other",
"metadata table in the publish tree. 2011/2012 straddle both the",
"wide-aggregate -> modern-leaf format boundary (exercised by",
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
"representation boundary (SB194); 2019/2020 retain the prior",
"per-capita/CPI regression anchors. Regenerated via",
"data-raw/regenerate_fixture_corpus.R."
),
data_vintage = source_manifest$data_vintage,
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+30 -10
View File
@@ -1,8 +1,8 @@
{
"schema_version": 6,
"built_at": "2026-07-27T13:04:05Z",
"pipeline_commit": "6098baf",
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.",
"built_at": "2026-07-30T20:01:56Z",
"pipeline_commit": "e64a046",
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
"data_vintage": {
"source_vintages": {
"2012": "10162019",
@@ -36,9 +36,9 @@
{
"year": 2011,
"path": "data/long/year=2011/part-0.parquet",
"sha256": "84302ab364dc9fc3b3fbbc3c3f8b826e3508b4d73ff7c42d094d3863cd1e37b5",
"row_count": 2864212,
"size_bytes": 3845911
"sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
"row_count": 496004,
"size_bytes": 2202455
},
{
"year": 2012,
@@ -74,9 +74,14 @@
"description": "canonical_fips_xwalk.parquet"
},
{
"path": "data/summary_categories.parquet",
"sha256": "0985b607f3f35a8dff62c0561261ab6922423b81d11c07b03bcb3e3461f85e33",
"description": "summary_categories.parquet"
"path": "data/census_collection_coverage.parquet",
"sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
"description": "census_collection_coverage.parquet"
},
{
"path": "data/code_set.parquet",
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
"description": "code_set.parquet"
},
{
"path": "data/harmonization_map.parquet",
@@ -88,10 +93,25 @@
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
"description": "harmonization_recipes.parquet"
},
{
"path": "data/lineage_events.parquet",
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
"description": "lineage_events.parquet"
},
{
"path": "data/representation.parquet",
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
"description": "representation.parquet"
},
{
"path": "data/series_breaks.parquet",
"sha256": "b0b6794b6887a4f300079adfa10029c2a77109faa4952fbff1c5a270793cc02b",
"sha256": "381090660c8b8a1bee852e7f63d29b9ecaf10f71870f41c63de92017c83b6f1f",
"description": "series_breaks.parquet"
},
{
"path": "data/summary_categories.parquet",
"sha256": "e4918abf8e9e6d1372d7ddc255dc199f9c303e250f4106449241b48c68abee67",
"description": "summary_categories.parquet"
}
]
},
+18 -4
View File
@@ -14,16 +14,16 @@
"basis_note": { "type": ["string", "null"] },
"expenditure_concept": {
"type": "string",
"enum": ["direct", "total"],
"description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments."
"enum": ["primary", "direct", "total"],
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
},
"expenditure_concept_note": {
"type": ["string", "null"],
"description": "How the intergovernmental leg was assembled; null for 'direct'."
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
},
"expenditure_concept_direct_suppressed": {
"type": "boolean",
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
},
"harmonization": { "type": "object" },
"recipe": { "type": ["object", "null"] },
@@ -33,6 +33,20 @@
"aggregate_fallback": { "type": ["object", "null"] },
"transformations":{ "type": "object" },
"series_break_refs": { "type": "array", "items": { "type": "string" } },
"completion": {
"type": "object",
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
"properties": {
"applied": { "type": "boolean" },
"rows_filled": { "type": "integer" },
"absence_means": { "type": "object" }
}
},
"corpus_break_refs": {
"type": "array",
"items": { "type": "string" },
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
},
"manifest": { "type": "object" },
"sql_query": { "type": "string" }
}
+7
View File
@@ -0,0 +1,7 @@
-- Category crosswalk. Numbered 11 (not with the other reference tables at
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
-- already exist when they register.
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+18 -1
View File
@@ -1,5 +1,22 @@
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
-- (summary_categories.category_type = 'expenditure'), never by item-code
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
-- every non-intergovernmental expenditure subtype: operations, capital,
-- assistance, interest, insurance_benefits.
--
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
-- 24-ig_long.sql).
CREATE OR REPLACE VIEW spending_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G')
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
)
AND NOT is_aggregate;
+11 -1
View File
@@ -1,5 +1,15 @@
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
-- Scope is Census General Revenue: every crosswalk revenue subtype EXCEPT
-- insurance_trust (Y01/Y02/Y04/Y11/Y12/Y51/Y52). Owner ruling 2026-07-30:
-- the default revenue concept stays general; surfacing insurance-trust
-- revenue through an explicit concept argument is uscogdata#12.
CREATE OR REPLACE VIEW revenue_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D')
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
AND revenue_subtype <> 'insurance_trust'
)
AND NOT is_aggregate;
+11 -1
View File
@@ -1,6 +1,16 @@
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
-- classification, applied to harmonized_code (the code the row is folded
-- onto) rather than the published item_code. Safe because the harmonized
-- space is leaf-only and every harmonized_code in the corpus is a
-- summary_categories member (verified at fixture regen; a code the
-- crosswalk cannot classify would be silently dropped here).
CREATE OR REPLACE VIEW spending_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G');
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
);
+8 -1
View File
@@ -1,6 +1,13 @@
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
-- classification (General Revenue = revenue minus insurance_trust), applied
-- to harmonized_code rather than the published item_code.
CREATE OR REPLACE VIEW revenue_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND LEFT(harmonized_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D');
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
AND revenue_subtype <> 'insurance_trust'
);
+12 -5
View File
@@ -1,4 +1,6 @@
-- Intergovernmental expenditure rows (M = to local govts, L = to state govts).
-- Intergovernmental expenditure rows: crosswalk spend_subtype =
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
-- = state payments to school systems -- uscogdata#11, finding F-017).
--
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
@@ -9,10 +11,15 @@
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
-- the pipeline's recipe joins use.
--
-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up
-- the L-NN codes, so including it would double-count.
-- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
-- rolls up the L-NN codes, so including it would double-count. The crosswalk
-- deliberately carries no `--` family-total codes, so membership excludes it
-- (guarded by "the IG leg never includes the L-- family total" in
-- tests/testthat/test-expenditure-concept.R).
CREATE OR REPLACE VIEW ig_long AS
SELECT *
FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L')
AND item_code NOT LIKE '%--';
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+9 -2
View File
@@ -8,8 +8,15 @@
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
-- while never dropping a row.
--
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
-- rather than the COALESCEd code: every IG harmonization target (M36) is
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
-- the column that exists on every row.
CREATE OR REPLACE VIEW ig_long_harmonized AS
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
FROM long
WHERE LEFT(item_code, 1) IN ('M', 'L')
AND item_code NOT LIKE '%--';
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
-3
View File
@@ -1,3 +0,0 @@
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW representation AS
SELECT *
FROM read_parquet('{url}data/representation.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW code_set AS
SELECT *
FROM read_parquet('{url}data/code_set.parquet');
+10 -1
View File
@@ -11,7 +11,8 @@ cog_find_peers(
same_state = FALSE,
pop_range = c(0.7, 1.3),
is_ratio = TRUE,
max_peers = 10L
max_peers = 10L,
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -34,6 +35,14 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
`pop_range` is interpreted as absolute population counts.}
\item{max_peers}{Integer cap on the number of peers returned.}
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
most recent census year with an observed population, so a cohort is not
built from a sample year in which most of the candidate universe is
absent. `"consistent"` needs a year range, which cohort selection does not
have, so it selects like `"all"` and is carried on the result as
`attr(x, "coverage")` for [cog_peer_compare()].}
}
\value{
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
+28 -7
View File
@@ -10,7 +10,8 @@ cog_geographic_rollup(
years,
per_capita = FALSE,
adjust_to_year = NULL,
expenditure_concept = c("direct", "total")
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -29,12 +30,32 @@ are excluded from the result.}
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
single-government queries but cannot be used here because combining Total
across multiple layers of government double-counts intergovernmental
transfers (a state's payment to a school district is the same dollar the
district reports as its own Direct spending).}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across multiple layers of
government double-counts intergovernmental transfers (a state's payment
to a school district is the same dollar the district reports as its own
Direct spending); `"primary"` and `"direct"` combine safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.}
}
\value{
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
+8 -4
View File
@@ -32,8 +32,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
}
\details{
* **Utility mode** (single `name`, the original behavior): returns all
rows whose `gov_name` matches the regex case-insensitively, sorted by
`population_acs` descending. Useful for exploratory lookups.
rows whose `gov_name` contains `name` as a **literal, case-insensitive
substring**, sorted by `population_acs` descending. Useful for
exploratory lookups. Regex metacharacters in `name` are escaped, so a
government is findable by its own complete name even when that name
contains parentheses or a period.
* **Basket mode** (`length(name) > 1`): resolves each input row to a
single canonical govid and returns a tibble in input order, suitable
for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -45,7 +48,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
2. **Exact pass:** case-insensitive equality against `gov_name`.
Single hit -> resolved. Multiple -> step 4.
3. **Substring fallback:** case-insensitive regex against `gov_name`.
3. **Substring fallback:** case-insensitive literal substring against
`gov_name` (metacharacters escaped).
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
`status = "no_match"`. Multiple hits -> step 4.
4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -58,7 +62,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
}
\examples{
\dontrun{
# Utility mode — exploratory regex lookup
# Utility mode — exploratory substring lookup
cog_gov_search("broward", state = "FL")
# Basket mode — resolve a known cohort
+62 -6
View File
@@ -11,7 +11,8 @@ cog_peer_compare(
years,
per_capita = TRUE,
adjust_to_year = NULL,
expenditure_concept = c("direct", "total")
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
)
}
\arguments{
@@ -29,10 +30,37 @@ population.}
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
single-government queries but cannot be used here because combining Total
across peer sets counts intergovernmental transfers twice.}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across peer sets counts
intergovernmental transfers twice; `"primary"` and `"direct"` combine
safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.
The comparison target is exempt from `"consistent"` balancing -- it is the
subject of the comparison, not a member of the cohort -- and the
`summary_*` quantiles are computed AFTER the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against
the cohort size: "3 of your 15 peers reported in FY2019".}
}
\value{
Tibble matching [cog_spending()]'s columns, plus a `role`
@@ -43,11 +71,39 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
`cohort_year`, and `cohort_govids`.
**The `summary_*` rows are per-category quantiles: they are not additive.**
Each one is computed **within each `(year, spend_subtype,
category)` cell** across the peer set, so a `summary_p50` row is *the
median peer's value in that one category*, not *the value of the median
peer's total*. The median peer for Police and the median peer for Fire
are usually different governments, so summing `summary_*` rows across
categories does not give any peer's total and misstates the band it
appears to describe — measured at −32.7% to +251.0% across 24 years on
one cohort, with a sign flip at FY2012.
Facet by `role` **and** `category` (the documented use, and what the
rows are built for). For a genuine "median peer's total spending" line,
sum each peer's own categories first and take the quantile of those
per-government totals:
```r
library(dplyr)
cmp |>
filter(role %in% c("target", "peer")) |>
group_by(year, role, canonical_govid) |>
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
filter(role == "peer") |>
group_by(year) |>
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
```
}
\description{
Pulls spending for the target plus a peer set (either a
[cog_find_peers()] result or a character vector of `canonical_govid`) and
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
`summary_p75`) so the result can be faceted by `role` in a single ggplot
call.
call. Those summary rows are quantiles **within each category**, not
quantiles of each peer's total — see the `@return` section before summing
them.
}
+27 -2
View File
@@ -11,7 +11,8 @@ cog_revenue(
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL
recipe = NULL,
complete = FALSE
)
}
\arguments{
@@ -55,12 +56,36 @@ argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
}
\description{
Mirror of [cog_spending()] for revenue categories. One row per
+69 -30
View File
@@ -12,7 +12,8 @@ cog_spending(
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
expenditure_concept = c("direct", "total")
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE
)
}
\arguments{
@@ -57,41 +58,79 @@ argument is ignored and the result's provenance reports
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{expenditure_concept}{`"direct"` (default) returns only the
government's own direct spending (item codes `E`/`F`/`G`), unchanged
from prior releases. `"total"` additionally UNIONs in the
intergovernmental leg -- payments to local governments (`M` codes) and
to the state government (`L` codes, excluding the `L--` family-total
rollup) -- so results gain rows with `spend_subtype ==
"intergovernmental"`. Requires the active corpus's `summary_categories`
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
`uscogdata_ig_categories_unsupported` on an older corpus rather than
silently under-reporting. Mutually exclusive with `recipe` (a recipe
already defines its own component codes). **Do not sum `"total"`
results across levels of government** (e.g. state + county + city):
a state's `M12` payment to a school district is the same dollar the
district reports as its own direct `E12`, so summing both double-counts
it. This matters in particular with [cog_geographic_rollup()], which
sums across exactly that kind of multi-layer government set.
\item{expenditure_concept}{Which spending concept to return. Concepts are
defined as sets of the crosswalk's `spend_subtype` values -- never as
item-code first letters, which cannot classify correctly (prefix `Y`
alone spans revenue, expenditure, and balance codes):
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
split), which the Direct leg excludes by construction but the IG leg
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
query, any (year, category) where this leaves intergovernmental rows
with NO Direct counterpart is flagged: the affected rows' `notes`
name the harmonization recipe that recovers the missing Direct
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG.}
* `"primary"` (default) -- the government's own service provision:
`operations` + `capital` + `assistance` subtypes.
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
`interest` (interest on debt) and `insurance_benefits` (insurance
trust benefit payments, e.g. pensions -- Census manual section
5.2.2.1 includes payments to retirees in Direct).
* `"total"` -- `direct` plus the intergovernmental leg: payments to
local governments (`M` codes), to the state government (`L` codes,
excluding the `L--` family-total rollup), and state payments to
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
`spend_subtype == "intergovernmental"`. Requires the active corpus's
`summary_categories` to carry M/L rows (added by cog_pipeline PR
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
older corpus rather than silently under-reporting. Mutually
exclusive with `recipe` (a recipe already defines its own component
codes).
**Do not sum `"total"` results across levels of government** (e.g.
state + county + city): a state's `M12` payment to a school district is
the same dollar the district reports as its own direct `E12`, so
summing both double-counts it. This matters in particular with
[cog_geographic_rollup()], which sums across exactly that kind of
multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
split), which the Direct leg excludes by construction but the IG leg
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
query, any (year, category) where this leaves intergovernmental rows
with NO Direct counterpart is flagged: the affected rows' `notes`
name the harmonization recipe that recovers the missing Direct
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
whose `completion` block reports `applied`, `rows_filled`, and the
per-year `absence_means` rule that was applied.
}
\description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
+63
View File
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
if (nzchar(p)) paste0(p, "/") else ""
}
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
# or "" when it isn't there.
#
# Tests that assert on documentation content have to read the sources, and the
# sources only exist when the suite runs from a checkout. Under R CMD check the
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
# shipped and `../../README.md` does not resolve -- so those tests must skip
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
# rcmdcheck, so the assertions are still enforced on every push; this only
# stops them from failing a context that structurally cannot satisfy them.
source_tree_path <- function(...) {
p <- testthat::test_path("..", "..", ...)
if (file.exists(p)) p else ""
}
# Skip unless every named source file is present (see source_tree_path()).
skip_if_no_source_tree <- function(...) {
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
character(1))
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
testthat::skip_if(
any(missing),
"package source tree not available (running against the installed package)"
)
invisible(paths)
}
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
skip_if_no_corpus <- function() {
p <- fixture_corpus_path()
@@ -58,6 +85,42 @@ with_doctored_schema_version <- function(version, code) {
force(code)
}
# Copy the bundled fixture to a temp dir with representation.parquet and
# code_set.parquet removed (and dropped from the manifest's metadata list),
# then run `code` against it. Models a corpus published BEFORE sparsification:
# schema_version is left alone deliberately, because it was never bumped for
# that change -- the pre-sparsification fixture this package shipped until
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
# is therefore the only honest signal, and this helper is what proves the
# package keys off it rather than off the version number.
with_corpus_missing_representation <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
dropped <- c("representation.parquet", "code_set.parquet")
file.remove(file.path(tmp, "data", dropped))
manifest_path <- file.path(tmp, "manifest.json")
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
m$files$metadata <- Filter(
function(f) !basename(f$path) %in% dropped, m$files$metadata
)
writeLines(
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
manifest_path
)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
# Copy the bundled fixture to a temp dir with summary_categories.parquet
# rewritten to drop every M/L (intergovernmental) row, then run `code`
# against it with a clean session (mirrors with_fixture_corpus()/
+11 -5
View File
@@ -17,7 +17,16 @@
# cog-api's llms.txt, which is silent on units).
test_that("returned amounts are documented as full US dollars where readers meet the package", {
testthat::skip("Blocked on uscogdata#15 (finding F-004)")
# README and vignettes ship only in the source tree, not in the installed
# package, so these assertions cannot run under R CMD check -- CI's earlier
# testthat::test_local() step is what enforces them. See
# skip_if_no_source_tree() in helper-fixture.R.
docs <- skip_if_no_source_tree(
"README.md",
c("vignettes", "total-spending.Rmd"),
c("vignettes", "population-denominators.Rmd")
)
says_units <- function(path) {
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
@@ -25,10 +34,7 @@ test_that("returned amounts are documented as full US dollars where readers meet
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
}
expect_true(says_units(testthat::test_path("..", "..", "README.md")))
expect_true(says_units(testthat::test_path("..", "..", "vignettes", "total-spending.Rmd")))
expect_true(says_units(testthat::test_path("..", "..", "vignettes",
"population-denominators.Rmd")))
for (path in docs) expect_true(says_units(path))
# Pin the documented claim to the actual behaviour, so the two cannot drift.
# The expected raw amount is read straight from the corpus's parquet
+23 -3
View File
@@ -8,14 +8,30 @@ test_that("cog_categories returns all categories grouped by subtype", {
expect_gt(nrow(r), 10L)
# corpus preserves Census-native "expenditure" vocabulary; the API takes
# "spending" as a friendlier alias.
expect_setequal(unique(r$category_type), c("expenditure", "revenue"))
#
# `balance` joined as a third category_type with the cash-and-security
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
# money verb, so it surfaces every category_type the corpus carries -- the
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
# return a balance row.
expect_setequal(unique(r$category_type),
c("expenditure", "revenue", "balance"))
})
test_that("cog_categories(type = 'spending') returns only expenditure rows", {
skip_if_no_corpus()
r <- cog_categories(type = "spending")
expect_true(all(r$category_type == "expenditure"))
expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental")))
# "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
# carrying dollars now maps to a category.
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
# joined with the I/Q/Y flow batch -- the last two characters of Census's
# expenditure taxonomy. `interest` is what makes the three-concept model
# computable: primary = direct minus debt service.
expect_true(all(r$subtype %in%
c("operations", "capital", "intergovernmental", "assistance",
"interest", "insurance_benefits")))
})
test_that("cog_categories surfaces the intergovernmental spending subtype", {
@@ -33,8 +49,12 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
skip_if_no_corpus()
r <- cog_categories(type = "revenue")
expect_true(all(r$category_type == "revenue"))
# `insurance_trust` (Y01/Y02/Y04/Y11/Y12/Y51/Y52) is deliberately NOT
# own_source: Census's "General Revenue" excludes insurance trust revenue,
# and Y01 alone is $1.31T corpus-wide.
expect_true(all(r$subtype %in%
c("own_source", "federal", "state", "local_aid")))
c("own_source", "federal", "state", "local_aid",
"insurance_trust")))
})
test_that("cog_categories(pattern = ...) filters case-insensitively", {
+193
View File
@@ -0,0 +1,193 @@
# tests/testthat/test-complete.R
#
# uscogdata#18. The published corpus no longer stores the wide era's explicit
# zeros (cog_pipeline#64, series break SB194), so absence means two different
# things:
#
# <= FY2011 (dense_source) : cell absent => Census published $0
# >= FY2012 (sparse_source): cell absent => not reported, unknown
#
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
# row's `value_source` so the two are distinguishable. Expected row sets here
# are built from the corpus parquet directly, never from the verb under test --
# verifying what a filter does through that same filter proves nothing.
# The (subtype, category) cells that SHOULD exist for one government-year:
# every code in force for that government's type, mapped through
# summary_categories, matching the verb's crosswalk subtype scope (the
# default concept, `primary`, is operations/capital/assistance -- see
# uscogdata#11) and excluding aggregate-flagged codes (which
# spending_long/revenue_long drop).
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
wt_raw_query(sprintf(
"SELECT DISTINCT c.%s AS subtype, c.category
FROM %s cs
JOIN %s x ON x.govs_type = cs.type
JOIN %s c ON c.item_code = cs.item_code
WHERE x.canonical_govid = '%s'
AND cs.year = %d
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%s IN (%s)",
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
q("summary_categories.parquet"), govid, year,
subtype_col, paste0("'", subtypes, "'", collapse = ",")
))
}
# The default expenditure concept's subtype scope, mirrored from
# R/spending.R's .spend_subtypes_primary.
primary_subtypes <- c("operations", "capital", "assistance")
test_that("complete = FALSE is the default and changes nothing", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
expect_equal(nrow(plain), nrow(explicit))
expect_false("value_source" %in% names(plain))
})
})
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2011 is dense_source: before sparsification this government carried a
# row for every code in force, most of them $0. complete = TRUE must
# reproduce that cell set exactly.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$spend_subtype, r$category),
key(expected$subtype, expected$category))
expect_gt(nrow(expected), 0L)
# Every filled cell in a dense-source year is a Census-published $0 --
# never "unknown", which is what the modern era's absences mean.
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
})
})
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
full <- cog_spending("121011212191", 2011L, complete = TRUE)
# Filling adds rows; it must never alter or drop one.
expect_gt(nrow(full), nrow(plain))
reported <- full[full$value_source == "reported", ]
expect_equal(nrow(reported), nrow(plain))
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
# ... and the total is unchanged, because every added cell is $0.
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
})
})
test_that("a sparse-source year's absences are unknown, not zero", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2019 is sparse_source: an absent cell means the government did not
# report, which is NOT a zero. Filling those with 0 would invent data --
# the exact error the representation contract exists to prevent.
r <- cog_spending("121011212191", 2019L, complete = TRUE)
filled <- r[r$value_source != "reported", ]
expect_gt(nrow(filled), 0L)
expect_true(all(filled$value_source == "not_reported"))
expect_true(all(is.na(filled$amt_nominal)))
expect_false(any(r$value_source == "census_zero"))
})
})
test_that("the fill is scoped to each government's own type", {
skip_if_no_corpus()
with_fixture_corpus({
# Filling against the union of all types would invent cells for codes a
# county can never report. Every filled category must be one that
# code_set puts in force for type 1 (county) specifically.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
county_cells <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
expect_true(all(r$category %in% county_cells$category))
})
})
test_that("complete = TRUE respects the category filter", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011L, category = "Police",
complete = TRUE)
expect_true(all(r$category == "Police"))
expect_true("value_source" %in% names(r))
})
})
test_that("cog_revenue() completes on its own flow", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
c("own_source", "federal", "state", "local_aid"),
"revenue_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$revenue_subtype, r$category),
key(expected$subtype, expected$category))
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
})
})
test_that("provenance records the completion and its absence rule", {
skip_if_no_corpus()
with_fixture_corpus({
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
"provenance")
expect_true(prov$completion$applied)
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
expect_gt(prov$completion$rows_filled, 0L)
off <- attr(cog_spending("121011212191", 2011L), "provenance")
expect_false(off$completion$applied)
expect_equal(off$completion$rows_filled, 0L)
})
})
test_that("complete = TRUE is refused where the fill would be guesswork", {
skip_if_no_corpus()
with_fixture_corpus({
# A recipe defines its own component codes and does not go through
# summary_categories at all, so there is no grid to fill from.
expect_error(
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
# The intergovernmental leg keeps aggregate rows by design
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
expect_error(
cog_spending("121011212191", 2011L, expenditure_concept = "total",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
})
})
test_that("complete = TRUE aborts on a corpus with no representation contract", {
skip_if_no_corpus()
# A corpus published before sparsification carries neither table, so there
# is nothing to fill from and no rule saying what an absence means. That
# must abort rather than guess.
with_corpus_missing_representation({
expect_error(
cog_spending("121011212191", 2011L, complete = TRUE),
class = "uscogdata_representation_unavailable"
)
# ... while an ordinary query on the same corpus still works.
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
})
})
+94
View File
@@ -0,0 +1,94 @@
# tests/testthat/test-corpus-breaks.R
#
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
# are caveats about the corpus itself rather than about one item code:
#
# SB085 1977 dollar precision across the 1976/1977 boundary
# SB087 2002 imputation exclusion FY2002-2006
# SB194 2012 dense -> sparse representation change
# SB086 2017 government ID scheme change
#
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
# and no row's item_code is ever the literal "ALL", so none of them could
# ever reach a user. They now travel in their own provenance field,
# `corpus_break_refs`, which keeps them distinguishable from the
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
# result, not one series).
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
skip_if_no_corpus()
with_fixture_corpus({
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
# DoD 4 intended to reach users.
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_true("SB194" %in% prov$corpus_break_refs)
})
})
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
skip_if_no_corpus()
with_fixture_corpus({
# 2019-2020 spans no catalogued corpus-wide break.
r <- cog_spending("121011212191", 2019:2020, "Police")
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
})
})
test_that("corpus_break_refs and series_break_refs stay disjoint", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_type(prov$series_break_refs, "character")
expect_type(prov$corpus_break_refs, "character")
# An ALL caveat must never masquerade as a break in a specific series.
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
expect_false("SB194" %in% prov$series_break_refs)
})
})
test_that(".build_corpus_break_refs matches on the break_year window alone", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
# series_breaks table is a full cross-vintage registry, so the matching
# logic is testable there even though no long partition covers it.
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1975:1980, schema_version = 6L
))
# ... and does not fire for a range that misses it, unlike a filter keyed
# on the era rather than the boundary.
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1978:1980, schema_version = 6L
))
# Unlike code-specific refs, these do not depend on which codes a result
# happens to contain -- that dependency is the whole defect.
expect_setequal(
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
"SB087"
)
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
expect_equal(
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
character(0)
)
})
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
out <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
expect_match(out, "SB194", fixed = TRUE)
})
})
+14 -3
View File
@@ -30,7 +30,6 @@ wt_coverage <- function(x) {
}
test_that("multi-government aggregates disclose reporting coverage on every result", {
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
# -- F-020: geographic rollups -------------------------------------------
# Wisconsin's city/village universe is 608 governments. On the bundled
@@ -49,10 +48,22 @@ test_that("multi-government aggregates disclose reporting coverage on every resu
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
# Cross-check against the raw partitions, scoped to the SAME universe the
# rollup was given -- the 608 govids above. Scoping instead on the long
# table's own `type`/`fips_state` asks a different question and answers 595:
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
# identity, when they were townships) while the xwalk lists them as
# govs_type = 2 (their present identity, as villages). Schema v6 made the
# long table's geography present-harmonized and moved as-of-year to the
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
# in R/manifest.R. n_units_reporting counts against the requested universe,
# so 597 is the number that answers "how many of the governments I asked
# about reported".
raw_2012 <- wt_raw_query(paste0(
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ",
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate"))
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
"AND canonical_govid IN (",
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
# -- F-023: peer cohorts --------------------------------------------------
+21 -4
View File
@@ -13,9 +13,13 @@ test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
}
})
test_that("expenditure_concept defaults to direct and preserves today's numbers", {
test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
gov <- "010000226085" # Alabama state government
base <- cog_spending(gov, years = 2019, category = "Police")
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
# Police maps only to operations/capital codes (E62/F62/G62), so the
# direct concept's extra subtypes (interest, insurance_benefits) cannot
# contribute and the two concepts must agree exactly here.
expl <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "direct")
expect_equal(base$amt_nominal, expl$amt_nominal)
@@ -59,7 +63,9 @@ test_that("the IG leg never includes the L-- family total", {
codes <- DBI::dbGetQuery(con,
"SELECT DISTINCT item_code FROM ig_long")$item_code
expect_false(any(grepl("--$", codes)))
expect_true(all(substr(codes, 1, 1) %in% c("M", "L")))
# Q joined the IG family with the crosswalk-membership rewrite
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
})
test_that("expenditure_concept rejects unknown values", {
@@ -246,9 +252,12 @@ test_that("both cross-government verbs still accept the direct default", {
})
test_that("provenance always records the expenditure concept", {
d <- cog_spending("010000226085", years = 2019, category = "Police")
p <- cog_spending("010000226085", years = 2019, category = "Police")
d <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
expect_equal(attr(t, "provenance")$expenditure_concept, "total")
# The note explains the non-obvious part: how legacy IG was assembled.
@@ -298,8 +307,16 @@ test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
# Direct/Total pairing. The flow-family gate in
# .attach_ig_counterparts() must keep ig_recipe_id NULL here.
#
# Anchored on FL state government, not AL. Coverage is presence-based: a
# recipe is only suggested when its component codes have rows for the
# requested government-year. AL state's only FY2011 B47 cell was an
# explicit zero, which the corpus no longer stores after sparsification
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
# FL state carries a real FY2011 B47 amount, so this exercises the guard
# against a suggestion that genuinely fires.
r <- suppressMessages(
cog_spending("010000226085", years = c(2005, 2011), category = "IG Federal")
cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
)
sugg <- attr(r, "provenance")$suggestions
expect_gt(length(sugg), 0L)
+43 -6
View File
@@ -19,8 +19,6 @@
# also check the FY2022 numbers above.
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
testthat::skip("Blocked on uscogdata#11 (findings F-012, F-017, F-018)")
mad <- "552025209777" # MADISON CITY, WI
wi_state <- "550000227544" # WISCONSIN (state government)
@@ -61,15 +59,54 @@ test_that("expenditure concepts classify on spend_type, not item-code prefix", {
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
# payments. All four share the first letter `Y` and the spend_type
# "Insurance Trust", so this pair of assertions is the concrete proof that
# classification is no longer keyed on the first letter.
# payments. All four share the first letter `Y`, so no first-letter allowlist
# can route them. The proof that classification is crosswalk-keyed:
# Y05 lands in `total` spending (insurance_benefits is inside `direct`),
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
# therefore can never appear in a spending result.
#
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
# DEFAULT stays Census General Revenue and so excludes insurance-trust
# revenue; Y01's revenue-side classification is asserted against the
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
# revenue concept argument is uscogdata#12.
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
spend_codes <- wt_codes_included(wi_total)
rev_codes <- wt_codes_included(wi_revenue)
expect_true("Y05" %in% spend_codes)
expect_false("Y05" %in% rev_codes)
expect_true("Y01" %in% rev_codes)
expect_false("Y01" %in% spend_codes)
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
con <- uscogdata:::.ensure_session()
y_class <- DBI::dbGetQuery(con,
"SELECT item_code, category_type, spend_subtype, revenue_subtype
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
})
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
# balances) share first letters with flow codes, so this could never be
# guaranteed under prefix classification; under crosswalk membership it
# falls out structurally -- asserted here at the verb level, on a
# government-year the fixture gives real balance rows (Wisconsin carries
# Y07/Y08/Y21/Y61-type balances in FY2019).
wi_state <- "550000227544"
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
expect_gt(nrow(balance), 0L)
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
rev <- cog_revenue(wi_state, 2019L)
expect_false(any(spend$category %in% balance$category))
expect_false(any(rev$category %in% balance$category))
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
})
+1 -1
View File
@@ -83,7 +83,7 @@ test_that("cog_explain prints the expenditure concept (I1)", {
capture.output(cog_explain(t)),
capture.output(cog_explain(t), type = "message")
), collapse = "\n")
expect_true(grepl("Concept: direct", txt_d))
expect_true(grepl("Concept: primary", txt_d))
expect_true(grepl("Concept: total", txt_t))
})
+112
View File
@@ -0,0 +1,112 @@
# tests/testthat/test-fixture-vintage.R
#
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
# every test in this package -- plus the whole cog-api suite -- runs against
# it. When the published corpus changes shape and the fixture does not, both
# suites stay green against a corpus that no longer exists (uscogdata#18).
#
# These tests pin the structural facts that distinguish the current published
# vintage from its predecessor, so a stale fixture fails loudly instead of
# passing quietly. They assert shape, never dollar values: re-running
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
# keep them green.
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
# not the package session: these assertions are about what the fixture
# CONTAINS, and routing them through the reader's own views would let a
# filter hide the very absence being checked.
fixture_query <- function(sql, ...) {
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
path <- function(rel) {
sprintf("read_parquet(%s)",
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
}
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
}
test_that("fixture ships every metadata table the publish tree does", {
skip_if_no_corpus()
# representation/code_set are what make a sparse corpus interpretable; a
# fixture without them predates sparsification (cog_pipeline#64).
expected <- c(
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet", "code_set.parquet",
"harmonization_map.parquet", "harmonization_recipes.parquet",
"lineage_events.parquet", "representation.parquet",
"series_breaks.parquet", "summary_categories.parquet"
)
on_disk <- basename(list.files(
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
))
expect_true(all(expected %in% on_disk))
# The manifest must list them too -- consumers read the manifest, not ls().
in_manifest <- with_fixture_corpus(
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
)
expect_true(all(expected %in% in_manifest))
})
test_that("fixture carries the dense/sparse representation contract", {
skip_if_no_corpus()
rep <- fixture_query(
"SELECT year, representation, absence_means FROM %s
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
"data/representation.parquet"
)
expect_equal(nrow(rep), 4L)
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
})
test_that("the fixture's wide era is sparse, not zero-padded", {
skip_if_no_corpus()
# FY2011 is a dense_source year: the corpus publishes only the cells Census
# reported non-zero, and an absent cell means Census published $0. Before
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
# zeros. A single explicit zero here means the fixture predates the change.
zeros_2011 <- fixture_query(
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
"data/long/year=2011/part-0.parquet"
)$n
expect_equal(zeros_2011, 0L)
# The modern era is a different regime: a reported zero there is real data
# (the government filed $0), so zeros legitimately survive and must not be
# asserted away.
expect_gt(
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
0L
)
})
test_that("code_set covers every fixture year with the reader-spec columns", {
skip_if_no_corpus()
cs <- fixture_query(
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
"data/code_set.parquet"
)
expect_true(all(
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
%in% names(cs)
))
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
})
test_that("every flow code carrying dollars has a category, J-prefix included", {
skip_if_no_corpus()
# The J (assistance/benefit) codes were uncategorised until the crosswalk
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
# duplication fix landed). Their absence is how a pre-crosswalk fixture
# gives itself away.
j <- fixture_query(
"SELECT item_code, category, category_type, spend_subtype FROM %s
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
"data/summary_categories.parquet"
)
expect_true("J19" %in% j$item_code)
expect_true(all(j$category_type == "expenditure"))
expect_true(all(j$spend_subtype == "assistance"))
expect_false(any(is.na(j$category)))
})
@@ -18,7 +18,6 @@
# semantics, not a row the fix makes findable.
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
testthat::skip("Blocked on uscogdata#16 (finding F-025)")
# -- correctness (1): a government must be findable by its own exact name ---
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
+9 -3
View File
@@ -13,10 +13,16 @@
# is unaffected, so the fix is documentation: one sentence in @return.
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
testthat::skip("Blocked on uscogdata#14 (finding F-021)")
rd <- paste(readLines(testthat::test_path("..", "..", "man", "cog_peer_compare.Rd"),
warn = FALSE), collapse = " ")
# man/ ships only in the source tree (the installed package carries a
# compiled help database instead), so the prose assertions below cannot run
# under R CMD check -- CI's earlier testthat::test_local() step enforces
# them. The numeric pin further down needs only the corpus, but it lives in
# the same test_that() as the sentence it protects, deliberately: they are
# one claim, and splitting them would let the prose drift while a separate
# test kept passing.
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
# The @return section must say the quantile is computed within each cell...
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
+5 -2
View File
@@ -80,8 +80,11 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
test_that("cog_geographic_rollup accepts data.frames per layer", {
skip_if_no_corpus()
fl_state <- cog_gov_search("^FLORIDA$", type = "state")
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
# Unanchored: utility mode matches literally now, so "^...$" would be
# searched for as characters rather than read as anchors (uscogdata#16).
# Both still resolve to exactly one row once scoped by type/state.
fl_state <- cog_gov_search("FLORIDA", type = "state")
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
r <- cog_geographic_rollup(
govids = list(state = fl_state, county = broward),
category = "Police", years = 2020L
+44 -20
View File
@@ -98,7 +98,11 @@ test_that("cog_spending rejects invalid inputs", {
test_that("cog_spending accepts a cog_gov_search result directly", {
skip_if_no_corpus()
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
# Unanchored: utility mode matches `name` as a literal substring now, so
# "^...$" would be searched for as those characters rather than read as
# anchors (uscogdata#16). Scoped by state and type, the bare name still
# resolves to exactly one row.
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
expect_gt(nrow(picks), 0L)
r <- cog_spending(picks, 2020L, "Corrections")
expect_equal(unique(r$canonical_govid), "121011212191")
@@ -244,24 +248,42 @@ test_that("basis defaults to 'harmonized' when not passed", {
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Corrections")
# FL state government. The harmonization block is scoped by government,
# year and flow prefix -- NOT by category -- so a Corrections query still
# counts every E/F/G-prefixed row the harmonized basis drops for having
# no harmonized_code. The three that apply here are E21/F21/G21
# (Education NEC, SB184-186, "discontinued_na", wide-era window ending
# FY2011); the other discontinued_na rulings live outside E/F/G.
# See docs/phase_r_harmonization_review.md § 1.3/1.4 and cog_pipeline
# data/harmonization_map.csv.
r <- cog_spending("120000226351", 2011:2012, "Corrections")
prov <- attr(r, "provenance")
expect_equal(prov$basis, "harmonized")
expect_true(prov$harmonization$applied)
expect_true(prov$harmonization$na_rows_excluded >= 0L)
expect_true(prov$harmonization$na_amount_excluded >= 0)
# Data-verified for the v6 fixture (corpus 2026-07-22). The Task 18 map
# extension added E/F/G-prefix discontinued_na rulings the earlier pin's
# comment predated: E21/F21/G21 (Education NEC local, SB184-186,
# "trivial; explicit-NA, full wide-era window"). Broward's 2011 legacy
# partition zero-pads exactly those three codes, so this query now
# excludes 3 NA-harmonized rows -- all with amt = 0, hence the excluded
# AMOUNT stays exactly zero. (The other discontinued_na rulings -- S74,
# Z61, X04, X06, the debt-detail family, L24 -- remain outside the
# E/F/G/K prefixes.) See docs/phase_r_harmonization_review.md § 1.3/1.4
# and cog_pipeline data/harmonization_map.csv E21/F21/G21 rows.
expect_equal(prov$harmonization$na_rows_excluded, 3L)
expect_equal(prov$harmonization$na_amount_excluded, 0)
# $2,825,439 thousands of FY2011 E21 + F21 + G21, reported in full USD.
# Pinning a non-zero amount is the point: the earlier Broward anchor's
# three rows were all explicit zeros, so the AMOUNT accounting was
# asserted only against 0 and could not have caught a bug.
expect_equal(prov$harmonization$na_amount_excluded, 2825439 * 1000)
})
})
test_that("sparsification removed the wide era's zero-pads from the exclusion count", {
skip_if_no_corpus()
with_fixture_corpus({
# Broward County FY2011 used to carry E21/F21/G21 rows of exactly $0 --
# the wide era stored every government x every code, zeros included. The
# published corpus no longer does (SB194, cog_pipeline#64), so there is
# now nothing for the harmonized basis to exclude. Absence in a
# dense_source year means Census published $0; it does not mean the
# exclusion machinery stopped working, which the FL state anchor above
# proves independently.
r <- cog_spending("121011212191", 2011:2012, "Corrections")
h <- attr(r, "provenance")$harmonization
expect_true(h$applied)
expect_equal(h$na_rows_excluded, 0L)
expect_equal(h$na_amount_excluded, 0)
})
})
@@ -303,11 +325,13 @@ test_that("provenance$series_break_refs is a populated-when-applicable character
r <- cog_spending("121011212191", 2020L, "Corrections")
refs <- attr(r, "provenance")$series_break_refs
expect_type(refs, "character")
# No catalogued series_breaks_pq row falls inside this fixture's
# 2011/2012/2019/2020 window for the codes this query touches (E04/G04)
# -- data-verified; the mechanism itself is what's under test here, via
# a query-shaped unit test in test-views.R since the fixture has no
# positive case to pin against.
# No catalogued code-specific series_breaks_pq row falls inside this
# fixture's 2011/2012/2019/2020 window for the codes this query touches
# (E04/G04) -- data-verified; the mechanism itself is what's under test
# here, via a query-shaped unit test in test-views.R since the fixture
# has no positive case to pin against. Corpus-wide ("ALL") entries never
# appear in this field by construction -- they travel in
# corpus_break_refs; see test-corpus-breaks.R.
expect_equal(refs, character(0))
})
})
+105 -32
View File
@@ -36,8 +36,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
# {url} exactly as .register_views() does, and executes them -- plus
# their 10-long.sql dependency -- against a synthetic hive-partitioned
# parquet tree written to a temp dir. A regression in any predicate (e.g.
# `NOT is_aggregate` dropped, the prefix list changed, the NULL guard
# removed) would change which of the rows below survive.
# `NOT is_aggregate` dropped, the crosswalk-membership subquery changed,
# the NULL guard removed) would change which of the rows below survive.
#
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
# PARQUET) rather than the arrow package: this package has no arrow
@@ -61,25 +61,45 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
-- 'S74' is outside BOTH flow families (E/F/G/K spending and
-- T/A/U/B/C/D revenue -- it mirrors the real corpus's own
-- non-flow-type codes like S74/Z61), so it can only leak into
-- EITHER view via the E/F/G/K or T/A/U/B/C/D prefix filter, never
-- both at once -- a prefix drawn from the other view's own family
-- (e.g. a real T-code for the spending row) would incorrectly
-- leak into the other view's assertion below and not discriminate
-- the predicate under test.
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by the E/F/G/K prefix filter
-- Revenue (T/A/U/B/C/D) rows, exercised against revenue_long_harmonized:
-- 'S74' and 'Z61' are classified `balance` in the synthetic
-- crosswalk below (mirroring the real corpus's own non-flow codes),
-- so each is excluded from its view ONLY by the crosswalk-membership
-- subquery -- the mechanism that replaced the prefix allowlists
-- (uscogdata#11) and keeps balance stocks out of both flows
-- (uscogdata#25).
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by crosswalk membership (balance)
-- Revenue rows, exercised against revenue_long_harmonized:
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by the T/A/U/B/C/D prefix filter
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by crosswalk membership (balance)
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The flow views classify by membership in summary_categories, so the
# synthetic corpus needs one too. Every flow code above is a member of its
# own flow (so is_aggregate / NULL-harmonized exclusions stay the SOLE
# excluder for those rows); S74/Z61 are members but classified balance, so
# membership itself is what excludes them.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('E36', 'Water Utilities', 'expenditure', 'operations', NULL),
('E38', 'Water Utilities', 'expenditure', 'operations', NULL),
('E05', 'Corrections', 'expenditure', 'operations', NULL),
('E99', 'Other', 'expenditure', 'operations', NULL),
('S74', 'Fund Balances', 'balance', NULL, NULL),
('U11', 'Interest Earnings','revenue', NULL, 'own_source'),
('U10', 'Interest Earnings','revenue', NULL, 'own_source'),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source'),
('T88', 'Other Taxes', 'revenue', NULL, 'own_source'),
('Z61', 'Fund Balances', 'balance', NULL, NULL)
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
@@ -89,6 +109,7 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
@@ -97,8 +118,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
GROUP BY item_code ORDER BY item_code"
)
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
# harmonized_code), and spend-E (wrong prefix family) must all be gone,
# and spend-A + spend-B must be folded together under E36.
# harmonized_code), and spend-E (balance, not an expenditure member) must
# all be gone, and spend-A + spend-B must be folded together under E36.
expect_equal(nrow(spend), 1L)
expect_equal(spend$item_code, "E36")
expect_equal(spend$amt, 150)
@@ -138,12 +159,27 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is
('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis
('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views
('ig-D', 'L--', 55555, false, 'L--'), -- family total: excluded from BOTH views
('ig-E', 'T29', 44444, false, 'T29') -- wrong prefix (revenue, not M/L): excluded from BOTH views
('ig-D', 'L--', 55555, false, 'L--'), -- family total: deliberately NOT a crosswalk member, excluded from BOTH views
('ig-E', 'T29', 44444, false, 'T29') -- revenue member, not intergovernmental: excluded from BOTH views
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The IG views classify by summary_categories membership
# (spend_subtype = 'intergovernmental'). L-- is deliberately absent --
# exactly as it is from the real crosswalk -- which is what excludes it.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('M04', 'Corrections', 'expenditure', 'intergovernmental', NULL),
('M38', 'Health', 'expenditure', 'intergovernmental', NULL),
('M36', 'Health', 'expenditure', 'intergovernmental', NULL),
('M47', 'IG Other', 'expenditure', 'intergovernmental', NULL),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source')
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
@@ -153,6 +189,7 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("24-ig_long.sql"))
DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql"))
@@ -160,8 +197,9 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
"SELECT item_code, SUM(amt) AS amt FROM ig_long
GROUP BY item_code ORDER BY item_code"
)
# L-- (family total) and T29 (wrong prefix) are gone; the aggregate row
# M47 survives -- proof `NOT is_aggregate` is absent from ig_long.
# L-- (family total, not a member) and T29 (revenue, not IG) are gone; the
# aggregate row M47 survives -- proof `NOT is_aggregate` is absent from
# ig_long.
expect_equal(raw$item_code, c("M04", "M38", "M47"))
expect_equal(raw$amt, c(100, 50, 99999))
@@ -177,11 +215,13 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
})
test_that(".build_series_break_refs matches fin_code + break_year window", {
# No series_breaks_pq row falls inside the bundled fixture's 2011-2020
# window (data-verified; see the "series_break_refs" test in
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
# 2011-2020 window (data-verified; see the "series_break_refs" test in
# test-spending.R), so this proves the matching logic itself against the
# live view + a synthetic year window that DOES hit a cataloged break
# (SB075, fin_code E62, break_year 2005).
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
# separate path with its own coverage -- SB194 does sit at 2012, inside
# the fixture window; see test-corpus-breaks.R.
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
@@ -320,14 +360,32 @@ test_that(".harmonization_view_files guard is necessary: registration against a
)
})
test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
test_that("spending_long carries exactly the non-IG expenditure crosswalk codes and excludes aggregates", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con,
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM spending_long"
)$pfx
expect_true(all(prefixes %in% c("E", "F", "G", "K")))
# Classification is crosswalk membership, not prefixes (uscogdata#11):
# every row's code must classify as expenditure and never as
# intergovernmental (which lives in ig_long).
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM spending_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'expenditure'
OR c.spend_subtype = 'intergovernmental'"
)$item_code
expect_length(stray, 0L)
# Balance codes are stocks, not flows -- they must never appear in a
# spending result (uscogdata#25). Prefix filtering could not guarantee
# this (W/X/Y/Z balance codes share letters with flow codes).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
@@ -335,14 +393,29 @@ test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
expect_equal(agg_count, 0)
})
test_that("revenue_long filters to T/A/U/B/C/D prefixes and excludes aggregates", {
test_that("revenue_long carries exactly the general-revenue crosswalk codes and excludes aggregates", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con,
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM revenue_long"
)$pfx
expect_true(all(prefixes %in% c("T", "A", "U", "B", "C", "D")))
# General Revenue scope: revenue crosswalk members minus insurance_trust
# (owner ruling 2026-07-30; an explicit wider concept is uscogdata#12).
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM revenue_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'revenue'
OR c.revenue_subtype = 'insurance_trust'"
)$item_code
expect_length(stray, 0L)
# No balance stock ever appears in a revenue result (uscogdata#25).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
+2
View File
@@ -13,6 +13,8 @@ knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
# Why per-year population matters
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
+101 -74
View File
@@ -1,8 +1,8 @@
---
title: "Total spending: Direct, Total, and when each is right"
title: "Total spending: Primary, Direct, Total, and when each is right"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}
%\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
@@ -17,17 +17,35 @@ knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
is about one government or several:
1. **"What did my county spend in total, a decade ago vs today?"** — one
government, tracked over time. Either `direct` or `total` spending answers
this correctly, as long as the same concept is used for both years.
government, tracked over time. Any concept answers this correctly, as
long as the same concept is used for both years.
2. **"How do all the counties in my state compare, a decade ago vs today,
against the neighboring state?"** — several governments, summed together.
Here only `direct` gives the right answer; summing `total` across
governments double-counts money that passes between them.
Here only a non-intergovernmental concept (`primary` or `direct`) gives
the right answer; summing `total` across governments double-counts money
that passes between them.
`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`)
controls which of these a query answers. This vignette walks through both
questions with code that actually runs against the package's bundled fixture
corpus, then explains why the second question refuses `"total"` outright.
`cog_spending()`'s `expenditure_concept` argument controls which of these a
query answers, via three nested concepts defined as sets of the crosswalk's
`spend_subtype` values (never item-code first letters — the letter `Y` alone
spans revenue, expenditure, and balance codes):
- `"primary"` (the default) — the government's own service provision:
`operations` + `capital` + `assistance`.
- `"direct"` — Census's published Direct Expenditure: `primary` plus
`interest` on debt and `insurance_benefits` (e.g. pension payments).
- `"total"` — `direct` plus the `intergovernmental` leg.
This vignette walks through both questions with code that actually runs
against the package's bundled fixture corpus, then explains why the second
question refuses `"total"` outright.
Before any of the numbers below: every amount column here — `amt_nominal`,
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
dollars**. The raw Census files report **thousands of dollars** and the
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
billion. Do not scale it again.
```{r}
library(uscogdata)
@@ -61,20 +79,23 @@ al_total <- cog_spending(
al_total
```
The `intergovernmental` rows are what `"total"` adds on top of `"direct"`
(`capital` + `operations`): Alabama's own payments out to counties and
cities for highway work. Because this query only ever concerns Alabama,
including that piece is safe -- there's no other government's number it
could be double-counted against.
The `intergovernmental` rows are what `"total"` adds on top of the
non-intergovernmental subtypes (here `capital` + `operations`): Alabama's
own payments out to counties and cities for highway work. Because this
query only ever concerns Alabama, including that piece is safe -- there's
no other government's number it could be double-counted against.
`"direct"` (the default) answers the same trend question just as validly:
`"primary"` (the default) answers the same trend question just as validly
(for Highways, which maps only to operations/capital codes, `"primary"` and
`"direct"` coincide -- there is no highway-specific interest or insurance
benefit to add):
```{r}
al_direct <- cog_spending(
al_primary <- cog_spending(
"010000226085", years = c(2012, 2020), category = "Highways"
# expenditure_concept = "direct" is the default; shown here for contrast
# expenditure_concept = "primary" is the default; shown here for contrast
)
al_direct
al_primary
```
Both are internally consistent series. What breaks the comparison is
@@ -86,8 +107,8 @@ every year in the series.
# Archetype 2: a cross-government rollup
`cog_geographic_rollup()` sums spending across state/county/city layers for
a place. Its default -- and, as shown below, its *only* accepted value for
`expenditure_concept` -- is `"direct"`:
a place. Its default is `"primary"`, and (as shown below) it accepts only
the non-intergovernmental concepts, `"primary"` and `"direct"`:
```{r}
fl_rollup <- cog_geographic_rollup(
@@ -138,15 +159,15 @@ shows up **twice** in the underlying corpus:
the county is the government that actually lets the contract and pays the
paving crew.
`direct` (item codes `E`/`F`/`G`) only ever counts the second of those --
the government that actually did the spending. `total` (Direct plus the
`M`/`L` intergovernmental legs) counts the first one *as well*, which is
exactly right for describing Alabama's own budget: Alabama's `total`
genuinely includes the $10M it committed to highways, whether it built the
road itself or paid the county to. But sum `total` across Alabama **and**
the county, and that $10M is counted twice -- once as Alabama's payment out,
once as the county's spending in -- reporting $20M of highway work for $10M
actually spent.
`primary` and `direct` (the crosswalk's non-intergovernmental expenditure
subtypes) only ever count the second of those -- the government that
actually did the spending. `total` (Direct plus the intergovernmental leg)
counts the first one *as well*, which is exactly right for describing
Alabama's own budget: Alabama's `total` genuinely includes the $10M it
committed to highways, whether it built the road itself or paid the county
to. But sum `total` across Alabama **and** the county, and that $10M is
counted twice -- once as Alabama's payment out, once as the county's
spending in -- reporting $20M of highway work for $10M actually spent.
This is exactly the shape of query `cog_geographic_rollup()` exists to run
(summing across layers of government), so it refuses `"total"` rather than
@@ -161,48 +182,51 @@ share of a government's own Direct spending is:
| Government type | Intergovernmental / Direct |
|---|---|
| State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) |
| County | 3.4%-5.1% (varies by year) |
| City | 2.6%-3.1% (varies by year) |
| State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) |
| County | 3.3%-4.8% (varies by year) |
| City | 2.4%-2.9% (varies by year) |
So the Direct/Total choice matters overwhelmingly for **state** governments
-- a state's Total genuinely differs from its Direct by a meaningful margin,
while for a county or city the two are close. The state range is also far
wider than a single flat figure would suggest: legacy wide-era years (2011:
48.4%) carry proportionally more intergovernmental spending than the modern
era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly
3x larger a decade earlier than it is today. That's also why the mistake
this vignette warns about is easy to make unnoticed at the county/city level
and costly at the state level: rolling up every government in a state using
`total` instead of `direct` overstates the true figure -- measured at 7.6%
for Alabama in FY2019, and 11.6% nationally.
So the Direct/Total choice matters overwhelmingly for **state**
governments -- a state's Total genuinely differs from its Direct by more
than a third, while for a county or city the two are close. (The state
share is much larger than pre-#11 measurements suggested, because the
intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state
payments to school systems -- for most states the single largest transfer
they make.) That's also why the mistake this vignette warns about is easy
to make unnoticed at the county/city level and costly at the state level:
rolling up every government using `total` instead of `primary`/`direct`
overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally.
# Why Total = Direct + M + L, not Direct + M
# Why Total = Direct + M + L + Q, not Direct + M
It's tempting to assume `total` only needs to add `M`. But `M` and `L` are
both money the queried government itself pays **out** -- they're not two
different accounts of a receiving government's revenue. `M` is what it
pays to other **local** governments (e.g. a county paying a city for a
shared paving contract); `L` is what it pays **up** to its **state**
government (e.g. a county's contribution to a state-administered program).
A local government's Total genuinely includes both legs, because both are
its own spending, just routed to a different kind of recipient. On the
bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for
state governments (a state has no "payments to the state government" leg of
its own) but is 43%-51% the size of `M` for counties (varies by year) and
144%-189% the size of `M` for cities (varies by year; 166% pooled across
all four) -- so a `total` that omitted `L` would silently undercount Total
specifically for local governments, and for cities `L` is often the
*larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes both legs (excluding
the `L--` family-total rollup row, which would double-count its own
components).
It's tempting to assume `total` only needs to add `M`. But the
intergovernmental leg has three families, all money the queried government
itself pays **out** -- they're not different accounts of a receiving
government's revenue. `M` is what it pays to other **local** governments
(e.g. a county paying a city for a shared paving contract); `L` is what it
pays **up** to its **state** government (e.g. a county's contribution to a
state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments
to **school systems** (K-12 and higher-ed aid -- for most states the
single largest transfer they make, and the piece the pre-#11 prefix
allowlist silently dropped, finding F-017). A government's Total genuinely
includes every leg it pays, because each is its own spending, just routed
to a different kind of recipient. On the bundled fixture corpus (all 50
states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has
no "payments to the state government" leg of its own) but is 43%-51% the
size of `M` for counties (varies by year) and 144%-189% the size of `M`
for cities (varies by year; 166% pooled across all four) -- so a `total`
that omitted `L` would silently undercount Total specifically for local
governments, and for cities `L` is often the *larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes every leg
(excluding the `L--` family-total rollup row, which would double-count its
own components).
# Composition rules
- `expenditure_concept` (whose spending counts -- Direct vs Direct plus
intergovernmental) is **orthogonal** to `basis` (which vintage of the
item-code space a query resolves against -- `"harmonized"` vs `"raw"`).
- `expenditure_concept` (whose spending counts -- Primary, Direct, or
Direct plus intergovernmental) is **orthogonal** to `basis` (which
vintage of the item-code space a query resolves against --
`"harmonized"` vs `"raw"`).
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
valid, meaningful query, and so is every other pairing.
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
@@ -218,12 +242,15 @@ components).
# Summary
- Comparing one government to itself over time: `"direct"` or `"total"`
both work -- pick one and hold it fixed across every year compared.
- Comparing one government to itself over time: any concept works -- pick
one and hold it fixed across every year compared.
- Comparing or summing across governments -- counties within a state, a
state against its neighbor, cities against counties: use `"direct"`.
`cog_geographic_rollup()` and `cog_peer_compare()` enforce this by
refusing `"total"`.
- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local
governments + `L` to the state government, excluding the `L--`
family-total row).
state against its neighbor, cities against counties: use `"primary"`
(the default) or `"direct"`. `cog_geographic_rollup()` and
`cog_peer_compare()` enforce this by refusing `"total"`.
- `"primary"` = operations + capital + assistance. `"direct"` = primary +
interest on debt + insurance trust benefits (Census's published Direct
Expenditure). `"total"` = direct + intergovernmental (`M` to local
governments, `L` to the state government excluding the `L--`
family-total row, and `Q11`/`Q12`/`Q18` state payments to school
systems).