Compare commits

...
Author SHA1 Message Date
jared 4b23dbd9f4 feat: revenue_concept = c("general", "total") off the crosswalk (#12)
R-CMD-check / check (push) Successful in 3m47s
R-CMD-check / check (pull_request) Successful in 3m27s
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.

`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:

  general = own_source + federal + state + local_aid   (the default)
  total   = general + utility + liquor_store + insurance_trust

The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.

- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
  subtype; the concept narrows in R via the existing subtype_scope
  machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
  printed `expenditure_concept` unconditionally, so a cog_revenue()
  caller was told "Concept: primary" -- a spending concept their result
  has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).

Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).

Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.

Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.

Closes #12
2026-07-30 20:49:49 -04:00
jared 93300ae0c1 feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00
jared 7d798b9937 chore: regenerate fixture corpus at pipeline_commit e64a046
R-CMD-check / check (pull_request) Successful in 3m18s
R-CMD-check / check (push) Successful in 4m6s
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.

Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.

test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.

Suite: 0 failures, 2 skips (the #11 and #12 blocks).
2026-07-30 16:04:35 -04:00
jared 915a4d0678 Merge pull request 'feat: coverage argument + always-on reporting-coverage metadata (#13)' (#24) from feat/coverage-disclosure-13 into main
R-CMD-check / check (push) Successful in 3m30s
2026-07-30 12:06:53 -04:00
jared 6f98d061a9 Merge pull request 'feat: complete = TRUE fills absent cells with their meaning (#18)' (#23) from feat/complete-argument-18 into main
R-CMD-check / check (push) Successful in 4m14s
Reviewed-on: #23
2026-07-30 12:04:27 -04:00
jared d95c9032c5 feat: coverage argument + always-on reporting-coverage metadata (#13)
R-CMD-check / check (pull_request) Successful in 3m13s
R-CMD-check / check (push) Successful in 3m18s
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".

On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.

Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.

Decisions worth stating:

  - n_units_expected is the universe the CALLER named, not the national one.
    That is what makes the ratio mean something: "597 of the 608 Wisconsin
    cities you asked about". For peers it is the cohort size, counted over
    peer rows only -- including the target would inflate every count by one
    and make a cohort that has entirely stopped reporting look non-empty.

  - The coverage table is built from the REQUESTED years, not the years
    present in the result, so a year in which nothing reported still appears
    with n_units_reporting = 0. A year that vanishes silently is precisely
    the disclosure failure at issue.

  - "census" filters years BEFORE the query, and aborts when the range holds
    no census year rather than returning an empty result for a query the
    caller believes they made.

  - "consistent" exempts the peer-comparison target: it is the subject of the
    comparison, not a member of the cohort being balanced, and dropping it
    would leave nothing to compare. The summary_* quantiles are computed
    AFTER the filter so they describe the cohort actually returned.

  - is_census_year is documented as a statement about the survey CALENDAR,
    never a claim of completeness -- FY1967 is a census year in which only 97
    of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
    that tells the truth.

On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().

One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.

Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
2026-07-30 11:57:11 -04:00
jared af85a23ea7 feat: complete = TRUE fills absent cells with their meaning (#18)
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
2026-07-30 11:47:51 -04:00
jared 8db944e4a0 Merge pull request 'fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)' (#22) from fix/kodor-batch-14-15-16 into main
R-CMD-check / check (push) Successful in 3m11s
Reviewed-on: #22
2026-07-30 11:37:08 -04:00
jared 2e8383b098 fix: let the doc-content tests survive R CMD check
R-CMD-check / check (push) Successful in 3m5s
R-CMD-check / check (pull_request) Successful in 3m11s
CI failed on the previous commit. testthat::test_local() from a checkout was
green, but rcmdcheck was not: under R CMD check the suite runs against the
INSTALLED package, where README.md, vignettes/ and man/ do not exist. Both
newly-activated tests read them through test_path("..", "..", ...) and died
on `cannot open the connection`.

The defect was latent in the committed tests, not introduced here -- they
shipped skip()ped, so CI had never executed either one. Removing the skips
is what exposed it, which is the mechanism working as intended.

Guarded with skip_if_no_source_tree(), so they skip in the installed-package
context that structurally cannot satisfy them. They are NOT thereby unchecked
in CI: the workflow runs testthat::test_local() from the checkout as its own
step before rcmdcheck, and there the paths resolve and the assertions run.

Deliberately not split: test-peer-summary-scope.R's numeric pin needs only
the corpus and would survive check on its own, but it exists to protect the
sentence above it. Separating them would let the prose drift while the pin
kept passing.

Verified locally: test_local 629 pass / 0 fail / 3 skip; rcmdcheck
0 errors / 0 warnings / 0 notes.
2026-07-30 11:31:28 -04:00
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00
jared ebac39e6de Merge pull request 'fix: surface ALL-scoped series breaks in provenance (#19)' (#21) from fix/all-scoped-series-breaks-19 into main
R-CMD-check / check (push) Successful in 3m23s
2026-07-30 10:33:24 -04:00
jared 47dc08c4b0 Merge pull request 'fix: regenerate the bundled fixture against the sparsified corpus (#18)' (#20) from fix/regen-fixture-corpus-18 into main
R-CMD-check / check (push) Successful in 3m7s
2026-07-30 10:32:40 -04:00
jared 1d553a788f fix: surface ALL-scoped series breaks in provenance (#19)
R-CMD-check / check (push) Successful in 3m1s
R-CMD-check / check (pull_request) Successful in 3m1s
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:

  SB085  1977  dollar precision across the 1976/1977 boundary
  SB087  2002  imputation exclusion FY2002-2006
  SB194  2012  dense -> sparse representation change
  SB086  2017  government id scheme change

SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.

Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.

On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.

Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.

Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
2026-07-30 10:27:48 -04:00
jared c375c55da7 fix: regenerate the bundled fixture against the sparsified corpus (#18)
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m51s
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.

Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.

Three test repairs, each a real consequence of sparsification rather than
a number to bump:

  test-categories.R          "assistance" joined the spending subtype
                             vocabulary with the J-prefix codes.

  test-spending.R            The harmonization block counts rows that
                             exist. Broward's E21/F21/G21 were zero-pads
                             and are gone, so the anchor moves to FL state,
                             whose three NA-mapped rows carry $2.83B --
                             the amount accounting was previously asserted
                             only against 0 and could not have caught a
                             bug. Broward keeps a test of its own, now
                             asserting the zero-pads are absent.

  test-expenditure-concept.R Coverage-gap suggestions are presence-based.
                             AL state's only FY2011 B47 cell was an
                             explicit zero, so ig_federal_b47_wide stopped
                             being a candidate there; FL state carries a
                             real amount, so the counterpart guard is
                             exercised against a suggestion that fires.

test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.

Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
2026-07-30 10:20:09 -04:00
jared 82acda6f93 Merge pull request 'test: add failing tests for Madison walkthrough findings' (#17) from test/walkthrough-findings into main
R-CMD-check / check (push) Successful in 3m6s
2026-07-29 10:35:31 -04:00
jared 9233c3d18e test: add failing tests for Madison walkthrough findings
R-CMD-check / check (push) Successful in 3m3s
R-CMD-check / check (pull_request) Successful in 2m58s
Six skipped tests, one per issue opened from the Madison walkthrough audit
(docs/walkthroughs/FINDINGS.md in cog_explorer). Each asserts the desired
behaviour, so it fails today and goes green when the fix lands; each is
guarded by a single skip() naming its issue and finding IDs, so the suite
stays green and activating a test is a one-line deletion.

  test-expenditure-concepts.R             #11  F-012, F-017, F-018
  test-revenue-concept-insurance-trust.R  #12  F-014
  test-coverage-disclosure.R              #13  F-020, F-023
  test-peer-summary-scope.R               #14  F-021
  test-amount-units-documented.R          #15  F-004
  test-gov-search-literal-match.R         #16  F-025

helper-walkthrough-raw.R reads the corpus's long parquet partitions directly,
bypassing uscogdata's SQL views. Every expected amount comes from there rather
than from the verb under test - verifying an absence through the filter that
creates it proves nothing, which was the most common defect in the audit itself.

Verified: with the skips removed all six fail (or error) against the bundled
fixture; with them in place the full suite is 576 pass / 0 fail / 6 skip.
2026-07-29 00:14:11 -04:00
jared 1f257812b6 Merge pull request 'expenditure_concept = direct|total in cog_spending(), refused in the cross-government verbs' (#10) from feat/expenditure-concept into main
R-CMD-check / check (push) Successful in 2m37s
Reviewed-on: #10
2026-07-27 13:13:01 -04:00
jaredandClaude Opus 5 d258cef8c5 fix: gate direct-suppressed flag/note on an actually-covering recipe
R-CMD-check / check (push) Successful in 2m53s
R-CMD-check / check (pull_request) Successful in 2m53s
.detect_direct_suppressed() equated "no Direct sibling row" with "Direct
was suppressed", but the dominant real cause is a government with
genuinely no direct spending in that category (e.g. a state funding K-12
entirely through school districts) -- correct, ordinary data, not
suppression. Measured: 32 of 50 states false-flagged on a clean FY2019
category = NULL total query, and all 141 flagged rows across 50 states x
{2011, 2019} fell back to "no covering recipe found" instead of naming one
-- including AL Corrections, which names corrections_combined correctly
when category is supplied explicitly.

Both the flag and its row note are now gated on a harmonization recipe
actually covering that exact (year, canonical_govid, category) triple, via
a new .covering_recipes() helper that runs the same generic recipe join
per-row regardless of whether the caller supplied a category filter.
.notes_column() takes the precomputed note vector directly instead of
searching a category-gated suggestions list; .direct_suppressed_note() is
removed (its "no recipe found" fallback no longer applies -- if no recipe
covers a triple, it isn't suppression).

Also recomputes two total-spending.Rmd figures the prior wave never
actually reconciled with its own "measured against the fixture" caption:
State IG/Direct (flat 17.2%, now 16.7%-48.4% varying by year) and City L/M
(flat 188.3%, now 144%-189% varying by year).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 13:05:58 -04:00
jared e7d3a7a310 docs: fix stale fixture description and Direct/Total vignette figures (M1, M2, M5)
M5: README.md described the bundled fixture as a "3.6 MB two-year slice
(2019 + 2020)"; it's now a 15 MB four-year slice (2011, 2012, 2019,
2020), matching the regenerated fixture and the vignette's own
description.

M2: total-spending.Rmd cited County 1.8% / City 0.8% intergovernmental-
to-Direct and County 91.6% L/M, all roughly 2x off against the bundled
fixture. Measured directly against the fixture (all 50 states, each of
its four years): County IG/Direct 3.4%-5.1%, City IG/Direct 2.6%-3.1%,
County L/M 43%-51% (all varying by year). State 17.2%, AL 7.6%,
national 11.6%, and City L/M 188.3% were re-checked and left as-is.

M1: the "Why Total = Direct + M + L" paragraph described money a local
government *receives* and the state "redistributing as M" -- backwards.
M and L are both the *queried* government's own payments *out*: M to
other local governments, L up to its state. Rewrote the explanation;
the conclusion and non-M2-flagged figures are unchanged.
2026-07-27 12:06:37 -04:00
jared a4eb80d823 feat: cog_explain() prints the expenditure concept and direct-suppressed flag (I1)
.print_provenance() printed "Basis:" but nothing about Direct vs Total
-- the most consequential switch this branch adds to cog_spending() was
invisible in the package's designated "what am I looking at" verb. Add
a "Concept: direct|total (<note>)" line next to Basis, and surface a
cli warning when provenance$expenditure_concept_direct_suppressed is
TRUE (see the C1 fix), so the suppressed-Direct case is visible in the
human-readable explain output, not just in the structured provenance.
2026-07-27 12:06:28 -04:00
jared aba7ffbac2 test: cover C1 direct-suppressed handling, C2 corpus guard, and I2 candidate filter
Adds regression coverage for the three preceding fixes:
- "total" on a legacy aggregate-only family (AL Corrections 2011) now
  fires recipe suggestions, flags provenance$expenditure_concept_direct_
  suppressed, and names a recovering recipe in the affected row's notes,
  plus a contrast test confirming the flag stays FALSE when the Direct
  leg is present.
- expenditure_concept = "total" aborts with class
  uscogdata_ig_categories_unsupported against a corpus whose
  summary_categories carries no M/L rows (new
  with_corpus_missing_ig_categories() fixture helper), is unaffected for
  "direct" on the same corpus, and still works on a corpus that does
  carry M/L rows.
- an M/L recipe (corrections_ig_local_combined) no longer appears as a
  raw suggestion for a Direct-flavored cog_spending() call.
2026-07-27 12:06:22 -04:00
jared c1c6b5a6ba fix: never suggest an intergovernmental (M/L) recipe as a Direct coverage-gap filler (I2)
.build_suggestions()'s candidate query picks recipes by component_code
matching the requested category's summary_categories rows, with no
flow-prefix filter. Task 1's M04/M05 category rows share the
"Corrections" category with the Direct-flavored E04/E05, so
corrections_ig_local_combined (entirely M-prefixed) became a raw
top-level candidate for a plain (Direct) cog_spending() call.
Following that hint would silently return intergovernmental dollars
under provenance$expenditure_concept = "direct".

Task 6's flow-family gate in .attach_ig_counterparts() already protects
the *counterpart* lookup (deciding whether a firing suggestion gets an
ig_recipe_id attached) but never touched the candidate list itself.
Exclude any recipe with an M/L-prefixed component from candidates
unconditionally -- an M/L recipe should never be a coverage-gap filler
for either verb, which is a stronger guarantee than the counterpart
gate's flow_prefixes check.

Confirmed via the full suite: before this fix, a Direct cog_spending()
call for category = "Corrections" printed "corrections_ig_local_combined
... re-run with recipe = 'corrections_ig_local_combined'" as its own
suggestion; after, it appears only as the "intergovernmental
counterpart" annotation on corrections_combined and its capital-outlay
siblings. The pre-existing "IG Federal" mis-scoped test (revenue-side
B-prefixed recipes) is unaffected -- those aren't M/L, so they remain
valid candidates with ig_recipe_id still gated to NULL.
2026-07-27 12:06:13 -04:00
jared c7260cb20c fix: scope total's coverage-gap detection to the Direct leg; require IG category rows (C1, C2)
C1: spending_long/spending_long_harmonized filter NOT is_aggregate but
ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a
legacy aggregate-only family (e.g. Corrections pre-2012) can survive on
the IG leg while Direct is suppressed. expenditure_concept = "total"
then UNIONs an IG-only figure that reads as a plausible Total, and the
coverage-gap suggestion machinery -- fed the UNION'd result -- saw the
surviving IG row as coverage and stayed silent.

  (a) .build_suggestions() is now fed a Direct-leg-only view of the
      result (IG rows filtered out before the gap-years computation),
      so the recipe hints fire for "total" exactly as they do for
      "direct".
  (b) Any row where IG has dollars but Direct has none for the same
      (year, canonical_govid, category) is now flagged: the row's
      `notes` name the recovering recipe (drawn from the Direct-leg
      suggestions), and provenance gains an explicit
      `expenditure_concept_direct_suppressed` boolean plus an appended
      warning on `expenditure_concept_note` -- both cheap for a
      downstream consumer (cog-api passes provenance through verbatim)
      to test, rather than silently asserting Direct + IG when that
      arithmetic didn't happen.

Measured before/after on AL state government, Corrections, 2011:
"total" already correctly returns the corpus's actual IG-only figure
($31,358,000, vs. true Direct of $521,651,000 via recipe =
"corrections_combined"), but before this fix it did so with 0
suggestions and an unqualified "Total = Direct + IG" note; after, it
fires 3 recipe hints and both the row notes and provenance say plainly
that Direct is unavailable through this basis.

C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59
with no schema_version bump, so schema_version can't gate "total" --
a pre-#59 corpus can report any supported schema_version and still
have zero M/L category rows, in which case ig_annotated's LEFT JOIN
silently produces NA category/spend_subtype (0 rows for a specific
category, or one invisible NA-subtype group for category = NULL). New
.require_ig_categories() checks summary_categories directly and aborts
with class uscogdata_ig_categories_unsupported, naming PR #59 and
directing the user to a newer corpus.

Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose
synthetic summary_categories carries only one E36 row) by asserting
the new guard fires against that same connection, rather than leaving
the two silently contradictory.
2026-07-27 12:06:03 -04:00
jared 54ece11867 Merge origin/main (PR #8: corpus URL trailing-slash normalization) into feat/expenditure-concept
Local main was stale at fa40266 when this branch was created, so it was
missing 748ca4a. Merging rather than rebasing to preserve the reviewed
commit SHAs recorded in the SDD ledger.
2026-07-27 11:26:17 -04:00
jared 3bd9b1f011 docs: total-spending vignette + README section on Direct vs Total
Task 7 (final) of the expenditure_concept plan. The vignette leads with the
two archetype questions -- a single government's own trend (either concept
works, held fixed across years) vs a cross-government rollup (direct only,
with the refusal error from cog_geographic_rollup() shown and explained) --
walked through with code that runs against the bundled fixture corpus
(years 2011/2012/2019/2020, substituting for "2017 vs today"). Explains the
double-counting mechanism (a state's M44 payment to a county is the same
dollar as the county's own E44/F44), why Total = Direct + M + L rather than
Direct + M, and the composition rules (expenditure_concept is orthogonal to
basis, mutually exclusive with recipe). README gets a short pointer section
with the one-line rule.
2026-07-27 11:23:31 -04:00
jared 24b2ff7d8c fix: gate IG-counterpart matching to the direct-expenditure flow family
Review found the suffix-set match alone is unsafe: revenue-side recipes
(ig_federal_b47_wide, ig_state_c47_wide, ig_local_d47_wide, and their *_89
siblings) coincidentally share exact suffix sets with M/L expenditure
recipes despite representing a different flow direction. Reachable today via
a mis-scoped cog_spending(category = "IG Federal") call, not just
cog_revenue(). Thread flow_prefixes (same parameter .build_harmonization_block
already uses) through .build_suggestions()/.attach_ig_counterparts() and
require a firing recipe's own prefixes to be both in the calling verb's flow
family and within {E,F,G} before searching the M/L catalog.
2026-07-27 11:08:57 -04:00
jared c28712f62f feat: name the intergovernmental counterpart in firing recipe suggestions
Closes uscogdata #6 item 4. Only extends suggestions that already fire -- a
concept hint on every healthy call would be noise.
2026-07-27 10:46:22 -04:00
jared 7913b0f664 feat: record expenditure_concept in provenance and its JSON schema
Always populated, never implicit, so a downstream artifact says which concept
produced it. cog-api passes provenance through verbatim.
2026-07-27 10:27:12 -04:00
jared 887acf7e81 test: add missing cog_peer_compare coverage to expenditure_concept tests
- 'both cross-government verbs still accept the direct default' now tests both verbs
- 'the refusal message names the fix and the reason' now asserts both functions name
  themselves correctly in their error messages (cog_geographic_rollup vs cog_peer_compare)

Addresses coordinator feedback to prevent test coverage gaps and ensure the helper's
verb name argument is pinned correctly.
2026-07-27 10:19:55 -04:00
jared 81fd1a5279 feat: refuse expenditure_concept = total in the cross-government verbs
Owner ruling R1. Combining Census Total across governments counts
intergovernmental transfers twice, and these results land in Tableau where a
warning would be invisible -- so this is a hard error whose message names the
fix and the reason.
2026-07-27 10:14:16 -04:00
jared e2088458e1 fix: address Task 3 code review (bool_or, invariant tests, guards, docs)
Nine review items on the expenditure_concept = direct|total feature:

- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
  bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
  dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
  wide-era data puts a $0 leaf row in the same group as the real aggregate
  row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
  across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
  aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
  scoped to the aggregate flag rather than bare code presence (M89/L89
  continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
  pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
  retained, NULL harmonized_code coalesced, L-- excluded), rather than
  relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
  necessary (not just incidental) against a corpus whose `long` genuinely
  lacks a harmonized_code column, and rewrote the misleading "v5-only
  parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
  test, guarded .verb_spendrev() against total on a non-spending
  view_base, added a roxygen caveat against summing total across levels
  of government, and replaced an uncheckable corpus-wide SQL comment
  figure with a fixture-verifiable one.

Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
2026-07-27 10:00:07 -04:00
jared fefd4fe969 feat: expenditure_concept = direct|total in cog_spending()
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
2026-07-27 09:28:17 -04:00
jared 7ed1da9b79 fix: drop the inert K prefix from the spending flow prefixes
K matches zero rows corpus-wide (audited pipeline-side). Numerically inert;
removed so the code stops implying a prefix the data never had.
2026-07-27 09:16:41 -04:00
jared 9240a18ea3 chore: regenerate fixture corpus with the intergovernmental category rows
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).

cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.

Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
2026-07-27 09:11:10 -04:00
jared 46fed3a241 docs: Task 1 amendment — cog_categories is a third consumer of summary_categories 2026-07-27 09:09:42 -04:00
jared c9d1a05d4f chore: gitignore .superpowers/sdd working artifacts, keep plans/ tracked 2026-07-27 09:02:54 -04:00
jared fcecd62a03 docs: implementation plan for uscogdata expenditure_concept (repo 2 of 3)
Seven TDD tasks. Records the two measured facts the design rests on: aggregate
IG rows carry no harmonized_code (so the IG leg must COALESCE item_code), and
aggregate IG codes are year-disjoint from their modern leaf components (so
skipping NOT is_aggregate cannot double-count). Baseline measured at PASS 467.

Plan lives under .superpowers/ because docs/ is the gitignored pkgdown output
dir; .Rbuildignore'd so it never ships in the package tarball.
2026-07-27 09:02:42 -04:00
jared e581e7360c Merge pull request 'fix(#3): normalize the corpus URL trailing slash at resolution' (#8) from fix/3-url-trailing-slash into main
R-CMD-check / check (push) Successful in 2m27s
Reviewed-on: #8
2026-07-25 19:03:11 -04:00
jared 748ca4a56e fix(#3): normalize the corpus URL's trailing slash at resolution
R-CMD-check / check (pull_request) Successful in 2m33s
R-CMD-check / check (push) Successful in 2m42s
Investigating "gov search doesn't work" (#3) turned up two separate things.

THE REPORTED SYMPTOM IS ALREADY FIXED.
#3 reported `cog_gov_search("Orange")` dying in jsonlite with
`lexical error: invalid char in json text. <html> <head>`. That was fixed the
same day the issue was filed, by 8743472 "fix(manifest): actionable errors when
USCOGDATA_URL is unset or returns non-JSON" (issue filed 2026-05-27 11:30;
commit 2026-05-27). The issue was simply never closed. Verified now: injecting
an HTML manifest.json raises a typed `uscogdata_invalid_manifest` condition
naming the likely causes, with the raw parse error demoted to a footnote, and
`cog_gov_search("Orange")` returns 62 rows against the live corpus.

THE ROOT CAUSE OF THAT HTML WAS STILL LIVE -- and is what this commit fixes.

Every consumer builds locations by CONCATENATION:
  manifest.R:95   paste0(url, "manifest.json")
  mirror.R:48,125 paste0(url, e$path)
  views.R         the parquet glob
and mirror.R:104 documents the invariant outright ('url ends in "/"'). The
error messages tell users to set `"<url-or-local-path>/"`. But `.resolve_url()`
was a bare `.cfg("url")` passthrough -- the invariant was assumed everywhere and
enforced nowhere.

So a URL entered without the slash failed silently and misleadingly:
  HTTPS -> ".../downloadmanifest.json"; the host answers with an HTML 404 page,
           which lands in the JSON parser as EXACTLY the #3 symptom -- and the
           guard then blames "login page / 404 / wrong share" when the real
           cause was one missing character.
  local -> ".../corpusdata/long/**/*.parquet" and a DuckDB "No files found".

Reproduced both: pointing USCOGDATA_URL at the bundled fixture without a
trailing slash gave
  No files found that match ".../fixture_corpusdata/long/**/*.parquet"

Normalizing once at resolution fixes every consumer at the same time, rather
than having each call site re-derive the invariant. An empty setting passes
through untouched so manifest.R's "not configured" guard still fires instead of
the value degrading into a bare "/" filesystem root.

RED->GREEN: 4 tests added, 2 failed first (append-missing-slash, local-path
normalization); the already-correct cases (slash present, empty setting) passed
throughout and pin them against regression. Same fixture path that produced the
DuckDB error above now returns 62 rows.

Suite: FAIL 0 | WARN 0 | SKIP 0 | PASS 471 (was 463; +8 = the new tests).
2026-07-25 18:37:42 -04:00
jared fa40266d07 Merge pull request 'Regenerate fixture corpus from the Option B (single-flavor aggregate) publish tree' (#7) from fix/fixture-option-b-aggregates into main
R-CMD-check / check (push) Successful in 2m34s
Reviewed-on: #7
2026-07-23 12:21:22 -04:00
jared 3583c05852 chore: regenerate fixture corpus from the Option B (single-flavor aggregate) publish tree
R-CMD-check / check (push) Successful in 2m49s
R-CMD-check / check (pull_request) Successful in 2m38s
Source: cog_pipeline publish_cache built 2026-07-23T16:06:45Z at 4f992a0
(pipeline PR #43, issue #28 Option B ruling: legacy aggregate families now
publish only the H2-designated Direct flavor; Census Total = code + M-code).

Fixture delta, verified against the prior partition: year=2011 loses
173,394 non-designated aggregate rows (3,037,606 -> 2,864,212; -05 max
rows/gov 2 -> 1), leaf rows byte-identical; 2012/2019/2020 partitions,
metadata parquets, and docs unchanged. manifest.json resyncs sha256 /
row_count / size_bytes for the changed partition.

Full suite vs the regenerated fixture: 467 PASS / 0 FAIL / 0 WARN / 0 SKIP
— zero pin adjudications needed (reader verbs filter is_aggregate rows and
no main test pins legacy aggregate counts).
2026-07-23 12:16:05 -04:00
jared bd53230ae7 Merge branch 'feat/schema-v6-support'
R-CMD-check / check (push) Successful in 19m38s
2026-07-22 13:16:11 -04:00
jaredandClaude Opus 4.8 e813ffd3aa feat: accept corpus schema v6 (FIPS geography harmonization); v6 fixture
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.

Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
  * test-manifest.R: v6 accepted; boundary rejection moves to v7.
  * test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
    extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
    Education NEC local, SB184-186). Broward's 2011 partition zero-pads
    exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
    excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 13:16:11 -04:00
jared b0df1ec668 Merge pull request 'Phase R2: basis= harmonized/raw, recipes, signposting (schema 4/5 dual-accept)' (#5) from feat/phase-r2-harmonization into main
R-CMD-check / check (push) Successful in 2m47s
Reviewed-on: #5
2026-07-19 11:05:05 -04:00
jared 77f48047b1 fix: exercise real harmonized-view SQL in tests; unambiguous recipe provenance
R-CMD-check / check (pull_request) Successful in 3m5s
R-CMD-check / check (push) Successful in 3m1s
test-views.R's harmonized-view test previously ran a hand-rolled REPLACE
query with no WHERE clause, so a regression in any of
inst/sql/22-spending_long_harmonized.sql / 23-revenue_long_harmonized.sql's
three predicates (NOT is_aggregate, harmonized_code IS NOT NULL, the
E/F/G/K or T/A/U/B/C/D prefix filter) would go uncaught. Replaced it with a
test that reads the real SQL files off disk, substitutes {url} exactly as
.register_views() does, and executes them (plus their 10-long.sql
dependency) against a synthetic hive-partitioned parquet tree written via
DuckDB's own COPY ... TO (FORMAT PARQUET) (no arrow dependency, matching
this package's existing convention). Ten rows are crafted so each predicate
is independently falsifiable by a specific row; manually broke each
predicate in turn to confirm the test fails exactly as expected, then
restored the SQL files (see the task report for the RED-phase transcript).

Also fixes a provenance ambiguity: a recipe= query bypasses
spending_annotated(_harmonized)/revenue_annotated(_harmonized) entirely
(.run_recipe() joins `long` directly), so basis= has no effect on it, but
provenance was still reporting basis = "harmonized"/"raw" (whatever the
argument resolved to) with harmonization$applied = FALSE alongside it --
misleading, since it looks like harmonization was evaluated and found
nothing to exclude rather than "not applicable here." Recipe results now
report basis = "recipe" with an inert harmonization block carrying an
explicit note, regardless of what basis= was passed.
2026-07-18 23:55:48 -04:00
jared 4de915b557 feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.

Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).

Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
2026-07-18 23:33:02 -04:00
jared 7818cd2b1a feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.

Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.

Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
2026-07-18 23:19:17 -04:00
jared 3b725770d2 Merge pull request 'Phase R1: cog_manifest() accessor + CPI coverage pins' (#4) from feat/phase-r1-forward into main
R-CMD-check / check (push) Successful in 2m40s
Reviewed-on: #4
2026-07-18 18:07:13 -04:00
jared 4d61692f05 feat: export cog_manifest() accessor + pin CPI coverage 1967-present
R-CMD-check / check (push) Successful in 2m47s
R-CMD-check / check (pull_request) Successful in 2m57s
2026-07-18 13:40:54 -04:00
jared 70cf553828 chore: regenerate fixture corpus from cog_pipeline Phase Q4 publish (a082b26)
R-CMD-check / check (push) Successful in 2m18s
Metadata tables refreshed from the Phase Q4 corpus (schema 4, pipeline_commit
a082b26): canonical_fips_xwalk.parquet now carries the extended population
bridge (pop_confidence exact 15.7% -> 97.3%); canonical_alias.parquet reflects
the Q3 rename continuations. 2019/2020 long partitions unchanged (continuations
remap only pre-2017 predecessor rows). Suite 336 PASS / 0 FAIL / 0 WARN.
2026-07-13 19:36:10 -04:00
jared 3c55447308 Merge feat/phase-p-schema-4: corpus schema_version 4 (Phase P canonical ids)
R-CMD-check / check (push) Successful in 2m14s
2026-07-11 14:11:53 -04:00
jared 92c9a7382e test: re-baseline canonical_govid literals to 12-char namespace
Swaps every hardcoded 9-char canonical_govid literal (Broward County,
Fort Lauderdale City, Florida/Alabama state govts, Bexar/Tarrant/Wayne
counties, San Diego/Oakland/Miami/Austin cities) for its 12-char Phase P
equivalent, resolved by name+type+state against the regenerated fixture
xwalk. Also updates two gov_name search patterns that no longer match
under Phase P canonical naming ("FLORIDA STATE GOVT" -> "FLORIDA"; the
"Miami" substring test now pins type = "city" since MIAMI-DADE COUNTY's
canonical name now also contains "Miami", which would otherwise make the
match ambiguous across govs_types instead of resolving via largest-pop).
Underlying per-year population figures for Broward County and Alabama
are unchanged, so no expected data-value literals needed recomputation.
Suite: 126 test blocks / 336 expectations, 0 FAIL / 0 WARN / 0 SKIP.
2026-07-11 09:32:02 -04:00
jared 570a9408a2 feat: regenerate fixture corpus from Phase P publish tree + committed regen script
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
2026-07-11 09:31:49 -04:00
jared e635a1fc9e feat!: require corpus schema_version 4 (Phase P canonical ids)
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
2026-07-11 09:31:33 -04:00
jared 874347242b fix(manifest): actionable errors when USCOGDATA_URL is unset or returns non-JSON
R-CMD-check / check (push) Successful in 2m5s
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:

1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
   when the resolved URL is empty or still contains the
   `REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
   `Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
   remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
   it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
   Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
   existing caches with non-JSON content are silently refetched instead
   of returning a parse error to the caller.

Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
2026-05-27 11:57:31 -04:00
jared 0dd3f15ada fix(ci): use knitr::rmarkdown vignette engine + ignore gitea/CLAUDE
R-CMD-check / check (push) Successful in 1m44s
R CMD build (run by rcmdcheck before R CMD check) rebuilds vignettes
from source regardless of --no-vignettes. Our vignette declared
%\VignetteEngine{knitr::knitr} which requires the 'markdown' package
that isn't in CI's dependency tree. Switch to %\VignetteEngine{knitr::rmarkdown},
which uses the already-Suggests-listed 'rmarkdown' package and matches
the output: rmarkdown::html_vignette directive in the YAML header.

Also add ^\.gitea$ and ^CLAUDE\.md$ to .Rbuildignore so R CMD check
stops emitting the "hidden file" / "non-standard top-level file" notes.

Local rcmdcheck (mirroring CI's exact args) now reports
0 errors / 0 warnings / 0 notes.

The pdflatex notice in the build log is unrelated — R CMD build prints
"Not building PDF manual" and continues; with --no-manual it's silenced
entirely.
2026-04-29 19:25:16 -04:00
jared 919548685b polish(per-year-pop): expand popyear in cog_explain + propagate pop_range
R-CMD-check / check (push) Failing after 1m30s
Final-review followups:

1. cog_explain rendered the popyear range as raw 2-digit values
   ("popyear range: 19-20"), which a user could read as the years 19-20.
   Added .expand_popyear() helper to format as 4-digit calendar years
   (2019-2020). Pivot at 70 to handle pre-2000 vintages if the corpus
   ever extends backward.

2. cog_peer_compare provenance was missing pop_range and is_ratio,
   omitted from the spec-required reproducibility metadata.
   cog_find_peers now stamps both as tibble attributes; cog_peer_compare
   reads them through to provenance$pop_range and provenance$is_ratio.

Test coverage extended: explain test asserts the 4-digit format and
rejects the old 2-digit form; peer-compare test asserts pop_range +
is_ratio propagate end-to-end.

326 PASS / 0 FAIL.
2026-04-29 19:16:48 -04:00
jared 716cfe25e5 build: ignore vignette build artifacts (doc/, Meta/)
Auto-added by devtools::document() during the per-year-population
documentation pass.
2026-04-29 19:10:15 -04:00
jared 33c0274727 docs(news): per-year population denominators (unreleased)
Summarizes the per-capita and peer-cohort behavior changes for users
upgrading from earlier 0.1 snapshots.
2026-04-29 19:09:00 -04:00
jared c46354f049 docs(vignette): population denominators rationale + usage
Explains the four population sources, why F-33 is the default, type-4/5
coverage gap, the popyear quirk, and how to build moving-window peer
cohorts manually.
2026-04-29 19:05:38 -04:00
jared 916212c327 feat(explain): render new per-capita provenance fields
cog_explain() now prints denominator_source, popyear_range, and
pop_source_counts under the Transformations section.
2026-04-29 19:02:37 -04:00
jared a2ced368f5 feat(provenance): record per-year denominator metadata
Updates transformations\$per_capita with the new denominator_source string,
popyear_range, and pop_source_counts. .attach_per_capita stashes
popyear_range on the result; .verb_spendrev strips the helper attr after
provenance is built.
2026-04-29 19:00:16 -04:00
jared b7ebb4cd88 feat(rollup): drop unavailable-pop rows + provenance audit
cog_geographic_rollup(per_capita = TRUE) now drops rows whose government
has no observed population for that year (pop_source == 'unavailable'),
matching the spec's exclusion rule. Records included/excluded govids in
provenance$rollup.
2026-04-29 17:50:59 -04:00
jared c334be7706 test(rollup): per-year denominator + provenance expectations (failing) 2026-04-29 17:50:07 -04:00
jared cadce8d528 docs(peers): document cohort_year on cog_peer_compare return 2026-04-29 17:45:05 -04:00
jared a92450ff76 feat(peers): stamp cohort_year on cog_peer_compare results
Reads attr(peers, 'cohort_year') when the caller passed a cog_find_peers()
tibble; NA when the caller passed a bare character vector. Stamped as a
constant column on the result and recorded in provenance alongside the
cohort govids.
2026-04-29 17:44:25 -04:00
jared 54dd40a61d feat(peers): cog_find_peers uses per-year population
Adds optional 'year' argument (defaults to most recent observed year for
the target). Filters and ranks candidates by gov_population_yearly.population
at that year. Returned column renamed population_acs -> population.
Cohort year attached as attr(x, 'cohort_year').

Adds .resolve_cohort_year() helper. Updates test assertions to use
'population' column name. Regenerates man/cog_find_peers.Rd.
2026-04-29 17:26:32 -04:00
jared 807ed35cb7 test(peers): per-year cohort expectations (failing)
Three RED tests that drive Task 6's cog_find_peers() rewrite:
- defaults year to most recent observed (expects cohort_year attr + population column)
- honors explicit year= argument (expects cohort_year attr)
- errors with "no observed population" for unobserved year
2026-04-29 17:17:52 -04:00
jared 9ae46746c0 refactor(notes): simplify aggregate-fallback predicate in .notes_column
The plan-supplied predicate combined three redundant checks
(is.null + any + %in% TRUE). Element-wise behavior was correct via
scalar recycling, but the form was confusing — a code-quality reviewer
misread it as a multi-row false-positive bug. Simplify to mirror the
parts[[2]] structure: gate on column presence, then element-wise
%in% TRUE check. Equivalent semantics, fewer ways to misread.
2026-04-29 17:08:56 -04:00
jared 9238b04b69 feat(notes): concatenate notes; flag unavailable population
.notes_column now joins multiple per-row notes with '; '. Adds the
'No population denominator available for this gov type' note when
pop_source is 'unavailable'. Gracefully handles absent pop_source
(per_capita = FALSE). Two new tests: one corpus-level (census_f33
branch) and one synthetic unit test covering multi-note concatenation.
2026-04-29 17:03:05 -04:00
jared cbc867bed1 docs(per-capita): refresh roxygen for per-year denominator + pop_source 2026-04-29 16:58:27 -04:00
jared 4ea0583d3a feat(per-capita): use per-year F-33 population in spending verbs
cog_spending(per_capita = TRUE) and cog_revenue(per_capita = TRUE) now
divide each year's amount by that gov-year's population from
gov_population_yearly (drawn from long.population) instead of a single
static ACS 2018-2022 value. Adds pop_source column with values
'census_f33' or 'unavailable'.
2026-04-29 16:36:22 -04:00
jared e4a105013e test(spending): document fixture-pop origin in per-year-denominator test 2026-04-29 16:29:24 -04:00
jared 21b3d66c0e test(spending): per-year denominator expectation (failing)
Adds a RED test asserting that cog_spending(per_capita = TRUE) divides
by the per-year F-33 population (Broward 2019: 1,935,878; 2020: 1,952,778)
rather than the static ACS value (1,940,907). Uses absolute-tolerance
expect_true(abs(...) < 1) instead of expect_equal(tolerance=1) because
testthat 3 treats the tolerance argument as relative.
2026-04-29 16:20:41 -04:00
jared df3fe3731b test(views): document hardcoded fixture-pop origin in gov_population_yearly test 2026-04-29 15:37:05 -04:00
jared ed9658d267 feat(sql): add gov_population_yearly view
Exposes one row per (year, canonical_govid) drawn from long.population.
Used by per-capita denominators and peer matching.
2026-04-29 15:12:53 -04:00
jared a28fb2e19b chore(fixture): regenerate against cog_pipeline Layer 1 + Layer 2 fixes
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:

- Salt Lake County now resolves to 20 distinct cities/townships instead
  of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
  townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
  reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
  a separate gap, documented in cog_pipeline).

manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.

All 283 uscogdata tests pass against the new fixture.
2026-04-29 15:05:38 -04:00
jared 24e4449be7 docs(plan): per-year population denominator implementation plan
15-task TDD plan covering: gov_population_yearly view, .attach_per_capita
per-year join, pop_source column + multi-note concatenation, cog_find_peers
year arg, cog_peer_compare cohort_year, rollup unavailable-pop exclusion,
provenance updates, cog_explain rendering, vignette, cog_pipeline data
dictionary, and NEWS entry.

Also corrects spec to match existing rollup semantics (side-by-side, not
summed) and adds /plans to .Rbuildignore.
2026-04-29 09:40:05 -04:00
jared cfcda04e0c docs(spec): correct cog_peer_compare signature in per-year-pop spec
Existing function takes a peers argument (caller supplies cohort) rather
than building one internally. Cohort year flows through via an attr on
the peers tibble produced by cog_find_peers.
2026-04-29 09:17:24 -04:00
jared a25ba5f348 docs: spec for per-year population denominators
Design doc for switching cog_spending / cog_revenue / cog_geographic_rollup
per-capita calculations from a static ACS 2018-2022 population to per-year
F-33 population already present in long.population. Covers shifting peer
matching to a user-selectable cohort year (defaulting to most recent
observed year), type-4/5 NA policy, provenance updates, and a new vignette
enumerating denominator sources for future extensibility.

Specs live in /specs (added to .Rbuildignore) since docs/ is reserved for
pkgdown output.
2026-04-29 09:12:32 -04:00
jared e7fa51eec7 Merge pull request 'feat(search): basket mode for cog_gov_search()' (#1) from feat/cog-gov-search-basket-mode into main
R-CMD-check / check (push) Successful in 1m34s
2026-04-28 15:09:44 -04:00
96 changed files with 9500 additions and 285 deletions
+7
View File
@@ -10,3 +10,10 @@
^\.gitignore$ ^\.gitignore$
\.gitkeep$ \.gitkeep$
^vignettes$ ^vignettes$
^specs$
^plans$
^doc$
^Meta$
^\.gitea$
^CLAUDE\.md$
^\.superpowers$
+3
View File
@@ -9,3 +9,6 @@ docs/
/Meta/ /Meta/
.DS_Store .DS_Store
/.quarto/ /.quarto/
# SDD working artifacts (ledger, briefs, review packages) — plans/ stays tracked
.superpowers/sdd/
File diff suppressed because it is too large Load Diff
+2 -2
View File
@@ -31,5 +31,5 @@ Suggests:
Config/testthat/edition: 3 Config/testthat/edition: 3
VignetteBuilder: knitr VignetteBuilder: knitr
RoxygenNote: 7.3.3 RoxygenNote: 7.3.3
MinCorpusSchema: 3 MinCorpusSchema: 4
MaxCorpusSchema: 3 MaxCorpusSchema: 5
+2
View File
@@ -7,7 +7,9 @@ export(cog_explain)
export(cog_find_peers) export(cog_find_peers)
export(cog_geographic_rollup) export(cog_geographic_rollup)
export(cog_gov_search) export(cog_gov_search)
export(cog_manifest)
export(cog_mirror) export(cog_mirror)
export(cog_peer_compare) export(cog_peer_compare)
export(cog_recipes)
export(cog_revenue) export(cog_revenue)
export(cog_spending) export(cog_spending)
+191
View File
@@ -1,5 +1,196 @@
# uscogdata 0.1.0 (development) # uscogdata 0.1.0 (development)
## Multi-government aggregates now disclose their reporting coverage
* The Census of Governments is a **complete census only in years ending in 2
and 7**; every other year is a sample, and the sample varies enormously. On
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
return value said nothing about, so a statewide total resting on a fifth of
the universe looked exactly like one resting on all of it.
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
`coverage`:
| value | effect |
|---|---|
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
* **Regardless of mode**, every result now carries `provenance$coverage` with
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
section. So the default mode can no longer mislead silently.
* `is_census_year` is a statement about the **survey calendar**, never a claim
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
cities report. `n_units_reporting` is the number that tells the truth.
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
it is the subject of the comparison, not a member of the cohort — and the
`summary_*` quantiles are computed after the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against the
cohort size.
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
is `NULL`: `"census"` snaps to the most recent census year with an observed
population, so a cohort is not built from a sample year in which most of the
candidate universe is absent.
## `complete = TRUE`: absent cells, labelled with why they are absent
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
(today's behaviour). With `complete = TRUE` the requested grid is filled
from the corpus's `code_set` table and every row carries a new
`value_source` column:
| `value_source` | meaning | `amt_nominal` |
|---|---|---|
| `reported` | the corpus carries this cell | as published |
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
The `NA` is deliberate and is the whole point: filling a modern absence
with `0` would invent data, which is precisely the error the corpus's
representation contract exists to prevent.
* This restores information the reader lost when the corpus was sparsified
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
had begun returning nothing at all — and improves on what came before it,
since the pre-sparsification corpus could not distinguish a published zero
from an unreported cell either.
* The grid is scoped to each government's **own type**, so a county is never
filled with cells only a state can report.
* Needs a corpus published from 2026-07-29 onward (when `representation` and
`code_set` began shipping); aborts with class
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
listing those tables rather than on `schema_version`, which was never
bumped for the change. Not available with `recipe` or
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
`absence_means` rule; `cog_explain()` prints a "Completion" section.
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
corpus as a whole rather than about one item code. `series_break_refs` is
built by matching `fin_code` against the item codes in the result, and no
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
representation change at FY2012) and `SB086` (the government id scheme
change at FY2017).
* Provenance gains `corpus_break_refs`, selected on the break-year window
alone and disjoint from `series_break_refs` by construction, so a consumer
can tell a whole-result caveat from a break in one series. `cog_explain()`
prints them under their own "Corpus-wide caveats" heading. cog-api passes
provenance through verbatim, so the field appears there without an API
change.
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
crosses the boundary where an absent cell stops meaning "Census published
`$0`" and starts meaning "not reported", and until now nothing said so.
## Bundled fixture regenerated against the sparsified corpus
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
2026-07-29 (`pipeline_commit 83f9715`, schema v6). The wide era no longer
stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of
which none are `$0`. **Absence now means two different things** — in a
`dense_source` year (≤ FY2011) an absent cell means Census published `$0`;
in a `sparse_source` year (≥ FY2012) it means not reported. The corpus
carries that rule in two new tables the fixture now ships,
`representation.parquet` and `code_set.parquet`, alongside
`census_collection_coverage.parquet` and `lineage_events.parquet`
(all ten publish-tree metadata tables, up from six). Catalogued upstream
as series break `SB194`.
* `cog_categories()` gains an `assistance` spending subtype: the J-prefix
aid/benefit codes (`J19`, `J67`, `J68`, `J85`) are categorised now that
the upstream crosswalk covers every flow code carrying dollars.
* Two consequences worth knowing about, both visible in provenance rather
than in returned dollars. The harmonization block's `na_rows_excluded`
counts only rows that exist, so wide-era codes that were zero-padded no
longer appear there. Coverage-gap `suggestions` are presence-based for the
same reason, so a recipe whose component codes were all `$0` for a given
government-year is no longer suggested for it.
* `tests/testthat/test-fixture-vintage.R` pins these structural facts, so a
fixture left behind by a future publish fails loudly instead of letting the
suite pass against a corpus that no longer exists.
## Breaking: corpus schema_version 4 (Phase P canonical ids)
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
`MaxCorpusSchema` in `DESCRIPTION` are both `4`); older corpora built
against schema 3 are rejected by `cog_open()` with a clear version-mismatch
error. `canonical_govid` is now uniformly 12 characters across every
vintage the corpus covers (previously a mix of 9-char legacy ids and
12-char FIPS ids depending on source year) — **every hardcoded
`canonical_govid` literal from a pre-Phase-P corpus is now invalid** and
must be re-resolved via `cog_gov_search()` or the new `canonical_alias`
lookup table. `canonical_fips_xwalk` gains four columns
(`legacy_govs_id`, `census_geoid`, `id_source`; `confidence` is renamed to
`pop_confidence`) and a companion `canonical_alias` table ships in the
corpus for mapping legacy/alternate ids onto the current canonical
namespace. The bundled fixture corpus (`inst/extdata/fixture_corpus/`) has
been regenerated against the Phase P publish tree, now ships the full
`canonical_fips_xwalk` and `canonical_alias` master tables alongside the
2019-2020 long partitions, and is reproducible via
`data-raw/regenerate_fixture_corpus.R`.
## Clearer errors when `USCOGDATA_URL` is unconfigured or returns non-JSON
* `cog_open()` now aborts with the `uscogdata_url_not_configured` error
class when the resolved corpus URL still contains the placeholder
`REPLACE_WITH_SHARE_TOKEN` sentinel (or is empty). The message lists both
remediation paths (`Sys.setenv(USCOGDATA_URL = ...)` and
`options(uscogdata.url = ...)`) and points at the bundled fixture for
offline testing. Previously the package proceeded to fetch the placeholder
URL, cached the resulting HTML welcome page, and failed downstream with a
cryptic `jsonlite` lexical-error.
* `.fetch_or_cache_manifest()` now parses the HTTP response body before
persisting it. Non-JSON responses (login pages, 404 HTML) raise
`uscogdata_invalid_manifest` with the URL, Content-Type, and underlying
parse error — and never write to the on-disk cache.
* Manifest cache writes are now atomic (write to `manifest.json.tmp.<pid>`
in `cache_dir`, then `file.rename` over the target), so an interrupted
fetch cannot replace a previously-good cache.
* Existing caches with non-JSON content (poisoned by the prior code path)
are silently refetched instead of returning a parse error to the caller.
* Local `USCOGDATA_URL` paths whose `manifest.json` is not valid JSON now
surface the same `uscogdata_invalid_manifest` class with file context.
## Per-capita denominators now use per-year Census F-33 population
* `cog_spending()` and `cog_revenue()` previously divided all years' amounts
by a single ACS 2018-2022 estimate (`canonical_fips_xwalk.population_acs`),
producing biased per-capita values for time-series analysis. They now
divide by the F-33 `population` recorded on each gov-year via the new
`gov_population_yearly` view. Result tibbles gain a `pop_source` column
with values `"census_f33"` or `"unavailable"`. `notes` is updated to
concatenate multiple notes with `"; "`.
## Peer cohorts can be set to a chosen year
* `cog_find_peers()` adds a `year` argument (default: most recent year for
which the target has an observed population in `gov_population_yearly`).
The returned column previously named `population_acs` is now `population`
and reflects the cohort year's vintage. The cohort year is attached to the
returned tibble as `attr(x, "cohort_year")`.
* `cog_peer_compare()` now stamps a `cohort_year` column on its result (read
from the peers tibble's attribute) and records `cohort_year` plus
`cohort_govids` in provenance. When the caller supplies a bare character
vector instead of a `cog_find_peers()` result, `cohort_year` is `NA`.
## Rollups exclude govs missing population
* `cog_geographic_rollup(per_capita = TRUE)` drops rows whose government has
`pop_source == "unavailable"` and records the dropped govids in
`provenance$rollup$excluded_govids`. This excludes special districts
(type 4) and school districts (type 5) from per-capita rollups by design.
## New: vignette and provenance metadata
* New vignette `population-denominators` covers the four population sources,
the type-4/5 coverage gap, the popyear quirk, and how to build moving-window
peer cohorts manually.
* Provenance gains `transformations$per_capita$popyear_range` and
`pop_source_counts`. `cog_explain()` renders both.
## New features ## New features
* `cog_gov_search()` gains a **basket mode**: passing vector `name` * `cog_gov_search()` gains a **basket mode**: passing vector `name`
+87
View File
@@ -0,0 +1,87 @@
# R/basis.R
# basis= resolution (harmonized/raw, with v4/v5 dual-accept) and the
# harmonization exclusion-count block attached to provenance.
#' Resolve the requested `basis` against the active corpus's schema_version.
#'
#' On a `schema_version >= 5` corpus, the requested basis is used as-is. On
#' an older (`schema_version == 4`) corpus, which has no harmonization
#' tables: a caller who left `basis` at its default (`"harmonized"`, so
#' `explicit` is `FALSE`) silently gets `"raw"` back, with a note recorded
#' for provenance; a caller who explicitly asked for
#' `basis = "harmonized"` gets a hard abort instead of a silent downgrade.
#'
#' @param basis `"harmonized"` or `"raw"` (already resolved via `match.arg`).
#' @param explicit `TRUE` if the caller passed `basis` explicitly (as
#' opposed to relying on the default `c("harmonized", "raw")`).
#' @param manifest The active session's parsed manifest list.
#' @return List with `basis` (the resolved value) and `note` (character or
#' `NA_character_`).
#' @noRd
.resolve_basis <- function(basis, explicit, manifest) {
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
if (schema_version >= 5L) {
return(list(basis = basis, note = NA_character_))
}
if (identical(basis, "harmonized") && explicit) {
cli::cli_abort(c(
"basis = \"harmonized\" requires corpus schema_version >= 5.",
x = "Active corpus has schema_version {schema_version}.",
i = "Use basis = \"raw\" (the default on this corpus), or point USCOGDATA_URL at a schema_version >= 5 corpus."
), class = "uscogdata_basis_unsupported")
}
list(
basis = "raw",
note = sprintf(
"basis resolved to \"raw\": corpus schema_version %d < 5 (harmonization tables unavailable)",
schema_version
)
)
}
#' Count + sum item-level rows that basis="harmonized" excludes because they
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
#' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
#' the same subtype-membership classification the verb SQL uses, never
#' item-code prefixes), govids, and years. Only meaningful when the resolved
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
#' basis never excludes rows this way).
#'
#' The intergovernmental leg is deliberately outside this count even for
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
#' @noRd
.build_harmonization_block <- function(con, govid, years, resolved,
subtype_col, subtype_scope) {
if (!identical(resolved$basis, "harmonized")) {
return(list(
applied = FALSE,
na_rows_excluded = 0L,
na_amount_excluded = 0,
note = resolved$note
))
}
sql <- sprintf(
"SELECT COUNT(*) AS n, COALESCE(SUM(amt), 0) * 1000.0 AS amt
FROM long
WHERE canonical_govid IN (%s) AND year IN (%s)
AND NOT is_aggregate AND harmonized_code IS NULL
AND item_code IN (
SELECT item_code FROM summary_categories WHERE %s IN (%s)
)",
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
subtype_col, .sql_lit_chr(subtype_scope)
)
na <- DBI::dbGetQuery(con, sql)
list(
applied = TRUE,
na_rows_excluded = as.integer(na$n),
na_amount_excluded = as.numeric(na$amt),
note = resolved$note
)
}
+149
View File
@@ -0,0 +1,149 @@
# R/complete.R
#
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
# cell the corpus does not carry still appears, labelled with WHY it is
# missing.
#
# The corpus stopped storing the wide era's explicit zeros
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
#
# <= FY2011 dense_source absent => Census published $0 (census_zero)
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
#
# Before sparsification a wide-era query whose cells were all $0 came back as
# explicit $0 rows; afterwards it came back empty, with nothing to say which
# of the two meanings applied. This restores that -- and improves on it,
# because the pre-sparsification corpus could not distinguish the two either.
#
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
# That difference is the entire point: writing 0 into a modern absence would
# invent data, which is the error the representation contract exists to stop.
#' @noRd
.abort_complete_unsupported <- function(reason, alternative) {
cli::cli_abort(c(
"{.code complete = TRUE} is not supported for this query.",
x = reason,
i = alternative
), class = "uscogdata_complete_unsupported")
}
#' @noRd
.require_representation <- function(con, manifest) {
needed <- c("representation.parquet", "code_set.parquet")
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
logical(1))]
if (length(missing) == 0L) return(invisible(TRUE))
cli::cli_abort(c(
"This corpus does not publish the representation contract.",
x = "Missing: {.file {missing}}.",
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
), class = "uscogdata_representation_unavailable")
}
#' The cells a government-year COULD carry: every code in force for that
#' government's own type, mapped through `summary_categories`, restricted to
#' the calling verb's crosswalk subtype scope (the same subtype-membership
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
#' operations/capital/assistance) and (when given) its category filter.
#'
#' Scoped by `govs_type` deliberately. Filling against the union of all types
#' would invent cells that the government can never report -- a county row for
#' "state IG transfer to school districts" -- and those inventions would then
#' be indistinguishable from real census zeros.
#'
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
#' aggregate rows. Without it the grid would offer cells the verb structurally
#' never returns, so every one of them would fill as a phantom $0.
#' @noRd
.completion_grid_sql <- function(subtype_col, govid, years, category,
subtype_scope) {
category_pred <- if (is.null(category)) {
""
} else {
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
}
sprintf(
"SELECT DISTINCT
cs.year,
x.canonical_govid,
x.gov_name,
c.%1$s AS subtype_value,
c.category,
r.absence_means
FROM code_set cs
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
JOIN summary_categories c ON c.item_code = cs.item_code
JOIN representation r ON r.year = cs.year
WHERE x.canonical_govid IN (%2$s)
AND cs.year IN (%3$s)
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%1$s IN (%4$s)
%5$s",
subtype_col, .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","),
.sql_lit_chr(subtype_scope), category_pred
)
}
#' Fill `result` out to the full grid, stamping `value_source` on every row.
#'
#' Returns the completed tibble with a `.completion` attribute carrying the
#' provenance block. Reported rows are passed through untouched -- filling
#' must never alter or drop what the corpus actually published.
#' @noRd
.complete_result <- function(result, con, subtype_col, govid, years, category,
subtype_scope) {
grid <- tibble::as_tibble(DBI::dbGetQuery(
con, .completion_grid_sql(subtype_col, govid, years, category, subtype_scope)
))
result$value_source <- rep("reported", nrow(result))
if (nrow(grid) == 0L) {
attr(result, ".completion") <- list(
applied = TRUE, rows_filled = 0L, absence_means = list()
)
return(result)
}
names(grid)[names(grid) == "subtype_value"] <- subtype_col
key <- function(d) {
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
}
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
if (nrow(missing) > 0L) {
filled <- tibble::tibble(
year = as.integer(missing$year),
canonical_govid = as.character(missing$canonical_govid),
gov_name = as.character(missing$gov_name),
category = as.character(missing$category),
# census_zero is a value Census published; not_reported is unknown and
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
amt_nominal = ifelse(missing$absence_means == "census_zero",
0, NA_real_),
codes_included = NA_character_,
aggregate_fallback = NA,
value_source = as.character(missing$absence_means)
)
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
if ("notes" %in% names(result)) filled$notes <- NA_character_
result <- dplyr::bind_rows(result, filled)
result <- result[order(result$year, result$canonical_govid,
result[[subtype_col]], result$category), ,
drop = FALSE]
}
rules <- unique(grid[, c("year", "absence_means")])
attr(result, ".completion") <- list(
applied = TRUE,
rows_filled = nrow(missing),
absence_means = stats::setNames(
as.list(as.character(rules$absence_means)), as.character(rules$year)
)
)
result
}
+24 -1
View File
@@ -21,7 +21,30 @@
.uscogdata_defaults[[key]] .uscogdata_defaults[[key]]
} }
.resolve_url <- function() .cfg("url") #' Resolve the corpus URL, guaranteeing the trailing slash the package assumes.
#'
#' Every consumer builds locations by CONCATENATION -- `paste0(url,
#' "manifest.json")` in manifest.R, `paste0(url, e$path)` in mirror.R, and the
#' parquet glob in views.R -- and mirror.R:104 documents the invariant outright
#' ('url ends in "/"'). Nothing enforced it, so a URL entered without the slash
#' failed silently and misleadingly:
#'
#' HTTPS -> ".../downloadmanifest.json"; the host answers with an HTML 404
#' page, which lands in the JSON parser as the lexical error
#' reported in issue #3 -- pointing the user at "login page / wrong
#' share" when the real cause was one missing character.
#' local -> ".../corpusdata/long/**/*.parquet" and a DuckDB "No files found".
#'
#' Normalizing here fixes every consumer at once, rather than each call site
#' re-deriving the same invariant. An empty setting is passed through
#' untouched so manifest.R's "not configured" guard still fires instead of the
#' value degrading into a bare "/" filesystem root.
#' @noRd
.resolve_url <- function() {
url <- .cfg("url")
if (is.null(url) || !nzchar(url) || grepl("/$", url)) return(url)
paste0(url, "/")
}
.resolve_cache_dir <- function() { .resolve_cache_dir <- function() {
v <- .cfg("cache_dir") v <- .cfg("cache_dir")
+107
View File
@@ -0,0 +1,107 @@
# R/coverage.R
#
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
# findings F-020 and F-023).
#
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
# 7. Every other year is a sample, and the sample varies enormously: on the
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
# what the verbs have always done -- correctly -- but the return value said
# nothing about it, so a statewide total resting on 18% of the universe looked
# exactly like one resting on 98%.
#
# Owner's settled design: a `coverage` argument selecting WHICH units to
# include, plus always-on metadata saying how many there were either way. The
# principle behind it: using these verbs correctly must not require the caller
# to know the survey calendar.
# Years ending in 2 or 7 are full censuses of every government; all others are
# samples.
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
#' @noRd
.is_census_year <- function(years) {
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
}
#' @noRd
.validate_coverage <- function(coverage) {
tryCatch(
match.arg(coverage, c("all", "census", "consistent")),
error = function(e) {
cli::cli_abort(
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
class = "uscogdata_invalid_coverage", parent = e
)
}
)
}
#' Restrict `years` to census years for `coverage = "census"`.
#'
#' Aborts rather than returning an empty result when the requested range holds
#' no census year: silently handing back zero rows for a query the caller
#' believes they made is the failure mode this whole issue is about.
#' @noRd
.apply_census_years <- function(years, coverage, verb) {
if (!identical(coverage, "census")) return(as.integer(years))
keep <- as.integer(years)[.is_census_year(years)]
if (length(keep) == 0L) {
cli::cli_abort(c(
"{.code coverage = \"census\"} leaves no years to query.",
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
), class = "uscogdata_no_census_years")
}
sort(keep)
}
#' Keep only units that report in EVERY requested year (a balanced panel).
#'
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
#' filter (the peer-comparison target, which is the subject of the comparison
#' rather than a member of the cohort being balanced).
#' @noRd
.filter_consistent <- function(result, years, id_col = "canonical_govid",
keep_ids = character(0)) {
years <- unique(as.integer(years))
if (nrow(result) == 0L || length(years) <= 1L) return(result)
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
present <- vapply(ids, function(g) {
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
}, logical(1))
consistent <- c(ids[present], keep_ids)
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
drop = FALSE]
}
#' Per-year coverage metadata, always attached regardless of mode.
#'
#' Built from the REQUESTED years rather than the years present in the result,
#' so a year in which nothing reported still appears -- with
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
#' missing year fails to make.
#'
#' `n_units_reporting` describes the result the caller actually received, so
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
#' `n_units_reporting` is the number that tells the truth.
#' @noRd
.coverage_table <- function(result, years, n_expected,
id_col = "canonical_govid", rows = NULL) {
years <- sort(unique(as.integer(years)))
src <- if (is.null(rows)) result else rows
reporting <- vapply(years, function(y) {
ids <- src[[id_col]][as.integer(src$year) == y]
length(unique(ids[!is.na(ids)]))
}, integer(1))
tibble::tibble(
year = years,
n_units_reporting = as.integer(reporting),
n_units_expected = rep(as.integer(n_expected), length(years)),
is_census_year = .is_census_year(years)
)
}
+152
View File
@@ -51,6 +51,38 @@ cog_explain <- function(result, format = c("print", "list")) {
cli::cli_text("Category: (all)") cli::cli_text("Category: (all)")
} }
if (!is.null(prov$basis)) {
note <- if (!is.null(prov$basis_note) && !is.na(prov$basis_note)) {
sprintf(" (%s)", prov$basis_note)
} else {
""
}
cli::cli_text("Basis: {prov$basis}{note}")
}
# Each verb reports its OWN concept. Both fields are always present (each
# defaults to its concept's default), so printing `expenditure_concept`
# unconditionally would tell a cog_revenue() caller "Concept: primary",
# which names a spending concept their result has nothing to do with.
if (identical(prov$verb, "cog_revenue")) {
if (!is.null(prov$revenue_concept)) {
cli::cli_text("Concept: {prov$revenue_concept} revenue")
}
} else if (!is.null(prov$expenditure_concept)) {
concept_note <- if (!is.null(prov$expenditure_concept_note) &&
!is.na(prov$expenditure_concept_note)) {
sprintf(" (%s)", prov$expenditure_concept_note)
} else {
""
}
cli::cli_text("Concept: {prov$expenditure_concept}{concept_note}")
if (isTRUE(prov$expenditure_concept_direct_suppressed)) {
cli::cli_alert_warning(
"Direct leg unavailable for at least one requested (year, category) -- affected rows report intergovernmental dollars alone, not Direct + IG. See each row's notes."
)
}
}
cli::cli_h2("Codes observed") cli::cli_h2("Codes observed")
codes <- prov$codes_summed$observed codes <- prov$codes_summed$observed
if (length(codes) == 0L) { if (length(codes) == 0L) {
@@ -66,6 +98,82 @@ cog_explain <- function(result, format = c("print", "list")) {
) )
} }
h <- prov$harmonization
if (!is.null(h) && isTRUE(h$applied)) {
cli::cli_h2("Harmonization")
cli::cli_text(
"Excluded {h$na_rows_excluded} row(s) with no harmonized_code (${format(h$na_amount_excluded, big.mark = ',')})"
)
}
rc <- prov$recipe
if (!is.null(rc)) {
cli::cli_h2("Recipe")
cli::cli_text("{rc$recipe_id}: {rc$label}")
comp_lines <- vapply(rc$components, function(x) {
sprintf("%s (%s, %s-%s, weight=%s)", x$component_code, x$gov_type_scope,
x$year_min, x$year_max, x$weight)
}, character(1))
cli::cli_ul(comp_lines)
}
if (length(prov$suggestions) > 0L) {
cli::cli_h2("Suggestions")
sugg_lines <- vapply(prov$suggestions, function(s) {
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
s$available_years[1], s$available_years[2], s$hint)
}, character(1))
cli::cli_ul(sugg_lines)
}
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
cli::cli_h2("Reporting coverage")
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
cov <- prov$coverage
cli::cli_ul(sprintf(
"%d: %d of %d units reporting (%.0f%%) -- %s year",
cov$year, cov$n_units_reporting, cov$n_units_expected,
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
ifelse(cov$is_census_year, "census", "sample")
))
if (any(!cov$is_census_year)) {
cli::cli_text(
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
)
}
}
if (isTRUE(prov$completion$applied)) {
cli::cli_h2("Completion")
cli::cli_text(
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
)
rules <- prov$completion$absence_means
if (length(rules) > 0L) {
cli::cli_ul(vapply(names(rules), function(y) {
sprintf("%s: an absent cell means %s", y,
if (identical(rules[[y]], "census_zero")) {
"Census published $0 (filled as 0)"
} else {
"the government did not report (filled as NA, not 0)"
})
}, character(1)))
}
}
if (length(prov$series_break_refs) > 0L) {
cli::cli_h2("Series breaks")
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
}
# Kept in a section of its own: these qualify the whole result, so folding
# them in with the per-code breaks above would invite reading them as a
# caveat about one series.
if (length(prov$corpus_break_refs) > 0L) {
cli::cli_h2("Corpus-wide caveats")
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
}
cli::cli_h2("Transformations") cli::cli_h2("Transformations")
uc <- prov$transformations$units_conversion uc <- prov$transformations$units_conversion
if (isTRUE(uc$applied)) { if (isTRUE(uc$applied)) {
@@ -74,6 +182,16 @@ cog_explain <- function(result, format = c("print", "list")) {
pc <- prov$transformations$per_capita pc <- prov$transformations$per_capita
if (isTRUE(pc$applied)) { if (isTRUE(pc$applied)) {
cli::cli_text("Per-capita denominator: {pc$denominator_source}") cli::cli_text("Per-capita denominator: {pc$denominator_source}")
if (length(pc$popyear_range) == 2L) {
lo <- .expand_popyear(pc$popyear_range[1])
hi <- .expand_popyear(pc$popyear_range[2])
cli::cli_text(" popyear range: {lo}-{hi}")
}
if (!is.null(pc$pop_source_counts)) {
cli::cli_text(
" pop_source counts: census_f33={pc$pop_source_counts$census_f33}, unavailable={pc$pop_source_counts$unavailable}"
)
}
} }
infl <- prov$transformations$inflation infl <- prov$transformations$inflation
if (isTRUE(infl$applied)) { if (isTRUE(infl$applied)) {
@@ -98,3 +216,37 @@ cog_explain <- function(result, format = c("print", "list")) {
invisible(NULL) invisible(NULL)
} }
# One "break-story" line per referenced break_id: "SB109 (2005): <join_advice>".
# Re-queries series_breaks_pq for the detail (break_year, join_advice) that
# provenance$series_break_refs deliberately doesn't carry (the schema keeps
# that field to a plain id array). Falls back to bare ids if no session is
# available (e.g. explaining a result after cog_close()) rather than
# erroring cog_explain() over a cosmetic detail.
#' @noRd
.series_break_story_lines <- function(break_ids) {
con <- tryCatch(.ensure_session(), error = function(e) NULL)
if (is.null(con) || !DBI::dbIsValid(con)) return(break_ids)
detail <- tryCatch(
DBI::dbGetQuery(con, sprintf(
"SELECT break_id, break_year, join_advice FROM series_breaks_pq
WHERE break_id IN (%s) ORDER BY break_id",
.sql_lit_chr(break_ids)
)),
error = function(e) NULL
)
if (is.null(detail) || nrow(detail) == 0L) return(break_ids)
sprintf("%s (%s): %s", detail$break_id, detail$break_year, detail$join_advice)
}
# Expand a 2-digit Census popyear (e.g. 19) to a 4-digit calendar year (2019).
# F-33 metadata stores popyear as 2 digits; pivot at 70 to handle a future
# corpus that ever spans pre-1970 vintages, though current scope is 2000+.
#' @noRd
.expand_popyear <- function(yy) {
yy <- as.integer(yy)
if (length(yy) == 0L || is.na(yy)) return(NA_integer_)
if (yy >= 100L) return(yy) # already 4-digit
if (yy < 70L) return(2000L + yy)
1900L + yy
}
+129 -12
View File
@@ -1,5 +1,66 @@
# R/manifest.R # R/manifest.R
# Sentinel substring baked into the placeholder default URL. If we see this
# in the resolved URL, the user hasn't configured USCOGDATA_URL yet.
.PLACEHOLDER_TOKEN <- "REPLACE_WITH_SHARE_TOKEN"
#' Abort with actionable guidance when the resolved corpus URL is still the
#' placeholder shipped with the package (or any URL containing the sentinel).
#' Called from `cog_open()` before any I/O so users see a clear message
#' instead of a downstream JSON parse error.
#' @noRd
.check_url_configured <- function(url) {
if (!is.character(url) || length(url) != 1L || !nzchar(url)) {
cli::cli_abort(c(
"USCOGDATA_URL is not configured.",
i = "Set the corpus location via one of:",
"*" = "{.code Sys.setenv(USCOGDATA_URL = \"<url-or-local-path>/\")}",
"*" = "{.code options(uscogdata.url = \"<url-or-local-path>/\")}",
i = "For an offline smoke test, use the bundled fixture: {.code system.file(\"extdata/fixture_corpus\", package = \"uscogdata\")}."
), class = "uscogdata_url_not_configured")
}
if (grepl(.PLACEHOLDER_TOKEN, url, fixed = TRUE)) {
sentinel <- .PLACEHOLDER_TOKEN
cli::cli_abort(c(
"USCOGDATA_URL is not configured (placeholder URL detected).",
x = "Current value contains the sentinel {.val {sentinel}}: {.url {url}}",
i = "Set the corpus location via one of:",
"*" = "{.code Sys.setenv(USCOGDATA_URL = \"<url-or-local-path>/\")}",
"*" = "{.code options(uscogdata.url = \"<url-or-local-path>/\")}",
i = "For an offline smoke test, use the bundled fixture: {.code system.file(\"extdata/fixture_corpus\", package = \"uscogdata\")}.",
i = "For the live Civilytics corpus, request the Nextcloud share URL from the package maintainer."
), class = "uscogdata_url_not_configured")
}
invisible(url)
}
#' Try to parse a JSON file. Returns parsed object on success, NULL on
#' any parse failure (so callers can decide whether to refetch).
#' @noRd
.try_parse_manifest_file <- function(path) {
tryCatch(
jsonlite::fromJSON(path, simplifyVector = FALSE),
error = function(e) NULL
)
}
#' Abort with a clear, classified error when a manifest payload (string or
#' file) cannot be parsed as JSON. Surfaces the URL, content-type if known,
#' and the underlying parse error.
#' @noRd
.abort_invalid_manifest <- function(source, content_type = NA_character_, parse_error = NULL) {
ct <- if (is.na(content_type) || !nzchar(content_type)) "<unknown>" else content_type
pmsg <- if (is.null(parse_error)) "" else conditionMessage(parse_error)
cli::cli_abort(c(
"Corpus manifest is not valid JSON.",
x = "Source: {source}",
i = "Content-Type: {ct}",
i = "Likely causes: USCOGDATA_URL points at a login page, a 404 HTML page, or the wrong share; or the corpus has not been published yet.",
i = "Set USCOGDATA_URL to a directory (local path or HTTPS) that serves manifest.json directly.",
if (nzchar(pmsg)) c(">" = "Parse error: {pmsg}") else NULL
), class = "uscogdata_invalid_manifest")
}
#' Fetch manifest.json from URL (or read from a local fixture path), #' Fetch manifest.json from URL (or read from a local fixture path),
#' cache locally, validate TTL. #' cache locally, validate TTL.
#' @noRd #' @noRd
@@ -11,23 +72,54 @@
if (!file.exists(local_manifest)) { if (!file.exists(local_manifest)) {
cli::cli_abort("Local fixture has no manifest.json at {local_manifest}") cli::cli_abort("Local fixture has no manifest.json at {local_manifest}")
} }
return(jsonlite::fromJSON(local_manifest, simplifyVector = FALSE)) return(tryCatch(
jsonlite::fromJSON(local_manifest, simplifyVector = FALSE),
error = function(e) .abort_invalid_manifest(source = local_manifest, parse_error = e)
))
} }
cache_path <- file.path(cache_dir, "manifest.json") cache_path <- file.path(cache_dir, "manifest.json")
ttl <- as.integer(.cfg("manifest_ttl_secs")) ttl <- as.integer(.cfg("manifest_ttl_secs"))
needs_fetch <- !file.exists(cache_path) || cache_fresh <- file.exists(cache_path) &&
difftime(Sys.time(), file.info(cache_path)$mtime, units = "secs") > ttl difftime(Sys.time(), file.info(cache_path)$mtime, units = "secs") <= ttl
if (needs_fetch) { # Honor a fresh cache only if its contents still parse as JSON. A previous
resp <- httr2::request(paste0(url, "manifest.json")) |> # version of this package could write HTML directly into the cache; treat
httr2::req_error(is_error = function(r) httr2::resp_status(r) >= 400) |> # such poisoned caches as if they were missing so the next call recovers.
httr2::req_perform() if (cache_fresh) {
writeLines(httr2::resp_body_string(resp), cache_path) parsed <- .try_parse_manifest_file(cache_path)
if (!is.null(parsed)) return(parsed)
} }
jsonlite::fromJSON(cache_path, simplifyVector = FALSE) resp <- httr2::request(paste0(url, "manifest.json")) |>
httr2::req_error(is_error = function(r) httr2::resp_status(r) >= 400) |>
httr2::req_perform()
body <- httr2::resp_body_string(resp)
# Parse BEFORE persisting. If the server returned HTML / a login page /
# any non-JSON body with a 2xx status, we must not write it to the cache.
parsed <- tryCatch(
jsonlite::fromJSON(body, simplifyVector = FALSE),
error = function(e) {
ct <- tryCatch(httr2::resp_content_type(resp), error = function(e2) NA_character_)
.abort_invalid_manifest(
source = paste0(url, "manifest.json"),
content_type = ct,
parse_error = e
)
}
)
# Atomic write: tmp file alongside cache_path (same filesystem -> no EXDEV)
# then rename. Ensures a partial write or interrupted process never
# replaces a previously-good cache.
if (!dir.exists(cache_dir)) dir.create(cache_dir, recursive = TRUE)
tmp <- paste0(cache_path, ".tmp.", Sys.getpid())
on.exit(if (file.exists(tmp)) unlink(tmp), add = TRUE)
writeLines(body, tmp)
file.rename(tmp, cache_path)
parsed
} }
#' @noRd #' @noRd
@@ -36,11 +128,21 @@
} }
#' @noRd #' @noRd
.validate_schema <- function(manifest, expected_version) { #' Schema v6 (FIPS geography harmonization, 2026-07-22) is accepted alongside
if (manifest$schema_version != expected_version) { #' 4/5. v6 renamed the long table's fips_state_code/fips_county_code to
#' fips_state_asof/fips_county_asof and added cog_legacy_state/
#' cog_legacy_county (26 -> 28 cols); this package references NONE of those
#' columns, so no code change was needed. NOTE the SILENT semantic change for
#' any consumer of the raw long table: long fips_state/fips_county are now
#' PRESENT/harmonized geography (current county identity carried back to every
#' year, matching canonical_fips_xwalk) rather than as-of-year; as-of-year
#' moved to the *_asof columns. This package's own geography always came from
#' the xwalk (already present-based), so behaviour is unchanged.
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L)) {
if (!manifest$schema_version %in% supported) {
cli::cli_abort(c( cli::cli_abort(c(
"Corpus schema version mismatch.", "Corpus schema version mismatch.",
x = "Package expects schema_version = {expected_version}; corpus has {manifest$schema_version}.", x = "Package supports schema_version in {paste(supported, collapse = ', ')}; corpus has {manifest$schema_version}.",
i = "Update uscogdata (install.packages or pak::pkg_install) or re-publish corpus." i = "Update uscogdata (install.packages or pak::pkg_install) or re-publish corpus."
)) ))
} }
@@ -54,3 +156,18 @@
} }
`%||%` <- function(a, b) if (is.null(a) || (length(a) == 1 && is.na(a))) b else a `%||%` <- function(a, b) if (is.null(a) || (length(a) == 1 && is.na(a))) b else a
#' Return the parsed corpus manifest for the active session.
#'
#' Opens a session (connecting to the configured corpus) if none is active,
#' then returns the manifest exactly as parsed from `manifest.json`. Useful
#' for consumers that need the published year range (`years` block, schema
#' v5+) or the partition list without issuing a data query.
#'
#' @return Named list: `schema_version`, `built_at`, `pipeline_commit`,
#' `data_vintage`, `scope`, `years` (schema v5+), `schema`, `files`.
#' @export
cog_manifest <- function() {
.ensure_session()
.uscogdata_env$manifest
}
+210 -38
View File
@@ -2,32 +2,43 @@
#' Find peer governments by similarity criteria #' Find peer governments by similarity criteria
#' #'
#' Selects peer governments from `canonical_fips_xwalk` by combinations of #' Selects peer governments by combinations of government type, state, and
#' government type, state, and population range. Peers are ordered by #' population range at a chosen `year`. Peers are ordered by `|log(pop_ratio)|`
#' `|log(pop_ratio)|` ascending (closest to the target's population first). #' ascending (closest to the target's population first).
#' #'
#' @param target_govid Character scalar — `canonical_govid` of the target. #' @param target_govid Character scalar — `canonical_govid` of the target.
#' @param year Integer scalar. Cohort vintage. When `NULL` (default), uses the
#' most recent year for which the target has an observed population in
#' `gov_population_yearly`.
#' @param same_type If `TRUE` (default) restrict peers to the target's #' @param same_type If `TRUE` (default) restrict peers to the target's
#' `govs_type`. #' `govs_type`.
#' @param same_state If `TRUE` restrict peers to the target's `fips_state`. #' @param same_state If `TRUE` restrict peers to the target's `fips_state`.
#' Default `FALSE`. #' Default `FALSE`.
#' @param pop_range Length-2 numeric vector giving lower/upper bounds. #' @param pop_range Length-2 numeric vector giving lower/upper bounds.
#' @param is_ratio If `TRUE` (default) `pop_range` is multiplied by the #' @param is_ratio If `TRUE` (default) `pop_range` is multiplied by the
#' target's `population_acs` to produce absolute bounds. If `FALSE`, #' target's population at `year` to produce absolute bounds. If `FALSE`,
#' `pop_range` is interpreted as absolute population counts. #' `pop_range` is interpreted as absolute population counts.
#' @param pop_year Reserved for future use (selecting ACS vintage). Currently
#' the corpus has a single snapshot so this argument has no effect.
#' @param max_peers Integer cap on the number of peers returned. #' @param max_peers Integer cap on the number of peers returned.
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
#' most recent census year with an observed population, so a cohort is not
#' built from a sample year in which most of the candidate universe is
#' absent. `"consistent"` needs a year range, which cohort selection does not
#' have, so it selects like `"all"` and is carried on the result as
#' `attr(x, "coverage")` for [cog_peer_compare()].
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`, #' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
#' `population_acs`, `pop_ratio`, `rank`. #' `population`, `pop_ratio`, `rank`. The cohort year is attached as
#' `attr(x, "cohort_year")`.
#' @export #' @export
cog_find_peers <- function(target_govid, cog_find_peers <- function(target_govid,
year = NULL,
same_type = TRUE, same_type = TRUE,
same_state = FALSE, same_state = FALSE,
pop_range = c(0.7, 1.3), pop_range = c(0.7, 1.3),
is_ratio = TRUE, is_ratio = TRUE,
pop_year = NULL, max_peers = 10L,
max_peers = 10L) { coverage = c("all", "census", "consistent")) {
coverage <- .validate_coverage(coverage)
if (!is.character(target_govid) || length(target_govid) != 1L) { if (!is.character(target_govid) || length(target_govid) != 1L) {
cli::cli_abort("`target_govid` must be a length-1 character string.") cli::cli_abort("`target_govid` must be a length-1 character string.")
} }
@@ -35,65 +46,127 @@ cog_find_peers <- function(target_govid,
pop_range[1] >= pop_range[2]) { pop_range[1] >= pop_range[2]) {
cli::cli_abort("`pop_range` must be a length-2 numeric with lo < hi.") cli::cli_abort("`pop_range` must be a length-2 numeric with lo < hi.")
} }
if (!is.null(year) &&
(!(is.numeric(year) || is.integer(year)) || length(year) != 1L)) {
cli::cli_abort("`year` must be NULL or a length-1 integer.")
}
con <- .ensure_session() con <- .ensure_session()
target_sql <- sprintf( # Confirm target exists in the xwalk and pull govs_type / fips_state.
"SELECT canonical_govid, gov_name, govs_type, fips_state, population_acs meta_sql <- sprintf(
"SELECT canonical_govid, gov_name, govs_type, fips_state
FROM canonical_fips_xwalk FROM canonical_fips_xwalk
WHERE canonical_govid = %s", WHERE canonical_govid = %s",
.sql_lit_chr(target_govid) .sql_lit_chr(target_govid)
) )
target <- DBI::dbGetQuery(con, target_sql) meta <- DBI::dbGetQuery(con, meta_sql)
if (nrow(target) == 0L) { if (nrow(meta) == 0L) {
cli::cli_abort(c( cli::cli_abort(c(
"govid {target_govid} not found in corpus.", "govid {target_govid} not found in corpus.",
i = "v0.1 covers types 0-3 only (state/county/city/township); see vignette('coverage-scope')." i = "v0.1 covers types 0-3 only (state/county/city/township); see vignette('coverage-scope')."
)) ))
} }
if (is.na(target$population_acs) || target$population_acs <= 0) {
cli::cli_abort("Target {target_govid} has missing or non-positive population; cannot build pop_ratio band.") cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
pop_sql <- sprintf(
"SELECT population FROM gov_population_yearly
WHERE canonical_govid = %s AND year = %d",
.sql_lit_chr(target_govid), as.integer(cohort_year)
)
target_pop <- DBI::dbGetQuery(con, pop_sql)$population
if (length(target_pop) == 0L || is.na(target_pop) || target_pop <= 0) {
cli::cli_abort(c(
"Target {target_govid} has no observed population in {cohort_year}.",
i = "Use a year for which population is observed; see gov_population_yearly."
))
} }
if (isTRUE(is_ratio)) { if (isTRUE(is_ratio)) {
lo <- target$population_acs * pop_range[1] lo <- target_pop * pop_range[1]
hi <- target$population_acs * pop_range[2] hi <- target_pop * pop_range[2]
} else { } else {
lo <- pop_range[1]; hi <- pop_range[2] lo <- pop_range[1]; hi <- pop_range[2]
} }
preds <- c( preds <- c(
sprintf("canonical_govid != %s", .sql_lit_chr(target_govid)), sprintf("p.canonical_govid != %s", .sql_lit_chr(target_govid)),
sprintf("population_acs BETWEEN %.6f AND %.6f", lo, hi) sprintf("p.year = %d", as.integer(cohort_year)),
sprintf("p.population BETWEEN %.6f AND %.6f", lo, hi)
) )
if (isTRUE(same_type)) preds <- c(preds, sprintf("govs_type = %d", target$govs_type)) if (isTRUE(same_type)) preds <- c(preds, sprintf("x.govs_type = %d", meta$govs_type))
if (isTRUE(same_state)) preds <- c(preds, sprintf("fips_state = %s", .sql_lit_chr(target$fips_state))) if (isTRUE(same_state)) preds <- c(preds, sprintf("x.fips_state = %s", .sql_lit_chr(meta$fips_state)))
peers_sql <- sprintf( peers_sql <- sprintf(
"SELECT canonical_govid, gov_name, fips_state, population_acs, "SELECT p.canonical_govid, x.gov_name, x.fips_state, p.population,
population_acs / %.6f AS pop_ratio p.population / %.6f AS pop_ratio
FROM canonical_fips_xwalk FROM gov_population_yearly p
JOIN canonical_fips_xwalk x USING (canonical_govid)
WHERE %s WHERE %s
ORDER BY ABS(LN(CAST(population_acs AS DOUBLE) / %.6f)) ORDER BY ABS(LN(CAST(p.population AS DOUBLE) / %.6f))
LIMIT %d", LIMIT %d",
target$population_acs, target_pop,
paste(preds, collapse = " AND "), paste(preds, collapse = " AND "),
target$population_acs, target_pop,
as.integer(max_peers) as.integer(max_peers)
) )
peers <- tibble::as_tibble(DBI::dbGetQuery(con, peers_sql)) peers <- tibble::as_tibble(DBI::dbGetQuery(con, peers_sql))
if (nrow(peers) > 0L) peers$rank <- seq_len(nrow(peers)) peers$rank <- if (nrow(peers) > 0L) seq_len(nrow(peers)) else integer(0)
else peers$rank <- integer(0) attr(peers, "cohort_year") <- as.integer(cohort_year)
attr(peers, "pop_range") <- as.numeric(pop_range)
attr(peers, "is_ratio") <- isTRUE(is_ratio)
attr(peers, "coverage") <- coverage
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
peers peers
} }
# `coverage` picks the cohort vintage when the caller did not name one.
# "census" snaps to the most recent CENSUS year with an observed population,
# so a cohort is not silently built from a sample year in which most of the
# candidate universe is absent. "consistent" is a comparison-time concept --
# it needs a year RANGE, which cohort selection does not have -- so it selects
# like "all" here and is carried on the result for cog_peer_compare().
#' @noRd
.resolve_cohort_year <- function(con, target_govid, year,
coverage = "all") {
if (!is.null(year)) return(as.integer(year))
if (identical(coverage, "census")) {
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
.sql_lit_chr(target_govid)
)
y <- DBI::dbGetQuery(con, sql)$y
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
cli::cli_abort(c(
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
), class = "uscogdata_no_census_years")
}
sql <- sprintf(
"SELECT MAX(year) AS y FROM gov_population_yearly
WHERE canonical_govid = %s",
.sql_lit_chr(target_govid)
)
y <- DBI::dbGetQuery(con, sql)$y
if (length(y) == 0L || is.na(y)) {
cli::cli_abort(
"Target {target_govid} has no observed population in any year."
)
}
as.integer(y)
}
#' Compare a target government against a peer set #' Compare a target government against a peer set
#' #'
#' Pulls spending for the target plus a peer set (either a #' Pulls spending for the target plus a peer set (either a
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and #' [cog_find_peers()] result or a character vector of `canonical_govid`) and
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`, #' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot #' `summary_p75`) so the result can be faceted by `role` in a single ggplot
#' call. #' call. Those summary rows are quantiles **within each category**, not
#' quantiles of each peer's total — see the `@return` section before summing
#' them.
#' #'
#' @param target_govid Character scalar. #' @param target_govid Character scalar.
#' @param peers A tibble from [cog_find_peers()] or a character vector of #' @param peers A tibble from [cog_find_peers()] or a character vector of
@@ -103,18 +176,92 @@ cog_find_peers <- function(target_govid,
#' @param per_capita Default `TRUE` — peer compare usually normalizes by #' @param per_capita Default `TRUE` — peer compare usually normalizes by
#' population. #' population.
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`. #' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across peer sets counts
#' intergovernmental transfers twice; `"primary"` and `"direct"` combine
#' safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#'
#' The comparison target is exempt from `"consistent"` balancing -- it is the
#' subject of the comparison, not a member of the cohort -- and the
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
#' cohort actually returned. `n_units_reporting` counts peers only, against
#' the cohort size: "3 of your 15 peers reported in FY2019".
#' @return Tibble matching [cog_spending()]'s columns, plus a `role` #' @return Tibble matching [cog_spending()]'s columns, plus a `role`
#' column taking values `"target"`, `"peer"`, `"summary_p25"`, #' column taking values `"target"`, `"peer"`, `"summary_p25"`,
#' `"summary_p50"`, or `"summary_p75"`, and `target_rank` (target's rank #' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
#' among target+peers at `max(years)`, NA for other rows). Provenance #' among target+peers at `max(years)`, NA for other rows), and
#' attribute reports `verb = "cog_peer_compare"` and `peer_count`. #' `cohort_year` (the year used to build the peer cohort, read from
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
#' `cohort_year`, and `cohort_govids`.
#'
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
#' Each one is computed **within each `(year, spend_subtype,
#' category)` cell** across the peer set, so a `summary_p50` row is *the
#' median peer's value in that one category*, not *the value of the median
#' peer's total*. The median peer for Police and the median peer for Fire
#' are usually different governments, so summing `summary_*` rows across
#' categories does not give any peer's total and misstates the band it
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
#' one cohort, with a sign flip at FY2012.
#'
#' Facet by `role` **and** `category` (the documented use, and what the
#' rows are built for). For a genuine "median peer's total spending" line,
#' sum each peer's own categories first and take the quantile of those
#' per-government totals:
#'
#' ```r
#' library(dplyr)
#' cmp |>
#' filter(role %in% c("target", "peer")) |>
#' group_by(year, role, canonical_govid) |>
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
#' filter(role == "peer") |>
#' group_by(year) |>
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
#' ```
#' @export #' @export
cog_peer_compare <- function(target_govid, peers, category, years, cog_peer_compare <- function(target_govid, peers, category, years,
per_capita = TRUE, adjust_to_year = NULL) { per_capita = TRUE, adjust_to_year = NULL,
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call() call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_peer_compare")
}
if (!is.character(target_govid) || length(target_govid) != 1L) { if (!is.character(target_govid) || length(target_govid) != 1L) {
cli::cli_abort("`target_govid` must be a length-1 character string.") cli::cli_abort("`target_govid` must be a length-1 character string.")
} }
cohort_year <- if (is.data.frame(peers)) {
ay <- attr(peers, "cohort_year")
if (is.null(ay)) NA_integer_ else as.integer(ay)
} else {
NA_integer_
}
pop_range <- if (is.data.frame(peers)) attr(peers, "pop_range") else NULL
is_ratio <- if (is.data.frame(peers)) attr(peers, "is_ratio") else NULL
peer_govids <- if (is.data.frame(peers)) { peer_govids <- if (is.data.frame(peers)) {
as.character(peers$canonical_govid) as.character(peers$canonical_govid)
} else { } else {
@@ -123,24 +270,49 @@ cog_peer_compare <- function(target_govid, peers, category, years,
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)] peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
all_govids <- unique(c(target_govid, peer_govids)) all_govids <- unique(c(target_govid, peer_govids))
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) years <- .apply_census_years(years, coverage, "cog_peer_compare")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer") r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
# The target is exempt from balancing: it is the subject of the comparison,
# not a member of the cohort being balanced, and dropping it would leave a
# peer comparison with nothing to compare. Filtering happens BEFORE the
# quantiles below, so a "consistent" cohort's summary rows describe that
# cohort rather than the unbalanced one.
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years, keep_ids = target_govid)
}
value_col <- .peer_value_col(per_capita, adjust_to_year) value_col <- .peer_value_col(per_capita, adjust_to_year)
summary_rows <- .peer_summary_rows(r, value_col) summary_rows <- .peer_summary_rows(r, value_col)
out <- dplyr::bind_rows(r, summary_rows) out <- dplyr::bind_rows(r, summary_rows)
rank_val <- .peer_target_rank(r, target_govid, years, value_col) rank_val <- .peer_target_rank(r, target_govid, years, value_col)
out$target_rank <- ifelse(out$role == "target", rank_val, NA_integer_) out$target_rank <- ifelse(out$role == "target", rank_val, NA_integer_)
out$cohort_year <- cohort_year
prov <- attr(r, "provenance") %||% list() prov <- attr(r, "provenance") %||% list()
prov$verb <- "cog_peer_compare" prov$verb <- "cog_peer_compare"
prov$call <- paste(deparse(call), collapse = " ") prov$call <- paste(deparse(call), collapse = " ")
prov$peer_count <- length(peer_govids) prov$peer_count <- length(peer_govids)
prov$cohort_year <- cohort_year
prov$cohort_govids <- peer_govids
prov$pop_range <- pop_range
prov$is_ratio <- is_ratio
prov$target <- list( prov$target <- list(
canonical_govid = target_govid, canonical_govid = target_govid,
gov_name = unique(r$gov_name[r$role == "target"]) gov_name = unique(r$gov_name[r$role == "target"])
) )
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
# reported in FY2019". Including the target would inflate every count by one
# and make a cohort that has entirely stopped reporting look non-empty.
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(
out, years, length(peer_govids),
rows = r[r$role == "peer", , drop = FALSE]
)
attr(out, "provenance") <- prov attr(out, "provenance") <- prov
out out
} }
+66 -3
View File
@@ -4,7 +4,15 @@
#' @noRd #' @noRd
.build_provenance <- function(verb, call, govid, years, category, .build_provenance <- function(verb, call, govid, years, category,
per_capita, adjust_to_year, result, sql, per_capita, adjust_to_year, result, sql,
subtype_col) { subtype_col, basis = NA_character_,
basis_note = NA_character_,
expenditure_concept = "primary",
expenditure_concept_note = NA_character_,
expenditure_concept_direct_suppressed = FALSE,
revenue_concept = "general",
harmonization = NULL, recipe = NULL,
suggestions = list(),
completion = NULL) {
manifest <- .uscogdata_env$manifest manifest <- .uscogdata_env$manifest
codes <- result[["codes_included"]] codes <- result[["codes_included"]]
@@ -29,6 +37,23 @@
unique(result$gov_name) unique(result$gov_name)
} }
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
con <- .uscogdata_env$con
have_con <- !is.null(con) && DBI::dbIsValid(con)
break_refs <- if (have_con) {
.build_series_break_refs(con, codes_observed, years, schema_version)
} else {
character(0)
}
# Corpus-wide caveats travel separately: they qualify the whole result
# rather than one series, and they do not depend on codes_observed (see
# .build_corpus_break_refs()).
corpus_refs <- if (have_con) {
.build_corpus_break_refs(con, years, schema_version)
} else {
character(0)
}
list( list(
verb = verb, verb = verb,
call = paste(deparse(call), collapse = " "), call = paste(deparse(call), collapse = " "),
@@ -38,6 +63,18 @@
), ),
years = as.integer(years), years = as.integer(years),
category = category, category = category,
basis = basis,
basis_note = basis_note,
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note,
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
revenue_concept = revenue_concept,
harmonization = harmonization %||% list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
note = NA_character_
),
recipe = recipe,
suggestions = suggestions,
scope = list( scope = list(
gov_types_included = as.integer(unlist(manifest$scope$gov_types_included)), gov_types_included = as.integer(unlist(manifest$scope$gov_types_included)),
gov_types_excluded = as.integer(unlist(manifest$scope$gov_types_excluded)), gov_types_excluded = as.integer(unlist(manifest$scope$gov_types_excluded)),
@@ -61,9 +98,27 @@
per_capita = list( per_capita = list(
applied = isTRUE(per_capita), applied = isTRUE(per_capita),
denominator_source = if (isTRUE(per_capita)) { denominator_source = if (isTRUE(per_capita)) {
"ACS 2018-2022 B01003_001 (population_acs from canonical_fips_xwalk)" "Census F-33 population (per-year, from long.population)"
} else { } else {
NA_character_ NA_character_
},
popyear_range = if (isTRUE(per_capita)) {
attr(result, ".popyear_range") %||% integer(0)
} else {
integer(0)
},
pop_source_counts = if (isTRUE(per_capita)) {
ps <- result[["pop_source"]]
if (is.null(ps) || length(ps) == 0L) {
list(census_f33 = 0L, unavailable = 0L)
} else {
list(
census_f33 = sum(ps == "census_f33", na.rm = TRUE),
unavailable = sum(ps == "unavailable", na.rm = TRUE)
)
}
} else {
NULL
} }
), ),
inflation = list( inflation = list(
@@ -72,7 +127,15 @@
index = if (is.null(adjust_to_year)) NA_character_ else "CPI-U (BLS CPIAUCSL annual average, bundled)" index = if (is.null(adjust_to_year)) NA_character_ else "CPI-U (BLS CPIAUCSL annual average, bundled)"
) )
), ),
series_break_refs = character(0), series_break_refs = break_refs,
corpus_break_refs = corpus_refs,
# What `complete = TRUE` filled, and the rule it filled by. Always
# present so a consumer can read `completion$applied` without testing
# for the key -- an absent block and applied = FALSE would otherwise be
# indistinguishable from an older reader version.
completion = completion %||% list(
applied = FALSE, rows_filled = 0L, absence_means = list()
),
manifest = list( manifest = list(
schema_version = as.integer(manifest$schema_version), schema_version = as.integer(manifest$schema_version),
pipeline_commit = manifest$pipeline_commit %||% NA_character_, pipeline_commit = manifest$pipeline_commit %||% NA_character_,
+167
View File
@@ -0,0 +1,167 @@
# R/recipes.R
# Harmonization recipes: multi-code, cross-vintage series built by summing a
# fixed set of component item codes with per-component weights and
# year/gov-type scoping (see the `harmonization_recipes` view, registered
# from data/harmonization_recipes.parquet, schema_version >= 5 only).
#
# Recipes exist because some cross-vintage series can't be expressed as a
# 1:1 harmonized_code mapping (basis = "harmonized"): the wide era (pre-2012)
# publishes only a combined aggregate row for these families (e.g.
# corrections functions 04+05), while the modern era splits them into leaf
# codes. A recipe's generic join sums whichever of its component codes are
# present for a given year, so the resulting series is continuous across
# that format boundary.
#' List available harmonization recipes
#'
#' Recipes are multi-code cross-vintage series (see [cog_spending()]'s
#' `recipe` argument) catalogued in the corpus's `harmonization_recipes`
#' table. Use this to discover valid `recipe` ids.
#'
#' @param pattern Optional regex matched case-insensitively against
#' `recipe_id` or `label`.
#' @return Tibble with columns `recipe_id`, `label`, `n_components`,
#' `year_min`, `year_max` (the min/max component year coverage), sorted by
#' `recipe_id`.
#' @export
cog_recipes <- function(pattern = NULL) {
if (!is.null(pattern) &&
(!is.character(pattern) || length(pattern) != 1L)) {
cli::cli_abort("`pattern` must be a length-1 character string or NULL.")
}
con <- .ensure_session()
.require_schema_v5(con, .uscogdata_env$manifest, "cog_recipes()")
where <- if (is.null(pattern)) {
""
} else {
sprintf(
"WHERE regexp_matches(recipe_id, %1$s, 'i') OR regexp_matches(label, %1$s, 'i')",
.sql_lit_chr(pattern)
)
}
sql <- paste(
"SELECT recipe_id, any_value(label) AS label,
COUNT(*) AS n_components,
MIN(year_min) AS year_min, MAX(year_max) AS year_max
FROM harmonization_recipes",
where,
"GROUP BY recipe_id
ORDER BY recipe_id"
)
out <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
out$year_min <- as.integer(out$year_min)
out$year_max <- as.integer(out$year_max)
out$n_components <- as.integer(out$n_components)
out
}
#' Abort unless the active corpus has schema_version >= 5.
#' @noRd
.require_schema_v5 <- function(con, manifest, what) {
sv <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
if (sv < 5L) {
cli::cli_abort(c(
sprintf("%s requires corpus schema_version >= 5.", what),
x = "Active corpus has schema_version {sv}.",
i = "Point USCOGDATA_URL at a schema_version >= 5 corpus to use harmonization recipes."
), class = "uscogdata_schema_unsupported")
}
invisible(sv)
}
#' Abort with the valid id list unless `recipe_id` exists in the catalog.
#' @noRd
.validate_recipe_id <- function(con, recipe_id) {
ids <- DBI::dbGetQuery(
con, "SELECT DISTINCT recipe_id FROM harmonization_recipes"
)$recipe_id
if (!recipe_id %in% ids) {
cli::cli_abort(c(
"Unknown recipe = {.val {recipe_id}}.",
i = "Valid ids: {paste(sort(ids), collapse = ', ')}",
i = "See cog_recipes() for labels and year coverage."
), class = "uscogdata_unknown_recipe")
}
invisible(TRUE)
}
#' Fetch the component rows for one recipe (label, component codes, scope,
#' year ranges, weights) -- both for running the recipe and for the
#' `recipe` provenance block.
#' @noRd
.recipe_components <- function(con, recipe_id) {
sql <- sprintf(
"SELECT recipe_id, label, component_code, gov_type_scope,
year_min, year_max, weight, source_break_ids, notes
FROM harmonization_recipes
WHERE recipe_id = %s
ORDER BY component_code",
.sql_lit_chr(recipe_id)
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
#' Run a recipe's generic join: sum `amt * weight` across whichever
#' component codes are present for each (year, canonical_govid), scoped by
#' gov_type_scope. Deliberately does NOT filter `NOT is_aggregate`: in the
#' wide era (<= 2011) these families' component codes exist ONLY as
#' aggregate rows (leaves first appear 2012), so excluding aggregates would
#' zero out the wide-era half of every recipe. This is safe by corpus
#' construction -- wide-era rows for these codes are aggregate-only, modern
#' rows are leaf-only, and every component row is year-scoped via
#' `year_min`/`year_max` -- so there is no double-counting. (Checkpoint
#' review docs/phase_r_harmonization_review.md § 0.2.)
#' @noRd
.run_recipe <- function(con, recipe_id, govid, years) {
sql <- sprintf(
"SELECT l.year, l.canonical_govid,
COALESCE(x.gov_name, l.gov_name) AS gov_name,
SUM(l.amt * r.weight) * 1000.0 AS amt_nominal,
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code) AS codes_included
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
WHERE r.recipe_id = %1$s
AND l.canonical_govid IN (%2$s)
AND l.year IN (%3$s)
GROUP BY 1, 2, 3
ORDER BY 1, 2",
.sql_lit_chr(recipe_id), .sql_lit_chr(govid),
paste(as.integer(years), collapse = ",")
)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
attr(result, "sql_query") <- sql
result
}
#' Shape a raw .run_recipe() result into the standard cog_spending()/
#' cog_revenue() column layout: subtype = "recipe", category = the recipe's
#' label, aggregate_fallback = FALSE (recipes resolve coverage gaps by
#' construction, not by falling back to an aggregate row).
#' @noRd
.shape_recipe_result <- function(result, subtype_col, label) {
sql_query <- attr(result, "sql_query")
n <- nrow(result)
result[[subtype_col]] <- rep("recipe", n)
result$category <- rep(label, n)
result$aggregate_fallback <- rep(FALSE, n)
result <- result[, c(
"year", "canonical_govid", "gov_name", subtype_col, "category",
"amt_nominal", "codes_included", "aggregate_fallback"
), drop = FALSE]
attr(result, "sql_query") <- sql_query
result
}
#' Turn a small data.frame into a list-of-lists (one list per row), the
#' shape used for the `recipe$components` provenance block.
#' @noRd
.df_to_row_list <- function(df) {
lapply(seq_len(nrow(df)), function(i) as.list(df[i, , drop = FALSE]))
}
+42 -4
View File
@@ -8,22 +8,60 @@
#' multiplies by 1000 and records the conversion in `provenance`). #' multiplies by 1000 and records the conversion in `provenance`).
#' #'
#' @inheritParams cog_spending #' @inheritParams cog_spending
#' @param revenue_concept Which of Census's two published revenue concepts to
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
#' values -- never as item-code first letters, which cannot classify
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
#' prefix `X` does the same):
#'
#' * `"general"` (default) -- Census General Revenue: `own_source` +
#' `federal` + `state` + `local_aid`. The manual defines this concept by
#' subtraction (section 4.3: *"General revenue comprises all revenue
#' except that classified as liquor store, utility, or insurance trust
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
#' insurance trust revenue are all excluded.
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
#' `general` plus utility, liquor store, and insurance trust revenue
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
#' `X01`/`X02`/`X05`/`X08`).
#'
#' The two are related by Census's own identity, `Total Revenue = General +
#' Utility + Liquor Store + Insurance Trust`.
#'
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
#' systems moved out of the annual finance file into the separate Annual
#' Survey of Public Pensions, so a `"total"` series steps down at the
#' FY2016/FY2017 seam for reasons that are about collection scope rather
#' than revenue (series breaks `SB197`-`SB202`).
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`, #' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, #' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' `codes_included`, `aggregate_fallback`, `notes`. #' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' and `value_source` when `complete = TRUE`.
#' @export #' @export
cog_revenue <- function(govid, years, category = NULL, cog_revenue <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL) { per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
# this verb's recipe families (see R/suggestions.R).
.verb_spendrev( .verb_spendrev(
verb = "cog_revenue", verb = "cog_revenue",
view = "revenue_annotated", view_base = "revenue_annotated",
subtype_col = "revenue_subtype", subtype_col = "revenue_subtype",
flow_prefixes = c("T", "A", "U", "B", "C", "D"),
call = match.call(), call = match.call(),
govid = govid, govid = govid,
years = years, years = years,
category = category, category = category,
per_capita = per_capita, per_capita = per_capita,
adjust_to_year = adjust_to_year adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe,
revenue_concept = revenue_concept,
complete = complete
) )
} }
+76 -9
View File
@@ -8,28 +8,67 @@
#' "place portraits" that compare a city to the surrounding county and #' "place portraits" that compare a city to the surrounding county and
#' containing state on one set of axes. #' containing state on one set of axes.
#' #'
#' When `per_capita = TRUE`, rows whose government has no observed
#' population in that year (`pop_source == "unavailable"`) are dropped from
#' the result. The dropped govids are recorded in
#' `provenance$rollup$excluded_govids`. This excludes special districts
#' (gov type 4) and school districts (gov type 5) from per-capita rollups
#' by design — see `vignette('population-denominators')`.
#'
#' @param govids Named list with any non-empty subset of elements named #' @param govids Named list with any non-empty subset of elements named
#' `state`, `county`, `city`. Each element is a character vector of #' `state`, `county`, `city`. Each element is a character vector of
#' `canonical_govid` values. At least one layer required. #' `canonical_govid` values. At least one layer required.
#' @param category Single category name or character vector (passed through #' @param category Single category name or character vector (passed through
#' to [cog_spending()]). #' to [cog_spending()]).
#' @param years Integer vector of years. #' @param years Integer vector of years.
#' @param per_capita If `TRUE`, per-capita uses each layer's own population #' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
#' from `canonical_fips_xwalk.population_acs`. #' population from `gov_population_yearly`. Govs with missing population
#' are excluded from the result.
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`. #' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
#' refused here because combining Total across multiple layers of
#' government double-counts intergovernmental transfers (a state's payment
#' to a school district is the same dollar the district reports as its own
#' Direct spending); `"primary"` and `"direct"` combine safely.
#' @param coverage How to handle the Census of Governments survey cycle,
#' which is a **complete census only in years ending in 2 and 7** -- every
#' other year is a sample, and the sample varies enormously (on the bundled
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
#' and 112 in FY2019).
#'
#' * `"all"` (default) -- every unit that reported that year. Unchanged
#' behaviour, so existing code keeps working.
#' * `"census"` -- census years only. Aborts if the requested range holds
#' none, rather than silently returning nothing.
#' * `"consistent"` -- only units reporting in *every* requested year, giving
#' a balanced panel.
#'
#' Regardless of mode, `provenance$coverage` always carries per-year
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
#' statement about the **survey calendar**, never a claim of completeness:
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
#' report. `n_units_reporting` is the number that tells the truth.
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` / #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
#' `amt_per_capita_nominal` / `amt_per_capita_real`, `codes_included`, #' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
#' `aggregate_fallback`, `scope_note`, `notes`. Carries a `provenance` #' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
#' attribute with `verb = "cog_geographic_rollup"` and `layers`. #' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
#' and `rollup$included_govids` / `rollup$excluded_govids`.
#' @export #' @export
cog_geographic_rollup <- function(govids, category, years, cog_geographic_rollup <- function(govids, category, years,
per_capita = FALSE, adjust_to_year = NULL) { per_capita = FALSE, adjust_to_year = NULL,
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")) {
call <- match.call() call <- match.call()
expenditure_concept <- match.arg(expenditure_concept)
coverage <- .validate_coverage(coverage)
if (identical(expenditure_concept, "total")) {
.abort_concept_not_aggregatable("cog_geographic_rollup")
}
.validate_rollup_layers(govids) .validate_rollup_layers(govids)
# Accept character vector OR a data.frame with canonical_govid per layer,
# so cog_gov_search() output can be piped into one of the layer slots.
govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]") govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]")
if (any(lengths(govids) == 0L)) { if (any(lengths(govids) == 0L)) {
cli::cli_abort("Each layer in `govids` must be non-empty after coercion.") cli::cli_abort("Each layer in `govids` must be non-empty after coercion.")
@@ -41,16 +80,44 @@ cog_geographic_rollup <- function(govids, category, years,
layer = rep(layer_names, lengths(govids)) layer = rep(layer_names, lengths(govids))
) )
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) # coverage = "census" drops non-census years BEFORE the query rather than
# after: a sample year's rows are not wanted at all, and fetching them only
# to discard them would also let them into the coverage table.
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
expenditure_concept = expenditure_concept)
r <- dplyr::left_join(r, layer_map, by = "canonical_govid", r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
relationship = "many-to-many") relationship = "many-to-many")
r$scope_note <- .rollup_scope_note(r$layer) r$scope_note <- .rollup_scope_note(r$layer)
if (identical(coverage, "consistent")) {
r <- .filter_consistent(r, years)
}
excluded <- character(0)
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
drop <- r$pop_source == "unavailable"
excluded <- unique(r$canonical_govid[drop])
r <- r[!drop, , drop = FALSE]
}
included <- unique(r$canonical_govid)
r <- .reorder_rollup_cols(r) r <- .reorder_rollup_cols(r)
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
prov$verb <- "cog_geographic_rollup" prov$verb <- "cog_geographic_rollup"
prov$call <- paste(deparse(call), collapse = " ") prov$call <- paste(deparse(call), collapse = " ")
prov$layers <- layer_names prov$layers <- layer_names
prov$rollup <- list(
included_govids = included,
excluded_govids = excluded
)
# n_units_expected is the universe the CALLER named -- the govids passed in
# -- not the national universe. That is what makes the ratio meaningful:
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
prov$coverage_mode <- coverage
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
attr(r, "provenance") <- prov attr(r, "provenance") <- prov
r r
+23 -10
View File
@@ -6,8 +6,11 @@
#' the cross-vintage canonical-government registry. Operates in two modes: #' the cross-vintage canonical-government registry. Operates in two modes:
#' #'
#' * **Utility mode** (single `name`, the original behavior): returns all #' * **Utility mode** (single `name`, the original behavior): returns all
#' rows whose `gov_name` matches the regex case-insensitively, sorted by #' rows whose `gov_name` contains `name` as a **literal, case-insensitive
#' `population_acs` descending. Useful for exploratory lookups. #' substring**, sorted by `population_acs` descending. Useful for
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
#' government is findable by its own complete name even when that name
#' contains parentheses or a period.
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a #' * **Basket mode** (`length(name) > 1`): resolves each input row to a
#' single canonical govid and returns a tibble in input order, suitable #' single canonical govid and returns a tibble in input order, suitable
#' for piping straight into [cog_spending()] / [cog_revenue()] / #' for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -19,7 +22,8 @@
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`. #' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
#' 2. **Exact pass:** case-insensitive equality against `gov_name`. #' 2. **Exact pass:** case-insensitive equality against `gov_name`.
#' Single hit -> resolved. Multiple -> step 4. #' Single hit -> resolved. Multiple -> step 4.
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`. #' 3. **Substring fallback:** case-insensitive literal substring against
#' `gov_name` (metacharacters escaped).
#' Single hit -> resolved (`match_method = "substring"`). Zero hits -> #' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
#' `status = "no_match"`. Multiple hits -> step 4. #' `status = "no_match"`. Multiple hits -> step 4.
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the #' 4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -48,7 +52,7 @@
#' [cog_spending()], [cog_revenue()]. #' [cog_spending()], [cog_revenue()].
#' @examples #' @examples
#' \dontrun{ #' \dontrun{
#' # Utility mode — exploratory regex lookup #' # Utility mode — exploratory substring lookup
#' cog_gov_search("broward", state = "FL") #' cog_gov_search("broward", state = "FL")
#' #'
#' # Basket mode — resolve a known cohort #' # Basket mode — resolve a known cohort
@@ -98,9 +102,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
if (!is.character(name) || length(name) != 1L) { if (!is.character(name) || length(name) != 1L) {
cli::cli_abort("`name` must be a length-1 character string.") cli::cli_abort("`name` must be a length-1 character string.")
} }
# Escaped, so `name` is a literal case-insensitive substring -- the same
# treatment basket mode has always given it. Interpolating it raw made a
# government unfindable by its own name whenever that name contains a
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
# match-everything wildcard, and let malformed pattern text reach the
# engine as an error -- which cog-api surfaced as a 500, reachable by
# typing a real name one character at a time (uscogdata#16, F-025).
preds <- c(preds, preds <- c(preds,
sprintf("regexp_matches(gov_name, %s, 'i')", sprintf("regexp_matches(gov_name, %s, 'i')",
.sql_lit_chr(name))) .sql_lit_chr(.escape_regex(name))))
} }
if (!is.null(state)) { if (!is.null(state)) {
st_fips <- .coerce_state_to_fips(state) st_fips <- .coerce_state_to_fips(state)
@@ -126,17 +137,19 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
canonical_govid = character(0), gov_name = character(0), canonical_govid = character(0), gov_name = character(0),
govs_type = integer(0), type_label = character(0), govs_type = integer(0), type_label = character(0),
fips_state = character(0), fips_county = character(0), fips_state = character(0), fips_county = character(0),
fips_place = character(0), first_year = integer(0), fips_place = character(0), legacy_govs_id = character(0),
last_year = integer(0), population_acs = integer(0), first_year = integer(0), last_year = integer(0),
confidence = character(0) census_geoid = character(0), population_acs = integer(0),
pop_confidence = character(0), id_source = character(0)
) )
} }
#' @noRd #' @noRd
.escape_regex <- function(x) { .escape_regex <- function(x) {
# Backslash-escape POSIX regex metacharacters so `name` is treated as a # Backslash-escape POSIX regex metacharacters so `name` is treated as a
# literal substring in the DuckDB regexp_matches call (substring fallback # literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
# only; utility-mode intentionally preserves regex behavior). # utility mode used to interpolate raw, which was a defect rather than a
# feature -- see the call site and uscogdata#16.
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE) gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
} }
+54
View File
@@ -0,0 +1,54 @@
# R/series_breaks.R
# Populates prov$series_break_refs (schema in inst/schemas/provenance-v1.json
# defines the field; it was always present but always empty pre-Phase-R2)
# with the ids of any catalogued series break whose fin_code appears among
# the result's observed item codes and whose break_year falls inside the
# requested year span -- the "break warnings in the provenance envelope"
# spec § 5 promises downstream consumers (cog-api passes provenance through
# verbatim). schema_version >= 5 only: series_breaks_pq isn't registered on
# an older corpus.
#' @noRd
.build_series_break_refs <- function(con, codes_observed, years, schema_version) {
if (schema_version < 5L || length(codes_observed) == 0L) return(character(0))
sql <- sprintf(
"SELECT DISTINCT break_id
FROM series_breaks_pq
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
AND break_year BETWEEN %d AND %d
ORDER BY break_id",
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
)
DBI::dbGetQuery(con, sql)$break_id
}
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
#' `"ALL"` rather than an item code. They qualify the whole result, so they
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
#' before uscogdata#19. Selection is on the break_year window alone: which
#' codes a result happens to contain is irrelevant to a caveat about the
#' corpus.
#'
#' All four catalogued entries are *boundary* caveats (dollar precision
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
#' representation change at 2012, the id scheme change at 2017), so the same
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
#' path uses is the right one -- a request that never crosses the boundary
#' is not affected by it.
#'
#' Returned separately from `series_break_refs` so a consumer can tell a
#' whole-result caveat from a break in one series; the two are disjoint by
#' construction.
#' @noRd
.build_corpus_break_refs <- function(con, years, schema_version) {
if (schema_version < 5L || length(years) == 0L) return(character(0))
sql <- sprintf(
"SELECT DISTINCT break_id
FROM series_breaks_pq
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
ORDER BY break_id",
min(as.integer(years)), max(as.integer(years))
)
DBI::dbGetQuery(con, sql)$break_id
}
+2 -1
View File
@@ -5,13 +5,14 @@
#' @noRd #' @noRd
cog_open <- function(url = .resolve_url(), cog_open <- function(url = .resolve_url(),
cache_dir = .resolve_cache_dir()) { cache_dir = .resolve_cache_dir()) {
.check_url_configured(url)
if (!dir.exists(cache_dir)) dir.create(cache_dir, recursive = TRUE) if (!dir.exists(cache_dir)) dir.create(cache_dir, recursive = TRUE)
con <- DBI::dbConnect(duckdb::duckdb()) con <- DBI::dbConnect(duckdb::duckdb())
DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;") DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;")
manifest <- .fetch_or_cache_manifest(url, cache_dir) manifest <- .fetch_or_cache_manifest(url, cache_dir)
.validate_schema(manifest, expected_version = 3L) .validate_schema(manifest, supported = c(4L, 5L, 6L))
.validate_scope(manifest) .validate_scope(manifest)
.register_views(con, url, manifest) .register_views(con, url, manifest)
+673 -33
View File
@@ -1,5 +1,57 @@
# R/spending.R # R/spending.R
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
# `spend_subtype` values. Classification is crosswalk membership, never
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
# route it (finding F-018).
#
# primary = operations + capital + assistance (the default)
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
# total = direct + intergovernmental (via the ig_* views)
#
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
# than intergovernmental -- including payments to retirees, i.e. insurance
# trust benefits. Verified against Census's own published FY2020 state
# aggregates (20statetypepu.txt): `total` reproduces the published
# expenditure sum to the dollar; omitting insurance benefits understates
# California's Direct by 10.9%.
.spend_subtypes_primary <- c("operations", "capital", "assistance")
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
#' @noRd
.expenditure_concept_subtypes <- function(concept) {
switch(concept,
primary = .spend_subtypes_primary,
# "total" = the direct subtypes here PLUS the intergovernmental leg,
# which travels through the ig_* views rather than this scope (see
# .build_verb_sql()).
direct = ,
total = .spend_subtypes_direct
)
}
# The two revenue concepts (uscogdata#12), again as crosswalk subtype sets.
# Census's manual section 4.3 defines the first by SUBTRACTING from the second
# -- "General revenue comprises all revenue except that classified as liquor
# store, utility, or insurance trust revenue" -- giving the identity
#
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
#
# Verified against Census's own computed concept fields (IndFin FY2012,
# Wisconsin state): 31,410,686 + 0 + 0 + 4,469,906 = 35,880,592, exact.
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
.revenue_subtypes_total <- c(.revenue_subtypes_general, "utility",
"liquor_store", "insurance_trust")
#' @noRd
.revenue_concept_subtypes <- function(concept) {
switch(concept,
general = .revenue_subtypes_general,
total = .revenue_subtypes_total
)
}
#' Summarized spending by category #' Summarized spending by category
#' #'
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are #' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
@@ -13,75 +65,416 @@
#' @param category Character vector of category names (from #' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories. #' `summary_categories.category`), or `NULL` for all categories.
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and #' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
#' `amt_per_capita_real` when `adjust_to_year` is set) using #' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
#' `population_acs` from the canonical xwalk. #' Census F-33 population from `gov_population_yearly`. Result also gains
#' a `pop_source` column with values `"census_f33"` or `"unavailable"`
#' (the latter for gov types 4/5 and any row whose population is missing
#' in that year).
#' @param adjust_to_year Integer base year for CPI-U real-dollar conversion, #' @param adjust_to_year Integer base year for CPI-U real-dollar conversion,
#' or `NULL` for nominal only. #' or `NULL` for nominal only.
#' @param basis `"harmonized"` (default) sums item codes through the
#' cross-vintage harmonization mapping (folding series-break-affected
#' codes onto a comparable target and excluding aggregate / discontinued
#' rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
#' reproduces the pre-Phase-R2 behavior (published item codes, no
#' folding). On a corpus with `schema_version < 5` (no harmonization
#' tables), `basis` silently resolves to `"raw"` when left at its default
#' and the resolution is recorded in the provenance; explicitly passing
#' `basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
#' is set (see below).
#' @param recipe Optional harmonization recipe id (see [cog_recipes()]) for
#' multi-code cross-vintage series that a 1:1 harmonized_code mapping
#' can't express (e.g. a wide-era aggregate that only splits into leaf
#' codes in the modern era). Mutually exclusive with `category`. The
#' result's subtype column reads `"recipe"` and `category` reads the
#' recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
#' `basis` entirely (it joins `long` directly rather than going through
#' the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
#' argument is ignored and the result's provenance reports
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
#' FALSE`, pointing at the `recipe` block instead) rather than a
#' possibly-misleading `"harmonized"`/`"raw"` value.
#' @param expenditure_concept Which spending concept to return. Concepts are
#' defined as sets of the crosswalk's `spend_subtype` values -- never as
#' item-code first letters, which cannot classify correctly (prefix `Y`
#' alone spans revenue, expenditure, and balance codes):
#'
#' * `"primary"` (default) -- the government's own service provision:
#' `operations` + `capital` + `assistance` subtypes.
#' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
#' `interest` (interest on debt) and `insurance_benefits` (insurance
#' trust benefit payments, e.g. pensions -- Census manual section
#' 5.2.2.1 includes payments to retirees in Direct).
#' * `"total"` -- `direct` plus the intergovernmental leg: payments to
#' local governments (`M` codes), to the state government (`L` codes,
#' excluding the `L--` family-total rollup), and state payments to
#' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
#' `spend_subtype == "intergovernmental"`. Requires the active corpus's
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
#' older corpus rather than silently under-reporting. Mutually
#' exclusive with `recipe` (a recipe already defines its own component
#' codes).
#'
#' **Do not sum `"total"` results across levels of government** (e.g.
#' state + county + city): a state's `M12` payment to a school district is
#' the same dollar the district reports as its own direct `E12`, so
#' summing both double-counts it. This matters in particular with
#' [cog_geographic_rollup()], which sums across exactly that kind of
#' multi-layer government set.
#'
#' In the legacy wide era (<= FY2011), some functions are published ONLY
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
#' split), which the Direct leg excludes by construction but the IG leg
#' deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
#' query, any (year, category) where this leaves intergovernmental rows
#' with NO Direct counterpart is flagged: the affected rows' `notes`
#' name the harmonization recipe that recovers the missing Direct
#' component (when one exists), and
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
#' figure in those rows is the intergovernmental leg alone, not Direct +
#' IG.
#' @param complete If `TRUE`, fill the requested grid so that a cell the
#' corpus does not carry still appears, labelled with **why** it is
#' missing, and add a `value_source` column to every row:
#'
#' * `"reported"` — the corpus carries this cell.
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
#' Census published `$0`. `amt_nominal` is `0`.
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
#' government did not report, and the value is unknown. `amt_nominal` is
#' `NA`, **not** `0` — writing a zero there would invent data.
#'
#' The grid comes from the corpus's `code_set` table, scoped to each
#' government's own type, so a county is never filled with cells only a
#' state can report. Reported rows are passed through untouched.
#'
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
#' when `representation`/`code_set` began shipping; aborts with class
#' `uscogdata_representation_unavailable` otherwise. Not available with
#' `recipe` or with `expenditure_concept = "total"` (class
#' `uscogdata_complete_unsupported`) — neither draws its cells from
#' `code_set`.
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`, #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, #' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' `codes_included`, `aggregate_fallback`, `notes`. Carries a `provenance` #' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' attribute matching `inst/schemas/provenance-v1.json`. #' and `value_source` when `complete = TRUE`.
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
#' whose `completion` block reports `applied`, `rows_filled`, and the
#' per-year `absence_means` rule that was applied.
#' @export #' @export
cog_spending <- function(govid, years, category = NULL, cog_spending <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL) { per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE) {
# flow_prefixes no longer classifies rows (crosswalk subtype membership
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
# machinery to this verb's recipe families (see R/suggestions.R; the
# catalog only has E/F/G-component direct-expenditure recipes).
.verb_spendrev( .verb_spendrev(
verb = "cog_spending", verb = "cog_spending",
view = "spending_annotated", view_base = "spending_annotated",
subtype_col = "spend_subtype", subtype_col = "spend_subtype",
flow_prefixes = c("E", "F", "G"),
call = match.call(), call = match.call(),
govid = govid, govid = govid,
years = years, years = years,
category = category, category = category,
per_capita = per_capita, per_capita = per_capita,
adjust_to_year = adjust_to_year adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe,
expenditure_concept = expenditure_concept,
complete = complete
) )
} }
#' @noRd #' @noRd
.verb_spendrev <- function(verb, view, subtype_col, call, .abort_concept_not_aggregatable <- function(verb) {
cli::cli_abort(c(
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
"*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
{.code \"direct\"} for any comparison or sum that spans more than \\
one government.",
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
money it hands to other governments. The receiving government reports \\
that same dollar again as its own Direct when it actually spends it, \\
so combining Total across governments double-counts intergovernmental \\
transfers.",
"i" = "For one government's own Total, use \\
{.code cog_spending(expenditure_concept = \"total\")}."
), class = "uscogdata_concept_not_aggregatable")
}
#' @noRd
.verb_spendrev <- function(verb, view_base, subtype_col, flow_prefixes, call,
govid, years, category, govid, years, category,
per_capita, adjust_to_year) { per_capita, adjust_to_year,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
revenue_concept = c("general", "total"),
complete = FALSE) {
basis_explicit <- length(basis) == 1L
basis <- match.arg(basis, c("harmonized", "raw"))
# match.arg() itself throws a base `simpleError`, not an rlang-classed
# condition; wrap it so an invalid expenditure_concept aborts consistently
# with the rest of this package's validation (cli::cli_abort -> rlang_error).
expenditure_concept <- tryCatch(
match.arg(expenditure_concept, c("primary", "direct", "total")),
error = function(e) {
cli::cli_abort(
"`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
class = "uscogdata_invalid_expenditure_concept",
parent = e
)
}
)
revenue_concept <- tryCatch(
match.arg(revenue_concept, c("general", "total")),
error = function(e) {
cli::cli_abort(
"`revenue_concept` must be one of {.val general} or {.val total}.",
class = "uscogdata_invalid_revenue_concept",
parent = e
)
}
)
# The concept's subtype scope. Every code path below -- the verb SQL, the
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
# by crosswalk subtype membership, never by item-code prefix. The
# expenditure "total" concept's extra intergovernmental leg is the one
# exception: it travels through the ig_* views rather than this scope,
# because its legacy rows are aggregate-flagged.
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
.expenditure_concept_subtypes(expenditure_concept)
} else {
.revenue_concept_subtypes(revenue_concept)
}
govid <- .coerce_govid_input(govid, arg = "govid") govid <- .coerce_govid_input(govid, arg = "govid")
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year) .validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
cli::cli_abort(c(
"`recipe` and `expenditure_concept = \"total\"` are mutually exclusive.",
i = "A recipe defines its own component codes; pass one or the other.",
i = "For a recipe's intergovernmental counterpart, use the matching IG recipe (e.g. `corrections_ig_local_combined`)."
), class = "uscogdata_recipe_concept_conflict")
}
# .verb_spendrev() is shared with cog_revenue(), which never exposes
# expenditure_concept and always resolves it to the default -- so nothing
# on the public API can reach this today. But it's a cheap guard against a
# future call (direct or via a modified cog_revenue()) that would UNION
# the IG leg's expenditure M/L/Q rows into a revenue result, which has no
# matching IG view and no sensible meaning.
if (identical(expenditure_concept, "total") &&
!identical(view_base, "spending_annotated")) {
cli::cli_abort(
paste0(
"`expenditure_concept = \"total\"` is only supported for spending ",
"(view_base = \"spending_annotated\"); got view_base = ",
"{.val {view_base}}."
),
class = "uscogdata_expenditure_concept_unsupported"
)
}
complete <- isTRUE(complete)
if (complete && !is.null(recipe)) {
.abort_complete_unsupported(
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
)
}
if (complete && identical(expenditure_concept, "total")) {
.abort_complete_unsupported(
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
)
}
years <- as.integer(years) years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
con <- .ensure_session() con <- .ensure_session()
manifest <- .uscogdata_env$manifest
scope <- .check_govids_in_scope(govid) scope <- .check_govids_in_scope(govid)
if (complete) .require_representation(con, manifest)
sql <- .build_verb_sql(view, subtype_col, govid, years, category) resolved <- .resolve_basis(basis, basis_explicit, manifest)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
recipe_block <- NULL
category_for_prov <- category
if (!is.null(recipe)) {
.require_schema_v5(con, manifest, "recipe =")
.validate_recipe_id(con, recipe)
comps <- .recipe_components(con, recipe)
recipe_label <- comps$label[[1]]
result <- .run_recipe(con, recipe, govid, years)
sql <- attr(result, "sql_query")
result <- .shape_recipe_result(result, subtype_col, recipe_label)
recipe_block <- list(
recipe_id = recipe, label = recipe_label,
components = .df_to_row_list(comps)
)
category_for_prov <- recipe_label
} else {
view <- .select_view(view_base, resolved$basis)
ig_view <- if (identical(expenditure_concept, "total")) {
.require_ig_categories(con)
.select_ig_view(resolved$basis)
} else {
NULL
}
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
subtype_scope)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
# Fill BEFORE per_capita / inflation so the added cells get the same
# treatment as reported ones: a census_zero stays $0 per capita and in real
# dollars, and a not_reported stays NA through both rather than becoming a
# spurious 0.
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
if (complete) {
result <- .complete_result(result, con, subtype_col, govid, years,
category, subtype_scope)
completion <- attr(result, ".completion")
attr(result, ".completion") <- NULL
}
if (per_capita) result <- .attach_per_capita(result, con, govid) if (per_capita) result <- .attach_per_capita(result, con, govid)
if (!is.null(adjust_to_year)) { if (!is.null(adjust_to_year)) {
result <- .attach_real_dollars(result, adjust_to_year, per_capita) result <- .attach_real_dollars(result, adjust_to_year, per_capita)
} }
result$notes <- .notes_column(result) # A recipe result doesn't go through spending_annotated(_harmonized) /
# revenue_annotated(_harmonized) at all -- .run_recipe()'s generic join
# reads `long` directly -- so `basis` and the `harmonization` exclusion
# count (which is itself computed from `long`, independent of which view
# a non-recipe query used) would describe a code path this result never
# took. Rather than report a technically-still-computed but misleading
# basis = "harmonized"/"raw" + harmonization$applied combo, recipe
# results report basis = "recipe" and an explicit, inert harmonization
# block pointing at the `recipe` block instead. Task 12 (cog-api) passes
# provenance through verbatim, so this needs to be unambiguous rather
# than technically-defensible-but-confusing.
if (!is.null(recipe)) {
basis_for_prov <- "recipe"
basis_note_for_prov <- NA_character_
harmonization <- list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
note = "basis/harmonization not applicable to recipe results; see the recipe block instead"
)
suggestions <- list()
} else {
basis_for_prov <- resolved$basis
basis_note_for_prov <- resolved$note
harmonization <- .build_harmonization_block(
con, govid, years, resolved, subtype_col, subtype_scope
)
# C1(a): gap detection must run against the Direct leg alone. `result`
# can also carry UNION'd intergovernmental rows (expenditure_concept =
# "total"), and the wide era (<= FY2011) routinely has legacy IG dollars
# surviving (ig_long deliberately keeps aggregate rows) for a
# (year, category) whose legacy Direct dollars were suppressed (spending_
# long/spending_long_harmonized both filter NOT is_aggregate). Passing
# the UNION'd result here would let a surviving IG row count as coverage
# and silently cancel the recipe-hint suggestion that should fire.
direct_leg_result <- if (identical(expenditure_concept, "total")) {
result[!(result[[subtype_col]] %in% "intergovernmental"), , drop = FALSE]
} else {
result
}
suggestions <- .build_suggestions(con, govid, years, category,
direct_leg_result,
resolved$basis, flow_prefixes)
}
# C1(b): when expenditure_concept = "total", flag any row where the IG
# leg has dollars but the Direct leg has none for that same (year,
# canonical_govid, category) AND a harmonization recipe actually recovers
# the missing Direct dollars for that exact triple -- see
# .detect_direct_suppressed() for why bare Direct-row absence alone is NOT
# sufficient (the dominant real cause is a government that simply has no
# direct spending in that category, which is correct, ordinary data). When
# a covering recipe is found, both the row-level notes and the provenance
# say so rather than pass silently as a plausible Total.
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
.detect_direct_suppressed(con, result, subtype_col)
} else {
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
}
direct_suppressed <- direct_suppressed_info$flag
direct_suppressed_flag <- isTRUE(any(direct_suppressed))
result$notes <- .notes_column(result, direct_suppressed_info$notes)
# Determine expenditure_concept_note: only non-empty for "total", explains
# how the IG leg was assembled from legacy-era aggregates. When the Direct
# leg is suppressed for at least one requested (year, category), append an
# explicit warning rather than let the base note's "Total = Direct + IG"
# framing stand unqualified for rows where that arithmetic didn't happen.
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
if (direct_suppressed_flag) {
paste0(
base_note,
" NOTE: for at least one requested (year, category) the Direct leg ",
"has NO rows in this corpus (a legacy aggregate-only family) -- the ",
"affected result rows report the intergovernmental leg alone, not ",
"Direct + IG. See `expenditure_concept_direct_suppressed` and each ",
"affected row's `notes`."
)
} else {
base_note
}
} else {
NA_character_
}
prov <- .build_provenance( prov <- .build_provenance(
verb = verb, verb = verb,
call = call, call = call,
govid = govid, govid = govid,
years = years, years = years,
category = category, category = category_for_prov,
per_capita = per_capita, per_capita = per_capita,
adjust_to_year = adjust_to_year, adjust_to_year = adjust_to_year,
result = result, result = result,
sql = sql, sql = sql,
subtype_col = subtype_col subtype_col = subtype_col,
basis = basis_for_prov,
basis_note = basis_note_for_prov,
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note_for_prov,
expenditure_concept_direct_suppressed = direct_suppressed_flag,
revenue_concept = revenue_concept,
harmonization = harmonization,
recipe = recipe_block,
suggestions = suggestions,
completion = completion
) )
prov$scope$govids_found <- scope$found prov$scope$govids_found <- scope$found
prov$scope$govids_missing <- scope$missing prov$scope$govids_missing <- scope$missing
attr(result, "provenance") <- prov attr(result, "provenance") <- prov
attr(result, ".popyear_range") <- NULL
if (length(suggestions) > 0L) .inform_suggestions(suggestions)
result result
} }
#' @noRd #' @noRd
.validate_verb_inputs <- function(govid, years, category, .validate_verb_inputs <- function(govid, years, category,
per_capita, adjust_to_year) { per_capita, adjust_to_year, recipe = NULL) {
if (!is.character(govid) || length(govid) == 0L) { if (!is.character(govid) || length(govid) == 0L) {
cli::cli_abort("`govid` must be a non-empty character vector.") cli::cli_abort("`govid` must be a non-empty character vector.")
} }
@@ -100,6 +493,59 @@ cog_spending <- function(govid, years, category = NULL,
cli::cli_abort("`adjust_to_year` must be NULL or a length-1 integer.") cli::cli_abort("`adjust_to_year` must be NULL or a length-1 integer.")
} }
} }
if (!is.null(recipe)) {
if (!is.character(recipe) || length(recipe) != 1L) {
cli::cli_abort("`recipe` must be NULL or a length-1 character string.")
}
if (!is.null(category)) {
cli::cli_abort(c(
"`recipe` and `category` are mutually exclusive.",
i = "Pass one or the other, not both."
), class = "uscogdata_recipe_category_conflict")
}
}
invisible(TRUE)
}
#' @noRd
.select_view <- function(view_base, basis) {
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
}
#' @noRd
.select_ig_view <- function(basis) {
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
}
#' Abort unless the active corpus's `summary_categories` actually carries
#' intergovernmental (M/L) rows.
#'
#' The 66 M/L category rows arrived via cog_pipeline PR #59 with NO
#' `schema_version` bump (`DESCRIPTION` still declares `MinCorpusSchema: 4`),
#' so `schema_version` alone cannot gate `expenditure_concept = "total"` --
#' a pre-#59 corpus can validly report schema_version 4, 5, or 6 and still
#' have zero M/L rows in `summary_categories`. Against such a corpus,
#' `ig_annotated`'s LEFT JOIN to `summary_categories` silently produces NA
#' `category`/`spend_subtype` for every IG row: with a `category` filter
#' this returns 0 rows (reads as "no intergovernmental spending" rather than
#' "can't tell"), and with `category = NULL` every IG dollar collapses into
#' one NA-subtype group that is invisible to the `spend_subtype ==
#' "intergovernmental"` filter this package's own tests, roxygen, and
#' vignette all rely on. Checking the data directly (rather than
#' schema_version) is the only reliable gate.
#' @noRd
.require_ig_categories <- function(con, what = "expenditure_concept = \"total\"") {
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')"
)$n
if (identical(as.integer(n), 0L)) {
cli::cli_abort(c(
sprintf("%s requires a corpus with intergovernmental category rows.", what),
x = "The active corpus's `summary_categories` has no M/L (intergovernmental) rows.",
i = "This corpus predates the intergovernmental category rows added by cog_pipeline PR #59.",
i = "Point USCOGDATA_URL at a newer corpus that includes the M/L summary_categories rows."
), class = "uscogdata_ig_categories_unsupported")
}
invisible(TRUE) invisible(TRUE)
} }
@@ -110,7 +556,8 @@ cog_spending <- function(govid, years, category = NULL,
} }
#' @noRd #' @noRd
.build_verb_sql <- function(view, subtype_col, govid, years, category) { .build_verb_sql <- function(view, subtype_col, govid, years, category,
ig_view = NULL, subtype_scope = NULL) {
govid_lit <- .sql_lit_chr(govid) govid_lit <- .sql_lit_chr(govid)
years_lit <- paste(as.integer(years), collapse = ",") years_lit <- paste(as.integer(years), collapse = ",")
category_pred <- if (is.null(category)) { category_pred <- if (is.null(category)) {
@@ -119,6 +566,39 @@ cog_spending <- function(govid, years, category = NULL,
sprintf("AND category IN (%s)", .sql_lit_chr(category)) sprintf("AND category IN (%s)", .sql_lit_chr(category))
} }
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
# The base views carry every subtype of their flow (spending_annotated has
# all five non-IG expenditure subtypes); the concept narrows here. For
# "total", the IG leg's rows are 'intergovernmental', so that value joins
# the allowlist exactly when ig_view is present.
subtype_pred <- if (is.null(subtype_scope)) {
""
} else {
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
}
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
# never UNION: the two legs are disjoint by crosswalk subtype (the direct
# view excludes 'intergovernmental'; the IG view is only that), so
# de-duplication would be pure cost, and a silent row-drop if two
# governments ever reported identical values.
source_expr <- if (is.null(ig_view)) {
view
} else {
sprintf("(SELECT * FROM %s UNION ALL SELECT * FROM %s)", view, ig_view)
}
# bool_or(), not bool_and(): a no-op for the Direct/revenue legs (those
# views filter NOT is_aggregate, so no row in any group is ever aggregate),
# but load-bearing for the IG leg, which deliberately keeps aggregate rows
# (see inst/sql/24-ig_long.sql). The wide era is dense -- every government
# has a row for every code in a family, most of them $0 -- so a $0 leaf
# commonly lands in the same (year, gov, subtype, category) group as the
# real aggregate row. bool_and() would then read FALSE for that group even
# though its dollars came entirely from an aggregate row, silently
# suppressing the "Aggregate fallback applied" note on exactly the rows
# this feature exists to surface.
sprintf( sprintf(
"SELECT "SELECT
year, year,
@@ -128,14 +608,15 @@ cog_spending <- function(govid, years, category = NULL,
category, category,
SUM(amt) * 1000.0 AS amt_nominal, SUM(amt) * 1000.0 AS amt_nominal,
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included, string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
bool_and(is_aggregate) AS aggregate_fallback bool_or(is_aggregate) AS aggregate_fallback
FROM %2$s FROM %2$s
WHERE canonical_govid IN (%3$s) WHERE canonical_govid IN (%3$s)
AND year IN (%4$s) AND year IN (%4$s)
%5$s %5$s
%6$s
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
ORDER BY year, canonical_govid, %1$s, category", ORDER BY year, canonical_govid, %1$s, category",
subtype_col, view, govid_lit, years_lit, category_pred subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
) )
} }
@@ -143,18 +624,32 @@ cog_spending <- function(govid, years, category = NULL,
.attach_per_capita <- function(result, con, govid) { .attach_per_capita <- function(result, con, govid) {
if (nrow(result) == 0L) { if (nrow(result) == 0L) {
result$amt_per_capita_nominal <- numeric(0) result$amt_per_capita_nominal <- numeric(0)
result$pop_source <- character(0)
attr(result, ".popyear_range") <- integer(0)
return(result) return(result)
} }
years_lit <- paste(unique(as.integer(result$year)), collapse = ",")
sql <- sprintf( sql <- sprintf(
"SELECT canonical_govid, population_acs "SELECT canonical_govid, year, population, popyear
FROM canonical_fips_xwalk FROM gov_population_yearly
WHERE canonical_govid IN (%s)", WHERE canonical_govid IN (%s)
.sql_lit_chr(govid) AND year IN (%s)",
.sql_lit_chr(govid), years_lit
) )
pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
result <- dplyr::left_join(result, pops, by = "canonical_govid") result <- dplyr::left_join(result, pops,
result$amt_per_capita_nominal <- result$amt_nominal / result$population_acs by = c("canonical_govid", "year"))
result$population_acs <- NULL result$amt_per_capita_nominal <- result$amt_nominal / result$population
result$pop_source <- ifelse(is.na(result$population),
"unavailable", "census_f33")
py <- result$popyear[!is.na(result$popyear)]
attr(result, ".popyear_range") <- if (length(py) > 0L) {
as.integer(c(min(py), max(py)))
} else {
integer(0)
}
result$population <- NULL
result$popyear <- NULL
result result
} }
@@ -174,12 +669,157 @@ cog_spending <- function(govid, years, category = NULL,
result result
} }
#' Detect rows where expenditure_concept = "total" is reporting the
#' intergovernmental leg with NO Direct counterpart in the same (year,
#' canonical_govid, category) group AND a harmonization recipe actually
#' recovers the missing Direct dollars for that exact (year, canonical_govid,
#' category) triple.
#'
#' Bare Direct-row absence is deliberately NOT sufficient on its own: the
#' dominant real cause of "no Direct sibling row" is a government that simply
#' has no direct spending in that category (e.g. a state that funds K-12
#' entirely through school districts), which is correct, ordinary data, not
#' suppression. Genuine suppression -- a legacy aggregate-only family whose
#' Direct-leg basis query excludes it by construction (spending_long/
#' spending_long_harmonized both filter NOT is_aggregate) -- always has a
#' covering harmonization recipe, because that is exactly what the recipe
#' catalog exists to recover (see R/suggestions.R and `cog_recipes()`). So
#' checking "does a recipe actually cover this triple" cleanly separates the
#' two cases instead of conflating them.
#'
#' Returns `list(flag, notes)`, both the same length as `result`: `flag` is
#' `TRUE` only for the `spend_subtype == "intergovernmental"` row(s) in a
#' suppressed group, and `notes` names the recovering recipe(s) for those
#' rows (`NA` everywhere else).
#' @noRd #' @noRd
.notes_column <- function(result) { .detect_direct_suppressed <- function(con, result, subtype_col) {
if (nrow(result) == 0L) return(character(0)) n <- nrow(result)
ifelse( empty_notes <- rep(NA_character_, n)
isTRUE(result$aggregate_fallback) | result$aggregate_fallback %in% TRUE, if (n == 0L) return(list(flag = logical(0), notes = character(0)))
"Aggregate fallback applied; see cog_explain()", is_ig <- result[[subtype_col]] %in% "intergovernmental"
"" if (!any(is_ig)) return(list(flag = rep(FALSE, n), notes = empty_notes))
)
key <- paste(result$year, result$canonical_govid, result$category, sep = "\r")
has_direct <- key %in% unique(key[!is_ig])
candidate <- is_ig & !has_direct
flag <- rep(FALSE, n)
notes <- empty_notes
if (!any(candidate)) return(list(flag = flag, notes = notes))
idx <- which(candidate)
rows <- unique(result[idx, c("year", "canonical_govid", "category")])
covering <- .covering_recipes(con, rows)
cov_key <- paste(covering$year, covering$canonical_govid, covering$category,
sep = "\r")
for (i in idx) {
k <- paste(result$year[i], result$canonical_govid[i], result$category[i],
sep = "\r")
m <- match(k, cov_key)
if (is.na(m)) next
ids <- covering$recipe_ids[[m]]
if (length(ids) == 0L) next
flag[i] <- TRUE
notes[i] <- sprintf(
"Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).",
paste(sort(unique(ids)), collapse = "', '")
)
}
list(flag = flag, notes = notes)
}
#' For each (year, canonical_govid, category) triple potentially affected by
#' a suppressed Direct leg, find the harmonization recipe(s) that (a) cover
#' this `category` (share a component item_code via `summary_categories`,
#' excluding any recipe that is itself entirely intergovernmental M/L -- the
#' same exclusion `.build_suggestions()` applies, see I2) and (b) actually
#' produce a `long` row for this exact (canonical_govid, year) via the same
#' generic join `.run_recipe()` uses (component year_min/year_max +
#' gov_type_scope, no is_aggregate filter -- a recipe's whole point is to
#' recover data that's aggregate-only). Adds a list-column `recipe_ids`
#' (possibly length-0) to `rows`.
#' @noRd
.covering_recipes <- function(con, rows) {
rows$recipe_ids <- vector("list", nrow(rows))
cats <- unique(rows$category[!is.na(rows$category)])
if (length(cats) == 0L) return(rows)
cand <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT sc.category, r.recipe_id
FROM harmonization_recipes r
JOIN summary_categories sc ON sc.item_code = r.component_code
WHERE sc.category IN (%s)
AND r.recipe_id NOT IN (
SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE LEFT(component_code, 1) IN ('M', 'L')
)",
.sql_lit_chr(cats)
))
if (nrow(cand) == 0L) return(rows)
recipe_ids_all <- unique(cand$recipe_id)
govids <- unique(rows$canonical_govid)
years <- unique(rows$year)
covered <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.canonical_govid, l.year
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%s)
AND l.canonical_govid IN (%s)
AND l.year IN (%s)",
.sql_lit_chr(recipe_ids_all), .sql_lit_chr(govids), paste(years, collapse = ",")
))
for (i in seq_len(nrow(rows))) {
cat_i <- rows$category[i]
if (is.na(cat_i)) next
cat_recipe_ids <- cand$recipe_id[cand$category == cat_i]
if (length(cat_recipe_ids) == 0L) next
sub <- covered[covered$canonical_govid == rows$canonical_govid[i] &
covered$year == rows$year[i] &
covered$recipe_id %in% cat_recipe_ids, ]
rows$recipe_ids[[i]] <- sort(unique(sub$recipe_id))
}
rows
}
#' @noRd
.notes_column <- function(result, direct_suppressed_notes = NULL) {
n <- nrow(result)
if (n == 0L) return(character(0))
parts <- vector("list", 3L)
agg <- result[["aggregate_fallback"]]
parts[[1]] <- if (!is.null(agg)) {
ifelse(agg %in% TRUE,
"Aggregate fallback applied; see cog_explain()",
NA_character_)
} else {
rep(NA_character_, n)
}
ps <- result[["pop_source"]]
parts[[2]] <- if (!is.null(ps)) {
ifelse(ps == "unavailable",
"No population denominator available for this gov type",
NA_character_)
} else {
rep(NA_character_, n)
}
parts[[3]] <- if (!is.null(direct_suppressed_notes)) {
direct_suppressed_notes
} else {
rep(NA_character_, n)
}
out <- character(n)
for (i in seq_len(n)) {
pieces <- vapply(parts, `[[`, character(1), i)
pieces <- pieces[!is.na(pieces)]
out[i] <- if (length(pieces) == 0L) "" else paste(pieces, collapse = "; ")
}
out
} }
+251
View File
@@ -0,0 +1,251 @@
# R/suggestions.R
# Recipe-component-driven signposting: when a basis = "harmonized" query for
# a category comes back with a coverage gap in some requested years (the
# result has no rows at all in that year) that a harmonization recipe would
# actually fill for this government, surface that recipe as a suggestion.
#
# This is deliberately keyed off the recipe catalog's component codes, not
# off harmonization_map rows: no live map row carries a non-blank
# suggested_recipe_id (the corpus's wide era exposes split families like
# corrections functions 04+05 ONLY as aggregate rows, which basis =
# "harmonized" excludes by construction -- there's no NA ruling to hang a
# suggestion off of, just a leaf-code absence a recipe happens to fill).
# See docs/phase_r_harmonization_review.md § 0.3.
#
# Scope is deliberately narrow: signposting only runs when the caller
# supplied a `category` (an un-scoped, all-categories query has no single
# coverage question to answer) and only flags a recipe when the ACTUAL
# result has zero rows in a requested year AND the candidate recipe's own
# generic join (same join .run_recipe() uses, including its wide-era
# aggregate rows) produces at least one row for this government in that
# year. Checking presence per-government (not corpus-wide) avoids false
# positives from ordinary reporting variance -- most governments don't use
# every sibling code in a multi-code category every year, and that is not
# a format-boundary gap worth signposting.
#
# C1(a): for expenditure_concept = "total" callers, `result` here must
# already be the Direct-leg subset (the caller filters out
# spend_subtype == "intergovernmental" rows before calling in). A gap year
# is "the requested year has no Direct rows", never "no rows at all" --
# an IG row surviving on a legacy aggregate that Direct excludes must not
# read as coverage and cancel the very suggestion that would recover it.
#' Build the `prov$suggestions` list for a (non-recipe) basis = "harmonized"
#' verb call: recipes whose generic join would fill a real gap in `result`.
#'
#' @param con Active DuckDB connection.
#' @param govid Character vector of canonical_govid values (the verb's raw
#' `govid`).
#' @param years Integer vector of requested years.
#' @param category `category` argument as passed to the verb (character
#' vector or `NULL`; suggestions are only computed when non-NULL).
#' @param result The verb's already-computed result tibble (post basis
#' query, pre per_capita/adjust_to_year), pre-filtered to the Direct leg
#' only when the caller's `expenditure_concept = "total"` (see C1(a)).
#' @param basis The *resolved* basis (`"harmonized"` or `"raw"`).
#' @param flow_prefixes The calling verb's own flow-type prefixes (e.g.
#' `c("E", "F", "G")` for `cog_spending()`, `c("T", "A", "U", "B", "C",
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
#' lookup scoped to the calling verb's own flow family.
#' @return List of `list(recipe_id, label, available_years, hint,
#' ig_recipe_id)`, possibly empty.
#' @noRd
.build_suggestions <- function(con, govid, years, category, result, basis,
flow_prefixes) {
if (!identical(basis, "harmonized") || is.null(category)) return(list())
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
# i.e. every one of its own component codes is M/L-prefixed. Without this,
# a category whose summary_categories rows span both a Direct family
# (e.g. E04/E05, "Corrections") and its M/L counterpart (M04/M05, same
# category since Task 1) makes the M/L recipe itself (e.g.
# `corrections_ig_local_combined`) a raw top-level candidate for a plain
# (Direct) cog_spending() call -- following that hint would silently
# return intergovernmental dollars under `expenditure_concept = "direct"`
# provenance. This is a stronger, unconditional exclusion than the
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
# should never be suggested as a coverage-gap filler for EITHER verb, not
# just kept from being named as the *counterpart* of another suggestion.
candidates <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE component_code IN (
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)
)
AND recipe_id NOT IN (
SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE LEFT(component_code, 1) IN ('M', 'L')
)",
.sql_lit_chr(category)
))$recipe_id
if (length(candidates) == 0L) return(list())
result_years <- if (is.null(result) || nrow(result) == 0L) {
integer(0)
} else {
unique(as.integer(result$year))
}
gap_years <- setdiff(as.integer(years), result_years)
if (length(gap_years) == 0L) return(list())
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
"SELECT recipe_id, any_value(label) AS label,
MIN(year_min) AS year_min, MAX(year_max) AS year_max
FROM harmonization_recipes
WHERE recipe_id IN (%s)
GROUP BY recipe_id",
.sql_lit_chr(candidates)
)))
# Which (recipe_id, year) pairs the recipe's own generic join actually
# covers for this government, restricted to the gap years -- the same
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
# no is_aggregate filter), just checking existence instead of summing.
covered <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.year
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%s)
AND l.canonical_govid IN (%s)
AND l.year IN (%s)",
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(gap_years, collapse = ",")
))
suggestions <- list()
for (rid in candidates) {
if (!rid %in% covered$recipe_id) next
m <- meta[meta$recipe_id == rid, ]
suggestions[[length(suggestions) + 1L]] <- list(
recipe_id = rid,
label = m$label[[1]],
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
hint = sprintf("re-run with recipe = '%s'", rid)
)
}
.attach_ig_counterparts(con, suggestions, flow_prefixes)
}
#' Attach `ig_recipe_id` to each suggestion: the intergovernmental-expenditure
#' recipe (an M-to-local or L-to-state recipe) whose component codes cover
#' exactly the same set of function suffixes as the firing recipe's own
#' components, e.g. `corrections_combined`'s {E04, E05} -> suffixes {"04",
#' "05"} matches `corrections_ig_local_combined`'s {M04, M05} -> the same
#' {"04", "05"}. `NULL` when no such recipe exists, which also covers the
#' case where the firing recipe already IS the IG recipe (self-matches are
#' excluded, so an IG recipe never names itself as its own counterpart).
#'
#' Matching is deliberately an exact set match, not "any suffix in common":
#' the two-digit suffix only means the same "function" across recipes that
#' share the underlying Census functional-classification scheme (E/F/G/L/M
#' all use "04"/"05" for corrections). M/L "combined other" codes (47/89/
#' 91-94) reuse digits for an unrelated catch-all construct, so e.g.
#' `general_gov_e89_wide`'s {E85, E89} -> {"85", "89"} must NOT match
#' `ige_local_m89_wide`'s {"89", "91", "92", "93"} on the shared "89" alone.
#' Checked by hand against the full harmonization_recipes catalog: only the
#' corrections family (E/F/G/M, suffixes 04/05) has an exact-set match in
#' this corpus.
#'
#' Exact-set suffix matching is NOT enough on its own, though: the same
#' reused-digit problem exists ACROSS the revenue-side IG families too.
#' `ig_local_d47_wide` (D47/D94, suffixes {"47","94"}) is an exact-set match
#' for `ige_local_m47_wide` (M47/M94, same suffixes) even though one is
#' intergovernmental REVENUE received from local governments and the other is
#' intergovernmental EXPENDITURE paid to local governments -- unrelated flows
#' that happen to reuse "47"/"94" for their own "transit/utilities" and
#' "other/combined" catch-alls. `ig_federal_b47_wide`, `ig_state_c47_wide`,
#' and their `*_89` siblings all collide the same way. None of this is
#' reachable via `cog_revenue()` in the bundled fixture today (its B/C/D
#' recipes never happen to have a covered gap year for any fixture govid),
#' but it IS reachable via a mis-scoped `cog_spending()` call on a
#' revenue-only category, e.g. `cog_spending(gov, category = "IG Federal")`
#' fires `ig_federal_b47_wide`/`ig_federal_b89_wide` for real in the fixture
#' -- so this is a live, not merely theoretical, gap.
#'
#' Two flow-family checks close this, both required (see
#' `tests/testthat/test-expenditure-concept.R`, "revenue-flavored ... never
#' receives an M/L counterpart" tests, for the pairwise verification):
#' 1. `own_prefix %in% flow_prefixes`: the firing recipe's own component
#' codes must belong to the calling verb's own flow family (the same
#' `flow_prefixes` `.build_harmonization_block()` uses, see
#' `R/basis.R`). This blocks a recipe surfaced through a mis-scoped
#' category from ever reaching the M/L search, e.g. `cog_spending()`'s
#' flow_prefixes are `c("E","F","G")`, which `ig_federal_b47_wide`'s own
#' `"B"` is not part of.
#' 2. `own_prefix %in% c("E","F","G")`: M/L only ever pairs with the
#' DIRECT-expenditure family, never with revenue (`cog_revenue()`'s
#' flow_prefixes already fold B/C/D in as ordinary revenue -- there is
#' no separate "Total" bolt-on for revenue the way `expenditure_concept`
#' adds one for spending) and never with ANOTHER M/L recipe (without
#' this check, `ige_local_m47_wide` would wrongly match sibling
#' `ige_state_l47_wide` on their shared {"47","94"} suffix set).
#' Condition 1 alone does not catch this: under `cog_revenue()`,
#' `ig_federal_b47_wide`'s own `"B"` IS inside revenue's own
#' `flow_prefixes`, so only this second, family-specific check blocks
#' the search.
#' @noRd
.attach_ig_counterparts <- function(con, suggestions, flow_prefixes) {
if (length(suggestions) == 0L) return(suggestions)
comp <- DBI::dbGetQuery(con,
"SELECT recipe_id, component_code FROM harmonization_recipes")
comp$prefix <- substr(comp$component_code, 1L, 1L)
comp$suffix <- substr(comp$component_code, 2L, nchar(comp$component_code))
suffix_sets <- lapply(split(comp$suffix, comp$recipe_id), function(x) sort(unique(x)))
prefix_sets <- lapply(split(comp$prefix, comp$recipe_id), function(x) sort(unique(x)))
ig_recipe_ids <- unique(comp$recipe_id[comp$prefix %in% c("M", "L")])
find_counterpart <- function(rid) {
own_prefix <- prefix_sets[[rid]]
own_suffix <- suffix_sets[[rid]]
if (is.null(own_prefix) || is.null(own_suffix)) return(NULL)
if (!all(own_prefix %in% flow_prefixes)) return(NULL)
if (!all(own_prefix %in% c("E", "F", "G"))) return(NULL)
for (cand in ig_recipe_ids) {
if (identical(cand, rid)) next
if (setequal(suffix_sets[[cand]], own_suffix)) return(cand)
}
NULL
}
lapply(suggestions, function(s) {
# `s$ig_recipe_id <- NULL` would DELETE the element rather than set it
# (standard R list-assignment gotcha), leaving no-match entries missing
# the key entirely instead of carrying it as NULL. Single-bracket
# assignment with a wrapped list preserves a NULL-valued element so the
# field is always present, per the brief's "NULL when there is none".
s["ig_recipe_id"] <- list(find_counterpart(s$recipe_id))
s
})
}
#' Emit the single cli::cli_inform() message summarizing all suggestions
#' for a verb call (the brief's "one message", not one per suggestion).
#' Bullet text is pre-formatted plain text (no cli/glue `{}` markup) since
#' recipe ids/labels are untrusted-ish data values, not literal call-site
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
#' continuation line is appended naming the intergovernmental counterpart
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
#' same bullet under cli, not a new bullet).
#' @noRd
.inform_suggestions <- function(suggestions) {
bullets <- vapply(suggestions, function(s) {
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
s$available_years[1], s$available_years[2], s$hint)
if (!is.null(s$ig_recipe_id)) {
bullet <- paste0(bullet, sprintf(
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
}
bullet
}, character(1))
cli::cli_inform(c(
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
stats::setNames(bullets, rep("*", length(bullets)))
))
}
+61 -1
View File
@@ -1,11 +1,71 @@
# R/views.R # R/views.R
# SQL files that cannot be registered unconditionally against a v4 corpus,
# for one of two distinct reasons -- both fail at CREATE VIEW time (DuckDB
# resolves a view's source schema eagerly, even though it defers execution),
# so a v4 corpus can't tolerate either unconditionally:
#
# (a) Missing FILE. 33-/34-/35- read_parquet() a v5-only parquet table
# (harmonization_map.parquet, harmonization_recipes.parquet,
# series_breaks.parquet) that doesn't exist at all on a v4 corpus --
# "IO Error: No files found".
#
# (b) Missing COLUMN. 22-/23-/25- reference `long.harmonized_code`, a
# column that does not exist on a v4 corpus's `long` table (harmonized
# space was introduced in schema v5) -- "Binder Error: Referenced
# column harmonized_code not found". 42-/43-/45- are on this list only
# because they SELECT s.* FROM the (a)/(b) views above, so they'd fail
# to resolve their own source view if it weren't already skipped.
#
# Registration is therefore gated on manifest$schema_version >= 5 for all of
# them; verb-level *usage* of the resulting views is separately gated by
# .resolve_basis() / .require_schema_v5().
.harmonization_view_files <- c(
"22-spending_long_harmonized.sql",
"23-revenue_long_harmonized.sql",
"25-ig_long_harmonized.sql",
"33-harmonization_map.sql",
"34-harmonization_recipes.sql",
"35-series_breaks_pq.sql",
"42-spending_annotated_harmonized.sql",
"43-revenue_annotated_harmonized.sql",
"45-ig_annotated_harmonized.sql"
)
# The representation contract (cog_pipeline#64): two parquet tables that say
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
# on schema_version, because the sparsification that introduced them did not
# bump the version -- the pre-sparsification corpus this package shipped
# against until 2026-07-30 was already schema v6 and carried neither table.
# Keying off the version number would therefore register a view over a file
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
# check exists to tolerate.
.representation_view_files <- c(
"36-representation.sql" = "representation.parquet",
"37-code_set.sql" = "code_set.parquet"
)
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
#' works identically for a local fixture and a remote share.
#' @noRd
.corpus_has_table <- function(manifest, file) {
paths <- vapply(manifest$files$metadata %||% list(),
function(f) as.character(f$path %||% ""), character(1))
file %in% basename(paths)
}
#' Register DuckDB views from inst/sql/ SQL files #' Register DuckDB views from inst/sql/ SQL files
#' @noRd #' @noRd
.register_views <- function(con, url, manifest) { .register_views <- function(con, url, manifest) {
sql_dir <- system.file("sql", package = "uscogdata") sql_dir <- system.file("sql", package = "uscogdata")
files <- list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE) files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
for (f in files) { for (f in files) {
base <- basename(f)
if (base %in% .harmonization_view_files && schema_version < 5L) next
if (base %in% names(.representation_view_files) &&
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
sql <- paste(readLines(f, warn = FALSE), collapse = "\n") sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
DBI::dbExecute(con, sql) DBI::dbExecute(con, sql)
+82 -3
View File
@@ -19,20 +19,99 @@ package implements.
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata") # pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
``` ```
## Amounts are in full US dollars
Every amount column this package returns — `amt_nominal`, `amt_real`,
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
The raw Census source files report **thousands of dollars**, and the corpus's
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
you never have to. The conversion is recorded in every result:
```r
r <- cog_spending("552025209777", 2020L)
attr(r, "provenance")$transformations$units_conversion
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
```
**Do not multiply again.** If you have read elsewhere that COG amounts are in
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
twice overstates every figure by 1000x, and the result looks plausible rather
than obviously wrong.
## Configuration ## Configuration
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash) - `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory - `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600) - `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
## Primary vs Direct vs Total spending
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
controls whose spending a result counts. Concepts are defined as sets of the
crosswalk's `spend_subtype` values — never item-code first letters, which
cannot classify correctly (the letter `Y` alone spans revenue, expenditure,
and balance codes):
- `"primary"` (the default) is the government's own service provision:
current operations, capital outlay, and assistance payments.
- `"direct"` is Census's published Direct Expenditure: `primary` plus
interest on debt and insurance trust benefit payments (e.g. pensions).
- `"total"` additionally adds the intergovernmental leg — money handed to
other governments to spend (`M`/`L` codes plus `Q11`/`Q12`/`Q18` state
payments to school systems) — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
**Rule of thumb: any figure that spans more than one government uses
`primary` or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()`
enforce this by refusing `expenditure_concept = "total"`. See
`vignette("total-spending", package = "uscogdata")` for the full
explanation with worked examples.
## General vs Total revenue
`cog_revenue(..., revenue_concept = c("general", "total"))` selects between
Census's two published revenue concepts, again defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:
- `"general"` (the default) is Census **General Revenue**: own-source
(taxes, charges, miscellaneous) plus federal, state and local
intergovernmental aid.
- `"total"` is Census **Total Revenue**: `general` plus utility revenue
(`A91`–`A94`), liquor store revenue (`A90`), and insurance trust revenue
(unemployment and workers' compensation `Y` codes plus the
employee-retirement `X` codes).
The manual defines the first by subtracting the other three from the second,
so the two are related by Census's own identity:
```
Total Revenue = General + Utility + Liquor Store + Insurance Trust
```
Two things worth knowing before switching to `"total"`:
- **Utility revenue is large for cities.** Measured on the bundled fixture,
utility plus liquor store revenue is 15.9% of city (type 2) revenue, versus
1.2% for states and 1.7% for counties. `general` excludes it by definition.
- **The employee-retirement (`X`) codes stop at FY2016**, when those systems
moved out of the annual finance file into the separate Annual Survey of
Public Pensions. A `"total"` series therefore steps down at the
FY2016/FY2017 seam for reasons of collection scope, not revenue (series
breaks `SB197`–`SB202`, in the corpus's `series_breaks` table).
## Developer notes ## Developer notes
### Testing ### Testing
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` — The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50 a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` at this all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
fixture, so the full test suite runs offline with no network dependency: at this fixture, so the full test suite runs offline with no network
dependency:
```r ```r
devtools::test() # uses bundled fixture, no credentials required devtools::test() # uses bundled fixture, no credentials required
+238
View File
@@ -0,0 +1,238 @@
# data-raw/regenerate_fixture_corpus.R
#
# Regenerate inst/extdata/fixture_corpus/ from a cog_pipeline publish tree.
#
# What this does:
# 1. Copies each requested year's long partition as-is (byte-for-byte)
# from <publish_cache>/data/long/ into the fixture. Default years are
# c(2011L, 2012L, 2019L, 2020L): 2011/2012 straddle the wide-aggregate
# -> modern-leaf format boundary (the harmonization/recipe seam), and
# 2019/2020 are the pre-existing per-capita/CPI regression anchors.
# Each partition is a full year (all states/govs) as published, so
# Broward County FL and every other previously-pinned government stay
# covered without any per-gov slicing logic.
# 2. Copies every metadata parquet the publish tree ships (see
# .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
# registries, not partitioned by year, so the fixture ships the complete
# tables rather than a year-scoped subset. representation.parquet and
# code_set.parquet are what make the sparse wide era interpretable --
# absence means "$0" in a dense_source year and "not reported" in a
# sparse_source one -- so a fixture without them cannot represent the
# published corpus.
# 3. Resyncs the four reference docs (data_dictionary.md,
# reader-specification.md, README.md, series_breaks.md) from the
# publish tree's docs/.
# 4. Hand-builds manifest.json for just the files the fixture ships,
# following the shape of the previous fixture manifest but with
# schema_version bumped to whatever the source manifest reports, and
# freshly computed sha256 / row_count / size_bytes for every fixture
# file (never copied from the source manifest, since paths and byte
# layout can differ subtly between a full corpus and a fixture).
#
# This is never a manual job: run it whenever cog_pipeline publishes a new
# corpus vintage that the fixture should track.
#
# Usage (from the uscogdata package root):
# Rscript data-raw/regenerate_fixture_corpus.R
# Rscript data-raw/regenerate_fixture_corpus.R /path/to/publish_cache
#
# Or from R:
# source("data-raw/regenerate_fixture_corpus.R")
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
# Every metadata parquet the publish tree ships, in the order they appear in
# the corpus manifest. Single source of truth for both the copy step and the
# fixture manifest, so the two can never drift apart.
.FIXTURE_METADATA_FILES <- c(
"canonical_alias.parquet",
"canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet",
"code_set.parquet",
"harmonization_map.parquet",
"harmonization_recipes.parquet",
"lineage_events.parquet",
"representation.parquet",
"series_breaks.parquet",
"summary_categories.parquet"
)
regenerate_fixture_corpus <- function(
publish_cache_dir = file.path(
"..", "cog_pipeline", "_targets", "publish_cache"
),
fixture_dir = file.path("inst", "extdata", "fixture_corpus"),
fixture_years = c(2011L, 2012L, 2019L, 2020L)) {
stopifnot(
requireNamespace("digest", quietly = TRUE),
requireNamespace("jsonlite", quietly = TRUE),
requireNamespace("duckdb", quietly = TRUE),
requireNamespace("DBI", quietly = TRUE)
)
publish_cache_dir <- normalizePath(publish_cache_dir, mustWork = TRUE)
if (!dir.exists(fixture_dir)) dir.create(fixture_dir, recursive = TRUE)
source_manifest <- jsonlite::fromJSON(
file.path(publish_cache_dir, "manifest.json"),
simplifyVector = TRUE
)
.copy_long_partitions(publish_cache_dir, fixture_dir, fixture_years)
.copy_metadata_parquets(publish_cache_dir, fixture_dir)
.copy_docs(publish_cache_dir, fixture_dir)
manifest <- .build_fixture_manifest(
fixture_dir, source_manifest, fixture_years
)
manifest_path <- file.path(fixture_dir, "manifest.json")
writeLines(
jsonlite::toJSON(manifest, auto_unbox = TRUE, pretty = TRUE, null = "null"),
manifest_path
)
size_bytes <- sum(file.info(
list.files(fixture_dir, recursive = TRUE, full.names = TRUE)
)$size)
message(sprintf(
"Fixture corpus regenerated at %s (%.2f MB total).",
fixture_dir, size_bytes / 1024^2
))
invisible(manifest)
}
# Copy each requested year's partition directory (just the parquet file
# inside it) from the publish tree into the fixture, as-is.
#' @noRd
.copy_long_partitions <- function(publish_cache_dir, fixture_dir, years) {
for (yr in years) {
part_rel <- file.path("data", "long", sprintf("year=%d", yr), "part-0.parquet")
src <- file.path(publish_cache_dir, part_rel)
dst <- file.path(fixture_dir, part_rel)
if (!file.exists(src)) {
stop(sprintf("Source partition missing: %s", src))
}
dir.create(dirname(dst), recursive = TRUE, showWarnings = FALSE)
ok <- file.copy(src, dst, overwrite = TRUE)
if (!ok) stop(sprintf("Failed to copy %s -> %s", src, dst))
}
invisible(NULL)
}
# Copy the full (not year-scoped) metadata tables listed in
# .FIXTURE_METADATA_FILES.
#' @noRd
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
for (f in .FIXTURE_METADATA_FILES) {
src <- file.path(publish_cache_dir, "data", f)
dst <- file.path(fixture_dir, "data", f)
if (!file.exists(src)) {
stop(sprintf("Source metadata file missing: %s", src))
}
dir.create(dirname(dst), recursive = TRUE, showWarnings = FALSE)
ok <- file.copy(src, dst, overwrite = TRUE)
if (!ok) stop(sprintf("Failed to copy %s -> %s", src, dst))
}
invisible(NULL)
}
# Resync the four reference docs shipped alongside the fixture.
#' @noRd
.copy_docs <- function(publish_cache_dir, fixture_dir) {
docs <- c(
"data_dictionary.md", "reader-specification.md",
"README.md", "series_breaks.md"
)
dst_dir <- file.path(fixture_dir, "docs")
dir.create(dst_dir, recursive = TRUE, showWarnings = FALSE)
for (f in docs) {
src <- file.path(publish_cache_dir, "docs", f)
if (!file.exists(src)) {
stop(sprintf("Source doc missing: %s", src))
}
ok <- file.copy(src, file.path(dst_dir, f), overwrite = TRUE)
if (!ok) stop(sprintf("Failed to copy doc %s", f))
}
invisible(NULL)
}
# Count rows in a parquet file via an ephemeral DuckDB connection.
#' @noRd
.parquet_row_count <- function(path) {
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbGetQuery(con, sprintf(
"SELECT COUNT(*) AS n FROM read_parquet(%s)",
.sql_quote(path)
))$n
}
#' @noRd
.sql_quote <- function(x) paste0("'", gsub("'", "''", x), "'")
# Hand-build manifest.json following the shape of the previous fixture
# manifest: schema_version / built_at / pipeline_commit / fixture_note /
# data_vintage / scope / schema / files.long_partitions / files.metadata /
# series_breaks_ref / reader_spec_ref. Every sha256 / row_count / size_bytes
# is freshly computed against the files actually written into fixture_dir.
#' @noRd
.build_fixture_manifest <- function(fixture_dir, source_manifest, years) {
long_partitions <- lapply(years, function(yr) {
rel <- file.path("data", "long", sprintf("year=%d", yr), "part-0.parquet")
path <- file.path(fixture_dir, rel)
list(
year = as.integer(yr),
path = gsub("\\\\", "/", rel),
sha256 = digest::digest(path, algo = "sha256", file = TRUE),
row_count = as.integer(.parquet_row_count(path)),
size_bytes = as.integer(file.info(path)$size)
)
})
metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
rel <- file.path("data", f)
path <- file.path(fixture_dir, rel)
list(
path = gsub("\\\\", "/", rel),
sha256 = digest::digest(path, algo = "sha256", file = TRUE),
description = f
)
})
list(
schema_version = as.integer(source_manifest$schema_version),
built_at = format(Sys.time(), "%Y-%m-%dT%H:%M:%SZ", tz = "UTC"),
pipeline_commit = source_manifest$pipeline_commit,
fixture_note = paste(
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
"corpus available via USCOGDATA_URL. Regenerated from the sparsified",
"schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
"zeros, so FY2011 absence means Census published $0 while FY2012+",
"absence means not reported. representation.parquet and",
"code_set.parquet carry that rule and ship in full, as do every other",
"metadata table in the publish tree. 2011/2012 straddle both the",
"wide-aggregate -> modern-leaf format boundary (exercised by",
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
"representation boundary (SB194); 2019/2020 retain the prior",
"per-capita/CPI regression anchors. Regenerated via",
"data-raw/regenerate_fixture_corpus.R."
),
data_vintage = source_manifest$data_vintage,
scope = source_manifest$scope,
schema = source_manifest$schema,
files = list(
long_partitions = long_partitions,
metadata = metadata
),
series_breaks_ref = source_manifest$series_breaks_ref,
reader_spec_ref = source_manifest$reader_spec_ref
)
}
if (identical(environment(), globalenv()) && sys.nframe() == 0L) {
args <- commandArgs(trailingOnly = TRUE)
if (length(args) >= 1L) {
regenerate_fixture_corpus(publish_cache_dir = args[[1]])
} else {
regenerate_fixture_corpus()
}
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+83 -49
View File
@@ -1,82 +1,116 @@
{ {
"schema_version": 3, "schema_version": 6,
"built_at": "2026-04-27T16:43:46Z", "built_at": "2026-07-31T00:47:27Z",
"pipeline_commit": "899af37", "pipeline_commit": "aadb46b",
"fixture_note": "Two-year (2019-2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL.", "fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
"data_vintage": { "data_vintage": {
"census_source_downloaded": "unknown", "source_vintages": {
"cpi_vintage": "FRED CPIAUCSL", "2012": "10162019",
"2013": "10162019",
"2014": "10162019",
"2015": "10162019",
"2016": "10162019",
"2017": "06102021",
"2018": "06102021",
"2019": "06102021",
"2020": "06122023",
"2021": "06122023",
"2022": "06052025",
"2023": "06052025"
},
"registry_rows": 148,
"acs_vintage": "ACS 2018-2022 5-year" "acs_vintage": "ACS 2018-2022 5-year"
}, },
"scope": { "scope": {
"gov_types_included": [ "gov_types_included": [0, 1, 2, 3],
0, "gov_types_excluded": [4, 5],
1,
2,
3
],
"gov_types_excluded": [
4,
5
],
"scope_note": "v0.1 covers state, county, city/municipality, and township governments. Special districts (type 4) and school districts (type 5) are excluded pending validation in a future cycle." "scope_note": "v0.1 covers state, county, city/municipality, and township governments. Special districts (type 4) and school districts (type 5) are excluded pending validation in a future cycle."
}, },
"schema": { "schema": {
"long_column_count": 24, "long_column_count": 28,
"long_columns": [ "long_columns": ["fips_state", "type", "fips_county", "govid", "gov_blank", "gov_name", "county_name", "fips_state_asof", "fips_county_asof", "cog_legacy_state", "cog_legacy_county", "fips_place_code", "population", "popyear", "enrollment", "enrollyear", "function_code", "sch_level_code", "fiscal_year_end", "srvy_year", "item_code", "amt", "srv_data", "impute_flag", "is_aggregate", "canonical_govid", "harmonized_code", "survey_weight"],
"fips_state",
"type",
"fips_county",
"govid",
"gov_blank",
"gov_name",
"county_name",
"fips_state_code",
"fips_county_code",
"fips_place_code",
"population",
"popyear",
"enrollment",
"enrollyear",
"function_code",
"sch_level_code",
"fiscal_year_end",
"srvy_year",
"item_code",
"amt",
"srv_data",
"impute_flag",
"is_aggregate",
"canonical_govid"
],
"data_dictionary": "docs/data_dictionary.md" "data_dictionary": "docs/data_dictionary.md"
}, },
"files": { "files": {
"long_partitions": [ "long_partitions": [
{
"year": 2011,
"path": "data/long/year=2011/part-0.parquet",
"sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
"row_count": 496004,
"size_bytes": 2202455
},
{
"year": 2012,
"path": "data/long/year=2012/part-0.parquet",
"sha256": "b82ac82d5e35f844b26c887445601f3748438c52c998ba4e403b025941a6f170",
"row_count": 1163338,
"size_bytes": 5929917
},
{ {
"year": 2019, "year": 2019,
"path": "data/long/year=2019/part-0.parquet", "path": "data/long/year=2019/part-0.parquet",
"sha256": "e1c9f426c6d7d3c51d06b3a652473b987b304836619c213f019cee4887714daa", "sha256": "5cbd4726dcc7d0dab5c2a05a64702e979533ae119ed0587073cd31c089e0d737",
"row_count": 318139, "row_count": 318139,
"size_bytes": 1424231 "size_bytes": 1719548
}, },
{ {
"year": 2020, "year": 2020,
"path": "data/long/year=2020/part-0.parquet", "path": "data/long/year=2020/part-0.parquet",
"sha256": "9b795853a848e8c955c80261b96b79630fc77394dcfb1a1ca288e2cd634053a3", "sha256": "ee548fec80bf1beda844fe03916ac145f10dd34c45968407cc330ec260935f00",
"row_count": 317500, "row_count": 317500,
"size_bytes": 1427150 "size_bytes": 1722918
} }
], ],
"metadata": [ "metadata": [
{
"path": "data/canonical_alias.parquet",
"sha256": "3f617051c23a99bea322889857f7106df0c92954564afeec181df7083ee6698e",
"description": "canonical_alias.parquet"
},
{ {
"path": "data/canonical_fips_xwalk.parquet", "path": "data/canonical_fips_xwalk.parquet",
"sha256": "86e53e04a35f6f90bb74bb1a273e053392afa782d6f518e3e3da9c976d47f7af", "sha256": "f98742f941269dacf8f7de5c273aa4dd4e75017a5bb70c054da35852a95a8d46",
"description": "canonical_fips_xwalk.parquet" "description": "canonical_fips_xwalk.parquet"
}, },
{
"path": "data/census_collection_coverage.parquet",
"sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
"description": "census_collection_coverage.parquet"
},
{
"path": "data/code_set.parquet",
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
"description": "code_set.parquet"
},
{
"path": "data/harmonization_map.parquet",
"sha256": "4cf32d0f817079ba4f28dc0ce65450d3247ebbf08d94c0c26c0d02af597bf812",
"description": "harmonization_map.parquet"
},
{
"path": "data/harmonization_recipes.parquet",
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
"description": "harmonization_recipes.parquet"
},
{
"path": "data/lineage_events.parquet",
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
"description": "lineage_events.parquet"
},
{
"path": "data/representation.parquet",
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
"description": "representation.parquet"
},
{
"path": "data/series_breaks.parquet",
"sha256": "06dcc995ff533e57cc65fa25086cc9bf83ba592c58bf7cc99269dc2576f69944",
"description": "series_breaks.parquet"
},
{ {
"path": "data/summary_categories.parquet", "path": "data/summary_categories.parquet",
"sha256": "60045e22bc2723318fa2cb73f8e5038250dc54d24b3447c6750dfe29035335b8", "sha256": "e3b0efa00ce713b8f45829b89cfde24b55333f26101f0495df82d85997d18d8e",
"description": "summary_categories.parquet" "description": "summary_categories.parquet"
} }
] ]
+38
View File
@@ -10,11 +10,49 @@
"target": { "type": "object" }, "target": { "type": "object" },
"years": { "type": "array", "items": { "type": "integer" } }, "years": { "type": "array", "items": { "type": "integer" } },
"category": { "type": ["string", "array", "null"] }, "category": { "type": ["string", "array", "null"] },
"basis": { "type": ["string", "null"] },
"basis_note": { "type": ["string", "null"] },
"expenditure_concept": {
"type": "string",
"enum": ["primary", "direct", "total"],
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
},
"expenditure_concept_note": {
"type": ["string", "null"],
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
},
"expenditure_concept_direct_suppressed": {
"type": "boolean",
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
},
"revenue_concept": {
"type": "string",
"enum": ["general", "total"],
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
},
"harmonization": { "type": "object" },
"recipe": { "type": ["object", "null"] },
"suggestions": { "type": "array" },
"scope": { "type": "object" }, "scope": { "type": "object" },
"codes_summed": { "type": "object" }, "codes_summed": { "type": "object" },
"aggregate_fallback": { "type": ["object", "null"] }, "aggregate_fallback": { "type": ["object", "null"] },
"transformations":{ "type": "object" }, "transformations":{ "type": "object" },
"series_break_refs": { "type": "array", "items": { "type": "string" } }, "series_break_refs": { "type": "array", "items": { "type": "string" } },
"completion": {
"type": "object",
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
"properties": {
"applied": { "type": "boolean" },
"rows_filled": { "type": "integer" },
"absence_means": { "type": "object" }
}
},
"corpus_break_refs": {
"type": "array",
"items": { "type": "string" },
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
},
"manifest": { "type": "object" }, "manifest": { "type": "object" },
"sql_query": { "type": "string" } "sql_query": { "type": "string" }
} }
+7
View File
@@ -0,0 +1,7 @@
-- Category crosswalk. Numbered 11 (not with the other reference tables at
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
-- already exist when they register.
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+18 -1
View File
@@ -1,5 +1,22 @@
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
-- (summary_categories.category_type = 'expenditure'), never by item-code
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
-- every non-intergovernmental expenditure subtype: operations, capital,
-- assistance, interest, insurance_benefits.
--
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
-- 24-ig_long.sql).
CREATE OR REPLACE VIEW spending_long AS CREATE OR REPLACE VIEW spending_long AS
SELECT * SELECT *
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G', 'K') WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
)
AND NOT is_aggregate; AND NOT is_aggregate;
+14 -1
View File
@@ -1,5 +1,18 @@
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
--
-- Carries EVERY revenue subtype. Which of Census's two published concepts a
-- query actually returns is decided per revenue_concept in R
-- (.verb_spendrev), exactly as expenditure_concept narrows spending_long:
-- general = own_source + federal + state + local_aid (the default)
-- total = general + utility + liquor_store + insurance_trust
-- Census defines the first by subtracting the other three from the second
-- (manual section 4.3), so both concepts need all four families present here.
CREATE OR REPLACE VIEW revenue_long AS CREATE OR REPLACE VIEW revenue_long AS
SELECT * SELECT *
FROM long FROM long
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D') WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
)
AND NOT is_aggregate; AND NOT is_aggregate;
+16
View File
@@ -0,0 +1,16 @@
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
-- classification, applied to harmonized_code (the code the row is folded
-- onto) rather than the published item_code. Safe because the harmonized
-- space is leaf-only and every harmonized_code in the corpus is a
-- summary_categories member (verified at fixture regen; a code the
-- crosswalk cannot classify would be silently dropped here).
CREATE OR REPLACE VIEW spending_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'expenditure'
AND spend_subtype <> 'intergovernmental'
);
+12
View File
@@ -0,0 +1,12 @@
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
-- classification (every revenue subtype; the concept narrows in R), applied
-- to harmonized_code rather than the published item_code.
CREATE OR REPLACE VIEW revenue_long_harmonized AS
SELECT * REPLACE (harmonized_code AS item_code)
FROM long
WHERE NOT is_aggregate
AND harmonized_code IS NOT NULL
AND harmonized_code IN (
SELECT item_code FROM summary_categories
WHERE category_type = 'revenue'
);
+25
View File
@@ -0,0 +1,25 @@
-- Intergovernmental expenditure rows: crosswalk spend_subtype =
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
-- = state payments to school systems -- uscogdata#11, finding F-017).
--
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
-- ONLY as aggregate-flagged rows -- filtering them would hide ~70% of legacy IG
-- dollars and make Total silently collapse to Direct. This is safe because the
-- aggregate codes and their modern leaf components are strictly year-disjoint
-- (M47 ends 2011 / M94 starts 2012; M89 is aggregate only <= 2011 and a leaf
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
-- the pipeline's recipe joins use.
--
-- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
-- rolls up the L-NN codes, so including it would double-count. The crosswalk
-- deliberately carries no `--` family-total codes, so membership excludes it
-- (guarded by "the IG leg never includes the L-- family total" in
-- tests/testthat/test-expenditure-concept.R).
CREATE OR REPLACE VIEW ig_long AS
SELECT *
FROM long
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
+22
View File
@@ -0,0 +1,22 @@
-- Harmonized-basis IG rows. Uses COALESCE(harmonized_code, item_code) rather
-- than harmonized_code alone: aggregate rows carry NO harmonized_code by
-- construction (harmonized space is leaf-only), so a plain
-- `harmonized_code IS NOT NULL` filter would drop every legacy IG aggregate --
-- in the bundled fixture corpus (year 2011; 2012+ all carry a harmonized_code)
-- that is $379,016,063k across 25,688 M rows and $2,277,458k across 19,266 L
-- rows (`SELECT year, LEFT(item_code,1), SUM(amt), COUNT(*) FROM ig_long
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
-- while never dropping a row.
--
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
-- rather than the COALESCEd code: every IG harmonization target (M36) is
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
-- the column that exists on every row.
CREATE OR REPLACE VIEW ig_long_harmonized AS
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
FROM long
WHERE item_code IN (
SELECT item_code FROM summary_categories
WHERE spend_subtype = 'intergovernmental'
);
-3
View File
@@ -1,3 +0,0 @@
CREATE OR REPLACE VIEW summary_categories AS
SELECT *
FROM read_parquet('{url}data/summary_categories.parquet');
+8
View File
@@ -0,0 +1,8 @@
CREATE OR REPLACE VIEW gov_population_yearly AS
SELECT DISTINCT
year,
canonical_govid,
population,
popyear
FROM long
WHERE population IS NOT NULL;
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW harmonization_map AS
SELECT *
FROM read_parquet('{url}data/harmonization_map.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW harmonization_recipes AS
SELECT *
FROM read_parquet('{url}data/harmonization_recipes.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW series_breaks_pq AS
SELECT *
FROM read_parquet('{url}data/series_breaks.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW representation AS
SELECT *
FROM read_parquet('{url}data/representation.parquet');
+3
View File
@@ -0,0 +1,3 @@
CREATE OR REPLACE VIEW code_set AS
SELECT *
FROM read_parquet('{url}data/code_set.parquet');
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW spending_annotated_harmonized AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.spend_subtype
FROM spending_long_harmonized s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW revenue_annotated_harmonized AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.revenue_subtype
FROM revenue_long_harmonized s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
+16
View File
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW ig_annotated AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.spend_subtype
FROM ig_long s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
+16
View File
@@ -0,0 +1,16 @@
CREATE OR REPLACE VIEW ig_annotated_harmonized AS
SELECT
s.*,
x.gov_name AS xwalk_gov_name,
x.govs_type,
x.type_label,
x.fips_state AS xwalk_fips_state,
x.fips_county AS xwalk_fips_county,
x.fips_place,
x.population_acs,
c.category,
c.category_type,
c.spend_subtype
FROM ig_long_harmonized s
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
LEFT JOIN summary_categories c USING (item_code);
+21 -10
View File
@@ -6,17 +6,22 @@
\usage{ \usage{
cog_find_peers( cog_find_peers(
target_govid, target_govid,
year = NULL,
same_type = TRUE, same_type = TRUE,
same_state = FALSE, same_state = FALSE,
pop_range = c(0.7, 1.3), pop_range = c(0.7, 1.3),
is_ratio = TRUE, is_ratio = TRUE,
pop_year = NULL, max_peers = 10L,
max_peers = 10L coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
\item{target_govid}{Character scalar — `canonical_govid` of the target.} \item{target_govid}{Character scalar — `canonical_govid` of the target.}
\item{year}{Integer scalar. Cohort vintage. When `NULL` (default), uses the
most recent year for which the target has an observed population in
`gov_population_yearly`.}
\item{same_type}{If `TRUE` (default) restrict peers to the target's \item{same_type}{If `TRUE` (default) restrict peers to the target's
`govs_type`.} `govs_type`.}
@@ -26,20 +31,26 @@ Default `FALSE`.}
\item{pop_range}{Length-2 numeric vector giving lower/upper bounds.} \item{pop_range}{Length-2 numeric vector giving lower/upper bounds.}
\item{is_ratio}{If `TRUE` (default) `pop_range` is multiplied by the \item{is_ratio}{If `TRUE` (default) `pop_range` is multiplied by the
target's `population_acs` to produce absolute bounds. If `FALSE`, target's population at `year` to produce absolute bounds. If `FALSE`,
`pop_range` is interpreted as absolute population counts.} `pop_range` is interpreted as absolute population counts.}
\item{pop_year}{Reserved for future use (selecting ACS vintage). Currently
the corpus has a single snapshot so this argument has no effect.}
\item{max_peers}{Integer cap on the number of peers returned.} \item{max_peers}{Integer cap on the number of peers returned.}
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
most recent census year with an observed population, so a cohort is not
built from a sample year in which most of the candidate universe is
absent. `"consistent"` needs a year range, which cohort selection does not
have, so it selects like `"all"` and is carried on the result as
`attr(x, "coverage")` for [cog_peer_compare()].}
} }
\value{ \value{
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`, Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
`population_acs`, `pop_ratio`, `rank`. `population`, `pop_ratio`, `rank`. The cohort year is attached as
`attr(x, "cohort_year")`.
} }
\description{ \description{
Selects peer governments from `canonical_fips_xwalk` by combinations of Selects peer governments by combinations of government type, state, and
government type, state, and population range. Peers are ordered by population range at a chosen `year`. Peers are ordered by `|log(pop_ratio)|`
`|log(pop_ratio)|` ascending (closest to the target's population first). ascending (closest to the target's population first).
} }
+45 -6
View File
@@ -9,7 +9,9 @@ cog_geographic_rollup(
category, category,
years, years,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL adjust_to_year = NULL,
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
@@ -22,17 +24,46 @@ to [cog_spending()]).}
\item{years}{Integer vector of years.} \item{years}{Integer vector of years.}
\item{per_capita}{If `TRUE`, per-capita uses each layer's own population \item{per_capita}{If `TRUE`, per-capita uses each gov's own per-year
from `canonical_fips_xwalk.population_acs`.} population from `gov_population_yearly`. Govs with missing population
are excluded from the result.}
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.} \item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across multiple layers of
government double-counts intergovernmental transfers (a state's payment
to a school district is the same dollar the district reports as its own
Direct spending); `"primary"` and `"direct"` combine safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.}
} }
\value{ \value{
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real` / `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
`amt_per_capita_nominal` / `amt_per_capita_real`, `codes_included`, `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
`aggregate_fallback`, `scope_note`, `notes`. Carries a `provenance` `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
attribute with `verb = "cog_geographic_rollup"` and `layers`. `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
and `rollup$included_govids` / `rollup$excluded_govids`.
} }
\description{ \description{
Wraps [cog_spending()], tags each row with its layer, and attaches a Wraps [cog_spending()], tags each row with its layer, and attaches a
@@ -41,3 +72,11 @@ human-readable `scope_note` documenting geographic-scope caveats (e.g.
"place portraits" that compare a city to the surrounding county and "place portraits" that compare a city to the surrounding county and
containing state on one set of axes. containing state on one set of axes.
} }
\details{
When `per_capita = TRUE`, rows whose government has no observed
population in that year (`pop_source == "unavailable"`) are dropped from
the result. The dropped govids are recorded in
`provenance$rollup$excluded_govids`. This excludes special districts
(gov type 4) and school districts (gov type 5) from per-capita rollups
by design — see `vignette('population-denominators')`.
}
+8 -4
View File
@@ -32,8 +32,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
} }
\details{ \details{
* **Utility mode** (single `name`, the original behavior): returns all * **Utility mode** (single `name`, the original behavior): returns all
rows whose `gov_name` matches the regex case-insensitively, sorted by rows whose `gov_name` contains `name` as a **literal, case-insensitive
`population_acs` descending. Useful for exploratory lookups. substring**, sorted by `population_acs` descending. Useful for
exploratory lookups. Regex metacharacters in `name` are escaped, so a
government is findable by its own complete name even when that name
contains parentheses or a period.
* **Basket mode** (`length(name) > 1`): resolves each input row to a * **Basket mode** (`length(name) > 1`): resolves each input row to a
single canonical govid and returns a tibble in input order, suitable single canonical govid and returns a tibble in input order, suitable
for piping straight into [cog_spending()] / [cog_revenue()] / for piping straight into [cog_spending()] / [cog_revenue()] /
@@ -45,7 +48,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`. 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
2. **Exact pass:** case-insensitive equality against `gov_name`. 2. **Exact pass:** case-insensitive equality against `gov_name`.
Single hit -> resolved. Multiple -> step 4. Single hit -> resolved. Multiple -> step 4.
3. **Substring fallback:** case-insensitive regex against `gov_name`. 3. **Substring fallback:** case-insensitive literal substring against
`gov_name` (metacharacters escaped).
Single hit -> resolved (`match_method = "substring"`). Zero hits -> Single hit -> resolved (`match_method = "substring"`). Zero hits ->
`status = "no_match"`. Multiple hits -> step 4. `status = "no_match"`. Multiple hits -> step 4.
4. **Disambiguation:** if matches share one `govs_type`, pick the 4. **Disambiguation:** if matches share one `govs_type`, pick the
@@ -58,7 +62,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
} }
\examples{ \examples{
\dontrun{ \dontrun{
# Utility mode — exploratory regex lookup # Utility mode — exploratory substring lookup
cog_gov_search("broward", state = "FL") cog_gov_search("broward", state = "FL")
# Basket mode — resolve a known cohort # Basket mode — resolve a known cohort
+18
View File
@@ -0,0 +1,18 @@
% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/manifest.R
\name{cog_manifest}
\alias{cog_manifest}
\title{Return the parsed corpus manifest for the active session.}
\usage{
cog_manifest()
}
\value{
Named list: `schema_version`, `built_at`, `pipeline_commit`,
`data_vintage`, `scope`, `years` (schema v5+), `schema`, `files`.
}
\description{
Opens a session (connecting to the configured corpus) if none is active,
then returns the manifest exactly as parsed from `manifest.json`. Useful
for consumers that need the published year range (`years` block, schema
v5+) or the partition list without issuing a data query.
}
+70 -5
View File
@@ -10,7 +10,9 @@ cog_peer_compare(
category, category,
years, years,
per_capita = TRUE, per_capita = TRUE,
adjust_to_year = NULL adjust_to_year = NULL,
expenditure_concept = c("primary", "direct", "total"),
coverage = c("all", "census", "consistent")
) )
} }
\arguments{ \arguments{
@@ -27,18 +29,81 @@ cog_peer_compare(
population.} population.}
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.} \item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
refused here because combining Total across peer sets counts
intergovernmental transfers twice; `"primary"` and `"direct"` combine
safely.}
\item{coverage}{How to handle the Census of Governments survey cycle,
which is a **complete census only in years ending in 2 and 7** -- every
other year is a sample, and the sample varies enormously (on the bundled
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
and 112 in FY2019).
* `"all"` (default) -- every unit that reported that year. Unchanged
behaviour, so existing code keeps working.
* `"census"` -- census years only. Aborts if the requested range holds
none, rather than silently returning nothing.
* `"consistent"` -- only units reporting in *every* requested year, giving
a balanced panel.
Regardless of mode, `provenance$coverage` always carries per-year
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
`provenance$coverage_mode` records the mode. `is_census_year` is a
statement about the **survey calendar**, never a claim of completeness:
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
report. `n_units_reporting` is the number that tells the truth.
The comparison target is exempt from `"consistent"` balancing -- it is the
subject of the comparison, not a member of the cohort -- and the
`summary_*` quantiles are computed AFTER the filter, so they describe the
cohort actually returned. `n_units_reporting` counts peers only, against
the cohort size: "3 of your 15 peers reported in FY2019".}
} }
\value{ \value{
Tibble matching [cog_spending()]'s columns, plus a `role` Tibble matching [cog_spending()]'s columns, plus a `role`
column taking values `"target"`, `"peer"`, `"summary_p25"`, column taking values `"target"`, `"peer"`, `"summary_p25"`,
`"summary_p50"`, or `"summary_p75"`, and `target_rank` (target's rank `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
among target+peers at `max(years)`, NA for other rows). Provenance among target+peers at `max(years)`, NA for other rows), and
attribute reports `verb = "cog_peer_compare"` and `peer_count`. `cohort_year` (the year used to build the peer cohort, read from
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
`cohort_year`, and `cohort_govids`.
**The `summary_*` rows are per-category quantiles: they are not additive.**
Each one is computed **within each `(year, spend_subtype,
category)` cell** across the peer set, so a `summary_p50` row is *the
median peer's value in that one category*, not *the value of the median
peer's total*. The median peer for Police and the median peer for Fire
are usually different governments, so summing `summary_*` rows across
categories does not give any peer's total and misstates the band it
appears to describe — measured at −32.7% to +251.0% across 24 years on
one cohort, with a sign flip at FY2012.
Facet by `role` **and** `category` (the documented use, and what the
rows are built for). For a genuine "median peer's total spending" line,
sum each peer's own categories first and take the quantile of those
per-government totals:
```r
library(dplyr)
cmp |>
filter(role %in% c("target", "peer")) |>
group_by(year, role, canonical_govid) |>
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
filter(role == "peer") |>
group_by(year) |>
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
```
} }
\description{ \description{
Pulls spending for the target plus a peer set (either a Pulls spending for the target plus a peer set (either a
[cog_find_peers()] result or a character vector of `canonical_govid`) and [cog_find_peers()] result or a character vector of `canonical_govid`) and
appends peer-distribution summary rows (`summary_p25`, `summary_p50`, appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
`summary_p75`) so the result can be faceted by `role` in a single ggplot `summary_p75`) so the result can be faceted by `role` in a single ggplot
call. call. Those summary rows are quantiles **within each category**, not
quantiles of each peer's total — see the `@return` section before summing
them.
} }
+22
View File
@@ -0,0 +1,22 @@
% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/recipes.R
\name{cog_recipes}
\alias{cog_recipes}
\title{List available harmonization recipes}
\usage{
cog_recipes(pattern = NULL)
}
\arguments{
\item{pattern}{Optional regex matched case-insensitively against
`recipe_id` or `label`.}
}
\value{
Tibble with columns `recipe_id`, `label`, `n_components`,
`year_min`, `year_max` (the min/max component year coverage), sorted by
`recipe_id`.
}
\description{
Recipes are multi-code cross-vintage series (see [cog_spending()]'s
`recipe` argument) catalogued in the corpus's `harmonization_recipes`
table. Use this to discover valid `recipe` ids.
}
+85 -4
View File
@@ -9,7 +9,11 @@ cog_revenue(
years, years,
category = NULL, category = NULL,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE
) )
} }
\arguments{ \arguments{
@@ -21,17 +25,94 @@ cog_revenue(
`summary_categories.category`), or `NULL` for all categories.} `summary_categories.category`), or `NULL` for all categories.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and \item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
`population_acs` from the canonical xwalk.} Census F-33 population from `gov_population_yearly`. Result also gains
a `pop_source` column with values `"census_f33"` or `"unavailable"`
(the latter for gov types 4/5 and any row whose population is missing
in that year).}
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion, \item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
or `NULL` for nominal only.} or `NULL` for nominal only.}
\item{basis}{`"harmonized"` (default) sums item codes through the
cross-vintage harmonization mapping (folding series-break-affected
codes onto a comparable target and excluding aggregate / discontinued
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
reproduces the pre-Phase-R2 behavior (published item codes, no
folding). On a corpus with `schema_version < 5` (no harmonization
tables), `basis` silently resolves to `"raw"` when left at its default
and the resolution is recorded in the provenance; explicitly passing
`basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
is set (see below).}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
multi-code cross-vintage series that a 1:1 harmonized_code mapping
can't express (e.g. a wide-era aggregate that only splits into leaf
codes in the modern era). Mutually exclusive with `category`. The
result's subtype column reads `"recipe"` and `category` reads the
recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
`basis` entirely (it joins `long` directly rather than going through
the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{revenue_concept}{Which of Census's two published revenue concepts to
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
values -- never as item-code first letters, which cannot classify
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
prefix `X` does the same):
* `"general"` (default) -- Census General Revenue: `own_source` +
`federal` + `state` + `local_aid`. The manual defines this concept by
subtraction (section 4.3: *"General revenue comprises all revenue
except that classified as liquor store, utility, or insurance trust
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
insurance trust revenue are all excluded.
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
`general` plus utility, liquor store, and insurance trust revenue
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
`X01`/`X02`/`X05`/`X08`).
The two are related by Census's own identity, `Total Revenue = General +
Utility + Liquor Store + Insurance Trust`.
Note that the employee-retirement (`X`) codes stop at FY2016, when those
systems moved out of the annual finance file into the separate Annual
Survey of Public Pensions, so a `"total"` series steps down at the
FY2016/FY2017 seam for reasons that are about collection scope rather
than revenue (series breaks `SB197`-`SB202`).}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
} }
\value{ \value{
Tibble with columns `year`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `canonical_govid`, `gov_name`,
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`, `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
`codes_included`, `aggregate_fallback`, `notes`. optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
} }
\description{ \description{
Mirror of [cog_spending()] for revenue categories. One row per Mirror of [cog_spending()] for revenue categories. One row per
+103 -5
View File
@@ -9,7 +9,11 @@ cog_spending(
years, years,
category = NULL, category = NULL,
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE
) )
} }
\arguments{ \arguments{
@@ -21,18 +25,112 @@ cog_spending(
`summary_categories.category`), or `NULL` for all categories.} `summary_categories.category`), or `NULL` for all categories.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and \item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
`population_acs` from the canonical xwalk.} Census F-33 population from `gov_population_yearly`. Result also gains
a `pop_source` column with values `"census_f33"` or `"unavailable"`
(the latter for gov types 4/5 and any row whose population is missing
in that year).}
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion, \item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
or `NULL` for nominal only.} or `NULL` for nominal only.}
\item{basis}{`"harmonized"` (default) sums item codes through the
cross-vintage harmonization mapping (folding series-break-affected
codes onto a comparable target and excluding aggregate / discontinued
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
reproduces the pre-Phase-R2 behavior (published item codes, no
folding). On a corpus with `schema_version < 5` (no harmonization
tables), `basis` silently resolves to `"raw"` when left at its default
and the resolution is recorded in the provenance; explicitly passing
`basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
is set (see below).}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
multi-code cross-vintage series that a 1:1 harmonized_code mapping
can't express (e.g. a wide-era aggregate that only splits into leaf
codes in the modern era). Mutually exclusive with `category`. The
result's subtype column reads `"recipe"` and `category` reads the
recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
`basis` entirely (it joins `long` directly rather than going through
the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{expenditure_concept}{Which spending concept to return. Concepts are
defined as sets of the crosswalk's `spend_subtype` values -- never as
item-code first letters, which cannot classify correctly (prefix `Y`
alone spans revenue, expenditure, and balance codes):
* `"primary"` (default) -- the government's own service provision:
`operations` + `capital` + `assistance` subtypes.
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
`interest` (interest on debt) and `insurance_benefits` (insurance
trust benefit payments, e.g. pensions -- Census manual section
5.2.2.1 includes payments to retirees in Direct).
* `"total"` -- `direct` plus the intergovernmental leg: payments to
local governments (`M` codes), to the state government (`L` codes,
excluding the `L--` family-total rollup), and state payments to
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
`spend_subtype == "intergovernmental"`. Requires the active corpus's
`summary_categories` to carry M/L rows (added by cog_pipeline PR
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
older corpus rather than silently under-reporting. Mutually
exclusive with `recipe` (a recipe already defines its own component
codes).
**Do not sum `"total"` results across levels of government** (e.g.
state + county + city): a state's `M12` payment to a school district is
the same dollar the district reports as its own direct `E12`, so
summing both double-counts it. This matters in particular with
[cog_geographic_rollup()], which sums across exactly that kind of
multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
split), which the Direct leg excludes by construction but the IG leg
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
query, any (year, category) where this leaves intergovernmental rows
with NO Direct counterpart is flagged: the affected rows' `notes`
name the harmonization recipe that recovers the missing Direct
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
} }
\value{ \value{
Tibble with columns `year`, `canonical_govid`, `gov_name`, Tibble with columns `year`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`, `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
`codes_included`, `aggregate_fallback`, `notes`. Carries a `provenance` optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
attribute matching `inst/schemas/provenance-v1.json`. and `value_source` when `complete = TRUE`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
whose `completion` block reports `applied`, `rows_filled`, and the
per-year `absence_means` rule that was applied.
} }
\description{ \description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,285 @@
# Per-year population denominators in uscogdata
**Date:** 2026-04-29
**Status:** Design — pending implementation
**Scope:** uscogdata 0.1 (pre-release; no version bump)
**Related:** cog_pipeline (data dictionary updates)
## Problem
`uscogdata::cog_spending(per_capita = TRUE)` and `cog_revenue(per_capita = TRUE)`
currently divide every year's nominal amount by a single static population value
— `canonical_fips_xwalk.population_acs`, the ACS 2018-2022 5-year estimate.
For a 24-year corpus (2000–2023) this introduces a systematic bias proportional
to each government's population change over that span. Fast-growing places have
their early-year per-capita numbers understated; shrinking places have theirs
overstated. The bias commonly exceeds 20% and can exceed 50% for cities like
Detroit. Provenance currently advertises this denominator explicitly, so the
error is visible to careful users — but the default behavior produces wrong
numbers.
`cog_geographic_rollup()` has the same bug. `cog_find_peers()` /
`cog_peer_compare()` use the same static value to define peer cohorts, which
is defensible for matching but is no longer necessary now that per-year
population is available.
## Background — population sources
| Source | What it is | Where it lives |
|---|---|---|
| **Census F-33 `population`** | Population value Census uses on each COG row to compute its own per-capita tables. Almost always a Population Estimates Program (PEP) estimate; sometimes lagged a year for fiscal-year alignment, recorded in `popyear` | `long.population`, `long.popyear` (per row) |
| **PEP** (raw) | Census Bureau's official annual intercensal estimates. Distinct from F-33 because F-33 sometimes uses a lagged vintage | Not in corpus; available via tidycensus |
| **ACS 5-year** | American Community Survey 5-year rolling average. Different methodology, includes margin of error, only available 2005-2009 onward | `canonical_fips_xwalk.population_acs` (one fixed vintage) |
| **Decennial** | Actual count, every 10 years | Not in corpus |
F-33 `population` is the right default: it's what Census itself uses, so per-
capita results published by uscogdata reconcile with Census's own published
tables.
## Approach
Use the per-row `population` already present in `long`, joined on
`(canonical_govid, year)`. No new external data dependency. Coverage:
- **Types 0–3** (state, county, city, township): observed every year by design
- **Types 4–5** (special districts, schools): always NA — masked in
`cog_pipeline/R/read_modern.R` because the F-33 schema does not carry a
population value for these gov types
Type-4 and type-5 govs return `NA` per-capita with a `pop_source = "unavailable"`
flag and a note. No silent substitution.
The architecture leaves the door open for future denominators (PEP, ACS,
decennial) by surfacing `pop_source` as a first-class result column. Adding a
new source later is a join change, not an API change.
## Detailed design
### New view: `gov_population_yearly`
```sql
-- inst/sql/32-gov_population_yearly.sql
CREATE OR REPLACE VIEW gov_population_yearly AS
SELECT DISTINCT
year,
canonical_govid,
population,
popyear
FROM long
WHERE population IS NOT NULL;
```
`SELECT DISTINCT` collapses the metadata column duplicated across each gov-year's
item rows. A test asserts `(year, canonical_govid)` is unique to catch any
future source-data divergence.
### `cog_spending()` and `cog_revenue()`
`.attach_per_capita()` (in `R/spending.R`) is rewritten to:
1. Query `gov_population_yearly` for the requested govids and years.
2. `LEFT JOIN` on `(canonical_govid, year)` so missing rows produce NA.
3. Compute `amt_per_capita_nominal = amt_nominal / population`. NA when
population is NA.
4. Drop `population` from the returned tibble (keep `pop_source` instead).
Result tibble gains one new column when `per_capita = TRUE`:
- `pop_source`: `"census_f33"` when a denominator was found, `"unavailable"`
when NA.
`notes` is extended: when `pop_source == "unavailable"`, append
`"No population denominator available for this gov type"`. The `notes` column
is updated to concatenate multiple notes with `"; "` (it currently holds at
most one).
`amt_per_capita_real` is NA whenever `amt_per_capita_nominal` is NA.
### `cog_geographic_rollup()`
The current implementation does **not** sum amounts within a layer — it returns
one row per `(year, canonical_govid, subtype, category)` tagged with its
layer, intended for side-by-side "place portrait" comparisons (a city, the
county containing it, the state containing both). That semantics is preserved.
The only behavior change in this work is per-row exclusion when `per_capita = TRUE`:
1. After `cog_spending()` returns with the per-row per-year denominator from
Task 3, drop rows where `pop_source == "unavailable"` so the result never
contains NA per-capita rows.
2. Record the dropped `canonical_govid`s in `provenance$rollup$excluded_govids`
and the kept ones in `provenance$rollup$included_govids`.
Documentation states explicitly: *Per-capita rollups include only governments
observed in both the finance and population panels for the given year. Special
districts and school districts (gov types 4 and 5) are therefore excluded from
per-capita rollups by design.*
Provenance gains:
- `rollup.included_govids` — `canonical_govid`s present in the result
- `rollup.excluded_govids` — `canonical_govid`s dropped for missing pop
### `cog_find_peers()`
Signature: `cog_find_peers(target_govid, year = NULL, pop_range = c(0.5, 2), ...)`
- `year` is a single integer. When `NULL`, defaults to the most recent year
present in `gov_population_yearly` for the target.
- Looks up target's `population` at `year`. Errors if NA, with a message
listing nearby years where target *is* observed.
- Filters candidates by `gov_population_yearly.population` at the same `year`,
within `pop_range[1] * target_pop` and `pop_range[2] * target_pop`.
- Orders by `|log(pop_ratio)|` ascending.
Returned columns: `canonical_govid`, `gov_name`, `govs_type`, `fips_state`,
`population`, `pop_ratio`, `rank`. The column previously named `population_acs`
is renamed to `population`.
The cohort year is attached as a tibble attribute: `attr(x, "cohort_year")`.
### `cog_peer_compare()`
Existing signature unchanged:
`cog_peer_compare(target_govid, peers, category, years, per_capita = TRUE, adjust_to_year = NULL)`.
The caller supplies `peers` (either a `cog_find_peers()` result tibble or a
character vector of `canonical_govid`). The cohort year is implicit in
whichever year the caller used to call `cog_find_peers()`.
Behavior changes:
- When `peers` is a tibble carrying `attr(peers, "cohort_year")`,
`cog_peer_compare()` reads it and stamps every result row with a constant
`cohort_year` column.
- When `peers` is a bare character vector, `cohort_year` in the result is `NA`.
- Provenance gets `cohort_year` (scalar or NA) and the cohort govids list.
Users who want time-varying cohorts call `cog_find_peers()` per year and
stitch the `cog_peer_compare()` results themselves — documented in the
vignette with a worked example.
### Provenance updates
`provenance$transformations$per_capita` becomes:
```r
list(
applied = TRUE,
denominator_source = "Census F-33 population (per-year, from long.population)",
popyear_range = c(<min>, <max>),
pop_source_counts = list(census_f33 = N1, unavailable = N2)
)
```
For peer compare results, additional provenance:
```r
list(
cohort_year = <int>,
cohort_govids = <character>,
pop_range = c(<lo>, <hi>)
)
```
For rollup results, additional provenance:
```r
list(
rollup = list(
included_govids = <character>,
excluded_govids = <character>
)
)
```
`R/explain.R` is updated to render the new fields.
### Documentation
**New vignette** `vignettes/population-denominators.Rmd`:
1. The four population sources explained
2. Why F-33 is the default — and how it reconciles with Census's own per-capita
tables
3. The `popyear` quirk: Census sometimes uses a lagged estimate for fiscal-year
alignment. Recorded in provenance, not in the result.
4. Worked example showing the bias from the old static-ACS approach versus
per-year F-33 (e.g., Detroit 2003 vs. 2023)
5. Worked example of a rolling-cohort peer comparison built by looping
`cog_peer_compare()` per year
6. Future direction: `pop_source` is structured so PEP, ACS time-series, or
decennial denominators can be added later without API changes
**`cog_pipeline/docs/data_dictionary.md`** entry for `long.population` and
`long.popyear`: definition, source (F-33 fixed-width files, byte ranges),
type-4/5 masking rule, relationship to PEP.
### Tests
- `gov_population_yearly` returns one row per `(year, canonical_govid)` (uniqueness)
- `cog_spending(per_capita = TRUE)` returns different denominators for
different years for a known gov in the fixture (use any gov whose population
changes between 2019 and 2020)
- Type-4 and type-5 govids in the fixture return `pop_source = "unavailable"`
and `NA` per-capita with the expected note
- `cog_geographic_rollup(per_capita = TRUE)` excludes missing-pop govs and
records them in provenance
- `cog_find_peers()` defaults `year` to the most recent year for a target
with known population history
- `cog_find_peers()` errors with a helpful message when target has no observed
population in the requested year
- `cog_peer_compare()` defaults `cohort_year` and produces a result with a
constant `cohort_year` column
- Provenance carries `denominator_source`, `popyear_range`, and
`pop_source_counts`
- Regression test against a fixed govid+year showing the new per-capita value
differs from the old (static-ACS) by exactly the ratio of `population_acs`
to `long.population` for that gov-year
### Migration
Pre-release; no version bump. `NEWS.md` Unreleased entry:
> **Per-capita denominators now use per-year Census F-33 population.**
> Previously, `cog_spending()` and `cog_revenue()` divided all years' amounts
> by a single ACS 2018-2022 population, producing biased per-capita values
> for time-series. They now divide by the F-33 `population` recorded for each
> gov-year. Type-4 (special districts) and type-5 (school districts) govs
> return `NA` per-capita with `pop_source = "unavailable"`.
>
> **Peer matching now uses per-year population.** `cog_find_peers()` gains a
> `year` argument (defaults to most recent observed year). `cog_peer_compare()`
> gains `cohort_year`. Cohorts are still fixed for a single peer-compare call;
> users wanting moving cohorts loop themselves.
>
> **Rollups exclude govs with missing population.** `cog_geographic_rollup()`
> per-capita totals include only govs where both the finance variable and
> population are observed in that year; excluded govids are recorded in
> provenance.
>
> Returned column `population_acs` from `cog_find_peers()` is renamed to
> `population` and reflects the cohort-year vintage.
### File impact
| File | Change |
|---|---|
| `inst/sql/32-gov_population_yearly.sql` | New |
| `R/spending.R` (`.attach_per_capita`, `.notes_column`) | Per-year join, `pop_source`, multi-note concat |
| `R/peers.R` (`cog_find_peers`, `cog_peer_compare`) | `year` / `cohort_year` args, query new view, column rename |
| `R/rollup.R` | Skip-with-record for missing-pop govs |
| `R/provenance.R` | New denominator/cohort/rollup fields |
| `R/explain.R` | Render new fields |
| `vignettes/population-denominators.Rmd` | New |
| `tests/testthat/` | Per-year denominator, type-4/5, rollup exclusion, peer cohort, provenance |
| `cog_pipeline/docs/data_dictionary.md` | Document `long.population`, `long.popyear`, masking |
| `NEWS.md` | Unreleased entry |
## Out of scope
- PEP/ACS/decennial denominators — architected for, not implemented
- `per_pupil` denominator using `long.enrollment` for type-5 — deferred
- Covering-county fallback for type-4 — deliberately not done
- Backfilling population for type-4/5 from any external source
- Changes to `cog_explorer` callers — separate follow-up, after this lands
+128
View File
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
if (nzchar(p)) paste0(p, "/") else "" if (nzchar(p)) paste0(p, "/") else ""
} }
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
# or "" when it isn't there.
#
# Tests that assert on documentation content have to read the sources, and the
# sources only exist when the suite runs from a checkout. Under R CMD check the
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
# shipped and `../../README.md` does not resolve -- so those tests must skip
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
# rcmdcheck, so the assertions are still enforced on every push; this only
# stops them from failing a context that structurally cannot satisfy them.
source_tree_path <- function(...) {
p <- testthat::test_path("..", "..", ...)
if (file.exists(p)) p else ""
}
# Skip unless every named source file is present (see source_tree_path()).
skip_if_no_source_tree <- function(...) {
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
character(1))
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
testthat::skip_if(
any(missing),
"package source tree not available (running against the installed package)"
)
invisible(paths)
}
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL). # Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
skip_if_no_corpus <- function() { skip_if_no_corpus <- function() {
p <- fixture_corpus_path() p <- fixture_corpus_path()
@@ -26,3 +53,104 @@ with_fixture_corpus <- function(code) {
}, add = TRUE) }, add = TRUE)
force(code) force(code)
} }
# Copy the bundled fixture to a temp dir with manifest.json's schema_version
# patched to `version`, then run `code` against it with a clean session
# (mirrors with_fixture_corpus()). Used to exercise the v4/v5 dual-accept
# path without a second physical fixture tree: a real v4 corpus has no
# harmonization_map/harmonization_recipes/series_breaks parquet files, but
# .register_views() only *reads* those when schema_version >= 5 (see
# R/views.R), so a doctored copy of the (v5) bundled fixture with the
# manifest's schema_version knocked down to 4 is a faithful stand-in.
with_doctored_schema_version <- function(version, code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
manifest_path <- file.path(tmp, "manifest.json")
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
m$schema_version <- as.integer(version)
writeLines(
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
manifest_path
)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
# Copy the bundled fixture to a temp dir with representation.parquet and
# code_set.parquet removed (and dropped from the manifest's metadata list),
# then run `code` against it. Models a corpus published BEFORE sparsification:
# schema_version is left alone deliberately, because it was never bumped for
# that change -- the pre-sparsification fixture this package shipped until
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
# is therefore the only honest signal, and this helper is what proves the
# package keys off it rather than off the version number.
with_corpus_missing_representation <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
dropped <- c("representation.parquet", "code_set.parquet")
file.remove(file.path(tmp, "data", dropped))
manifest_path <- file.path(tmp, "manifest.json")
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
m$files$metadata <- Filter(
function(f) !basename(f$path) %in% dropped, m$files$metadata
)
writeLines(
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
manifest_path
)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
# Copy the bundled fixture to a temp dir with summary_categories.parquet
# rewritten to drop every M/L (intergovernmental) row, then run `code`
# against it with a clean session (mirrors with_fixture_corpus()/
# with_doctored_schema_version()). Models a real pre-cog_pipeline-PR#59
# corpus: the 66 M/L category rows shipped with NO schema_version bump (see
# C2 in the expenditure-concept review), so schema_version is left
# untouched here -- only the category data itself is rolled back.
with_corpus_missing_ig_categories <- function(code) {
src <- fixture_corpus_path()
tmp <- withr::local_tempdir(.local_envir = parent.frame())
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf(
"COPY (SELECT * FROM read_parquet(%s) WHERE LEFT(item_code, 1) NOT IN ('M', 'L'))
TO %s (FORMAT PARQUET)",
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
))
file.remove(cats_path)
file.rename(filtered_path, cats_path)
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
uscogdata:::cog_close()
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
on.exit({
uscogdata:::cog_close()
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
}, add = TRUE)
force(code)
}
+44
View File
@@ -0,0 +1,44 @@
# Helper for the Madison-walkthrough finding tests (uscogdata #11-#16).
#
# Those tests all assert something about what a `cog_*` verb includes or
# excludes. The expected amounts must therefore come from the RAW corpus, never
# from the verb under test: verifying an absence through the filter that creates
# it proves nothing. `wt_raw_*()` opens its own DuckDB connection straight onto
# the corpus's `long` parquet partitions, bypassing uscogdata's SQL views (and
# therefore its `flow_prefixes` filtering) entirely.
wt_corpus_glob <- function() {
url <- Sys.getenv("USCOGDATA_URL")
if (!nzchar(url)) testthat::skip("USCOGDATA_URL is not set")
paste0(sub("/$", "", url), "/data/long/**/*.parquet")
}
wt_raw_query <- function(sql) {
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbGetQuery(con, sql)
}
# Sum of `amt` (in $1,000s, as the corpus stores it) for one government-year,
# restricted either to an explicit set of item codes or to a set of first-letter
# prefixes. Aggregate rows are excluded, matching every published verb.
wt_raw_amt <- function(govid, year, codes = NULL, prefixes = NULL) {
stopifnot(xor(is.null(codes), is.null(prefixes)))
filter_sql <- if (!is.null(codes)) {
paste0("item_code IN (", paste0("'", codes, "'", collapse = ", "), ")")
} else {
paste0("LEFT(item_code, 1) IN (", paste0("'", prefixes, "'", collapse = ", "), ")")
}
out <- wt_raw_query(paste0(
"SELECT COALESCE(SUM(amt), 0) AS amt FROM read_parquet('", wt_corpus_glob(), "') ",
"WHERE canonical_govid = '", govid, "' AND year = ", year,
" AND NOT is_aggregate AND ", filter_sql
))
out$amt[[1]]
}
# The item codes a verb reports having summed, flattened out of the
# comma-separated `codes_included` column.
wt_codes_included <- function(df) {
sort(unique(trimws(unlist(strsplit(stats::na.omit(df$codes_included), ",")))))
}
+7
View File
@@ -41,3 +41,10 @@ test_that(".inflate preserves NA amounts", {
expect_true(is.na(result[2])) expect_true(is.na(result[2]))
expect_false(any(is.na(result[c(1, 3)]))) expect_false(any(is.na(result[c(1, 3)])))
}) })
test_that("bundled CPI covers the full 1967+ corpus era through this year", {
cpi <- .cpi_table()
expect_lte(min(cpi$year), 1967L)
expect_gte(max(cpi$year), as.integer(format(Sys.Date(), "%Y")))
expect_false(any(is.na(cpi$cpi)))
})
@@ -0,0 +1,52 @@
# Madison walkthrough audit -- finding F-004. Tracked as uscogdata#15.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# The raw Census files report thousands of dollars; this package multiplies by
# 1000 and returns full US dollars. That is the friendlier choice and is not
# wrong -- but cog_explorer's CLAUDE.md states "All raw `amt` values are in
# $1,000s", so a reader who applies that rule to amt_nominal overstates every
# figure by 1000x, and gets a plausible-looking number rather than an obvious
# error. The audit rates this the highest-consequence definitional gap it found.
#
# Deliberately NOT asserted here: man/cog_spending.Rd and man/cog_revenue.Rd,
# which ALREADY carry the statement in their @return sections (verified
# 2026-07-29), as does cog-api's data-dictionary.md (since 2b71b41). The gap is
# in the surfaces a reader meets first and in cog_explorer's own conventions
# doc -- see uscogdata#15 for the full surface-by-surface table and for the two
# secondary tasks (cog_explorer/CLAUDE.md, which has no git remote, and
# cog-api's llms.txt, which is silent on units).
test_that("returned amounts are documented as full US dollars where readers meet the package", {
# README and vignettes ship only in the source tree, not in the installed
# package, so these assertions cannot run under R CMD check -- CI's earlier
# testthat::test_local() step is what enforces them. See
# skip_if_no_source_tree() in helper-fixture.R.
docs <- skip_if_no_source_tree(
"README.md",
c("vignettes", "total-spending.Rmd"),
c("vignettes", "population-denominators.Rmd")
)
says_units <- function(path) {
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
grepl("full US dollars|full U\\.S\\. dollars", txt, ignore.case = TRUE) &&
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
}
for (path in docs) expect_true(says_units(path))
# Pin the documented claim to the actual behaviour, so the two cannot drift.
# The expected raw amount is read straight from the corpus's parquet
# partitions -- never through cog_spending(), which is the thing being
# described. Madison FY2020: E/F/G = 623,347 ($1,000s) -> $623,347,000.
raw_thousands <- wt_raw_amt("552025209777", 2020L, prefixes = c("E", "F", "G"))
expect_equal(raw_thousands, 623347)
returned <- cog_spending(govid = "552025209777", years = 2020L)
expect_equal(sum(returned$amt_nominal), raw_thousands * 1000)
units <- attr(returned, "provenance")$transformations$units_conversion
expect_true(units$applied)
expect_equal(units$multiplier, 1000)
})
+36 -3
View File
@@ -8,22 +8,55 @@ test_that("cog_categories returns all categories grouped by subtype", {
expect_gt(nrow(r), 10L) expect_gt(nrow(r), 10L)
# corpus preserves Census-native "expenditure" vocabulary; the API takes # corpus preserves Census-native "expenditure" vocabulary; the API takes
# "spending" as a friendlier alias. # "spending" as a friendlier alias.
expect_setequal(unique(r$category_type), c("expenditure", "revenue")) #
# `balance` joined as a third category_type with the cash-and-security
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
# money verb, so it surfaces every category_type the corpus carries -- the
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
# return a balance row.
expect_setequal(unique(r$category_type),
c("expenditure", "revenue", "balance"))
}) })
test_that("cog_categories(type = 'spending') returns only expenditure rows", { test_that("cog_categories(type = 'spending') returns only expenditure rows", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories(type = "spending") r <- cog_categories(type = "spending")
expect_true(all(r$category_type == "expenditure")) expect_true(all(r$category_type == "expenditure"))
expect_true(all(r$subtype %in% c("operations", "capital"))) # "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
# carrying dollars now maps to a category.
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
# joined with the I/Q/Y flow batch -- the last two characters of Census's
# expenditure taxonomy. `interest` is what makes the three-concept model
# computable: primary = direct minus debt service.
expect_true(all(r$subtype %in%
c("operations", "capital", "intergovernmental", "assistance",
"interest", "insurance_benefits")))
})
test_that("cog_categories surfaces the intergovernmental spending subtype", {
skip_if_no_corpus()
r <- cog_categories(type = "spending")
expect_true("intergovernmental" %in% r$subtype)
# IG rows reuse the existing functional categories -- they add a subtype,
# not new category values.
ig_cats <- sort(unique(r$category[r$subtype == "intergovernmental"]))
direct_cats <- sort(unique(r$category[r$subtype != "intergovernmental"]))
expect_true(all(ig_cats %in% c(direct_cats, "Other Education")))
}) })
test_that("cog_categories(type = 'revenue') returns only revenue rows", { test_that("cog_categories(type = 'revenue') returns only revenue rows", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_categories(type = "revenue") r <- cog_categories(type = "revenue")
expect_true(all(r$category_type == "revenue")) expect_true(all(r$category_type == "revenue"))
# The four non-general subtypes are deliberately NOT own_source: Census's
# General Revenue excludes insurance trust (Y01 alone is $1.31T corpus-wide,
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
# (A90) revenue by definition, which is what makes both of its published
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
expect_true(all(r$subtype %in% expect_true(all(r$subtype %in%
c("own_source", "federal", "state", "local_aid"))) c("own_source", "federal", "state", "local_aid",
"insurance_trust", "utility", "liquor_store")))
}) })
test_that("cog_categories(pattern = ...) filters case-insensitively", { test_that("cog_categories(pattern = ...) filters case-insensitively", {
+193
View File
@@ -0,0 +1,193 @@
# tests/testthat/test-complete.R
#
# uscogdata#18. The published corpus no longer stores the wide era's explicit
# zeros (cog_pipeline#64, series break SB194), so absence means two different
# things:
#
# <= FY2011 (dense_source) : cell absent => Census published $0
# >= FY2012 (sparse_source): cell absent => not reported, unknown
#
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
# row's `value_source` so the two are distinguishable. Expected row sets here
# are built from the corpus parquet directly, never from the verb under test --
# verifying what a filter does through that same filter proves nothing.
# The (subtype, category) cells that SHOULD exist for one government-year:
# every code in force for that government's type, mapped through
# summary_categories, matching the verb's crosswalk subtype scope (the
# default concept, `primary`, is operations/capital/assistance -- see
# uscogdata#11) and excluding aggregate-flagged codes (which
# spending_long/revenue_long drop).
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
wt_raw_query(sprintf(
"SELECT DISTINCT c.%s AS subtype, c.category
FROM %s cs
JOIN %s x ON x.govs_type = cs.type
JOIN %s c ON c.item_code = cs.item_code
WHERE x.canonical_govid = '%s'
AND cs.year = %d
AND NOT cs.is_aggregate
AND c.category IS NOT NULL
AND c.%s IN (%s)",
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
q("summary_categories.parquet"), govid, year,
subtype_col, paste0("'", subtypes, "'", collapse = ",")
))
}
# The default expenditure concept's subtype scope, mirrored from
# R/spending.R's .spend_subtypes_primary.
primary_subtypes <- c("operations", "capital", "assistance")
test_that("complete = FALSE is the default and changes nothing", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
expect_equal(nrow(plain), nrow(explicit))
expect_false("value_source" %in% names(plain))
})
})
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2011 is dense_source: before sparsification this government carried a
# row for every code in force, most of them $0. complete = TRUE must
# reproduce that cell set exactly.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$spend_subtype, r$category),
key(expected$subtype, expected$category))
expect_gt(nrow(expected), 0L)
# Every filled cell in a dense-source year is a Census-published $0 --
# never "unknown", which is what the modern era's absences mean.
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
})
})
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
skip_if_no_corpus()
with_fixture_corpus({
plain <- cog_spending("121011212191", 2011L)
full <- cog_spending("121011212191", 2011L, complete = TRUE)
# Filling adds rows; it must never alter or drop one.
expect_gt(nrow(full), nrow(plain))
reported <- full[full$value_source == "reported", ]
expect_equal(nrow(reported), nrow(plain))
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
# ... and the total is unchanged, because every added cell is $0.
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
})
})
test_that("a sparse-source year's absences are unknown, not zero", {
skip_if_no_corpus()
with_fixture_corpus({
# FY2019 is sparse_source: an absent cell means the government did not
# report, which is NOT a zero. Filling those with 0 would invent data --
# the exact error the representation contract exists to prevent.
r <- cog_spending("121011212191", 2019L, complete = TRUE)
filled <- r[r$value_source != "reported", ]
expect_gt(nrow(filled), 0L)
expect_true(all(filled$value_source == "not_reported"))
expect_true(all(is.na(filled$amt_nominal)))
expect_false(any(r$value_source == "census_zero"))
})
})
test_that("the fill is scoped to each government's own type", {
skip_if_no_corpus()
with_fixture_corpus({
# Filling against the union of all types would invent cells for codes a
# county can never report. Every filled category must be one that
# code_set puts in force for type 1 (county) specifically.
r <- cog_spending("121011212191", 2011L, complete = TRUE)
county_cells <- raw_expected_cells("121011212191", 2011L,
primary_subtypes, "spend_subtype")
expect_true(all(r$category %in% county_cells$category))
})
})
test_that("complete = TRUE respects the category filter", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011L, category = "Police",
complete = TRUE)
expect_true(all(r$category == "Police"))
expect_true("value_source" %in% names(r))
})
})
test_that("cog_revenue() completes on its own flow", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
expected <- raw_expected_cells("121011212191", 2011L,
c("own_source", "federal", "state", "local_aid"),
"revenue_subtype")
key <- function(sub, cat) paste(sub, cat, sep = "|")
expect_setequal(key(r$revenue_subtype, r$category),
key(expected$subtype, expected$category))
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
})
})
test_that("provenance records the completion and its absence rule", {
skip_if_no_corpus()
with_fixture_corpus({
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
"provenance")
expect_true(prov$completion$applied)
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
expect_gt(prov$completion$rows_filled, 0L)
off <- attr(cog_spending("121011212191", 2011L), "provenance")
expect_false(off$completion$applied)
expect_equal(off$completion$rows_filled, 0L)
})
})
test_that("complete = TRUE is refused where the fill would be guesswork", {
skip_if_no_corpus()
with_fixture_corpus({
# A recipe defines its own component codes and does not go through
# summary_categories at all, so there is no grid to fill from.
expect_error(
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
# The intergovernmental leg keeps aggregate rows by design
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
expect_error(
cog_spending("121011212191", 2011L, expenditure_concept = "total",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
})
})
test_that("complete = TRUE aborts on a corpus with no representation contract", {
skip_if_no_corpus()
# A corpus published before sparsification carries neither table, so there
# is nothing to fill from and no rule saying what an absence means. That
# must abort rather than guess.
with_corpus_missing_representation({
expect_error(
cog_spending("121011212191", 2011L, complete = TRUE),
class = "uscogdata_representation_unavailable"
)
# ... while an ordinary query on the same corpus still works.
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
})
})
+38
View File
@@ -26,3 +26,41 @@ test_that(".resolve_cache_dir falls back to R_user_dir", {
}) })
}) })
}) })
# ---------------------------------------------------------------------------
# Trailing-slash normalization (uscogdata #3 follow-up).
#
# EVERY consumer builds paths by concatenation: paste0(url, "manifest.json")
# (manifest.R), paste0(url, e$path) (mirror.R), and the parquet glob in
# views.R. mirror.R:104 even comments 'url ends in "/"' -- an assumption the
# package documents and relies on but never enforced.
#
# A URL missing its trailing slash therefore fails SILENTLY and confusingly:
# HTTPS -> ".../downloadmanifest.json" -> the host answers with an HTML 404
# page -> the jsonlite lexical error that issue #3 reported;
# local -> ".../corpusdata/long/**/*.parquet" -> DuckDB "No files found".
# Neither message points at the real cause. Normalize once, at resolution.
# ---------------------------------------------------------------------------
test_that(".resolve_url appends a missing trailing slash", {
withr::local_envvar(USCOGDATA_URL = "https://example.org/s/TOKEN/download")
expect_equal(.resolve_url(), "https://example.org/s/TOKEN/download/")
})
test_that(".resolve_url leaves an existing trailing slash alone", {
withr::local_envvar(USCOGDATA_URL = "https://example.org/s/TOKEN/download/")
expect_equal(.resolve_url(), "https://example.org/s/TOKEN/download/")
})
test_that(".resolve_url normalizes a local path without a trailing slash", {
withr::local_envvar(USCOGDATA_URL = "/tmp/corpus")
expect_equal(.resolve_url(), "/tmp/corpus/")
})
test_that(".resolve_url does not invent a slash for an empty setting", {
# An unset/empty URL must stay empty so the "not configured" guard in
# manifest.R still fires, rather than degrading into a bare "/" root.
withr::local_envvar(USCOGDATA_URL = "")
withr::local_options(uscogdata.url = "")
expect_equal(.resolve_url(), "")
})
+94
View File
@@ -0,0 +1,94 @@
# tests/testthat/test-corpus-breaks.R
#
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
# are caveats about the corpus itself rather than about one item code:
#
# SB085 1977 dollar precision across the 1976/1977 boundary
# SB087 2002 imputation exclusion FY2002-2006
# SB194 2012 dense -> sparse representation change
# SB086 2017 government ID scheme change
#
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
# and no row's item_code is ever the literal "ALL", so none of them could
# ever reach a user. They now travel in their own provenance field,
# `corpus_break_refs`, which keeps them distinguishable from the
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
# result, not one series).
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
skip_if_no_corpus()
with_fixture_corpus({
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
# DoD 4 intended to reach users.
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_true("SB194" %in% prov$corpus_break_refs)
})
})
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
skip_if_no_corpus()
with_fixture_corpus({
# 2019-2020 spans no catalogued corpus-wide break.
r <- cog_spending("121011212191", 2019:2020, "Police")
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
})
})
test_that("corpus_break_refs and series_break_refs stay disjoint", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
prov <- attr(r, "provenance")
expect_type(prov$series_break_refs, "character")
expect_type(prov$corpus_break_refs, "character")
# An ALL caveat must never masquerade as a break in a specific series.
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
expect_false("SB194" %in% prov$series_break_refs)
})
})
test_that(".build_corpus_break_refs matches on the break_year window alone", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
# series_breaks table is a full cross-vintage registry, so the matching
# logic is testable there even though no long partition covers it.
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1975:1980, schema_version = 6L
))
# ... and does not fire for a range that misses it, unlike a filter keyed
# on the era rather than the boundary.
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
con, years = 1978:1980, schema_version = 6L
))
# Unlike code-specific refs, these do not depend on which codes a result
# happens to contain -- that dependency is the whole defect.
expect_setequal(
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
"SB087"
)
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
expect_equal(
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
character(0)
)
})
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2011:2012, "Police")
out <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
expect_match(out, "SB194", fixed = TRUE)
})
})
+103
View File
@@ -0,0 +1,103 @@
# Madison walkthrough audit -- findings F-020 and F-023. Tracked as uscogdata#13.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# The owner's settled design (2026-07-28): a `coverage` argument on
# cog_geographic_rollup(), cog_find_peers()/cog_peer_compare() and their
# cog-api equivalents --
# "all" every unit that reported that year (today's behaviour, DEFAULT)
# "census" census years only (years ending 2 or 7)
# "consistent" only units reporting in every requested year (balanced panel)
# -- PLUS always-on coverage metadata on every result regardless of mode:
# n_units_reporting, n_units_expected, is_census_year.
#
# Motivating principle: using these verbs correctly must not require the user to
# know that the Census of Governments is a complete census only in years ending
# in 2 and 7.
#
# The helper below accepts that metadata either as columns on the returned
# tibble or as a per-year table in provenance$coverage -- the design fixes the
# three field names and that they reach the caller, not the container.
wt_coverage <- function(x) {
prov <- attr(x, "provenance")
cov <- prov$coverage
if (is.null(cov)) {
needed <- c("year", "n_units_reporting", "n_units_expected", "is_census_year")
expect_true(all(needed %in% names(x)))
cov <- unique(x[, needed])
}
cov[order(cov$year), ]
}
test_that("multi-government aggregates disclose reporting coverage on every result", {
# -- F-020: geographic rollups -------------------------------------------
# Wisconsin's city/village universe is 608 governments. On the bundled
# fixture, FY2012 (a census year) has 597 of them reporting while FY2019 and
# FY2020 (sample years) have 112 and 114 -- an 18%-98% swing that today's
# return value says nothing about. Counts cross-checked against the raw
# corpus, not through cog_geographic_rollup(), which is under test.
wi <- cog_gov_search(name = NULL, state = "WI", type = "city")
expect_equal(nrow(wi), 608L)
roll <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
category = NULL, years = c(2011L, 2012L, 2019L, 2020L))
cov <- wt_coverage(roll)
expect_equal(cov$n_units_expected, rep(608L, 4L))
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
# Cross-check against the raw partitions, scoped to the SAME universe the
# rollup was given -- the 608 govids above. Scoping instead on the long
# table's own `type`/`fips_state` asks a different question and answers 595:
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
# identity, when they were townships) while the xwalk lists them as
# govs_type = 2 (their present identity, as villages). Schema v6 made the
# long table's geography present-harmonized and moved as-of-year to the
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
# in R/manifest.R. n_units_reporting counts against the requested universe,
# so 597 is the number that answers "how many of the governments I asked
# about reported".
raw_2012 <- wt_raw_query(paste0(
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
"AND canonical_govid IN (",
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
# -- F-023: peer cohorts --------------------------------------------------
# CHILTON CITY, WI (ACS population 4,017): a 15-peer cohort fixed at FY2012
# reports 15 of 15 in FY2012 and only 3 of 15 in FY2019 and FY2020. Nothing
# in cog_peer_compare()'s return distinguishes those years today.
chilton <- "552015177095"
peers <- cog_find_peers(chilton, year = 2012L, max_peers = 15L)
expect_equal(nrow(peers), 15L)
cmp <- cog_peer_compare(target_govid = chilton, peers = peers, category = NULL,
years = c(2012L, 2019L, 2020L), per_capita = TRUE)
cov_peers <- wt_coverage(cmp)
expect_equal(cov_peers$n_units_expected, rep(15L, 3L))
expect_equal(cov_peers$n_units_reporting, c(15L, 3L, 3L))
expect_equal(cov_peers$is_census_year, c(TRUE, FALSE, FALSE))
# -- the three coverage modes --------------------------------------------
expect_equal(attr(cog_peer_compare(target_govid = chilton, peers = peers,
category = NULL, years = c(2012L, 2019L, 2020L),
per_capita = TRUE),
"provenance")$coverage_mode, "all") # unchanged default
consistent <- cog_peer_compare(target_govid = chilton, peers = peers,
category = NULL, years = c(2012L, 2019L, 2020L),
per_capita = TRUE, coverage = "consistent")
n_by_year <- tapply(consistent$canonical_govid[consistent$role == "peer"],
consistent$year[consistent$role == "peer"],
function(g) length(unique(g)))
expect_equal(unname(as.integer(n_by_year)), c(3L, 3L, 3L)) # balanced panel
census_only <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
category = NULL,
years = c(2011L, 2012L, 2019L, 2020L),
coverage = "census")
expect_equal(sort(unique(census_only$year)), 2012)
})
+553
View File
@@ -0,0 +1,553 @@
test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
con <- .ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM long WHERE LEFT(item_code, 1) = 'K'")$n
expect_equal(n, 0)
sql_files <- c("20-spending_long.sql", "22-spending_long_harmonized.sql")
for (f in sql_files) {
txt <- paste(readLines(system.file("sql", f, package = "uscogdata")),
collapse = " ")
expect_false(grepl("'K'", txt, fixed = TRUE),
label = paste(f, "must not reference the inert K prefix"))
}
})
test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
gov <- "010000226085" # Alabama state government
base <- cog_spending(gov, years = 2019, category = "Police")
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
# Police maps only to operations/capital codes (E62/F62/G62), so the
# direct concept's extra subtypes (interest, insurance_benefits) cannot
# contribute and the two concepts must agree exactly here.
expl <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "direct")
expect_equal(base$amt_nominal, expl$amt_nominal)
expect_false("intergovernmental" %in% base$spend_subtype)
})
test_that("expenditure_concept = 'total' adds an intergovernmental subtype", {
gov <- "010000226085"
d <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending(gov, years = 2019, category = "Police",
expenditure_concept = "total")
expect_true("intergovernmental" %in% t$spend_subtype)
# Direct rows are untouched; Total only ever ADDS. Use %in% rather than
# != : a category = NULL result can contain a NULL-subtype group (codes
# with no summary_categories row, e.g. E16/E21/E85/F16/F85/G16/G21/G85),
# and `NA != "intergovernmental"` is NA, not TRUE, which would silently
# smuggle an all-NA phantom row into dt.
dt <- t[!(t$spend_subtype %in% "intergovernmental"), ]
expect_equal(sort(dt$amt_nominal), sort(d$amt_nominal))
expect_gt(sum(t$amt_nominal), sum(d$amt_nominal))
})
test_that("legacy-era Total does not collapse to Direct (the is_aggregate trap)", {
# In the wide era the IG dollars live almost entirely on aggregate-flagged
# rows. A Total leg that inherited the Direct leg's NOT is_aggregate filter
# would silently return Total == Direct here.
gov <- "010000226085"
d <- cog_spending(gov, years = 2011, category = "Education K-12",
expenditure_concept = "direct")
t <- cog_spending(gov, years = 2011, category = "Education K-12",
expenditure_concept = "total")
expect_true("intergovernmental" %in% t$spend_subtype)
ig <- sum(t$amt_nominal[t$spend_subtype == "intergovernmental"])
expect_gt(ig, 0)
expect_gt(sum(t$amt_nominal), sum(d$amt_nominal))
})
test_that("the IG leg never includes the L-- family total", {
con <- .ensure_session()
codes <- DBI::dbGetQuery(con,
"SELECT DISTINCT item_code FROM ig_long")$item_code
expect_false(any(grepl("--$", codes)))
# Q joined the IG family with the crosswalk-membership rewrite
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
})
test_that("expenditure_concept rejects unknown values", {
expect_error(
cog_spending("010000226085", years = 2019, expenditure_concept = "gross"),
class = "rlang_error"
)
})
test_that("total composes with basis = 'raw' and basis = 'harmonized'", {
gov <- "010000226085"
h <- cog_spending(gov, years = 2011, category = "Education K-12",
expenditure_concept = "total", basis = "harmonized")
r <- cog_spending(gov, years = 2011, category = "Education K-12",
expenditure_concept = "total", basis = "raw")
ig_h <- sum(h$amt_nominal[h$spend_subtype == "intergovernmental"])
ig_r <- sum(r$amt_nominal[r$spend_subtype == "intergovernmental"])
# The only IG harmonization rule is M38 -> M36 (year-disjoint), so the IG
# total must agree between bases even though the code labels may differ.
expect_equal(ig_h, ig_r)
})
test_that("recipe = and expenditure_concept = 'total' together aborts", {
expect_error(
cog_spending("121011212191", 2020L, recipe = "corrections_combined",
expenditure_concept = "total"),
class = "uscogdata_recipe_concept_conflict"
)
})
test_that("aggregate-sourced IG dollars are flagged aggregate_fallback = TRUE (bool_or, not bool_and)", {
# Regression test: .build_verb_sql() originally used bool_and(is_aggregate)
# for aggregate_fallback, which is correct for the Direct leg (a group can
# never mix aggregate and non-aggregate rows there -- spending_long filters
# NOT is_aggregate) but wrong for the IG leg. The wide era is dense -- every
# government has a $0 row for every code in a family -- so a $0 leaf sits in
# the same (year, gov, subtype, category) group as the real aggregate row
# and flips bool_and() to FALSE. Measured: AL state 2011 had $5,740,775,000
# of aggregate-sourced IG dollars (Corrections $31,358,000 + Education K-12
# $5,152,385,000 + General Government $557,032,000) reporting
# aggregate_fallback = FALSE under bool_and(), with the only TRUE row being
# Transit Utilities at $0. bool_or() reports all of them correctly.
gov <- "010000226085"
t <- cog_spending(gov, years = 2011, category = "Education K-12",
expenditure_concept = "total")
ig <- t[t$spend_subtype == "intergovernmental", ]
expect_equal(nrow(ig), 1L)
expect_true(ig$aggregate_fallback)
expect_true(nzchar(ig$notes))
expect_match(ig$notes, "Aggregate fallback applied", fixed = TRUE)
})
test_that("legacy aggregate IG codes are year-disjoint from their modern leaf components", {
# The safety of ig_long's deliberate omission of `NOT is_aggregate` (see
# inst/sql/24-ig_long.sql) rests entirely on each legacy code's AGGREGATE
# instance being year-disjoint from the modern leaf codes it rolls up --
# if a future corpus rebuild ever back-filled a leaf into a year where the
# code is still flagged aggregate, `total` would silently double-count and
# this suite would still pass. This test fails loudly if that ever
# happens.
#
# Note the invariant is scoped to the AGGREGATE flag, not bare code
# presence: M89/L89 do NOT disappear after the wide era the way M47/L47
# do -- they continue past 2011 as their OWN independent leaf line item
# (is_aggregate = FALSE) alongside M91-93/L91-93, which is fine because a
# non-aggregate M89/L89 no longer represents a rollup of those codes.
# (Verified in the fixture: M89/L89 are is_aggregate = TRUE only in 2011,
# when M91-93/L91-93 don't exist yet; from 2012 on M89/L89 are
# is_aggregate = FALSE leaves coexisting with M91-93/L91-93.)
#
# Pairs are the M/L-prefixed components (this package's ig_long only
# covers M/L; other prefixes in the same rollup, e.g. N/O/P/Q/R, fall
# outside its domain and are irrelevant here) enumerated in
# cog_pipeline's data/wide_to_long_xwalk.csv `full_desc` column (read
# once at authoring time, not at test time -- this test stays offline):
# M47 "To local governments, total (includes N47, O47, P47, R47, and M94)"
# M89 "To local governments, total (incl N89, O89, P89, R89, M91, M92, and M93)"
# L47 "To state government (includes L94)"
# L89 "To state government (includes L91, L92, and L93)"
con <- .ensure_session()
pairs <- list(
list(aggregate = "M47", components = "M94"),
list(aggregate = "M89", components = c("M91", "M92", "M93")),
list(aggregate = "L47", components = "L94"),
list(aggregate = "L89", components = c("L91", "L92", "L93"))
)
agg_years_by_code <- DBI::dbGetQuery(con,
"SELECT DISTINCT year, item_code FROM ig_long WHERE is_aggregate")
codes_by_year <- DBI::dbGetQuery(con, "SELECT DISTINCT year, item_code FROM ig_long")
for (p in pairs) {
agg_years <- agg_years_by_code$year[agg_years_by_code$item_code == p$aggregate]
for (yr in agg_years) {
codes_yr <- codes_by_year$item_code[codes_by_year$year == yr]
has_component <- any(p$components %in% codes_yr)
expect_false(
has_component,
label = sprintf(
"year %s has aggregate-flagged %s co-occurring with a modern component (%s)",
yr, p$aggregate, paste(p$components, collapse = ",")
)
)
}
}
})
test_that(".verb_spendrev rejects expenditure_concept = 'total' for a non-spending view_base", {
# cog_revenue() never exposes expenditure_concept and always resolves it
# to the "direct" default, so there is no revenue codepath that reaches
# this today -- but .verb_spendrev() is shared, and nothing else stops a
# future caller from passing expenditure_concept = "total" alongside
# view_base = "revenue_annotated", which would UNION expenditure M/L rows
# into a revenue result. Exercise the internal helper directly.
expect_error(
uscogdata:::.verb_spendrev(
verb = "cog_revenue_test", view_base = "revenue_annotated",
subtype_col = "revenue_subtype",
flow_prefixes = c("T", "A", "U", "B", "C", "D"),
call = quote(cog_revenue_test()),
govid = "010000226085", years = 2019L, category = NULL,
per_capita = FALSE, adjust_to_year = NULL, basis = "raw",
recipe = NULL, expenditure_concept = "total"
),
class = "uscogdata_expenditure_concept_unsupported"
)
})
test_that("cog_geographic_rollup refuses expenditure_concept = 'total'", {
expect_error(
cog_geographic_rollup(
govids = list(state = "010000226085"),
category = "Police", years = 2019,
expenditure_concept = "total"
),
class = "uscogdata_concept_not_aggregatable"
)
})
test_that("cog_peer_compare refuses expenditure_concept = 'total'", {
expect_error(
cog_peer_compare(
target_govid = "010000226085", peers = "010000226085",
category = "Police", years = 2019,
expenditure_concept = "total"
),
class = "uscogdata_concept_not_aggregatable"
)
})
test_that("the refusal message names the fix and the reason", {
err <- tryCatch(
cog_geographic_rollup(govids = list(state = "010000226085"),
category = "Police", years = 2019,
expenditure_concept = "total"),
condition = function(e) e
)
msg <- paste(conditionMessage(err), collapse = " ")
expect_match(msg, "direct")
expect_match(msg, "double-count|double count")
expect_match(msg, "cog_geographic_rollup")
# Test that cog_peer_compare's message names its own function
err2 <- tryCatch(
cog_peer_compare(target_govid = "010000226085", peers = "010000226085",
category = "Police", years = 2019,
expenditure_concept = "total"),
condition = function(e) e
)
msg2 <- paste(conditionMessage(err2), collapse = " ")
expect_match(msg2, "direct")
expect_match(msg2, "double-count|double count")
expect_match(msg2, "cog_peer_compare")
})
test_that("both cross-government verbs still accept the direct default", {
expect_no_error(
cog_geographic_rollup(govids = list(state = "010000226085"),
category = "Police", years = 2019)
)
expect_no_error(
cog_peer_compare(target_govid = "010000226085", peers = "010000226085",
category = "Police", years = 2019)
)
})
test_that("provenance always records the expenditure concept", {
p <- cog_spending("010000226085", years = 2019, category = "Police")
d <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
expect_equal(attr(t, "provenance")$expenditure_concept, "total")
# The note explains the non-obvious part: how legacy IG was assembled.
expect_true(nzchar(attr(t, "provenance")$expenditure_concept_note))
expect_true(is.na(attr(d, "provenance")$expenditure_concept_note) ||
!nzchar(attr(d, "provenance")$expenditure_concept_note))
})
test_that("the provenance schema documents expenditure_concept", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE
)
expect_true("expenditure_concept" %in% names(sch$properties))
})
test_that("a firing suggestion names the intergovernmental counterpart recipe", {
# Corrections has no legacy leaf rows, so the coverage-gap suggestion fires;
# corrections_ig_local_combined is its IG counterpart.
r <- suppressMessages(
cog_spending("010000226085", years = c(2005, 2011), category = "Corrections")
)
sugg <- attr(r, "provenance")$suggestions
expect_gt(length(sugg), 0L)
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
expect_true("corrections_combined" %in% ids)
ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id))
expect_true("corrections_ig_local_combined" %in% ig)
})
test_that("no suggestion fires for a healthy query", {
r <- cog_spending("010000226085", years = 2019, category = "Police")
expect_length(attr(r, "provenance")$suggestions, 0L)
})
test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to a revenue-flavored recipe", {
# "IG Federal" is a revenue-only category (summary_categories maps it to
# B-prefixed component codes only; its recipes are ig_federal_b47_wide /
# ig_federal_b89_wide). A cog_spending() call scoped to it returns zero
# spending rows for every requested year -- there is no spending
# component in this category at all -- so the coverage-gap machinery
# fires for real (not hypothetically) even though this isn't the kind of
# format-boundary gap the recipe catalog is meant to signpost. This is
# exactly the live-corpus risk flagged in review: ig_federal_b47_wide's
# own component codes (B47/B94, suffixes {"47","94"}) are an EXACT
# suffix-set match for the expenditure recipe ige_local_m47_wide
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
# Direct/Total pairing. The flow-family gate in
# .attach_ig_counterparts() must keep ig_recipe_id NULL here.
#
# Anchored on FL state government, not AL. Coverage is presence-based: a
# recipe is only suggested when its component codes have rows for the
# requested government-year. AL state's only FY2011 B47 cell was an
# explicit zero, which the corpus no longer stores after sparsification
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
# FL state carries a real FY2011 B47 amount, so this exercises the guard
# against a suggestion that genuinely fires.
r <- suppressMessages(
cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
)
sugg <- attr(r, "provenance")$suggestions
expect_gt(length(sugg), 0L)
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
expect_true("ig_federal_b47_wide" %in% ids)
ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id))
expect_length(ig, 0L)
})
test_that("C1: 'total' on a legacy aggregate-only family reports the IG-only figure honestly, not as Direct + IG", {
# AL state government, Corrections, 2011. Measured pre-fix: 'total'
# returned $31,358,000 (the IG leg alone, on an aggregate-flagged M04/M05
# row) with 0 suggestions (the surviving IG row made the gap-detection
# machinery think the Direct leg was covered) and a note asserting
# "Total = Direct + intergovernmental" with no caveat. True Direct (via
# recipe = "corrections_combined") is $521,651,000 -- the IG-only figure
# is ~6% of it.
gov <- "010000226085"
d <- cog_spending(gov, years = 2011, category = "Corrections",
expenditure_concept = "direct")
expect_equal(nrow(d), 0L)
t <- suppressMessages(cog_spending(
gov, years = 2011, category = "Corrections", expenditure_concept = "total"
))
expect_equal(nrow(t), 1L)
expect_equal(t$spend_subtype, "intergovernmental")
expect_equal(t$amt_nominal, 31358000)
r <- cog_spending(gov, years = 2011, recipe = "corrections_combined")
expect_equal(r$amt_nominal, 521651000)
# C1(a): the recipe hints must fire for "total" exactly as they do for
# "direct" -- the surviving IG row must not be mistaken for Direct
# coverage.
prov <- attr(t, "provenance")
expect_gt(length(prov$suggestions), 0L)
ids <- vapply(prov$suggestions, function(s) s$recipe_id %||% "", character(1))
expect_true("corrections_combined" %in% ids)
# C1(b): the affected row's notes name a recovering recipe rather than
# staying silent, and the provenance carries a flag a downstream consumer
# (e.g. cog-api, which passes provenance through verbatim) can test.
expect_true(nzchar(t$notes))
expect_match(t$notes, "unavailable", fixed = TRUE)
expect_match(t$notes, "corrections_combined", fixed = TRUE)
expect_true(prov$expenditure_concept_direct_suppressed)
# The base "Total = Direct + IG" note must NOT stand unqualified when that
# arithmetic didn't actually happen for this row.
expect_match(prov$expenditure_concept_note, "NOTE", fixed = TRUE)
expect_match(prov$expenditure_concept_note,
"expenditure_concept_direct_suppressed", fixed = TRUE)
})
test_that("C1(b): expenditure_concept_direct_suppressed is FALSE when the Direct leg is present", {
d <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
expect_false(isTRUE(attr(d, "provenance")$expenditure_concept_direct_suppressed))
expect_false(isTRUE(attr(t, "provenance")$expenditure_concept_direct_suppressed))
expect_false(any(nzchar(t$notes[t$spend_subtype == "intergovernmental"]) &
grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"])))
})
# M/I fix: .detect_direct_suppressed() was equating "no Direct sibling row"
# with "Direct was suppressed", but the dominant real cause is a government
# that simply has no direct spending in that category -- correct, ordinary
# data. The fix gates the flag (and its row note) on a harmonization recipe
# ACTUALLY covering that exact (year, canonical_govid, category) triple.
test_that("M/I: true positive, category supplied explicitly (unchanged behavior)", {
al <- "010000226085"
t_cat <- suppressMessages(cog_spending(
al, years = 2011, category = "Corrections", expenditure_concept = "total"
))
expect_true(attr(t_cat, "provenance")$expenditure_concept_direct_suppressed)
expect_match(t_cat$notes, "corrections_combined", fixed = TRUE)
expect_match(t_cat$notes, "unavailable", fixed = TRUE)
})
test_that("M/I: true positive, category = NULL now also names the recipe (was the fallback bug)", {
# Root bug: .build_suggestions() short-circuits to list() when category is
# NULL, so the note previously always hit its "no covering recipe found"
# fallback here even though corrections_combined genuinely covers this row.
al <- "010000226085"
t_null <- suppressMessages(cog_spending(
al, years = 2011, category = NULL, expenditure_concept = "total"
))
corr_row <- t_null[t_null$category %in% "Corrections", ]
expect_equal(nrow(corr_row), 1L)
expect_true(attr(t_null, "provenance")$expenditure_concept_direct_suppressed)
expect_match(corr_row$notes, "corrections_combined", fixed = TRUE)
expect_match(corr_row$notes, "unavailable", fixed = TRUE)
expect_false(grepl("no covering recipe found", corr_row$notes, fixed = TRUE))
})
test_that("M/I: false positive -- Virginia Education K-12 FY2019 total is NOT flagged", {
# States fund K-12 through school districts, so the Direct leg (E12/F12/
# G12) is genuinely, correctly zero -- not suppressed. Must not be flagged
# and must carry no suppression note.
va <- "510000227542"
t_va <- suppressMessages(cog_spending(
va, years = 2019, category = "Education K-12", expenditure_concept = "total"
))
expect_equal(nrow(t_va), 1L)
expect_equal(t_va$spend_subtype, "intergovernmental")
expect_equal(t_va$amt_nominal, 8028179000)
expect_false(isTRUE(attr(t_va, "provenance")$expenditure_concept_direct_suppressed))
expect_false(nzchar(t_va$notes) && grepl("unavailable", t_va$notes))
})
test_that("M/I: false positive by construction -- 'Other Education' has no E/F/G code, never flagged", {
# "Other Education" maps only to M21/L21 in summary_categories -- there is
# no E/F/G code for it in this corpus at all, so no Direct-recovering
# recipe can exist and it must never be flagged, in any fixture year.
con <- uscogdata:::.ensure_session()
years_all <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
states <- DBI::dbGetQuery(con,
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
oe <- suppressMessages(cog_spending(
states, years = years_all, category = "Other Education",
expenditure_concept = "total"
))
expect_false(isTRUE(attr(oe, "provenance")$expenditure_concept_direct_suppressed))
expect_false(any(nzchar(oe$notes) & grepl("unavailable", oe$notes)))
})
test_that("M/I: a clean FY2019 category = NULL total query flags far fewer than the pre-fix 32/50 states", {
con <- uscogdata:::.ensure_session()
states <- DBI::dbGetQuery(con,
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
r <- suppressMessages(cog_spending(
states, years = 2019, category = NULL, expenditure_concept = "total"
))
ig <- r[r$spend_subtype == "intergovernmental", ]
flagged <- ig[nzchar(ig$notes) & grepl("unavailable", ig$notes), ]
expect_lt(length(unique(flagged$canonical_govid)), 32L)
# Every remaining flagged row must actually name a covering recipe --
# never the old no-recipe-found fallback.
expect_true(all(grepl("recipe = '", flagged$notes, fixed = TRUE)))
expect_false(any(grepl("no covering recipe found", flagged$notes, fixed = TRUE)))
})
test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", {
with_corpus_missing_ig_categories({
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')"
)$n
expect_equal(n, 0)
err <- tryCatch(
cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total"),
condition = function(e) e
)
expect_s3_class(err, "uscogdata_ig_categories_unsupported")
msg <- conditionMessage(err)
expect_match(msg, "PR #59|predates", perl = TRUE)
})
# 'direct' is unaffected on the same corpus -- the guard is scoped to
# expenditure_concept = "total" only.
with_corpus_missing_ig_categories({
expect_no_error(
cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "direct")
)
})
})
test_that("C2: expenditure_concept = 'total' still works on a corpus that DOES carry M/L category rows", {
expect_no_error(
cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
)
})
test_that("I2: an intergovernmental (M/L) recipe never appears as its own top-level suggestion", {
# Task 1's M04/M05 category rows share the "Corrections" summary_categories
# category with the Direct-flavored E04/E05, so `corrections_ig_local_
# combined` (entirely M-prefixed) becomes a raw *candidate* in
# .build_suggestions()'s component_code-driven query. Following a
# "re-run with recipe = 'corrections_ig_local_combined'" hint on a plain
# cog_spending() call would silently return intergovernmental dollars
# under provenance$expenditure_concept = "direct". Task 6's gate
# (.attach_ig_counterparts()) already protects the *counterpart* lookup;
# this exercises that the candidate list itself is filtered too.
r <- suppressMessages(
cog_spending("010000226085", years = c(2005, 2011), category = "Corrections")
)
sugg <- attr(r, "provenance")$suggestions
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
expect_true("corrections_combined" %in% ids)
expect_false("corrections_ig_local_combined" %in% ids)
})
test_that(".attach_ig_counterparts() never pairs a revenue-side recipe with its coincidental M/L suffix twin", {
# Broader version of the case above, run at the matching-helper level
# (the same level code review's pairwise enumeration was done at) rather
# than end-to-end: the fixture has no (govid, year) combination where
# cog_revenue() itself produces a covered gap for any B/C/D recipe, so an
# end-to-end repro for THIS specific set of recipes isn't reachable
# today. Each of these six recipes shares an exact suffix set with an
# M/L expenditure recipe purely by reused-digit coincidence:
# ig_federal_b47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
# ig_federal_b89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
# ig_state_c47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
# ig_state_c89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
# ig_local_d47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
# ig_local_d89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
# None of them may receive an ig_recipe_id under cog_revenue()'s own
# flow_prefixes, since M/L only ever pairs with the direct-expenditure
# (E/F/G) family.
con <- uscogdata:::.ensure_session()
fake_suggestion <- function(rid) {
list(recipe_id = rid, label = "x", available_years = c(1967L, 2023L),
hint = "h")
}
fake_suggestions <- lapply(
c("ig_federal_b47_wide", "ig_federal_b89_wide",
"ig_state_c47_wide", "ig_state_c89_wide",
"ig_local_d47_wide", "ig_local_d89_wide"),
fake_suggestion
)
out <- uscogdata:::.attach_ig_counterparts(
con, fake_suggestions, c("T", "A", "U", "B", "C", "D")
)
ig <- unlist(lapply(out, function(s) s$ig_recipe_id))
expect_length(ig, 0L)
})
+112
View File
@@ -0,0 +1,112 @@
# Madison walkthrough audit -- findings F-012, F-017, F-018.
# Tracked as uscogdata#11. See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# The owner's settled three-concept model (2026-07-28):
# total = primary + interest + intergovernmental transfers
# direct = primary + interest (Census's published Direct Expenditure)
# primary = direct minus debt service (the NEW DEFAULT)
# implemented by reclassifying on the crosswalk's `spend_type` column, NOT on
# item-code first letters -- F-018 shows prefix `Y` carries both revenue
# (Y01/Y02) and expenditure (Y05/Y06) codes, so no first-letter allowlist can
# route them correctly.
#
# Fixture reproducibility: the finding's headline reconciliation is Madison
# FY2022, where the corpus carries I89 = 46,609 (thousands) and Census's
# published Direct Expenditure is $654,893,000 against cog_spending()'s
# $608,284,000 (-7.1%). FY2022 is outside the bundled fixture's year window
# (2011/2012/2019/2020), so the same invariant is asserted on FY2020, where the
# fixture carries I89 = 27,704. Anyone running against the full corpus should
# also check the FY2022 numbers above.
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
mad <- "552025209777" # MADISON CITY, WI
wi_state <- "550000227544" # WISCONSIN (state government)
# -- F-012: `primary` is the new default, and equals today's E/F/G figure ---
primary <- cog_spending(govid = mad, years = 2020L)
expect_equal(attr(primary, "provenance")$expenditure_concept, "primary")
expect_equal(sum(primary$amt_nominal), 623347000)
# -- F-012: `direct` adds interest on long-term debt ------------------------
# Expected interest read from the RAW corpus, never through cog_spending(),
# which is the filter under test.
interest <- wt_raw_amt(mad, 2020L, prefixes = "I")
expect_equal(interest, 27704) # I89, in $1,000s
direct <- cog_spending(govid = mad, years = 2020L, expenditure_concept = "direct")
expect_equal(sum(direct$amt_nominal), 651051000) # 623,347 + 27,704 thousands
expect_equal(sum(direct$amt_nominal) - sum(primary$amt_nominal), interest * 1000)
expect_true("I89" %in% wt_codes_included(direct))
# -- F-017: `total` carries Q12/Q18, state IG transfers to school districts --
# Wisconsin FY2019: Q12 = 6,431,530 and Q18 = 533,391 (thousands). Today
# neither verb's flow_prefixes contains "Q", so both are dropped from the one
# concept that is supposed to include intergovernmental transfers.
ig_expected <- wt_raw_amt(wi_state, 2019L, prefixes = c("M", "L", "Q"))
expect_equal(ig_expected, 11609814) # M 4,644,893 + Q 6,964,921
wi_direct <- cog_spending(govid = wi_state, years = 2019L,
expenditure_concept = "direct")
wi_total <- cog_spending(govid = wi_state, years = 2019L,
expenditure_concept = "total")
# total - direct is exactly the intergovernmental component. Asserted as a
# delta rather than a grand total so this stays correct however the J and Y
# families land inside `primary`.
expect_equal(sum(wi_total$amt_nominal) - sum(wi_direct$amt_nominal),
ig_expected * 1000)
expect_true(all(c("Q12", "Q18") %in% wt_codes_included(wi_total)))
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
# payments. All four share the first letter `Y`, so no first-letter allowlist
# can route them. The proof that classification is crosswalk-keyed:
# Y05 lands in `total` spending (insurance_benefits is inside `direct`),
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
# therefore can never appear in a spending result.
#
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
# DEFAULT stays Census General Revenue and so excludes insurance-trust
# revenue; Y01's revenue-side classification is asserted against the
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
# revenue concept argument is uscogdata#12.
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
spend_codes <- wt_codes_included(wi_total)
rev_codes <- wt_codes_included(wi_revenue)
expect_true("Y05" %in% spend_codes)
expect_false("Y05" %in% rev_codes)
expect_false("Y01" %in% spend_codes)
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
con <- uscogdata:::.ensure_session()
y_class <- DBI::dbGetQuery(con,
"SELECT item_code, category_type, spend_subtype, revenue_subtype
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
})
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
# balances) share first letters with flow codes, so this could never be
# guaranteed under prefix classification; under crosswalk membership it
# falls out structurally -- asserted here at the verb level, on a
# government-year the fixture gives real balance rows (Wisconsin carries
# Y07/Y08/Y21/Y61-type balances in FY2019).
wi_state <- "550000227544"
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
expect_gt(nrow(balance), 0L)
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
rev <- cog_revenue(wi_state, 2019L)
expect_false(any(spend$category %in% balance$category))
expect_false(any(rev$category %in% balance$category))
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
})
+92 -4
View File
@@ -1,6 +1,6 @@
test_that("cog_explain prints verb header and target", { test_that("cog_explain prints verb header and target", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections") r <- cog_spending("121011212191", 2020L, "Corrections")
# cli writes to stderr; capture both stdout and message streams. # cli writes to stderr; capture both stdout and message streams.
txt <- paste(c( txt <- paste(c(
capture.output(cog_explain(r)), capture.output(cog_explain(r)),
@@ -8,19 +8,19 @@ test_that("cog_explain prints verb header and target", {
), collapse = "\n") ), collapse = "\n")
expect_true(grepl("cog_spending", txt)) expect_true(grepl("cog_spending", txt))
expect_true(grepl("Corrections", txt)) expect_true(grepl("Corrections", txt))
expect_true(grepl("101006006", txt)) expect_true(grepl("121011212191", txt))
}) })
test_that("cog_explain format='list' returns structured provenance", { test_that("cog_explain format='list' returns structured provenance", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections") r <- cog_spending("121011212191", 2020L, "Corrections")
prov <- cog_explain(r, format = "list") prov <- cog_explain(r, format = "list")
expect_identical(prov, attr(r, "provenance")) expect_identical(prov, attr(r, "provenance"))
}) })
test_that("cog_explain returns result invisibly for chaining", { test_that("cog_explain returns result invisibly for chaining", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections") r <- cog_spending("121011212191", 2020L, "Corrections")
res <- withVisible(cog_explain(r)) res <- withVisible(cog_explain(r))
expect_false(res$visible) expect_false(res$visible)
expect_identical(res$value, r) expect_identical(res$value, r)
@@ -30,3 +30,91 @@ test_that("cog_explain errors on non-verb input", {
df <- tibble::tibble(a = 1) df <- tibble::tibble(a = 1)
expect_error(cog_explain(df), "provenance") expect_error(cog_explain(df), "provenance")
}) })
test_that("cog_explain prints basis + harmonization block", {
skip_if_no_corpus()
r <- cog_spending("121011212191", 2020L, "Corrections")
txt <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_true(grepl("Basis: harmonized", txt))
expect_true(grepl("Harmonization", txt))
expect_true(grepl("Excluded 0 row", txt))
})
test_that("cog_explain prints a Recipe section for recipe = results", {
skip_if_no_corpus()
r <- cog_spending("121011212191", c(2011L, 2012L), recipe = "corrections_combined")
txt <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_true(grepl("Recipe", txt))
expect_true(grepl("corrections_combined", txt))
expect_true(grepl("E04", txt))
expect_true(grepl("E05", txt))
})
test_that("cog_explain prints a Suggestions section when the provenance has one", {
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("121011212191", c(2011L, 2012L), category = "Corrections")
)
txt <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_true(grepl("Suggestions", txt))
expect_true(grepl("corrections_combined", txt))
expect_true(grepl("re-run with recipe", txt))
})
test_that("cog_explain prints the expenditure concept (I1)", {
skip_if_no_corpus()
d <- cog_spending("010000226085", years = 2019, category = "Police")
t <- cog_spending("010000226085", years = 2019, category = "Police",
expenditure_concept = "total")
txt_d <- paste(c(
capture.output(cog_explain(d)),
capture.output(cog_explain(d), type = "message")
), collapse = "\n")
txt_t <- paste(c(
capture.output(cog_explain(t)),
capture.output(cog_explain(t), type = "message")
), collapse = "\n")
expect_true(grepl("Concept: primary", txt_d))
expect_true(grepl("Concept: total", txt_t))
})
test_that("cog_explain surfaces the C1(b) direct-suppressed flag as a warning", {
skip_if_no_corpus()
t <- suppressMessages(cog_spending(
"010000226085", years = 2011, category = "Corrections",
expenditure_concept = "total"
))
expect_true(attr(t, "provenance")$expenditure_concept_direct_suppressed)
txt <- paste(c(
capture.output(cog_explain(t)),
capture.output(cog_explain(t), type = "message")
), collapse = "\n")
expect_true(grepl("Direct leg unavailable", txt))
})
test_that("cog_explain prints denominator + popyear_range + counts", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", years = 2019:2020,
category = "Police", per_capita = TRUE)
out <- paste(c(
capture.output(cog_explain(r)),
capture.output(cog_explain(r), type = "message")
), collapse = "\n")
expect_true(grepl("Census F-33", out))
expect_true(grepl("popyear", out, ignore.case = TRUE))
expect_true(grepl("census_f33", out))
# popyear_range should render as 4-digit calendar years, not raw 2-digit
expect_true(grepl("2019-2020", out))
expect_false(grepl("popyear range: 19-20", out, fixed = TRUE))
})
})
+112
View File
@@ -0,0 +1,112 @@
# tests/testthat/test-fixture-vintage.R
#
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
# every test in this package -- plus the whole cog-api suite -- runs against
# it. When the published corpus changes shape and the fixture does not, both
# suites stay green against a corpus that no longer exists (uscogdata#18).
#
# These tests pin the structural facts that distinguish the current published
# vintage from its predecessor, so a stale fixture fails loudly instead of
# passing quietly. They assert shape, never dollar values: re-running
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
# keep them green.
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
# not the package session: these assertions are about what the fixture
# CONTAINS, and routing them through the reader's own views would let a
# filter hide the very absence being checked.
fixture_query <- function(sql, ...) {
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
path <- function(rel) {
sprintf("read_parquet(%s)",
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
}
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
}
test_that("fixture ships every metadata table the publish tree does", {
skip_if_no_corpus()
# representation/code_set are what make a sparse corpus interpretable; a
# fixture without them predates sparsification (cog_pipeline#64).
expected <- c(
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
"census_collection_coverage.parquet", "code_set.parquet",
"harmonization_map.parquet", "harmonization_recipes.parquet",
"lineage_events.parquet", "representation.parquet",
"series_breaks.parquet", "summary_categories.parquet"
)
on_disk <- basename(list.files(
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
))
expect_true(all(expected %in% on_disk))
# The manifest must list them too -- consumers read the manifest, not ls().
in_manifest <- with_fixture_corpus(
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
)
expect_true(all(expected %in% in_manifest))
})
test_that("fixture carries the dense/sparse representation contract", {
skip_if_no_corpus()
rep <- fixture_query(
"SELECT year, representation, absence_means FROM %s
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
"data/representation.parquet"
)
expect_equal(nrow(rep), 4L)
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
})
test_that("the fixture's wide era is sparse, not zero-padded", {
skip_if_no_corpus()
# FY2011 is a dense_source year: the corpus publishes only the cells Census
# reported non-zero, and an absent cell means Census published $0. Before
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
# zeros. A single explicit zero here means the fixture predates the change.
zeros_2011 <- fixture_query(
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
"data/long/year=2011/part-0.parquet"
)$n
expect_equal(zeros_2011, 0L)
# The modern era is a different regime: a reported zero there is real data
# (the government filed $0), so zeros legitimately survive and must not be
# asserted away.
expect_gt(
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
0L
)
})
test_that("code_set covers every fixture year with the reader-spec columns", {
skip_if_no_corpus()
cs <- fixture_query(
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
"data/code_set.parquet"
)
expect_true(all(
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
%in% names(cs)
))
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
})
test_that("every flow code carrying dollars has a category, J-prefix included", {
skip_if_no_corpus()
# The J (assistance/benefit) codes were uncategorised until the crosswalk
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
# duplication fix landed). Their absence is how a pre-crosswalk fixture
# gives itself away.
j <- fixture_query(
"SELECT item_code, category, category_type, spend_subtype FROM %s
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
"data/summary_categories.parquet"
)
expect_true("J19" %in% j$item_code)
expect_true(all(j$category_type == "expenditure"))
expect_true(all(j$spend_subtype == "assistance"))
expect_false(any(is.na(j$category)))
})
@@ -0,0 +1,57 @@
# Madison walkthrough audit -- finding F-025. Tracked as uscogdata#16.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# cog_gov_search()'s UTILITY mode interpolates `name` into
# regexp_matches(gov_name, <name>, 'i')
# unescaped (R/search.R:102), while BASKET mode in the same file already routes
# it through .escape_regex() (R/search.R:307) with the comment "so `name` is
# treated as a literal substring". Two failure modes result:
# correctness -- a real government cannot be found by its own exact name, and
# a single "." matches everything (HTTP 200 both ways via the API);
# robustness -- malformed regex reaches the engine and errors, which cog-api
# surfaces as a 500, reachable by typing a real name one
# character at a time.
#
# NOT asserted here: the finding's `q=St. Louis` example. Under correct literal
# matching that search still returns 0 rows, because the stored name is
# "ST LOUIS CITY" with no period -- it demonstrates today's over-matching
# semantics, not a row the fix makes findable.
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
# -- correctness (1): a government must be findable by its own exact name ---
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
# grouping, so its own complete name matches nothing.
fredonia <- cog_gov_search(name = "FREDONIA (BRISCOE) CITY")
expect_equal(nrow(fredonia), 1L)
expect_equal(fredonia$canonical_govid, "052117184386")
expect_equal(cog_gov_search(name = "FREDONIA (BRISCOE)")$canonical_govid,
"052117184386")
# -- correctness (2): a metacharacter must not become a wildcard ------------
# No Wisconsin city or village name contains a literal period -- established
# against the raw registry below, NOT through the verb under test. A literal
# search for "." must therefore return nothing; today it returns all 608.
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
xwalk <- paste0(sub("/$", "", Sys.getenv("USCOGDATA_URL")),
"/data/canonical_fips_xwalk.parquet")
with_dot <- DBI::dbGetQuery(con, paste0(
"SELECT COUNT(*) n FROM read_parquet('", xwalk, "') ",
"WHERE fips_state = '55' AND govs_type = 2 AND gov_name LIKE '%.%'"))
expect_equal(as.integer(with_dot$n[[1]]), 0L)
expect_equal(nrow(cog_gov_search(name = ".", state = "WI", type = "city")), 0L)
expect_equal(nrow(cog_gov_search(name = "M.dison", state = "WI", type = "city")), 0L)
expect_equal(nrow(cog_gov_search(name = "Mad(i|o)son", state = "WI", type = "city")), 0L)
# A metacharacter-free name still resolves exactly as before.
expect_equal(nrow(cog_gov_search(name = "Madison", state = "WI", type = "city")), 1L)
# -- robustness: malformed pattern text returns no rows, and does not error --
# "[" alone, and "Athens-Clarke County (bal" -- an in-progress substring of
# ATHENS-CLARKE COUNTY (BALANCE), a real government -- both currently raise
# (DuckDB: "Invalid Input Error: missing ]").
expect_equal(nrow(cog_gov_search(name = "[")), 0L)
expect_equal(nrow(cog_gov_search(name = "Athens-Clarke County (bal")), 0L)
})
+155
View File
@@ -0,0 +1,155 @@
# tests/testthat/test-manifest.R
#
# Tests for the guards on .fetch_or_cache_manifest() and cog_open() that
# protect users from silent failures when USCOGDATA_URL is misconfigured
# or returns non-JSON content.
test_that("cog_open aborts with actionable error when URL is the placeholder default", {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
withr::with_envvar(c(USCOGDATA_URL = placeholder), {
expect_error(
uscogdata:::cog_open(),
class = "uscogdata_url_not_configured"
)
})
})
test_that("placeholder guard fires for any URL containing the sentinel token", {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
# Sentinel detection should be substring-based — covers any host that still
# has REPLACE_WITH_SHARE_TOKEN baked in (default or partial user edit).
withr::with_envvar(c(USCOGDATA_URL = "https://other.example/s/REPLACE_WITH_SHARE_TOKEN/x/"), {
expect_error(
uscogdata:::cog_open(),
class = "uscogdata_url_not_configured"
)
})
})
test_that("placeholder guard error names both env var and option as remediation", {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
withr::with_envvar(c(USCOGDATA_URL = placeholder), {
msg <- tryCatch(uscogdata:::cog_open(), error = conditionMessage)
expect_match(msg, "USCOGDATA_URL", fixed = TRUE)
expect_match(msg, "uscogdata.url", fixed = TRUE)
})
})
test_that("local manifest containing HTML produces uscogdata_invalid_manifest, not raw parse error", {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
tmp <- withr::local_tempdir()
writeLines(
c("<html>", " <head><title>Welcome to our server</title></head>", "</html>"),
file.path(tmp, "manifest.json")
)
withr::with_envvar(c(USCOGDATA_URL = paste0(tmp, "/")), {
err <- expect_error(
uscogdata:::cog_open(),
class = "uscogdata_invalid_manifest"
)
expect_match(conditionMessage(err), "manifest", ignore.case = TRUE)
})
})
test_that("remote manifest fetch does not poison cache when response is HTML", {
uscogdata:::cog_close()
on.exit(uscogdata:::cog_close(), add = TRUE)
tmp_cache <- withr::local_tempdir()
cache_path <- file.path(tmp_cache, "manifest.json")
# Pretend the cache already exists with stale-but-fresh-by-mtime HTML
# (simulating a previous poisoned write from the old behavior). When the
# fetcher sees invalid JSON in the cache, it must refetch rather than
# silently returning a parse error to the caller.
writeLines("<html>poisoned</html>", cache_path)
Sys.setFileTime(cache_path, Sys.time()) # ensure within TTL
# We don't have a live HTTP fixture here, so the refetch will fail at the
# network layer — but the failure should NOT be a jsonlite parse error on
# the cached HTML; it should be a network-level httr2 error. The cache
# file itself must remain untouched (no atomic-write half-states).
withr::with_envvar(
c(
USCOGDATA_URL = "https://invalid.localhost.uscogdata.test/",
USCOGDATA_CACHE_DIR = tmp_cache
),
{
err <- tryCatch(uscogdata:::cog_open(), error = identity)
expect_s3_class(err, "error")
# Must not be a JSON lexical error on HTML.
expect_false(grepl("lexical error", conditionMessage(err), fixed = TRUE))
}
)
# Atomic write contract: no stray tmp files left behind in cache_dir.
expect_length(
list.files(tmp_cache, pattern = "manifest\\.json\\.tmp"),
0L
)
})
test_that("cog_manifest returns the active session's parsed manifest", {
with_fixture_corpus({
m <- cog_manifest()
expect_type(m, "list")
expect_true(m$schema_version >= 4L)
yrs <- vapply(m$files$long_partitions, function(p) as.integer(p$year),
integer(1))
expect_setequal(yrs, c(2011L, 2012L, 2019L, 2020L))
})
})
test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others", {
expect_silent(uscogdata:::.validate_schema(list(schema_version = 4L)))
expect_silent(uscogdata:::.validate_schema(list(schema_version = 5L)))
# v6 = FIPS geography harmonization (2026-07-22): _code -> _asof rename +
# cog_legacy_* columns (26 -> 28 cols). This package references none of the
# renamed columns and its geography comes from the xwalk, so v6 is accepted
# without behavioural change -- see .validate_schema()'s note.
expect_silent(uscogdata:::.validate_schema(list(schema_version = 6L)))
expect_error(
uscogdata:::.validate_schema(list(schema_version = 3L)),
"schema_version"
)
expect_error(
uscogdata:::.validate_schema(list(schema_version = 7L)),
"schema_version"
)
})
test_that("cog_open succeeds against a doctored schema_version 4 corpus (dual-accept)", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
con <- cog_open()
expect_true(DBI::dbIsValid(con))
expect_equal(as.integer(cog_manifest()$schema_version), 4L)
# Core (pre-Phase-R2) views must still register on a v4 corpus.
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)$table_name
expect_true(all(c("spending_annotated", "revenue_annotated") %in% views))
# Schema-v5-only harmonization views must NOT register on a v4 corpus:
# their parquet sources don't exist there and DuckDB's read_parquet()
# errors eagerly at CREATE VIEW time for a missing file/glob, so
# .register_views() gates these on manifest$schema_version >= 5.
expect_false(any(c(
"spending_long_harmonized", "spending_annotated_harmonized",
"harmonization_recipes", "harmonization_map", "series_breaks_pq"
) %in% views))
})
})
+2 -2
View File
@@ -61,7 +61,7 @@ test_that("cog_mirror reads back via a fresh session against the mirror", {
cog_close() cog_close()
options(uscogdata.url = paste0(normalizePath(tmp), "/")) options(uscogdata.url = paste0(normalizePath(tmp), "/"))
r <- cog_spending("101006006", 2020L, "Corrections") r <- cog_spending("121011212191", 2020L, "Corrections")
expect_gt(nrow(r), 0L) expect_gt(nrow(r), 0L)
expect_equal(unique(r$canonical_govid), "101006006") expect_equal(unique(r$canonical_govid), "121011212191")
}) })
+51
View File
@@ -0,0 +1,51 @@
# Madison walkthrough audit -- finding F-021. Tracked as uscogdata#14.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# .peer_summary_rows() computes stats::quantile() separately INSIDE each
# (year, spend_subtype, category) cell. A summary_p50 row is therefore "the
# median peer's value in that one category", not "the value of the median
# peer's total". Summing those rows across categories -- the obvious move for a
# caller who wants one peer-median total line and reads only the column names --
# misstated a total-spending band by -32.7% to +251.0% across the 24 years the
# audit tested, with a sign flip at FY2012.
#
# The verb is not wrong and its documented use (faceting by role AND category)
# is unaffected, so the fix is documentation: one sentence in @return.
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
# man/ ships only in the source tree (the installed package carries a
# compiled help database instead), so the prose assertions below cannot run
# under R CMD check -- CI's earlier testthat::test_local() step enforces
# them. The numeric pin further down needs only the corpus, but it lives in
# the same test_that() as the sentence it protects, deliberately: they are
# one claim, and splitting them would let the prose drift while a separate
# test kept passing.
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
# The @return section must say the quantile is computed within each cell...
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
# ...and must warn that the rows are not additive across category.
expect_match(rd, "not additive|do(es)? not sum|cannot be summed", ignore.case = TRUE)
# ...naming the grouping explicitly.
expect_match(rd, "spend_subtype", fixed = TRUE)
# Pin the mechanism numerically so a future refactor that quietly changes the
# quantile grouping fails here rather than silently invalidating the sentence
# above. Fixture: Madison, 10 peers found at FY2020, category = NULL.
peers <- cog_find_peers("552025209777", year = 2020L, max_peers = 10L)
cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers,
category = NULL, years = 2020L, per_capita = TRUE)
naive <- sum(cmp$amt_per_capita_nominal[cmp$role == "summary_p50"], na.rm = TRUE)
peer_rows <- cmp[cmp$role == "peer", ]
per_gov <- tapply(peer_rows$amt_per_capita_nominal, peer_rows$canonical_govid,
sum, na.rm = TRUE)
correct <- unname(stats::quantile(per_gov, 0.5, na.rm = TRUE))
expect_equal(round(naive), 6180) # summing the built-in summary rows
expect_equal(round(correct), 2043) # quantile of each peer's OWN total
expect_gt(naive / correct, 2) # a +200% misstatement on this cohort
})
+64 -16
View File
@@ -1,29 +1,29 @@
test_that("cog_find_peers returns same-type peers in the default pop band", { test_that("cog_find_peers returns same-type peers in the default pop band", {
skip_if_no_corpus() skip_if_no_corpus()
peers <- cog_find_peers("101006006") # Broward County peers <- cog_find_peers("121011212191") # Broward County
expect_s3_class(peers, "tbl_df") expect_s3_class(peers, "tbl_df")
expected_cols <- c("canonical_govid", "gov_name", "fips_state", expected_cols <- c("canonical_govid", "gov_name", "fips_state",
"population_acs", "pop_ratio", "rank") "population", "pop_ratio", "rank")
expect_true(all(expected_cols %in% names(peers))) expect_true(all(expected_cols %in% names(peers)))
expect_true(all(peers$pop_ratio >= 0.7 & peers$pop_ratio <= 1.3)) expect_true(all(peers$pop_ratio >= 0.7 & peers$pop_ratio <= 1.3))
expect_false("101006006" %in% peers$canonical_govid) expect_false("121011212191" %in% peers$canonical_govid)
expect_equal(peers$rank, seq_len(nrow(peers))) expect_equal(peers$rank, seq_len(nrow(peers)))
}) })
test_that("cog_find_peers respects same_state restriction", { test_that("cog_find_peers respects same_state restriction", {
skip_if_no_corpus() skip_if_no_corpus()
peers <- cog_find_peers("101006006", same_state = TRUE, peers <- cog_find_peers("121011212191", same_state = TRUE,
pop_range = c(0.1, 10)) pop_range = c(0.1, 10))
expect_true(all(peers$fips_state == "12")) expect_true(all(peers$fips_state == "12"))
}) })
test_that("cog_find_peers absolute pop range works", { test_that("cog_find_peers absolute pop range works", {
skip_if_no_corpus() skip_if_no_corpus()
peers <- cog_find_peers("101006006", peers <- cog_find_peers("121011212191",
pop_range = c(1.5e6, 2.5e6), pop_range = c(1.5e6, 2.5e6),
is_ratio = FALSE, max_peers = 20L) is_ratio = FALSE, max_peers = 20L)
expect_true(all(peers$population_acs >= 1.5e6 & expect_true(all(peers$population >= 1.5e6 &
peers$population_acs <= 2.5e6)) peers$population <= 2.5e6))
}) })
test_that("cog_find_peers errors cleanly on unknown govid", { test_that("cog_find_peers errors cleanly on unknown govid", {
@@ -33,8 +33,8 @@ test_that("cog_find_peers errors cleanly on unknown govid", {
test_that("cog_peer_compare accepts a cog_find_peers result directly", { test_that("cog_peer_compare accepts a cog_find_peers result directly", {
skip_if_no_corpus() skip_if_no_corpus()
peers <- cog_find_peers("101006006", max_peers = 4L) peers <- cog_find_peers("121011212191", max_peers = 4L)
r <- cog_peer_compare("101006006", peers, "Police", years = 2020L) r <- cog_peer_compare("121011212191", peers, "Police", years = 2020L)
expect_s3_class(r, "tbl_df") expect_s3_class(r, "tbl_df")
expect_true("role" %in% names(r)) expect_true("role" %in% names(r))
expect_setequal( expect_setequal(
@@ -47,8 +47,8 @@ test_that("cog_peer_compare accepts a cog_find_peers result directly", {
test_that("cog_peer_compare accepts a character vector of govids", { test_that("cog_peer_compare accepts a character vector of govids", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_peer_compare( r <- cog_peer_compare(
"101006006", "121011212191",
peers = c("441015015", "441220220"), # Bexar, Tarrant peers = c("481029175853", "481439135072"), # Bexar, Tarrant
category = "Police", years = 2020L category = "Police", years = 2020L
) )
expect_true("peer" %in% r$role) expect_true("peer" %in% r$role)
@@ -58,8 +58,8 @@ test_that("cog_peer_compare accepts a character vector of govids", {
test_that("cog_peer_compare summary rows use real per-capita when requested", { test_that("cog_peer_compare summary rows use real per-capita when requested", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_peer_compare( r <- cog_peer_compare(
"101006006", "121011212191",
peers = c("441015015", "441220220", "231082082"), peers = c("481029175853", "481439135072", "261163166615"),
category = "Police", years = 2019:2020, category = "Police", years = 2019:2020,
per_capita = TRUE, adjust_to_year = 2022L per_capita = TRUE, adjust_to_year = 2022L
) )
@@ -72,8 +72,8 @@ test_that("cog_peer_compare summary rows use real per-capita when requested", {
test_that("cog_peer_compare provenance reports the outer verb + peer count", { test_that("cog_peer_compare provenance reports the outer verb + peer count", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_peer_compare("101006006", r <- cog_peer_compare("121011212191",
peers = c("441015015", "441220220"), peers = c("481029175853", "481439135072"),
category = "Police", years = 2020L) category = "Police", years = 2020L)
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
expect_equal(prov$verb, "cog_peer_compare") expect_equal(prov$verb, "cog_peer_compare")
@@ -82,9 +82,57 @@ test_that("cog_peer_compare provenance reports the outer verb + peer count", {
test_that("cog_peer_compare handles zero peers gracefully", { test_that("cog_peer_compare handles zero peers gracefully", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_peer_compare("101006006", r <- cog_peer_compare("121011212191",
peers = character(0), peers = character(0),
category = "Police", years = 2020L) category = "Police", years = 2020L)
expect_true(all(r$role == "target")) expect_true(all(r$role == "target"))
expect_equal(sum(grepl("^summary_", r$role)), 0L) expect_equal(sum(grepl("^summary_", r$role)), 0L)
}) })
test_that("cog_find_peers defaults `year` to most recent observed year for target", {
skip_if_no_corpus()
peers <- cog_find_peers("121011212191")
expect_equal(attr(peers, "cohort_year"), 2020L)
# Returned column is now `population`, not `population_acs`
expect_true("population" %in% names(peers))
expect_false("population_acs" %in% names(peers))
})
test_that("cog_find_peers honors an explicit `year`", {
skip_if_no_corpus()
peers <- cog_find_peers("121011212191", year = 2019L)
expect_equal(attr(peers, "cohort_year"), 2019L)
})
test_that("cog_find_peers errors when target has no observed pop in `year`", {
skip_if_no_corpus()
expect_error(
cog_find_peers("121011212191", year = 1999L),
"no observed population"
)
})
test_that("cog_peer_compare stamps cohort_year from peers attribute", {
skip_if_no_corpus()
peers <- cog_find_peers("121011212191", year = 2019L, max_peers = 4L,
pop_range = c(0.5, 1.5))
r <- cog_peer_compare("121011212191", peers, "Police", years = 2020L)
expect_true("cohort_year" %in% names(r))
expect_true(all(r$cohort_year == 2019L))
prov <- attr(r, "provenance")
expect_equal(prov$cohort_year, 2019L)
expect_equal(length(prov$cohort_govids), nrow(peers))
# pop_range and is_ratio propagate from cog_find_peers attrs
expect_equal(prov$pop_range, c(0.5, 1.5))
expect_true(prov$is_ratio)
})
test_that("cog_peer_compare cohort_year is NA for bare character peers", {
skip_if_no_corpus()
r <- cog_peer_compare(
"121011212191",
peers = c("481029175853", "481439135072"),
category = "Police", years = 2020L
)
expect_true(all(is.na(r$cohort_year)))
})
+216
View File
@@ -0,0 +1,216 @@
# tests/testthat/test-recipes.R
#
# cog_recipes(), recipe = in cog_spending()/cog_revenue(), and the
# recipe-component-driven signposting in prov$suggestions (Phase R2 /
# Task 11, schema_version 5).
test_that("cog_recipes lists the curated catalog including corrections_combined", {
skip_if_no_corpus()
r <- cog_recipes()
expect_s3_class(r, "tbl_df")
expect_equal(names(r), c("recipe_id", "label", "n_components", "year_min", "year_max"))
expect_equal(nrow(r), 24L)
expect_true("corrections_combined" %in% r$recipe_id)
expect_true("t19_selective_sales_wide" %in% r$recipe_id)
expect_true("ig_federal_b89_wide" %in% r$recipe_id)
expect_true("rents_royalties_u4_wide" %in% r$recipe_id)
expect_true("higher_ed_e18_wide" %in% r$recipe_id)
expect_true("cash_securities_z77_wide" %in% r$recipe_id)
# Superseded id from the pre-curation brief text must NOT be present.
expect_false("corrections_judicial_combined" %in% r$recipe_id)
})
test_that("cog_recipes(pattern=) filters by recipe_id or label", {
skip_if_no_corpus()
r <- cog_recipes("corrections")
expect_true(nrow(r) >= 1L)
expect_true(all(grepl("corrections", r$recipe_id, ignore.case = TRUE) |
grepl("corrections", r$label, ignore.case = TRUE)))
})
test_that("cog_recipes requires schema_version >= 5", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
expect_error(cog_recipes(), class = "uscogdata_schema_unsupported")
})
})
# --- recipe = : generic join, no is_aggregate filter -----------------------
test_that("recipe = 'corrections_combined' is continuous across the 2011->2012 seam", {
skip_if_no_corpus()
r <- cog_spending("121011212191", years = c(2011L, 2012L),
recipe = "corrections_combined")
expect_equal(nrow(r), 2L)
expect_true(all(c("year", "canonical_govid", "gov_name", "spend_subtype",
"category", "amt_nominal", "codes_included",
"aggregate_fallback", "notes") %in% names(r)))
expect_equal(unique(r$spend_subtype), "recipe")
expect_equal(unique(r$category), "Corrections (functions 04+05 combined)")
expect_false(any(r$aggregate_fallback))
r2011 <- r$amt_nominal[r$year == 2011L]
r2012 <- r$amt_nominal[r$year == 2012L]
# 2011: E05 only exists as a wide-era AGGREGATE row (is_aggregate = TRUE)
# for Broward -- data-verified $216,088,000. Since .run_recipe() does NOT
# filter is_aggregate (amendment: the recipe join must not, because these
# families exist ONLY as aggregate rows in the wide era), the recipe
# correctly picks this up.
expect_equal(r2011, 216088000)
# 2012: modern E04 leaf ($213,056,000); Broward reports no E05 leaf that
# year, so the recipe total equals E04 alone -- still continuous with the
# 2011 aggregate, proving the wide-aggregate -> modern-leaf handoff.
expect_equal(r2012, 213056000)
expect_true(all(grepl("E04|E05", r$codes_included)))
})
test_that("recipe result carries a recipe provenance block with component rows", {
skip_if_no_corpus()
r <- cog_spending("121011212191", years = c(2011L, 2012L),
recipe = "corrections_combined")
prov <- attr(r, "provenance")
expect_equal(prov$basis, "recipe")
expect_equal(prov$category, "Corrections (functions 04+05 combined)")
expect_type(prov$recipe, "list")
expect_equal(prov$recipe$recipe_id, "corrections_combined")
expect_equal(prov$recipe$label, "Corrections (functions 04+05 combined)")
expect_length(prov$recipe$components, 2L)
comp_codes <- vapply(prov$recipe$components, function(x) x$component_code, character(1))
expect_setequal(comp_codes, c("E04", "E05"))
# A recipe query resolves its own coverage; it should never also carry
# suggestions for itself.
expect_length(prov$suggestions, 0L)
})
test_that("recipe results report an unambiguous basis/harmonization, ignoring basis=", {
skip_if_no_corpus()
# A recipe query bypasses spending_annotated(_harmonized) entirely --
# .run_recipe() joins `long` directly -- so `basis` must never read
# "harmonized"/"raw" (which would describe a code path this query never
# took) regardless of what the caller passed for `basis`. Task 12
# consumes provenance verbatim, so this needs to be unambiguous.
r_default <- cog_spending("121011212191", years = c(2011L, 2012L),
recipe = "corrections_combined")
r_raw <- cog_spending("121011212191", years = c(2011L, 2012L),
recipe = "corrections_combined", basis = "raw")
r_harm <- cog_spending("121011212191", years = c(2011L, 2012L),
recipe = "corrections_combined", basis = "harmonized")
for (r in list(r_default, r_raw, r_harm)) {
prov <- attr(r, "provenance")
expect_equal(prov$basis, "recipe")
expect_true(is.na(prov$basis_note))
expect_false(prov$harmonization$applied)
expect_equal(prov$harmonization$na_rows_excluded, 0L)
expect_match(prov$harmonization$note, "recipe", ignore.case = TRUE)
}
# basis= truly has zero effect on a recipe query's actual numbers.
expect_equal(r_raw$amt_nominal, r_harm$amt_nominal)
expect_equal(r_default$amt_nominal, r_raw$amt_nominal)
})
test_that("recipe = 't19_selective_sales_wide' sums the local T11/T14 legs when present", {
skip_if_no_corpus()
# Westminster City, CA (canonical_govid 082001211654): T11 = 0 in 2011,
# T11 = 568 (T14 = 0/absent) in 2012 -- a real, data-verified equality/
# inequality pair inside the amended fixture window (2011-2012), standing
# in for the brief's original 2004/2005 example (out of scope per the
# amended fixture years; the underlying local-tax-split boundary is
# nationally FY2005, but this government's own T11 reporting activates
# within our 2011-2012 window).
r <- cog_revenue("082001211654", years = c(2011L, 2012L),
recipe = "t19_selective_sales_wide")
# Raw, single-code T19 total (not the "Other Taxes" category total, which
# would also sum in T11/T14/T21/T23/T27/T29/T53/T99 -- queried directly to
# isolate exactly the code the brief's equality/inequality check is about).
con <- uscogdata:::.ensure_session()
raw_t19 <- DBI::dbGetQuery(con, "
SELECT year, SUM(amt) * 1000.0 AS amt
FROM revenue_long
WHERE canonical_govid = '082001211654' AND item_code = 'T19'
AND year IN (2011, 2012)
GROUP BY year ORDER BY year
")
raw_t19_2011 <- raw_t19$amt[raw_t19$year == 2011L]
raw_t19_2012 <- raw_t19$amt[raw_t19$year == 2012L]
expect_equal(raw_t19_2011, 2231000)
expect_equal(raw_t19_2012, 2365000)
recipe_2011 <- r$amt_nominal[r$year == 2011L]
recipe_2012 <- r$amt_nominal[r$year == 2012L]
expect_equal(recipe_2011, raw_t19_2011) # equality: no local T11/T14 yet
expect_gt(recipe_2012, raw_t19_2012) # inequality: local T11 joins in
expect_equal(recipe_2012, raw_t19_2012 + 568000)
})
test_that("recipe = and category = together aborts", {
skip_if_no_corpus()
expect_error(
cog_spending("121011212191", 2020L, category = "Corrections",
recipe = "corrections_combined"),
class = "uscogdata_recipe_category_conflict"
)
})
test_that("unknown recipe id aborts and lists valid ids", {
skip_if_no_corpus()
err <- tryCatch(
cog_spending("121011212191", 2020L, recipe = "does_not_exist"),
error = identity
)
expect_s3_class(err, "uscogdata_unknown_recipe")
expect_match(conditionMessage(err), "corrections_combined")
})
test_that("recipe = requires schema_version >= 5", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
expect_error(
cog_spending("121011212191", 2020L, recipe = "corrections_combined"),
class = "uscogdata_schema_unsupported"
)
})
})
# --- signposting -------------------------------------------------------
test_that("signposting suggests corrections_combined across the 2011->2012 gap", {
skip_if_no_corpus()
expect_message(
r <- cog_spending("121011212191", years = c(2011L, 2012L),
category = "Corrections"),
"recipe"
)
prov <- attr(r, "provenance")
expect_true(length(prov$suggestions) >= 1L)
ids <- vapply(prov$suggestions, function(s) s$recipe_id, character(1))
expect_true("corrections_combined" %in% ids)
hit <- prov$suggestions[[which(ids == "corrections_combined")]]
expect_equal(hit$hint, "re-run with recipe = 'corrections_combined'")
expect_equal(hit$available_years, c(1967L, 2023L))
})
test_that("no signposting when the result already has full year coverage", {
skip_if_no_corpus()
r <- cog_spending("121011212191", years = 2019:2020, category = "Corrections")
prov <- attr(r, "provenance")
expect_length(prov$suggestions, 0L)
})
test_that("no signposting when category is NULL (unscoped query)", {
skip_if_no_corpus()
r <- cog_spending("121011212191", years = c(2011L, 2012L))
prov <- attr(r, "provenance")
expect_length(prov$suggestions, 0L)
})
test_that("no signposting under basis = 'raw'", {
skip_if_no_corpus()
r <- cog_spending("121011212191", years = c(2011L, 2012L),
category = "Corrections", basis = "raw")
prov <- attr(r, "provenance")
expect_length(prov$suggestions, 0L)
})
@@ -0,0 +1,139 @@
# Madison walkthrough audit -- finding F-014. Tracked as uscogdata#12.
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
#
# cog_revenue()'s flow_prefixes = c("T","A","U","B","C","D") never returns
# item-code prefix X (Employee Retirement) or Y (other Insurance Trust). Per
# Census's standard identity, Total Revenue = General + Utility + Liquor Store +
# Insurance Trust Revenue, and Employee Retirement System contributions and
# earnings ARE the Insurance Trust Revenue component -- so prefix X sits inside
# a published Census revenue concept exactly the way I89 sits inside Census's
# Direct Expenditure concept (finding F-012).
#
# RULED 2026-07-30. `revenue_concept = c("general", "total")` mirrors
# `expenditure_concept`, and the two values are Census's two published revenue
# concepts, related by the manual's own identity (section 4.3, which defines
# the first by SUBTRACTING from the second):
#
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
#
# so `general` is the four general subtypes (own_source/federal/state/
# local_aid) and `total` is every revenue subtype. Naming utility (A91-A94)
# and liquor store (A90) separately is what makes BOTH computable -- before
# cog_pipeline#79 they sat in own_source, so the default was really
# "General + Utility + Liquor", a concept Census does not publish.
#
# Fixture reproducibility: Madison's own X-prefix revenue (FY1970-FY1986,
# $15,098,000 nominal, $0 thereafter) is outside the bundled fixture's year
# window (2011/2012/2019/2020), so the same invariant is asserted on Wisconsin
# state government FY2012, where the fixture carries nonzero X01/X02/X05/X08.
test_that("cog_revenue() can return Census Total Revenue including Insurance Trust (prefix X)", {
wi_state <- "550000227544" # WISCONSIN (state government)
# Revenue-shaped Employee Retirement codes, read from the RAW corpus rather
# than through cog_revenue(), which is the filter under test:
# X01/X02 employee contributions, X05 contributions from other governments,
# X08 total earnings on investments.
#
# X04 and X06 are deliberately NOT in this set, though an earlier draft of
# this test included X04. Both are exhibit codes for INTRAgovernmental
# transfers (the administering government paying into its own fund), which
# X05's own definition excludes by name. Census agrees: its computed "Total
# Emp Ret Rev" for this government-year is exactly the four codes below.
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X02", "X05", "X08"))
expect_equal(x_revenue, 2283883) # 615,835 + 245,083 + 560,382 + 862,583
# The Y-prefix insurance trust revenue (unemployment + workers comp), which
# is the other half of the same Census concept.
y_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("Y01", "Y11"))
expect_equal(y_revenue, 1259785)
general <- cog_revenue(govid = wi_state, years = 2012L)
expect_equal(attr(general, "provenance")$revenue_concept, "general")
expect_equal(sum(general$amt_nominal), 31338293000)
total <- cog_revenue(govid = wi_state, years = 2012L, revenue_concept = "total")
expect_equal(attr(total, "provenance")$revenue_concept, "total")
# total - general is the whole insurance trust leg, X and Y together.
# Asserted as a delta as well as a level so this stays correct however the
# utility/liquor families land (both are $0 for WI state in FY2012).
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
(x_revenue + y_revenue) * 1000)
expect_equal(sum(total$amt_nominal), 34881961000)
expect_true(all(c("X01", "X02", "X05", "X08") %in% wt_codes_included(total)))
# Sibling codes under the SAME first letter must stay out: X11/X12 are
# benefit payments (an expenditure) and X21/X30/X47 are cash and securities
# holdings (a balance-sheet stock). This is the F-018 point restated on the
# revenue side -- the split comes from the crosswalk, not from the letter X.
expect_false(any(c("X11", "X12", "X21", "X30", "X47") %in% wt_codes_included(total)))
# Every returned row still resolves to a category (cog_pipeline#79 added the
# X crosswalk rows; relaxing a prefix filter alone would have produced
# category = NA rows).
expect_false(any(is.na(total$category)))
})
test_that("revenue_concept = 'general' is the default and is strict Census General Revenue", {
wi_state <- "550000227544"
default <- cog_revenue(govid = wi_state, years = 2012L)
explicit <- cog_revenue(govid = wi_state, years = 2012L,
revenue_concept = "general")
expect_equal(sum(default$amt_nominal), sum(explicit$amt_nominal))
# General Revenue excludes utility, liquor store AND insurance trust
# revenue. WI state carries $0 of utility/liquor in FY2012, so the level
# assertion above cannot see those two -- assert the subtype scope directly.
#
# A subset, not setequal: `state` means "intergovernmental revenue FROM the
# state government" (the C codes), which a STATE government does not receive
# from itself, so it is legitimately absent here.
expect_true(all(default$revenue_subtype %in%
c("own_source", "federal", "state", "local_aid")))
expect_false(any(c("utility", "liquor_store", "insurance_trust") %in%
default$revenue_subtype))
})
test_that("utility and liquor store revenue are inside `total` and outside `general`", {
# A city, where utility revenue is material: this is the case the WI state
# baseline structurally cannot exercise. Measured on the fixture, utility +
# liquor is 15.9% of what cog_revenue() returned for type-2 governments
# before the general/total split, so this is the largest behaviour change
# the concept split introduces.
con <- uscogdata:::.ensure_session()
gov <- DBI::dbGetQuery(con,
"SELECT canonical_govid, SUM(amt) amt FROM long
WHERE year = 2012 AND type = 2 AND NOT is_aggregate
AND item_code IN ('A91','A92','A93','A94')
GROUP BY 1 ORDER BY amt DESC LIMIT 1")$canonical_govid
util_raw <- wt_raw_amt(gov, 2012L, codes = c("A90", "A91", "A92", "A93", "A94"))
expect_gt(util_raw, 0)
general <- cog_revenue(govid = gov, years = 2012L)
total <- cog_revenue(govid = gov, years = 2012L, revenue_concept = "total")
expect_false(any(c("utility", "liquor_store") %in% general$revenue_subtype))
expect_true("utility" %in% total$revenue_subtype)
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
util_raw * 1000 +
wt_raw_amt(gov, 2012L, codes = c("Y01", "Y11", "X01", "X02",
"X05", "X08")) * 1000)
})
test_that("revenue_concept rejects unknown values and never returns a balance row", {
expect_error(
cog_revenue("550000227544", years = 2012L, revenue_concept = "gross"),
class = "uscogdata_invalid_revenue_concept"
)
# uscogdata#25 restated for the widest revenue concept: stocks are not
# flows, and `total` must not quietly admit the X/Y/W/Z balance families.
con <- uscogdata:::.ensure_session()
balance <- DBI::dbGetQuery(con,
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
total <- cog_revenue("550000227544", years = 2012L, revenue_concept = "total")
expect_false(any(total$category %in% balance$category))
expect_length(intersect(wt_codes_included(total), balance$item_code), 0L)
})
+22 -5
View File
@@ -1,24 +1,24 @@
test_that("cog_revenue returns expected shape for Broward Property Tax 2020", { test_that("cog_revenue returns expected shape for Broward Property Tax 2020", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_revenue("101006006", years = 2020L, category = "Property Tax") r <- cog_revenue("121011212191", years = 2020L, category = "Property Tax")
expect_s3_class(r, "tbl_df") expect_s3_class(r, "tbl_df")
expected_cols <- c("year", "canonical_govid", "gov_name", "revenue_subtype", expected_cols <- c("year", "canonical_govid", "gov_name", "revenue_subtype",
"category", "amt_nominal", "codes_included", "category", "amt_nominal", "codes_included",
"aggregate_fallback", "notes") "aggregate_fallback", "notes")
expect_true(all(expected_cols %in% names(r))) expect_true(all(expected_cols %in% names(r)))
expect_equal(unique(r$canonical_govid), "101006006") expect_equal(unique(r$canonical_govid), "121011212191")
expect_equal(unique(r$year), 2020L) expect_equal(unique(r$year), 2020L)
}) })
test_that("cog_revenue with no category filter returns multiple categories", { test_that("cog_revenue with no category filter returns multiple categories", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_revenue("101006006", years = 2020L) r <- cog_revenue("121011212191", years = 2020L)
expect_gt(length(unique(r$category)), 1L) expect_gt(length(unique(r$category)), 1L)
}) })
test_that("cog_revenue with per_capita + adjust_to_year adds all columns", { test_that("cog_revenue with per_capita + adjust_to_year adds all columns", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_revenue("101006006", 2020L, r <- cog_revenue("121011212191", 2020L,
per_capita = TRUE, adjust_to_year = 2022L) per_capita = TRUE, adjust_to_year = 2022L)
expect_true(all(c("amt_nominal", "amt_real", expect_true(all(c("amt_nominal", "amt_real",
"amt_per_capita_nominal", "amt_per_capita_real") %in% "amt_per_capita_nominal", "amt_per_capita_real") %in%
@@ -27,7 +27,7 @@ test_that("cog_revenue with per_capita + adjust_to_year adds all columns", {
test_that("cog_revenue result has provenance attribute", { test_that("cog_revenue result has provenance attribute", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_revenue("101006006", 2020L) r <- cog_revenue("121011212191", 2020L)
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
expect_equal(prov$verb, "cog_revenue") expect_equal(prov$verb, "cog_revenue")
expect_true(grepl("revenue_annotated", prov$sql_query)) expect_true(grepl("revenue_annotated", prov$sql_query))
@@ -36,3 +36,20 @@ test_that("cog_revenue result has provenance attribute", {
test_that("cog_revenue rejects invalid inputs", { test_that("cog_revenue rejects invalid inputs", {
expect_error(cog_revenue(list(), 2020L), "character|data frame") expect_error(cog_revenue(list(), 2020L), "character|data frame")
}) })
test_that("cog_revenue basis = 'harmonized' (default) matches 'raw' in this fixture window", {
skip_if_no_corpus()
r_raw <- cog_revenue("121011212191", 2019:2020, basis = "raw")
r_harm <- cog_revenue("121011212191", 2019:2020, basis = "harmonized")
expect_equal(attr(r_raw, "provenance")$basis, "raw")
expect_equal(attr(r_harm, "provenance")$basis, "harmonized")
expect_equal(sum(r_raw$amt_nominal), sum(r_harm$amt_nominal))
})
test_that("cog_revenue provenance carries the harmonization block", {
skip_if_no_corpus()
r <- cog_revenue("121011212191", 2020L)
h <- attr(r, "provenance")$harmonization
expect_true(h$applied)
expect_true(h$na_rows_excluded >= 0L)
})
+53 -12
View File
@@ -2,9 +2,9 @@ test_that("cog_geographic_rollup aggregates state + county + city layers", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list( govids = list(
state = "100000000", # Florida state govt state = "120000226351", # Florida state govt
county = "101006006", # Broward County county = "121011212191", # Broward County
city = "102006004" # Fort Lauderdale City city = "122011161585" # Fort Lauderdale City
), ),
category = "Police", category = "Police",
years = 2019:2020 years = 2019:2020
@@ -23,7 +23,7 @@ test_that("cog_geographic_rollup aggregates state + county + city layers", {
test_that("cog_geographic_rollup respects per_capita + adjust_to_year", { test_that("cog_geographic_rollup respects per_capita + adjust_to_year", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list(county = "101006006", city = "102006004"), govids = list(county = "121011212191", city = "122011161585"),
category = "Police", category = "Police",
years = 2020L, years = 2020L,
per_capita = TRUE, per_capita = TRUE,
@@ -43,8 +43,8 @@ test_that("cog_geographic_rollup respects per_capita + adjust_to_year", {
test_that("cog_geographic_rollup scope_notes describe each layer", { test_that("cog_geographic_rollup scope_notes describe each layer", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list(state = "100000000", county = "101006006", govids = list(state = "120000226351", county = "121011212191",
city = "102006004"), city = "122011161585"),
category = "Police", years = 2020L category = "Police", years = 2020L
) )
state_notes <- unique(r$scope_note[r$layer == "state"]) state_notes <- unique(r$scope_note[r$layer == "state"])
@@ -58,7 +58,7 @@ test_that("cog_geographic_rollup scope_notes describe each layer", {
test_that("cog_geographic_rollup single-layer call works", { test_that("cog_geographic_rollup single-layer call works", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list(county = c("101006006")), govids = list(county = c("121011212191")),
category = "Corrections", category = "Corrections",
years = 2020L years = 2020L
) )
@@ -69,7 +69,7 @@ test_that("cog_geographic_rollup single-layer call works", {
test_that("cog_geographic_rollup provenance reports the outer verb", { test_that("cog_geographic_rollup provenance reports the outer verb", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list(state = "100000000", county = "101006006"), govids = list(state = "120000226351", county = "121011212191"),
category = "Police", years = 2020L category = "Police", years = 2020L
) )
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
@@ -80,8 +80,11 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
test_that("cog_geographic_rollup accepts data.frames per layer", { test_that("cog_geographic_rollup accepts data.frames per layer", {
skip_if_no_corpus() skip_if_no_corpus()
fl_state <- cog_gov_search("^FLORIDA STATE GOVT$", type = "state") # Unanchored: utility mode matches literally now, so "^...$" would be
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county") # searched for as characters rather than read as anchors (uscogdata#16).
# Both still resolve to exactly one row once scoped by type/state.
fl_state <- cog_gov_search("FLORIDA", type = "state")
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
r <- cog_geographic_rollup( r <- cog_geographic_rollup(
govids = list(state = fl_state, county = broward), govids = list(state = fl_state, county = broward),
category = "Police", years = 2020L category = "Police", years = 2020L
@@ -92,9 +95,47 @@ test_that("cog_geographic_rollup accepts data.frames per layer", {
test_that("cog_geographic_rollup rejects invalid inputs", { test_that("cog_geographic_rollup rejects invalid inputs", {
expect_error(cog_geographic_rollup(list(), "Police", 2020L), "length") expect_error(cog_geographic_rollup(list(), "Police", 2020L), "length")
expect_error(cog_geographic_rollup(c("101006006"), "Police", 2020L), "list") expect_error(cog_geographic_rollup(c("121011212191"), "Police", 2020L), "list")
expect_error( expect_error(
cog_geographic_rollup(list(planet = "100000000"), "Police", 2020L), cog_geographic_rollup(list(planet = "120000226351"), "Police", 2020L),
"state|county|city" "state|county|city"
) )
}) })
test_that("cog_geographic_rollup per-capita uses summed per-year populations", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_geographic_rollup(
govids = list(state = "010000226085",
county = "121011212191"),
category = "Police",
years = 2019:2020,
per_capita = TRUE
)
state_ops <- r[r$layer == "state" & r$spend_subtype == "operations", ]
county_ops <- r[r$layer == "county" & r$spend_subtype == "operations", ]
state_implied <- state_ops$amt_nominal / state_ops$amt_per_capita_nominal
county_implied <- county_ops$amt_nominal /
county_ops$amt_per_capita_nominal
# Per-year, per-layer denominator is the layer's own per-year population
expect_equal(state_implied[state_ops$year == 2019], 4874747, tolerance = 1)
expect_equal(state_implied[state_ops$year == 2020], 4903185, tolerance = 1)
expect_equal(county_implied[county_ops$year == 2019], 1935878, tolerance = 1)
})
})
test_that("cog_geographic_rollup records included/excluded govids in provenance", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_geographic_rollup(
govids = list(county = "121011212191"),
category = "Police",
years = 2019:2020,
per_capita = TRUE
)
prov <- attr(r, "provenance")
expect_true("rollup" %in% names(prov))
expect_true("121011212191" %in% prov$rollup$included_govids)
expect_true(is.character(prov$rollup$excluded_govids))
})
})
+22 -14
View File
@@ -146,7 +146,7 @@ test_that(".resolve_basket_row exact match returns one row", {
expect_equal(out$match_method, "exact") expect_equal(out$match_method, "exact")
expect_equal(out$n_candidates, 1L) expect_equal(out$n_candidates, 1L)
expect_equal(nrow(out$row), 1L) expect_equal(nrow(out$row), 1L)
expect_equal(out$row$canonical_govid, "101006006") expect_equal(out$row$canonical_govid, "121011212191")
expect_equal(out$row$gov_name, "BROWARD COUNTY") expect_equal(out$row$gov_name, "BROWARD COUNTY")
}) })
@@ -157,7 +157,7 @@ test_that(".resolve_basket_row exact match is case-insensitive", {
) )
expect_equal(out$status, "resolved") expect_equal(out$status, "resolved")
expect_equal(out$match_method, "exact") expect_equal(out$match_method, "exact")
expect_equal(out$row$canonical_govid, "101006006") expect_equal(out$row$canonical_govid, "121011212191")
}) })
test_that(".resolve_basket_row exact match honors per-row type", { test_that(".resolve_basket_row exact match honors per-row type", {
@@ -166,7 +166,7 @@ test_that(".resolve_basket_row exact match honors per-row type", {
name = "SAN DIEGO CITY", state = "CA", type = "city", con = con name = "SAN DIEGO CITY", state = "CA", type = "city", con = con
) )
expect_equal(out$status, "resolved") expect_equal(out$status, "resolved")
expect_equal(out$row$canonical_govid, "052037010") expect_equal(out$row$canonical_govid, "062073207598")
}) })
test_that(".resolve_basket_row substring fallback resolves single match", { test_that(".resolve_basket_row substring fallback resolves single match", {
@@ -177,7 +177,7 @@ test_that(".resolve_basket_row substring fallback resolves single match", {
expect_equal(out$status, "resolved") expect_equal(out$status, "resolved")
expect_equal(out$match_method, "substring") expect_equal(out$match_method, "substring")
expect_equal(out$n_candidates, 1L) expect_equal(out$n_candidates, 1L)
expect_equal(out$row$canonical_govid, "101006006") expect_equal(out$row$canonical_govid, "121011212191")
}) })
test_that(".resolve_basket_row no_match returns 0-row tibble", { test_that(".resolve_basket_row no_match returns 0-row tibble", {
@@ -207,15 +207,19 @@ test_that(".resolve_basket_row treats empty/whitespace name as no_match", {
test_that(".resolve_basket_row largest_pop within single type", { test_that(".resolve_basket_row largest_pop within single type", {
# FL Miami substring matches 10 cities (all govs_type = 2), largest pop # FL Miami substring matches 10 cities (all govs_type = 2), largest pop
# is MIAMI CITY at 443665. # is MIAMI CITY at 443665. Under Phase P canonical naming, MIAMI-DADE
# COUNTY (govs_type = 1) also contains "Miami", so `type = "city"` pins
# the match set to a single type (as the query docs promise it will for
# per-row `type`), keeping this test's original intent: multiple
# same-type name matches resolve to the largest-population row.
con <- uscogdata:::.ensure_session() con <- uscogdata:::.ensure_session()
out <- uscogdata:::.resolve_basket_row( out <- uscogdata:::.resolve_basket_row(
name = "Miami", state = "FL", type = NA_character_, con = con name = "Miami", state = "FL", type = "city", con = con
) )
expect_equal(out$status, "largest_pop") expect_equal(out$status, "largest_pop")
expect_equal(out$match_method, "substring") expect_equal(out$match_method, "substring")
expect_gte(out$n_candidates, 2L) expect_gte(out$n_candidates, 2L)
expect_equal(out$row$canonical_govid, "102013013") expect_equal(out$row$canonical_govid, "122086194757")
expect_equal(out$row$gov_name, "MIAMI CITY") expect_equal(out$row$gov_name, "MIAMI CITY")
}) })
@@ -241,7 +245,7 @@ test_that(".resolve_basket_row resolves with type override on ambiguous case", {
) )
expect_equal(out$status, "resolved") expect_equal(out$status, "resolved")
expect_equal(out$match_method, "substring") expect_equal(out$match_method, "substring")
expect_equal(out$row$canonical_govid, "052037010") expect_equal(out$row$canonical_govid, "062073207598")
}) })
# ---- basket mode public surface ---- # ---- basket mode public surface ----
@@ -254,7 +258,7 @@ test_that("cog_gov_search basket mode resolves clean inputs in input order", {
) )
expect_s3_class(basket, "tbl_df") expect_s3_class(basket, "tbl_df")
expect_equal(nrow(basket), 3L) expect_equal(nrow(basket), 3L)
expect_equal(basket$canonical_govid, c("101006006", "052037010", "442227001")) expect_equal(basket$canonical_govid, c("121011212191", "062073207598", "482453176394"))
expect_equal(basket$gov_name, c("BROWARD COUNTY", "SAN DIEGO CITY", "AUSTIN CITY")) expect_equal(basket$gov_name, c("BROWARD COUNTY", "SAN DIEGO CITY", "AUSTIN CITY"))
}) })
@@ -284,7 +288,7 @@ test_that("cog_gov_search basket mode skips ambiguous and no_match rows", {
)) ))
# Broward resolves; San Diego ambiguous; Notarealplace no_match. # Broward resolves; San Diego ambiguous; Notarealplace no_match.
expect_equal(nrow(basket), 1L) expect_equal(nrow(basket), 1L)
expect_equal(basket$canonical_govid, "101006006") expect_equal(basket$canonical_govid, "121011212191")
res <- attr(basket, "resolution") res <- attr(basket, "resolution")
expect_equal(nrow(res), 3L) expect_equal(nrow(res), 3L)
expect_equal(res$status, c("resolved", "ambiguous", "no_match")) expect_equal(res$status, c("resolved", "ambiguous", "no_match"))
@@ -306,20 +310,24 @@ test_that("cog_gov_search basket mode recycles single state", {
state = "CA" state = "CA"
) )
expect_equal(nrow(basket), 2L) expect_equal(nrow(basket), 2L)
expect_equal(basket$canonical_govid, c("052037010", "052001009")) expect_equal(basket$canonical_govid, c("062073207598", "062001123093"))
}) })
test_that("cog_gov_search basket mode within-type largest_pop records candidates", { test_that("cog_gov_search basket mode within-type largest_pop records candidates", {
skip_if_no_corpus() skip_if_no_corpus()
# `type = "city"` for the Miami row pins the match set to govs_type = 2;
# under Phase P canonical naming MIAMI-DADE COUNTY also contains "Miami"
# and would otherwise make this an ambiguous (cross-type) match.
basket <- suppressMessages(cog_gov_search( basket <- suppressMessages(cog_gov_search(
name = c("Miami", "OAKLAND CITY"), name = c("Miami", "OAKLAND CITY"),
state = c("FL", "CA") state = c("FL", "CA"),
type = c("city", NA)
)) ))
expect_equal(nrow(basket), 2L) expect_equal(nrow(basket), 2L)
res <- attr(basket, "resolution") res <- attr(basket, "resolution")
miami_row <- res[res$query_name == "Miami", ] miami_row <- res[res$query_name == "Miami", ]
expect_equal(miami_row$status, "largest_pop") expect_equal(miami_row$status, "largest_pop")
expect_equal(miami_row$canonical_govid, "102013013") expect_equal(miami_row$canonical_govid, "122086194757")
expect_gte(miami_row$n_candidates, 2L) expect_gte(miami_row$n_candidates, 2L)
expect_gte(nrow(miami_row$candidates[[1]]), 2L) expect_gte(nrow(miami_row$candidates[[1]]), 2L)
}) })
@@ -396,7 +404,7 @@ test_that("cog_gov_search basket mode skips per-row excluded type without aborti
)) ))
# Broward should resolve; the special_district row should be no_match. # Broward should resolve; the special_district row should be no_match.
expect_equal(nrow(basket), 1L) expect_equal(nrow(basket), 1L)
expect_equal(basket$canonical_govid, "101006006") expect_equal(basket$canonical_govid, "121011212191")
res <- attr(basket, "resolution") res <- attr(basket, "resolution")
expect_equal(res$status, c("resolved", "no_match")) expect_equal(res$status, c("resolved", "no_match"))
# query_type should record what the user passed for the excluded-type row # query_type should record what the user passed for the excluded-type row
+229 -13
View File
@@ -1,12 +1,12 @@
test_that("cog_spending returns expected shape for Broward Corrections 2020", { test_that("cog_spending returns expected shape for Broward Corrections 2020", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", years = 2020L, category = "Corrections") r <- cog_spending("121011212191", years = 2020L, category = "Corrections")
expect_s3_class(r, "tbl_df") expect_s3_class(r, "tbl_df")
expected_cols <- c("year", "canonical_govid", "gov_name", "spend_subtype", expected_cols <- c("year", "canonical_govid", "gov_name", "spend_subtype",
"category", "amt_nominal", "codes_included", "category", "amt_nominal", "codes_included",
"aggregate_fallback", "notes") "aggregate_fallback", "notes")
expect_true(all(expected_cols %in% names(r))) expect_true(all(expected_cols %in% names(r)))
expect_equal(unique(r$canonical_govid), "101006006") expect_equal(unique(r$canonical_govid), "121011212191")
expect_equal(unique(r$year), 2020L) expect_equal(unique(r$year), 2020L)
expect_equal(unique(r$category), "Corrections") expect_equal(unique(r$category), "Corrections")
expect_true(all(r$spend_subtype %in% c("operations", "capital"))) expect_true(all(r$spend_subtype %in% c("operations", "capital")))
@@ -15,7 +15,7 @@ test_that("cog_spending returns expected shape for Broward Corrections 2020", {
test_that("cog_spending vectorised years + categories", { test_that("cog_spending vectorised years + categories", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2019:2020, r <- cog_spending("121011212191", 2019:2020,
category = c("Corrections", "Police")) category = c("Corrections", "Police"))
expect_true(all(r$year %in% 2019:2020)) expect_true(all(r$year %in% 2019:2020))
expect_true(all(r$category %in% c("Corrections", "Police"))) expect_true(all(r$category %in% c("Corrections", "Police")))
@@ -24,7 +24,7 @@ test_that("cog_spending vectorised years + categories", {
test_that("cog_spending with per_capita adds per-capita nominal column", { test_that("cog_spending with per_capita adds per-capita nominal column", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections", per_capita = TRUE) r <- cog_spending("121011212191", 2020L, "Corrections", per_capita = TRUE)
expect_true("amt_per_capita_nominal" %in% names(r)) expect_true("amt_per_capita_nominal" %in% names(r))
expect_false("amt_real" %in% names(r)) expect_false("amt_real" %in% names(r))
expect_false("amt_per_capita_real" %in% names(r)) expect_false("amt_per_capita_real" %in% names(r))
@@ -34,7 +34,7 @@ test_that("cog_spending with per_capita adds per-capita nominal column", {
test_that("cog_spending with adjust_to_year adds real column", { test_that("cog_spending with adjust_to_year adds real column", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2019:2020, "Corrections", r <- cog_spending("121011212191", 2019:2020, "Corrections",
adjust_to_year = 2022L) adjust_to_year = 2022L)
expect_true("amt_real" %in% names(r)) expect_true("amt_real" %in% names(r))
r2019 <- dplyr::filter(r, year == 2019L) r2019 <- dplyr::filter(r, year == 2019L)
@@ -43,7 +43,7 @@ test_that("cog_spending with adjust_to_year adds real column", {
test_that("cog_spending with per_capita + adjust_to_year adds all columns", { test_that("cog_spending with per_capita + adjust_to_year adds all columns", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections", r <- cog_spending("121011212191", 2020L, "Corrections",
per_capita = TRUE, adjust_to_year = 2022L) per_capita = TRUE, adjust_to_year = 2022L)
expect_true(all(c("amt_nominal", "amt_real", expect_true(all(c("amt_nominal", "amt_real",
"amt_per_capita_nominal", "amt_per_capita_real") %in% "amt_per_capita_nominal", "amt_per_capita_real") %in%
@@ -68,16 +68,16 @@ test_that("cog_spending for unknown govid returns empty tibble + informs", {
test_that("cog_spending records found + missing govids in provenance", { test_that("cog_spending records found + missing govids in provenance", {
skip_if_no_corpus() skip_if_no_corpus()
suppressMessages( suppressMessages(
r <- cog_spending(c("101006006", "XXXINVALID"), 2020L, "Corrections") r <- cog_spending(c("121011212191", "XXXINVALID"), 2020L, "Corrections")
) )
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
expect_equal(sort(prov$scope$govids_found), "101006006") expect_equal(sort(prov$scope$govids_found), "121011212191")
expect_equal(sort(prov$scope$govids_missing), "XXXINVALID") expect_equal(sort(prov$scope$govids_missing), "XXXINVALID")
}) })
test_that("cog_spending result has provenance attribute matching schema", { test_that("cog_spending result has provenance attribute matching schema", {
skip_if_no_corpus() skip_if_no_corpus()
r <- cog_spending("101006006", 2020L, "Corrections") r <- cog_spending("121011212191", 2020L, "Corrections")
prov <- attr(r, "provenance") prov <- attr(r, "provenance")
expect_type(prov, "list") expect_type(prov, "list")
expect_equal(prov$verb, "cog_spending") expect_equal(prov$verb, "cog_spending")
@@ -93,20 +93,24 @@ test_that("cog_spending result has provenance attribute matching schema", {
test_that("cog_spending rejects invalid inputs", { test_that("cog_spending rejects invalid inputs", {
expect_error(cog_spending(list(), 2020L), "character|data frame") expect_error(cog_spending(list(), 2020L), "character|data frame")
expect_error(cog_spending("101006006", "2020"), "years") expect_error(cog_spending("121011212191", "2020"), "years")
}) })
test_that("cog_spending accepts a cog_gov_search result directly", { test_that("cog_spending accepts a cog_gov_search result directly", {
skip_if_no_corpus() skip_if_no_corpus()
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county") # Unanchored: utility mode matches `name` as a literal substring now, so
# "^...$" would be searched for as those characters rather than read as
# anchors (uscogdata#16). Scoped by state and type, the bare name still
# resolves to exactly one row.
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
expect_gt(nrow(picks), 0L) expect_gt(nrow(picks), 0L)
r <- cog_spending(picks, 2020L, "Corrections") r <- cog_spending(picks, 2020L, "Corrections")
expect_equal(unique(r$canonical_govid), "101006006") expect_equal(unique(r$canonical_govid), "121011212191")
}) })
test_that("cog_spending accepts a cog_find_peers result directly", { test_that("cog_spending accepts a cog_find_peers result directly", {
skip_if_no_corpus() skip_if_no_corpus()
peers <- cog_find_peers("101006006", max_peers = 3L) peers <- cog_find_peers("121011212191", max_peers = 3L)
r <- cog_spending(peers, 2020L, "Police") r <- cog_spending(peers, 2020L, "Police")
expect_setequal(unique(r$canonical_govid), expect_setequal(unique(r$canonical_govid),
sort(peers$canonical_govid)) sort(peers$canonical_govid))
@@ -128,3 +132,215 @@ test_that("cog_spending accepts a basket-mode cog_gov_search result", {
expect_s3_class(spending, "tbl_df") expect_s3_class(spending, "tbl_df")
expect_setequal(unique(spending$canonical_govid), basket$canonical_govid) expect_setequal(unique(spending$canonical_govid), basket$canonical_govid)
}) })
test_that("per_capita denominator is the per-year F-33 population", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", years = 2019:2020,
category = "Police", per_capita = TRUE)
r_ops <- r[r$spend_subtype == "operations", ]
# Implied denominator from amt_nominal / amt_per_capita_nominal
implied_pop <- r_ops$amt_nominal / r_ops$amt_per_capita_nominal
names(implied_pop) <- r_ops$year
# Use absolute tolerance: within 1 person of per-year F-33 values.
# Hardcoded values are Broward County's per-year Census F-33 population
# from the bundled fixture (regenerated 2026-07-11 against cog_pipeline
# publish tree, pipeline_commit 1a00925, Phase P schema_version 4).
# 1,940,907 is the static ACS 2018-2022 5-year value the legacy
# implementation would use; we assert it is NOT what we get.
expect_true(abs(implied_pop[["2019"]] - 1935878) < 1)
expect_true(abs(implied_pop[["2020"]] - 1952778) < 1)
expect_false(all(abs(implied_pop - 1940907) < 1))
})
})
test_that("pop_source = 'census_f33' does not produce unavailable-pop note", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", years = 2019L,
category = "Police", per_capita = TRUE)
expect_true(all(r$pop_source == "census_f33"))
expect_true(all(is.na(r$notes) | r$notes == "" |
!grepl("No population denominator", r$notes)))
})
})
test_that("aggregate fallback + unavailable pop produce concatenated notes", {
# Unit-level test of .notes_column with a synthetic data frame so we don't
# depend on having a type-4/5 gov in the fixture.
result <- tibble::tibble(
aggregate_fallback = c(FALSE, TRUE, TRUE),
pop_source = c("census_f33", "census_f33", "unavailable")
)
notes <- uscogdata:::.notes_column(result)
expect_equal(notes[1], "")
expect_equal(notes[2], "Aggregate fallback applied; see cog_explain()")
expect_equal(notes[3],
"Aggregate fallback applied; see cog_explain(); No population denominator available for this gov type")
})
test_that("provenance records per-year denominator metadata", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", years = 2019:2020,
category = "Police", per_capita = TRUE)
pc <- attr(r, "provenance")$transformations$per_capita
expect_true(pc$applied)
expect_match(pc$denominator_source, "Census F-33", fixed = FALSE)
expect_match(pc$denominator_source, "per-year", fixed = TRUE)
expect_equal(pc$pop_source_counts$census_f33, nrow(r))
expect_equal(pc$pop_source_counts$unavailable, 0L)
expect_equal(length(pc$popyear_range), 2L)
})
})
# --- basis = "harmonized" / "raw" (Phase R2, schema v5) --------------------
test_that("basis = 'raw' reproduces the pre-harmonization Broward Police totals", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", years = 2019:2020, category = "Police",
basis = "raw")
# Regression pin captured against the schema v5 fixture (2026-07-18,
# pipeline_commit ece9b32) before basis = "harmonized" existed as a
# concept; these are the same totals the pre-Phase-R2 default query
# returned (spending_annotated is untouched by the harmonized views).
ops <- r$amt_nominal[r$year == 2019L & r$spend_subtype == "operations"]
cap <- r$amt_nominal[r$year == 2020L & r$spend_subtype == "capital"]
expect_equal(ops, 483560000)
expect_equal(cap, 26693000)
expect_equal(attr(r, "provenance")$basis, "raw")
})
})
test_that("basis = 'harmonized' (default) matches 'raw' when no harmonization rule applies", {
skip_if_no_corpus()
with_fixture_corpus({
# Every `method = "collapse"` mapping in the curated harmonization_map
# ends by FY2004 for codes inside the spending/revenue flow-type
# prefixes (E/F/G/K, T/A/U/B/C/D); the one collapse extending to FY2011
# (L38/M38 -> L36/M36) is intergovernmental-transfer (L/M prefix) codes
# that were never part of spending_long/revenue_long to begin with. So
# for the fixture's 2011-2020 window, basis = "harmonized" is a
# data-verified no-op vs "raw" for in-scope codes -- this is the
# positive-control counterpart to the synthetic REPLACE-mechanism test
# in test-views.R, which proves the fold itself works when data exists.
r_raw <- cog_spending("121011212191", c(2011L, 2012L, 2019L, 2020L),
"Police", basis = "raw")
r_harm <- cog_spending("121011212191", c(2011L, 2012L, 2019L, 2020L),
"Police", basis = "harmonized")
expect_equal(attr(r_harm, "provenance")$basis, "harmonized")
expect_equal(
r_harm$amt_nominal[order(r_harm$year, r_harm$spend_subtype)],
r_raw$amt_nominal[order(r_raw$year, r_raw$spend_subtype)]
)
})
})
test_that("basis defaults to 'harmonized' when not passed", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2020L, "Police")
expect_equal(attr(r, "provenance")$basis, "harmonized")
})
})
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
skip_if_no_corpus()
with_fixture_corpus({
# FL state government. The harmonization block is scoped by government,
# year and flow prefix -- NOT by category -- so a Corrections query still
# counts every E/F/G-prefixed row the harmonized basis drops for having
# no harmonized_code. The three that apply here are E21/F21/G21
# (Education NEC, SB184-186, "discontinued_na", wide-era window ending
# FY2011); the other discontinued_na rulings live outside E/F/G.
# See docs/phase_r_harmonization_review.md § 1.3/1.4 and cog_pipeline
# data/harmonization_map.csv.
r <- cog_spending("120000226351", 2011:2012, "Corrections")
prov <- attr(r, "provenance")
expect_equal(prov$basis, "harmonized")
expect_true(prov$harmonization$applied)
expect_equal(prov$harmonization$na_rows_excluded, 3L)
# $2,825,439 thousands of FY2011 E21 + F21 + G21, reported in full USD.
# Pinning a non-zero amount is the point: the earlier Broward anchor's
# three rows were all explicit zeros, so the AMOUNT accounting was
# asserted only against 0 and could not have caught a bug.
expect_equal(prov$harmonization$na_amount_excluded, 2825439 * 1000)
})
})
test_that("sparsification removed the wide era's zero-pads from the exclusion count", {
skip_if_no_corpus()
with_fixture_corpus({
# Broward County FY2011 used to carry E21/F21/G21 rows of exactly $0 --
# the wide era stored every government x every code, zeros included. The
# published corpus no longer does (SB194, cog_pipeline#64), so there is
# now nothing for the harmonized basis to exclude. Absence in a
# dense_source year means Census published $0; it does not mean the
# exclusion machinery stopped working, which the FL state anchor above
# proves independently.
r <- cog_spending("121011212191", 2011:2012, "Corrections")
h <- attr(r, "provenance")$harmonization
expect_true(h$applied)
expect_equal(h$na_rows_excluded, 0L)
expect_equal(h$na_amount_excluded, 0)
})
})
test_that("basis = 'raw' never populates the harmonization exclusion block", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2020L, "Corrections", basis = "raw")
h <- attr(r, "provenance")$harmonization
expect_false(h$applied)
expect_equal(h$na_rows_excluded, 0L)
})
})
test_that("v4 corpus: basis silently resolves to raw (default) with a provenance note", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
r <- cog_spending("121011212191", 2019L, "Police")
prov <- attr(r, "provenance")
expect_equal(prov$basis, "raw")
expect_match(prov$basis_note, "raw", fixed = TRUE)
expect_match(prov$basis_note, "schema_version", fixed = TRUE)
expect_false(prov$harmonization$applied)
})
})
test_that("v4 corpus: explicit basis = 'harmonized' aborts", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
expect_error(
cog_spending("121011212191", 2019L, "Police", basis = "harmonized"),
class = "uscogdata_basis_unsupported"
)
})
})
test_that("provenance$series_break_refs is a populated-when-applicable character vector", {
skip_if_no_corpus()
with_fixture_corpus({
r <- cog_spending("121011212191", 2020L, "Corrections")
refs <- attr(r, "provenance")$series_break_refs
expect_type(refs, "character")
# No catalogued code-specific series_breaks_pq row falls inside this
# fixture's 2011/2012/2019/2020 window for the codes this query touches
# (E04/G04) -- data-verified; the mechanism itself is what's under test
# here, via a query-shaped unit test in test-views.R since the fixture
# has no positive case to pin against. Corpus-wide ("ALL") entries never
# appear in this field by construction -- they travel in
# corpus_break_refs; see test-corpus-breaks.R.
expect_equal(refs, character(0))
})
})
test_that("v4 corpus: explicit basis = 'raw' still works", {
skip_if_no_corpus()
with_doctored_schema_version(4L, {
r <- cog_spending("121011212191", 2019L, "Police", basis = "raw")
expect_equal(attr(r, "provenance")$basis, "raw")
expect_gt(nrow(r), 0L)
})
})
+423 -11
View File
@@ -9,19 +9,383 @@ test_that("all expected views register on session open", {
expected <- c( expected <- c(
"long", "spending_long", "revenue_long", "long", "spending_long", "revenue_long",
"canonical_fips_xwalk", "summary_categories", "canonical_fips_xwalk", "summary_categories",
"spending_annotated", "revenue_annotated" "spending_annotated", "revenue_annotated",
"ig_long", "ig_annotated",
"ig_long_harmonized", "ig_annotated_harmonized"
) )
expect_true(all(expected %in% views$table_name)) expect_true(all(expected %in% views$table_name))
}) })
test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", { test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (real SQL text, synthetic parquet)", {
# spending_long_harmonized / revenue_long_harmonized are three-predicate
# views:
# SELECT * REPLACE (harmonized_code AS item_code)
# FROM long
# WHERE NOT is_aggregate
# AND harmonized_code IS NOT NULL
# AND LEFT(harmonized_code, 1) IN (<flow prefixes>)
# None of the curated harmonization_map's `collapse` rulings land inside
# the bundled fixture's 2011-2020 window for spending/revenue-prefixed
# codes (see the "basis = 'harmonized' (default) matches 'raw'" test in
# test-spending.R and docs/phase_r_harmonization_review.md § 0.2/§ 2), so
# there is no real fixture row that exercises a nonzero fold or a
# predicate-excluded row. Rather than re-implement the WHERE clause by
# hand against an in-memory VALUES table (which would only prove the SQL
# *pattern* works, not that the deployed inst/sql/22-/23- text actually
# applies it), this test reads the real SQL files off disk, substitutes
# {url} exactly as .register_views() does, and executes them -- plus
# their 10-long.sql dependency -- against a synthetic hive-partitioned
# parquet tree written to a temp dir. A regression in any predicate (e.g.
# `NOT is_aggregate` dropped, the crosswalk-membership subquery changed,
# the NULL guard removed) would change which of the rows below survive.
#
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
# PARQUET) rather than the arrow package: this package has no arrow
# dependency (CLAUDE.md "No arrow dependency -- DuckDB reads parquet
# natively"), and DuckDB can round-trip its own parquet writer/reader
# without adding one for tests either.
skip_if_no_corpus()
tmp <- withr::local_tempdir()
part_dir <- file.path(tmp, "data", "long", "year=2004")
dir.create(part_dir, recursive = TRUE)
part_path <- file.path(part_dir, "part-0.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
-- Spending (E/F/G/K) rows, exercised against spending_long_harmonized:
('spend-A', 'E36', 100, false, 'E36'), -- control: passes every predicate as-is
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
-- 'S74' and 'Z61' are classified `balance` in the synthetic
-- crosswalk below (mirroring the real corpus's own non-flow codes),
-- so each is excluded from its view ONLY by the crosswalk-membership
-- subquery -- the mechanism that replaced the prefix allowlists
-- (uscogdata#11) and keeps balance stocks out of both flows
-- (uscogdata#25).
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by crosswalk membership (balance)
-- Revenue rows, exercised against revenue_long_harmonized:
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by crosswalk membership (balance)
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The flow views classify by membership in summary_categories, so the
# synthetic corpus needs one too. Every flow code above is a member of its
# own flow (so is_aggregate / NULL-harmonized exclusions stay the SOLE
# excluder for those rows); S74/Z61 are members but classified balance, so
# membership itself is what excludes them.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('E36', 'Water Utilities', 'expenditure', 'operations', NULL),
('E38', 'Water Utilities', 'expenditure', 'operations', NULL),
('E05', 'Corrections', 'expenditure', 'operations', NULL),
('E99', 'Other', 'expenditure', 'operations', NULL),
('S74', 'Fund Balances', 'balance', NULL, NULL),
('U11', 'Interest Earnings','revenue', NULL, 'own_source'),
('U10', 'Interest Earnings','revenue', NULL, 'own_source'),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source'),
('T88', 'Other Taxes', 'revenue', NULL, 'own_source'),
('Z61', 'Fund Balances', 'balance', NULL, NULL)
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
}
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
spend <- DBI::dbGetQuery(con,
"SELECT item_code, SUM(amt) AS amt FROM spending_long_harmonized
GROUP BY item_code ORDER BY item_code"
)
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
# harmonized_code), and spend-E (balance, not an expenditure member) must
# all be gone, and spend-A + spend-B must be folded together under E36.
expect_equal(nrow(spend), 1L)
expect_equal(spend$item_code, "E36")
expect_equal(spend$amt, 150)
rev <- DBI::dbGetQuery(con,
"SELECT item_code, SUM(amt) AS amt FROM revenue_long_harmonized
GROUP BY item_code ORDER BY item_code"
)
expect_equal(nrow(rev), 1L)
expect_equal(rev$item_code, "U11")
expect_equal(rev$amt, 225)
})
test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmonized_code, and exclude the L-- family total (real SQL text, synthetic parquet)", {
# ig_long / ig_long_harmonized have the subtlest predicates in the package:
# a deliberately ABSENT `NOT is_aggregate` (unlike every other *_long view),
# and COALESCE(harmonized_code, item_code) instead of a plain
# `harmonized_code IS NOT NULL` filter. The only end-to-end guard on this
# today is bound to AL state / 2011 / Education K-12, where M12 happens to
# be the sole IG code present -- regenerate the fixture without that one
# row and the guard would die silently while staying green. As with the
# 22-/23- test above, this reads the real inst/sql/24-/25- text off disk
# and executes it against a synthetic hive-partitioned parquet tree, so a
# regression in either predicate changes which rows survive.
skip_if_no_corpus()
tmp <- withr::local_tempdir()
part_dir <- file.path(tmp, "data", "long", "year=2004")
dir.create(part_dir, recursive = TRUE)
part_path <- file.path(part_dir, "part-0.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is
('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis
('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views
('ig-D', 'L--', 55555, false, 'L--'), -- family total: deliberately NOT a crosswalk member, excluded from BOTH views
('ig-E', 'T29', 44444, false, 'T29') -- revenue member, not intergovernmental: excluded from BOTH views
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
# The IG views classify by summary_categories membership
# (spend_subtype = 'intergovernmental'). L-- is deliberately absent --
# exactly as it is from the real crosswalk -- which is what excludes it.
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('M04', 'Corrections', 'expenditure', 'intergovernmental', NULL),
('M38', 'Health', 'expenditure', 'intergovernmental', NULL),
('M36', 'Health', 'expenditure', 'intergovernmental', NULL),
('M47', 'IG Other', 'expenditure', 'intergovernmental', NULL),
('T29', 'Other Taxes', 'revenue', NULL, 'own_source')
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
}
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
DBI::dbExecute(con, .read_view_sql("24-ig_long.sql"))
DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql"))
raw <- DBI::dbGetQuery(con,
"SELECT item_code, SUM(amt) AS amt FROM ig_long
GROUP BY item_code ORDER BY item_code"
)
# L-- (family total, not a member) and T29 (revenue, not IG) are gone; the
# aggregate row M47 survives -- proof `NOT is_aggregate` is absent from
# ig_long.
expect_equal(raw$item_code, c("M04", "M38", "M47"))
expect_equal(raw$amt, c(100, 50, 99999))
harmonized <- DBI::dbGetQuery(con,
"SELECT item_code, SUM(amt) AS amt FROM ig_long_harmonized
GROUP BY item_code ORDER BY item_code"
)
# M38 folds to M36 (real harmonized_code present); M47 keeps its raw code
# via COALESCE(NULL, 'M47') -- proof the aggregate row is NOT dropped by
# a plain `harmonized_code IS NOT NULL` filter. L-- and T29 stay excluded.
expect_equal(harmonized$item_code, c("M04", "M36", "M47"))
expect_equal(harmonized$amt, c(100, 50, 99999))
})
test_that(".build_series_break_refs matches fin_code + break_year window", {
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
# 2011-2020 window (data-verified; see the "series_break_refs" test in
# test-spending.R), so this proves the matching logic itself against the
# live view + a synthetic year window that DOES hit a cataloged break
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
# separate path with its own coverage -- SB194 does sit at 2012, inside
# the fixture window; see test-corpus-breaks.R.
skip_if_no_corpus() skip_if_no_corpus()
con <- cog_open() con <- cog_open()
on.exit(cog_close()) on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con, refs <- uscogdata:::.build_series_break_refs(
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM spending_long" con, codes_observed = c("E62", "E04"), years = c(2003L, 2006L),
)$pfx schema_version = 5L
expect_true(all(prefixes %in% c("E", "F", "G", "K"))) )
expect_true("SB075" %in% refs)
expect_true("SB071" %in% refs)
# Gated on schema_version >= 5 even when the codes/years would otherwise match.
refs_v4 <- uscogdata:::.build_series_break_refs(
con, codes_observed = c("E62"), years = c(2003L, 2006L), schema_version = 4L
)
expect_equal(refs_v4, character(0))
})
test_that("schema v5 harmonization views register when the corpus supports them", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
manifest <- uscogdata:::.uscogdata_env$manifest
skip_if(as.integer(manifest$schema_version) < 5L, "fixture is schema_version < 5")
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'"
)
expected_v5 <- c(
"spending_long_harmonized", "revenue_long_harmonized",
"spending_annotated_harmonized", "revenue_annotated_harmonized",
"harmonization_map", "harmonization_recipes", "series_breaks_pq"
)
expect_true(all(expected_v5 %in% views$table_name))
})
test_that(".harmonization_view_files guard is necessary: registration against a v4-shaped corpus (no harmonized_code column at all) succeeds only because the harmonized views are skipped", {
# with_doctored_schema_version() (used elsewhere in this suite) only
# rewrites manifest.json's schema_version -- the underlying `long` parquet
# is still the bundled v6 fixture, which DOES have a harmonized_code
# column, so it only proves the skip *happens*, not that it is *required*.
# This test builds a genuinely v4-shaped corpus: `long` has no
# harmonized_code column at all, matching a real pre-Phase-R2 publish
# tree, and then shows two things: (1) the real .register_views(), gated
# on manifest$schema_version, registers cleanly against it; (2) the exact
# SQL text of a gated file (25-ig_long_harmonized.sql), executed directly
# against the same corpus without the gate, fails -- proving the gate is
# load-bearing, not incidental.
tmp <- withr::local_tempdir()
part_dir <- file.path(tmp, "data", "long", "year=2004")
dir.create(part_dir, recursive = TRUE)
part_path <- file.path(part_dir, "part-0.parquet")
write_con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('gov-1', 'E36', 100, false, 500000, 2020)
) AS t(canonical_govid, item_code, amt, is_aggregate, population, popyear)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(part_path)))
xwalk_path <- file.path(tmp, "data", "canonical_fips_xwalk.parquet")
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('gov-1', 'Test Gov', 1, 'County', '01', '001', NULL, 500000)
) AS t(canonical_govid, gov_name, govs_type, type_label, fips_state,
fips_county, fips_place, population_acs)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(xwalk_path)))
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
DBI::dbExecute(write_con, sprintf("
COPY (
SELECT * FROM (VALUES
('E36', 'Test Category', 'expenditure', 'direct', NULL)
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
) TO %s (FORMAT PARQUET)
", uscogdata:::.sql_lit_chr(cats_path)))
# Confirm the synthetic `long` genuinely lacks harmonized_code (not just
# NULL values -- the column itself must be absent) before trusting the
# rest of this test.
cols <- DBI::dbGetQuery(write_con, sprintf(
"DESCRIBE SELECT * FROM read_parquet(%s)", uscogdata:::.sql_lit_chr(part_path)
))$column_name
expect_false("harmonized_code" %in% cols)
url <- paste0(tmp, "/")
# (1) Full .register_views() against this v4-shaped corpus must succeed --
# this is the behavior the guard exists to protect.
con <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
expect_no_error(
uscogdata:::.register_views(con, url, manifest = list(schema_version = 4L))
)
views <- DBI::dbGetQuery(con,
"SELECT table_name FROM information_schema.tables
WHERE table_schema = 'main' AND table_type = 'VIEW'")$table_name
expect_true(all(c("ig_long", "ig_annotated", "spending_annotated") %in% views))
expect_false(any(c("ig_long_harmonized", "ig_annotated_harmonized",
"spending_long_harmonized") %in% views))
# (2) Prove the gate is load-bearing: the exact SQL text of the skipped
# file, executed directly (bypassing .register_views()'s schema_version
# check) against the SAME corpus, fails because it references
# long.harmonized_code, a column this corpus's `long` does not have.
sql_dir <- system.file("sql", package = "uscogdata")
.read_view_sql <- function(filename) {
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
gsub("\\{url\\}", url, txt, fixed = FALSE)
}
con2 <- DBI::dbConnect(duckdb::duckdb())
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
DBI::dbExecute(con2, .read_view_sql("10-long.sql"))
expect_error(DBI::dbExecute(con2, .read_view_sql("25-ig_long_harmonized.sql")))
# Reconciling this test with the C2 guard (expenditure-concept review):
# `ig_annotated`/`spending_annotated` registering cleanly above proves
# only that CREATE VIEW binds against a `summary_categories` with no M/L
# rows at all (this synthetic corpus's own summary_categories has a
# single E36 row, see the COPY above) -- a LEFT JOIN never fails to
# resolve regardless of what the joined-to table contains. It does NOT
# mean querying expenditure_concept = "total" against this shape is safe:
# exactly this corpus (schema_version reported as supported, but
# summary_categories predates the M/L rows cog_pipeline PR #59 added) is
# what .require_ig_categories() exists to catch at the *verb* level,
# since PR #59 shipped those rows with no schema_version bump. Confirm
# the new runtime guard actually fires against this same `con`.
expect_error(
uscogdata:::.require_ig_categories(con),
class = "uscogdata_ig_categories_unsupported"
)
})
test_that("spending_long carries exactly the non-IG expenditure crosswalk codes and excludes aggregates", {
skip_if_no_corpus()
con <- cog_open()
on.exit(cog_close())
# Classification is crosswalk membership, not prefixes (uscogdata#11):
# every row's code must classify as expenditure and never as
# intergovernmental (which lives in ig_long).
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM spending_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'expenditure'
OR c.spend_subtype = 'intergovernmental'"
)$item_code
expect_length(stray, 0L)
# Balance codes are stocks, not flows -- they must never appear in a
# spending result (uscogdata#25). Prefix filtering could not guarantee
# this (W/X/Y/Z balance codes share letters with flow codes).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con, agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate" "SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
@@ -29,14 +393,29 @@ test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
expect_equal(agg_count, 0) expect_equal(agg_count, 0)
}) })
test_that("revenue_long filters to T/A/U/B/C/D prefixes and excludes aggregates", { test_that("revenue_long carries exactly the revenue crosswalk codes and excludes aggregates", {
skip_if_no_corpus() skip_if_no_corpus()
con <- cog_open() con <- cog_open()
on.exit(cog_close()) on.exit(cog_close())
prefixes <- DBI::dbGetQuery(con,
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM revenue_long" # The view carries EVERY revenue subtype; which of Census's two published
)$pfx # concepts a query returns is decided per `revenue_concept` in R
expect_true(all(prefixes %in% c("T", "A", "U", "B", "C", "D"))) # (uscogdata#12), exactly as `expenditure_concept` narrows spending_long.
stray <- DBI::dbGetQuery(con,
"SELECT DISTINCT s.item_code
FROM revenue_long s
LEFT JOIN summary_categories c USING (item_code)
WHERE c.category_type IS DISTINCT FROM 'revenue'"
)$item_code
expect_length(stray, 0L)
# No balance stock ever appears in a revenue result (uscogdata#25).
balance_n <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE item_code IN (
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
)"
)$n
expect_equal(balance_n, 0)
agg_count <- DBI::dbGetQuery(con, agg_count <- DBI::dbGetQuery(con,
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate" "SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
@@ -57,3 +436,36 @@ test_that("spending_annotated carries category + xwalk columns", {
expect_true(nm %in% names(row), info = paste("missing column:", nm)) expect_true(nm %in% names(row), info = paste("missing column:", nm))
} }
}) })
test_that("gov_population_yearly exposes one row per (year, canonical_govid)", {
skip_if_no_corpus()
with_fixture_corpus({
con <- uscogdata:::.ensure_session()
df <- DBI::dbGetQuery(
con,
"SELECT year, canonical_govid, population, popyear
FROM gov_population_yearly
WHERE canonical_govid = '121011212191'
ORDER BY year"
)
expect_setequal(df$year, c(2011L, 2012L, 2019L, 2020L))
expect_equal(nrow(df), 4L)
expect_true(all(!is.na(df$population)))
# Hardcoded values are from the bundled fixture (regenerated 2026-07-18
# against cog_pipeline publish tree, pipeline_commit ece9b32, Phase R2
# schema_version 5, years 2011/2012/2019/2020). Update if the fixture is
# rebuilt against a different source vintage.
expect_equal(df$population[df$year == 2011L], 1759591L)
expect_equal(df$population[df$year == 2012L], 1819773L)
expect_equal(df$population[df$year == 2019L], 1935878L)
expect_equal(df$population[df$year == 2020L], 1952778L)
# Uniqueness on (year, canonical_govid) across the whole view.
dup <- DBI::dbGetQuery(
con,
"SELECT year, canonical_govid, COUNT(*) AS n
FROM gov_population_yearly
GROUP BY year, canonical_govid HAVING n > 1"
)
expect_equal(nrow(dup), 0L)
})
})
+64
View File
@@ -0,0 +1,64 @@
---
title: "Population denominators"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Population denominators}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
```{r setup, include = FALSE}
knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
```
# Why per-year population matters
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
# The four population sources
| Source | What it is | Default in uscogdata? |
|---|---|---|
| Census F-33 `population` | Population value Census used on each COG row to compute its published per-capita tables. Almost always a Population Estimates Program (PEP) estimate; sometimes lagged a year for fiscal-year alignment, recorded in `popyear`. | **Yes — default for `cog_spending(per_capita = TRUE)` etc.** |
| PEP (raw) | Census Bureau's official annual intercensal estimates, distinct from F-33 because F-33 sometimes uses a lagged vintage. | No (not in corpus) |
| ACS 5-year | American Community Survey 5-year rolling average. Different methodology, has margin of error, only available 2005-2009 onward. | Used by `cog_find_peers()` historically; replaced in 0.1 by per-year F-33. Still available in `canonical_fips_xwalk.population_acs` for non-time-series uses. |
| Decennial count | Actual count, every 10 years. | No (not in corpus) |
The F-33 denominator is preferred because it's the same value Census uses internally — so `uscogdata` per-capita numbers reconcile with Census's own published tables.
# Coverage
F-33 `population` is observed for gov types 0–3 (state, county, city, township). Gov types 4 (special districts) and 5 (school districts) have `population` masked to NA in the F-33 schema. uscogdata returns:
- `pop_source = "census_f33"` and a numeric `amt_per_capita_*` for types 0–3.
- `pop_source = "unavailable"` and `NA` per-capita for types 4–5, with a corresponding entry in `notes`.
`cog_geographic_rollup(per_capita = TRUE)` excludes unavailable-pop rows from the result; the dropped govids are listed in `provenance\$rollup\$excluded_govids`.
# The popyear quirk
Census sometimes uses a population estimate from one year prior to the fiscal year being reported (e.g., FY2018 paired with a 2017 PEP estimate) so the denominator is available before the fiscal year closes. `popyear` records which vintage was paired; `cog_spending()` returns the popyear range in `provenance\$transformations\$per_capita\$popyear_range` rather than as a per-row column.
# Time-varying peer cohorts
`cog_find_peers(target, year = Y)` builds a cohort matched on each candidate's population at year `Y`. The cohort is fixed once chosen; `cog_peer_compare()` then runs that cohort across whatever `years` you ask for. To run a moving-window comparison, build cohorts year-by-year and stitch the results:
```r
years <- 2010:2023
out <- purrr::map_dfr(years, function(y) {
peers <- cog_find_peers("261163166615", year = y, max_peers = 10L)
cog_peer_compare("261163166615", peers,
category = "Police", years = y,
per_capita = TRUE)
})
```
Each row in `out` has `cohort_year == year`, so a faceted plot shows cohort drift directly.
# Future direction
`pop_source` is a column on the result, not a fixed value, so adding a new denominator (PEP from tidycensus, decennial counts, ACS time-series) is a join change rather than an API change. A future release may add `cog_spending(..., pop_source = "pep")` for users who need a single externally-audited series.
+256
View File
@@ -0,0 +1,256 @@
---
title: "Total spending: Primary, Direct, Total, and when each is right"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
```{r setup, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```
# Two questions that sound the same but aren't
"Total spending" means two different things depending on whether the question
is about one government or several:
1. **"What did my county spend in total, a decade ago vs today?"** — one
government, tracked over time. Any concept answers this correctly, as
long as the same concept is used for both years.
2. **"How do all the counties in my state compare, a decade ago vs today,
against the neighboring state?"** — several governments, summed together.
Here only a non-intergovernmental concept (`primary` or `direct`) gives
the right answer; summing `total` across governments double-counts money
that passes between them.
`cog_spending()`'s `expenditure_concept` argument controls which of these a
query answers, via three nested concepts defined as sets of the crosswalk's
`spend_subtype` values (never item-code first letters — the letter `Y` alone
spans revenue, expenditure, and balance codes):
- `"primary"` (the default) — the government's own service provision:
`operations` + `capital` + `assistance`.
- `"direct"` — Census's published Direct Expenditure: `primary` plus
`interest` on debt and `insurance_benefits` (e.g. pension payments).
- `"total"` — `direct` plus the `intergovernmental` leg.
This vignette walks through both questions with code that actually runs
against the package's bundled fixture corpus, then explains why the second
question refuses `"total"` outright.
Before any of the numbers below: every amount column here — `amt_nominal`,
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
dollars**. The raw Census files report **thousands of dollars** and the
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
billion. Do not scale it again.
```{r}
library(uscogdata)
# Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50
# states) so this vignette knits without network access. In real use,
# USCOGDATA_URL is instead set to the published corpus URL -- see README.md.
Sys.setenv(USCOGDATA_URL = paste0(
system.file("extdata/fixture_corpus", package = "uscogdata"), "/"
))
```
The fixture doesn't carry 2017 or the present year, so the examples below use
the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs
today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published
corpus and swap in real years; the mechanics are identical.
# Archetype 1: one government's own trend
For a single government, `total` is a legitimate way to describe "everything
this government spent, including money it handed to other governments to
spend on its behalf":
```{r}
al_total <- cog_spending(
"010000226085", # Alabama, the state government
years = c(2012, 2020),
category = "Highways",
expenditure_concept = "total"
)
al_total
```
The `intergovernmental` rows are what `"total"` adds on top of the
non-intergovernmental subtypes (here `capital` + `operations`): Alabama's
own payments out to counties and cities for highway work. Because this
query only ever concerns Alabama, including that piece is safe -- there's
no other government's number it could be double-counted against.
`"primary"` (the default) answers the same trend question just as validly
(for Highways, which maps only to operations/capital codes, `"primary"` and
`"direct"` coincide -- there is no highway-specific interest or insurance
benefit to add):
```{r}
al_primary <- cog_spending(
"010000226085", years = c(2012, 2020), category = "Highways"
# expenditure_concept = "primary" is the default; shown here for contrast
)
al_primary
```
Both are internally consistent series. What breaks the comparison is
**switching concepts between the two years being compared** -- e.g. `direct`
for 2012 and `total` for 2020 -- which manufactures a trend that isn't
really there. Pick one concept for a given question and hold it fixed across
every year in the series.
# Archetype 2: a cross-government rollup
`cog_geographic_rollup()` sums spending across state/county/city layers for
a place. Its default is `"primary"`, and (as shown below) it accepts only
the non-intergovernmental concepts, `"primary"` and `"direct"`:
```{r}
fl_rollup <- cog_geographic_rollup(
govids = list(
state = "120000226351", # Florida
county = c("121011212191", "121099101897") # Broward + Palm Beach
),
category = "Highways",
years = c(2012, 2020)
)
fl_rollup
```
For the neighboring state, the comparison is a single government, so it's a
plain `cog_spending()` call rather than a rollup:
```{r}
ga_state <- cog_spending(
"130000226087", years = c(2012, 2020), category = "Highways" # Georgia
)
ga_state
```
Now the same rollup, but asking for `expenditure_concept = "total"`:
```{r, error = TRUE}
cog_geographic_rollup(
govids = list(state = "120000226351", county = "121011212191"),
category = "Highways",
years = 2020,
expenditure_concept = "total"
)
```
`cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason)
refuses `"total"` outright rather than silently returning an inflated
number. The next section is why.
# The mechanism
Suppose Alabama gives a county $10M toward a highway project. That $10M
shows up **twice** in the underlying corpus:
- Once on Alabama's own record, coded `M44` ("to local governments,
Highways") -- Alabama's intergovernmental leg.
- Again on the county's record, coded `E44` / `F44` ("Highways, current
operations" / "capital outlay") -- the county's direct spending, because
the county is the government that actually lets the contract and pays the
paving crew.
`primary` and `direct` (the crosswalk's non-intergovernmental expenditure
subtypes) only ever count the second of those -- the government that
actually did the spending. `total` (Direct plus the intergovernmental leg)
counts the first one *as well*, which is exactly right for describing
Alabama's own budget: Alabama's `total` genuinely includes the $10M it
committed to highways, whether it built the road itself or paid the county
to. But sum `total` across Alabama **and** the county, and that $10M is
counted twice -- once as Alabama's payment out, once as the county's
spending in -- reporting $20M of highway work for $10M actually spent.
This is exactly the shape of query `cog_geographic_rollup()` exists to run
(summing across layers of government), so it refuses `"total"` rather than
silently overstating every multi-layer figure it produces.
# How big is the risk in practice
Intergovernmental transfers aren't evenly distributed by government type.
Measured against the bundled fixture corpus (all 50 states, each of its
four years -- 2011, 2012, 2019, 2020), intergovernmental spending as a
share of a government's own Direct spending is:
| Government type | Intergovernmental / Direct |
|---|---|
| State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) |
| County | 3.3%-4.8% (varies by year) |
| City | 2.4%-2.9% (varies by year) |
So the Direct/Total choice matters overwhelmingly for **state**
governments -- a state's Total genuinely differs from its Direct by more
than a third, while for a county or city the two are close. (The state
share is much larger than pre-#11 measurements suggested, because the
intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state
payments to school systems -- for most states the single largest transfer
they make.) That's also why the mistake this vignette warns about is easy
to make unnoticed at the county/city level and costly at the state level:
rolling up every government using `total` instead of `primary`/`direct`
overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally.
# Why Total = Direct + M + L + Q, not Direct + M
It's tempting to assume `total` only needs to add `M`. But the
intergovernmental leg has three families, all money the queried government
itself pays **out** -- they're not different accounts of a receiving
government's revenue. `M` is what it pays to other **local** governments
(e.g. a county paying a city for a shared paving contract); `L` is what it
pays **up** to its **state** government (e.g. a county's contribution to a
state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments
to **school systems** (K-12 and higher-ed aid -- for most states the
single largest transfer they make, and the piece the pre-#11 prefix
allowlist silently dropped, finding F-017). A government's Total genuinely
includes every leg it pays, because each is its own spending, just routed
to a different kind of recipient. On the bundled fixture corpus (all 50
states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has
no "payments to the state government" leg of its own) but is 43%-51% the
size of `M` for counties (varies by year) and 144%-189% the size of `M`
for cities (varies by year; 166% pooled across all four) -- so a `total`
that omitted `L` would silently undercount Total specifically for local
governments, and for cities `L` is often the *larger* of the two legs.
`cog_spending(expenditure_concept = "total")` includes every leg
(excluding the `L--` family-total rollup row, which would double-count its
own components).
# Composition rules
- `expenditure_concept` (whose spending counts -- Primary, Direct, or
Direct plus intergovernmental) is **orthogonal** to `basis` (which
vintage of the item-code space a query resolves against --
`"harmonized"` vs `"raw"`).
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
valid, meaningful query, and so is every other pairing.
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
recipe already defines its own component codes (some recipes have their
own matching intergovernmental counterpart recipe instead -- see
`cog_recipes()` and the "firing suggestion" notes surfaced in
`cog_spending()`'s provenance), so layering a second, generic `total`
union on top of a recipe query has no well-defined meaning. Passing both
together aborts with an error naming the conflict.
- `expenditure_concept` is a **spending-only** concept: `cog_revenue()`
doesn't expose it (revenue's own intergovernmental codes are a different
axis -- see `?cog_revenue`).
# Summary
- Comparing one government to itself over time: any concept works -- pick
one and hold it fixed across every year compared.
- Comparing or summing across governments -- counties within a state, a
state against its neighbor, cities against counties: use `"primary"`
(the default) or `"direct"`. `cog_geographic_rollup()` and
`cog_peer_compare()` enforce this by refusing `"total"`.
- `"primary"` = operations + capital + assistance. `"direct"` = primary +
interest on debt + insurance trust benefits (Census's published Direct
Expenditure). `"total"` = direct + intergovernmental (`M` to local
governments, `L` to the state government excluding the `L--`
family-total row, and `Q11`/`Q12`/`Q18` state payments to school
systems).