diff --git a/.Rbuildignore b/.Rbuildignore index f06fe76..7205877 100644 --- a/.Rbuildignore +++ b/.Rbuildignore @@ -16,3 +16,4 @@ ^Meta$ ^\.gitea$ ^CLAUDE\.md$ +^\.superpowers$ diff --git a/.gitignore b/.gitignore index 7beed34..f03f1e7 100644 --- a/.gitignore +++ b/.gitignore @@ -9,3 +9,6 @@ docs/ /Meta/ .DS_Store /.quarto/ + +# SDD working artifacts (ledger, briefs, review packages) — plans/ stays tracked +.superpowers/sdd/ diff --git a/.superpowers/plans/2026-07-27-expenditure-concept-reader.md b/.superpowers/plans/2026-07-27-expenditure-concept-reader.md new file mode 100644 index 0000000..7db25a6 --- /dev/null +++ b/.superpowers/plans/2026-07-27-expenditure-concept-reader.md @@ -0,0 +1,1008 @@ +# `expenditure_concept` in uscogdata — Implementation Plan (Repo 2 of 3) + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add `expenditure_concept = c("direct","total")` to `cog_spending()`, so a +user can ask for Census "Total" (Direct + intergovernmental) for a single +government — and cannot accidentally get it inside the cross-government verbs, +where summing Total double-counts intergovernmental transfers. + +**Architecture:** The pipeline now publishes 66 intergovernmental (`M`/`L`) codes +in `summary_categories.parquet` under `spend_subtype = "intergovernmental"` +(merged as pipeline PR #59). `"direct"` keeps today's behavior exactly — the +existing `spending_annotated{,_harmonized}` views. `"total"` adds a second leg, +a new `ig_annotated{,_harmonized}` view, UNION'd in. The IG leg deliberately does +**not** reuse the Direct leg's `NOT is_aggregate` filter, because in the legacy +era the IG dollars live almost entirely on aggregate-flagged rows. + +**Tech Stack:** R, DuckDB, testthat (3rd ed), roxygen2, `cli`. + +## Global Constraints + +- **Repo root:** `/home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata`. Always `cd` there in the same command as any git or R invocation. It is NESTED inside the `cog_explorer` git repo, so a bare `git` command can answer from the WRONG repo — confirm with `git rev-parse --show-toplevel` before believing any alarming git result. +- **Branch:** `feat/expenditure-concept` (already created off `main` @ `e581e73`). +- **Interpreter:** `/usr/bin/Rscript`, absolute paths. Conda shadows system R and lacks arrow/duckdb; IDE "package not installed" diagnostics are NOISE. +- **Commits:** `git -c commit.gpgsign=false commit` — a plain commit fails on a GPG "signing failed: Timeout". +- **Measured baseline on `main` (2026-07-27):** `FAIL 0 | WARN 0 | SKIP 0 | PASS 467`. (The earlier handoff's "471 with PR #8 / 463 without" figures are both wrong; 467 is measured.) +- **Tests are offline.** `tests/testthat/setup.R` points `USCOGDATA_URL` at the bundled fixture. Never write a test that requires the live corpus or a network fetch. +- **Do NOT merge the PR.** One PR per repo, stop for owner review. +- **`expenditure_concept` is spending-only.** Do NOT add it to `cog_revenue()` — the Direct/Total distinction is an expenditure concept and the parameter name says so. Revenue's intergovernmental codes (`B`/`C`/`D`) are already categorized by source and are a different axis. +- **Ruling reference:** `../cog_pipeline/docs/superpowers/specs/2026-07-25-total-spending-semantics-design.md` (rulings R1–R4). + +## The two facts this design rests on (both measured against the published corpus) + +**1. Aggregate IG rows carry NO `harmonized_code`.** Harmonized space is +leaf-only by construction: + +| prefix | is_aggregate | harmonized_code NULL | rows | Σ amt | +|---|---|---|---:|---:| +| M | TRUE | **yes** | 2,918,024 | 6,434,679,266 | +| L | TRUE | **yes** | 2,188,518 | 307,079,345 | +| L | FALSE | yes (`L21`,`L24`) | 1,380,245 | 516,301 | + +So the IG leg **cannot** go through `harmonized_code` alone — it would drop +almost all legacy IG. It joins on `COALESCE(harmonized_code, item_code)`, which +keeps the one real IG collapse rule (`M38 → M36`, SB012 — year-disjoint, 1967–2011 +vs 2012+) while never dropping an aggregate row. + +**2. Aggregate IG codes are year-disjoint from their components, so ignoring +`is_aggregate` cannot double-count.** Measured per year: + +| year | `M47` | `M94` | `M89`(agg) | `M91-93` | `L89`(agg) | `L91-93` | `M05`(agg) | `M04` | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| 2010 | 6408 | 0 | 6408 | 0 | 6408 | 0 | 6408 | 0 | +| 2011 | 6422 | 0 | 6422 | 0 | 6422 | 0 | 6422 | 0 | +| 2012 | 0 | 224 | 0 | 329 | 0 | 86 | 0 | 598 | +| 2023 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 35 | + +Every aggregate/component pair is perfectly disjoint — the same argument the +pipeline used to let recipe joins skip the filter +(`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2). The **only** row that +must be excluded is `L--`, the IG-to-state family total (Σ 248,812,372), which +genuinely does roll up the `L-NN` codes. + +**What the rule is worth:** Σ IG by era — + +| era | rule (`NOT LIKE '%--'`) | naive `NOT is_aggregate` | visible under naive | +|---|---:|---:|---:| +| legacy (≤2011) | 9,259,744,623 | 2,766,798,384 | **29.9%** | +| modern (2012+) | 3,648,584,806 | 3,648,584,806 | 100% | + +--- + +## File Structure + +| File | Responsibility | Change | +|---|---|---| +| `inst/extdata/fixture_corpus/` | Offline test corpus | **Regenerate** (Task 1) — currently 194 category rows, 0 IG | +| `inst/sql/24-ig_long.sql` | IG rows, raw item_code | **Create** | +| `inst/sql/25-ig_long_harmonized.sql` | IG rows, COALESCE'd code | **Create** | +| `inst/sql/44-ig_annotated.sql` | IG + xwalk + categories | **Create** | +| `inst/sql/45-ig_annotated_harmonized.sql` | same, harmonized leg | **Create** | +| `R/views.R` | Registers `inst/sql/` files | **Verify** — likely globs, may need no change | +| `R/spending.R` | `cog_spending()`, `.verb_spendrev()`, `.build_verb_sql()` | **Modify** — new arg, UNION leg, drop `"K"` | +| `R/rollup.R` | `cog_geographic_rollup()` | **Modify** — arg + hard error | +| `R/peers.R` | `cog_peer_compare()` | **Modify** — arg + hard error | +| `R/provenance.R` | `.build_provenance()` | **Modify** — record the concept | +| `R/suggestions.R` | `.build_suggestions()` | **Modify** — name the IG counterpart recipe | +| `inst/schemas/provenance-v1.json` | Provenance contract | **Modify** — document the field | +| `tests/testthat/test-expenditure-concept.R` | The feature's tests | **Create** | +| `README.md`, `vignettes/total-spending.Rmd` | The two archetype questions | **Modify / Create** | + +--- + +### Task 1: Regenerate the fixture corpus + +Everything downstream tests against the fixture, and today's fixture has **194** +category rows and **zero** `M`/`L` codes — so no IG test can pass until this +lands. `data-raw/regenerate_fixture_corpus.R` is the maintained, scripted job for +exactly this. + +**Files:** +- Modify: `inst/extdata/fixture_corpus/**` (regenerated, not hand-edited) +- Test: existing suite (this task must not change any assertion's meaning) + +**Interfaces:** +- Consumes: the merged pipeline publish tree at + `/home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/cog_pipeline/_targets/publish_cache`. +- Produces: a fixture whose `data/summary_categories.parquet` has **260** rows + including **66** with `spend_subtype = "intergovernmental"` (34 `M`, 32 `L`), + and whose `docs/` copies carry the amended `Total = Direct + M + L` text. + +- [ ] **Step 1: Record the pre-regeneration state** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'p <- arrow::read_parquet("inst/extdata/fixture_corpus/data/summary_categories.parquet"); cat("BEFORE rows:", nrow(p), " IG:", sum(grepl("^[ML]", p$item_code)), "\n")' && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); r <- as.data.frame(testthat::test_local(reporter="silent")); cat(sprintf("BEFORE suite: PASS %d FAIL %d SKIP %d\n", sum(r$passed), sum(r$failed), sum(r$skipped)))' +``` + +Expected: `BEFORE rows: 194 IG: 0`, and `PASS 467 FAIL 0 SKIP 0`. + +- [ ] **Step 2: Regenerate** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript data-raw/regenerate_fixture_corpus.R \ + /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/cog_pipeline/_targets/publish_cache +``` + +- [ ] **Step 3: Verify the fixture picked up the IG rows** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e ' +p <- arrow::read_parquet("inst/extdata/fixture_corpus/data/summary_categories.parquet") +cat("AFTER rows:", nrow(p), "\n"); print(table(p$spend_subtype, useNA="ifany")) +cat("M:", sum(startsWith(p$item_code,"M")), " L:", sum(startsWith(p$item_code,"L")), "\n") +stopifnot(nrow(p)==260L, sum(p$spend_subtype %in% "intergovernmental")==66L) +cat("FIXTURE HAS IG ROWS\n")' +``` + +Expected: 260 rows; `capital 70 / intergovernmental 66 / operations 37 / NA 87`; M 34, L 32. + +- [ ] **Step 4: Run the full suite and adjudicate every change** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); r <- as.data.frame(testthat::test_local(reporter="silent")); cat(sprintf("AFTER suite: PASS %d FAIL %d SKIP %d\n", sum(r$passed), sum(r$failed), sum(r$skipped))); bad <- r[r$failed>0|r$error, c("file","test")]; if(nrow(bad)) print(bad) else cat("NO FAILURES\n")' +``` + +**Expected: still PASS 467, FAIL 0.** The IG rows are inert until Task 3 — no +existing view selects `M`/`L` prefixes, so nothing should move. + +**If any test fails, STOP and report rather than editing the assertion.** A +failure here means the fixture regeneration changed something beyond the +category table (e.g. a year partition or the manifest), and that needs +adjudicating against the pipeline change, exactly as `e813ffd` and PR #7 were +handled. Do not re-baseline a pinned value to make it pass. + +### Task 1 AMENDMENT (added 2026-07-27, after Step 4 came back red) + +Step 4 predicted "still PASS 467" on the reasoning that the IG rows are inert +until Task 3. **That was wrong, and the implementer correctly stopped rather than +re-baselining.** Measured: `PASS 466 | FAIL 1`, failing at +`tests/testthat/test-categories.R:18`. + +**Why:** `cog_categories()` (`R/categories.R:50-57`) is a **third consumer** of +`summary_categories` that the plan's file-structure table missed. It groups by +`(category, category_type, COALESCE(spend_subtype, revenue_subtype))` with **no +item-code prefix filter**, so the 66 new IG rows flow straight through to an +already-exported verb. The test's closed enumeration +`expect_true(all(r$subtype %in% c("operations", "capital")))` no longer holds — +`intergovernmental` is now a third spending subtype. + +**Adjudication (controller, 2026-07-27): the new behavior is CORRECT; update the +test.** `cog_categories()` is a *discovery* verb — its documented purpose is +"use this to discover valid `category` values". After Task 3, users will see +`spend_subtype = "intergovernmental"` in `cog_spending(expenditure_concept = +"total")` results. A discovery verb that hid a subtype the package can return +would misreport the data model. Note the `category` values themselves are +unchanged: IG rows reuse the existing functional categories, so only the +*subtype* enumeration grows. + +This is an adjudicated behavior change, not a pinned value bent to get green — +the distinction the brief asked you to protect. Record it in the commit message. + +**Required changes in `tests/testthat/test-categories.R`:** + +1. Widen the spending-subtype assertion at `:18`: + +```r + expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental"))) +``` + +2. Add positive coverage immediately after that test, so the new subtype is + asserted rather than merely tolerated: + +```r +test_that("cog_categories surfaces the intergovernmental spending subtype", { + skip_if_no_corpus() + r <- cog_categories(type = "spending") + expect_true("intergovernmental" %in% r$subtype) + # IG rows reuse the existing functional categories -- they add a subtype, + # not new category values. + ig_cats <- sort(unique(r$category[r$subtype == "intergovernmental"])) + direct_cats <- sort(unique(r$category[r$subtype != "intergovernmental"])) + expect_true(all(ig_cats %in% c(direct_cats, "Other Education"))) +}) +``` + +`"Other Education"` is allowed through because the pipeline ruled it a new +category carrying IG dollars and, for now, zero Direct dollars (pipeline #58). + +**Do not touch `R/categories.R`.** Its behavior is right as-is. + +After these edits the suite must be **FAIL 0**, with PASS at 467 minus the +widened assertion's arithmetic plus the new test's expectations — report the +real number. + +- [ ] **Step 5: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add inst/extdata/fixture_corpus tests/testthat/test-categories.R && \ +git -c commit.gpgsign=false commit -m "chore: regenerate fixture corpus with the intergovernmental category rows + +Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under +spend_subtype = intergovernmental (194 -> 260 rows). Inert until the +expenditure_concept views land; suite unchanged at PASS 467." +``` + +--- + +### Task 2: Drop the inert `"K"` prefix + +`flow_prefixes = c("E","F","G","K")` includes `K`, which matches **zero** rows +corpus-wide — audited pipeline-side (`docs/phase_r_harmonization_review.md` § 6: +*"Every raw FinEstDAT 2012–2023 contains zero K-prefix rows … Absence in the +corpus = absence in the product."*). Removing it changes no number. + +**Files:** +- Modify: `R/spending.R:58` (`flow_prefixes`), `inst/sql/20-spending_long.sql`, `inst/sql/22-spending_long_harmonized.sql` +- Test: `tests/testthat/test-expenditure-concept.R` (create) + +**Interfaces:** +- Produces: `cog_spending()`'s Direct leg selects prefixes `E`/`F`/`G` only. + Every existing result is numerically unchanged. + +- [ ] **Step 1: Write the failing test** + +Create `tests/testthat/test-expenditure-concept.R`: + +```r +test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", { + con <- .ensure_session() + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM long WHERE LEFT(item_code, 1) = 'K'")$n + expect_equal(n, 0) + + sql_files <- c("20-spending_long.sql", "22-spending_long_harmonized.sql") + for (f in sql_files) { + txt <- paste(readLines(system.file("sql", f, package = "uscogdata")), + collapse = " ") + expect_false(grepl("'K'", txt, fixed = TRUE), + label = paste(f, "must not reference the inert K prefix")) + } +}) +``` + +- [ ] **Step 2: Run it to verify it fails** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); testthat::test_file("tests/testthat/test-expenditure-concept.R")' +``` + +Expected: the `COUNT(*)` expectation PASSES (K really is absent) and the two +`expect_false` expectations FAIL (the SQL still says `'K'`). + +- [ ] **Step 3: Remove `K` from all three places** + +`R/spending.R:58` — `flow_prefixes = c("E", "F", "G"),` + +`inst/sql/20-spending_long.sql` — `WHERE LEFT(item_code, 1) IN ('E', 'F', 'G')` + +`inst/sql/22-spending_long_harmonized.sql` — `AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G')` + +- [ ] **Step 4: Verify green and numerically inert** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); r <- as.data.frame(testthat::test_local(reporter="silent")); cat(sprintf("PASS %d FAIL %d SKIP %d\n", sum(r$passed), sum(r$failed), sum(r$skipped)))' +``` + +Expected: FAIL 0, and PASS = 467 + the new expectations. Any *changed* number in +a pre-existing test would mean `K` was not inert — stop and report. + +- [ ] **Step 5: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add R/spending.R inst/sql/20-spending_long.sql inst/sql/22-spending_long_harmonized.sql tests/testthat/test-expenditure-concept.R && \ +git -c commit.gpgsign=false commit -m "fix: drop the inert K prefix from the spending flow prefixes + +K matches zero rows corpus-wide (audited pipeline-side). Numerically inert; +removed so the code stops implying a prefix the data never had." +``` + +--- + +### Task 3: `expenditure_concept` in `cog_spending()` + +**Files:** +- Create: `inst/sql/24-ig_long.sql`, `inst/sql/25-ig_long_harmonized.sql`, `inst/sql/44-ig_annotated.sql`, `inst/sql/45-ig_annotated_harmonized.sql` +- Modify: `R/spending.R` (`cog_spending()`, `.verb_spendrev()`, `.build_verb_sql()`, `.select_view()`) +- Verify: `R/views.R` registers the new files +- Test: `tests/testthat/test-expenditure-concept.R` + +**Interfaces:** +- Consumes: Task 1's fixture, Task 2's prefix list. +- Produces: `cog_spending(govid, years, category = NULL, per_capita = FALSE, + adjust_to_year = NULL, basis = c("harmonized","raw"), recipe = NULL, + expenditure_concept = c("direct","total"))`. With `"total"`, results gain rows + whose `spend_subtype` is `"intergovernmental"`; `"direct"` is byte-for-byte + today's behavior. Later tasks call it with `expenditure_concept` passed through. + +- [ ] **Step 1: Write the failing tests** + +Append to `tests/testthat/test-expenditure-concept.R`: + +```r +test_that("expenditure_concept defaults to direct and preserves today's numbers", { + gov <- "010000226085" # Alabama state government + base <- cog_spending(gov, years = 2019, category = "Police") + expl <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "direct") + expect_equal(base$amt_nominal, expl$amt_nominal) + expect_false("intergovernmental" %in% base$spend_subtype) +}) + +test_that("expenditure_concept = 'total' adds an intergovernmental subtype", { + gov <- "010000226085" + d <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "direct") + t <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "total") + expect_true("intergovernmental" %in% t$spend_subtype) + # Direct rows are untouched; Total only ever ADDS. + dt <- t[t$spend_subtype != "intergovernmental", ] + expect_equal(sort(dt$amt_nominal), sort(d$amt_nominal)) + expect_gt(sum(t$amt_nominal), sum(d$amt_nominal)) +}) + +test_that("legacy-era Total does not collapse to Direct (the is_aggregate trap)", { + # In the wide era the IG dollars live almost entirely on aggregate-flagged + # rows. A Total leg that inherited the Direct leg's NOT is_aggregate filter + # would silently return Total == Direct here. + gov <- "010000226085" + d <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "direct") + t <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total") + expect_true("intergovernmental" %in% t$spend_subtype) + ig <- sum(t$amt_nominal[t$spend_subtype == "intergovernmental"]) + expect_gt(ig, 0) + expect_gt(sum(t$amt_nominal), sum(d$amt_nominal)) +}) + +test_that("the IG leg never includes the L-- family total", { + con <- .ensure_session() + codes <- DBI::dbGetQuery(con, + "SELECT DISTINCT item_code FROM ig_long")$item_code + expect_false(any(grepl("--$", codes))) + expect_true(all(substr(codes, 1, 1) %in% c("M", "L"))) +}) + +test_that("expenditure_concept rejects unknown values", { + expect_error( + cog_spending("010000226085", years = 2019, expenditure_concept = "gross"), + class = "rlang_error" + ) +}) + +test_that("total composes with basis = 'raw' and basis = 'harmonized'", { + gov <- "010000226085" + h <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total", basis = "harmonized") + r <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total", basis = "raw") + ig_h <- sum(h$amt_nominal[h$spend_subtype == "intergovernmental"]) + ig_r <- sum(r$amt_nominal[r$spend_subtype == "intergovernmental"]) + # The only IG harmonization rule is M38 -> M36 (year-disjoint), so the IG + # total must agree between bases even though the code labels may differ. + expect_equal(ig_h, ig_r) +}) +``` + +- [ ] **Step 2: Run to verify they fail** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); testthat::test_file("tests/testthat/test-expenditure-concept.R")' +``` + +Expected: failures on `unused argument (expenditure_concept = ...)` and on the +missing `ig_long` view. + +- [ ] **Step 3: Create the four SQL views** + +`inst/sql/24-ig_long.sql`: + +```sql +-- Intergovernmental expenditure rows (M = to local govts, L = to state govts). +-- +-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the +-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published +-- ONLY as aggregate-flagged rows -- filtering them would hide ~70% of legacy IG +-- dollars and make Total silently collapse to Direct. This is safe because the +-- aggregate codes and their modern leaf components are strictly year-disjoint +-- (M47 ends 2011 / M94 starts 2012; M89 is aggregate only <= 2011 and a leaf +-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument +-- the pipeline's recipe joins use. +-- +-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up +-- the L-NN codes, so including it would double-count. +CREATE OR REPLACE VIEW ig_long AS +SELECT * +FROM long +WHERE LEFT(item_code, 1) IN ('M', 'L') + AND item_code NOT LIKE '%--'; +``` + +`inst/sql/25-ig_long_harmonized.sql`: + +```sql +-- Harmonized-basis IG rows. Uses COALESCE(harmonized_code, item_code) rather +-- than harmonized_code alone: aggregate rows carry NO harmonized_code by +-- construction (harmonized space is leaf-only), so a plain +-- `harmonized_code IS NOT NULL` filter would drop every legacy IG aggregate -- +-- 6.4e9 of M and 3.1e8 of L in corpus units. COALESCE keeps the one real IG +-- collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+) while +-- never dropping a row. +CREATE OR REPLACE VIEW ig_long_harmonized AS +SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code) +FROM long +WHERE LEFT(item_code, 1) IN ('M', 'L') + AND item_code NOT LIKE '%--'; +``` + +`inst/sql/44-ig_annotated.sql`: + +```sql +CREATE OR REPLACE VIEW ig_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.spend_subtype +FROM ig_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); +``` + +`inst/sql/45-ig_annotated_harmonized.sql`: identical, but `FROM ig_long_harmonized s` +and `CREATE OR REPLACE VIEW ig_annotated_harmonized AS`. + +- [ ] **Step 4: Register the harmonized views against the schema guard** + +`.register_views()` (`R/views.R:25-35`) globs `inst/sql/*.sql` and executes them +in **sorted filename order**, so the four new files register automatically — the +numbering above is chosen so `ig_long` (24/25) precedes `ig_annotated` (44/45), +which in turn follows `canonical_fips_xwalk` (30) and `summary_categories` (31). + +But there is a guard at `R/views.R:30`: any file listed in +`.harmonization_view_files` is **skipped** when `schema_version < 5`. The two new +harmonized views read `harmonized_code`, which does not exist on a pre-v5 corpus, +so they MUST be added to that vector (`R/views.R:13-21`) or `cog_open()` will +abort against an older corpus: + +```r +.harmonization_view_files <- c( + "22-spending_long_harmonized.sql", + "23-revenue_long_harmonized.sql", + "25-ig_long_harmonized.sql", + "33-harmonization_map.sql", + "34-harmonization_recipes.sql", + "35-series_breaks_pq.sql", + "42-spending_annotated_harmonized.sql", + "43-revenue_annotated_harmonized.sql", + "45-ig_annotated_harmonized.sql" +) +``` + +Note the asymmetry that follows: on a pre-v5 corpus `ig_annotated_harmonized` +does not exist, so `.select_ig_view()` must resolve to the raw `ig_annotated` +whenever the *resolved* basis is `"raw"` — which the existing `.resolve_basis()` +already guarantees for pre-v5 corpora. Pass `resolved$basis`, never the +user's raw `basis` argument, exactly as `.select_view()` does. + +Verify registration: + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); con <- .ensure_session(); print(DBI::dbGetQuery(con, "SELECT COUNT(*) n, COUNT(DISTINCT item_code) codes FROM ig_long"))' +``` + +Expected: a nonzero row count and 66 distinct codes (the fixture ships the full +metadata tables, but only the fixture's years of `long`, so `n` will be smaller +than the full corpus — the **codes** figure is the one that matters and may be +below 66 if a code is absent from the fixture's years; report what you see). + +- [ ] **Step 5: Wire the argument through** + +In `R/spending.R`, add the parameter to `cog_spending()` (roxygen `@param` too): + +```r +cog_spending <- function(govid, years, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, + basis = c("harmonized", "raw"), recipe = NULL, + expenditure_concept = c("direct", "total")) { + .verb_spendrev( + verb = "cog_spending", + view_base = "spending_annotated", + subtype_col = "spend_subtype", + flow_prefixes = c("E", "F", "G"), + call = match.call(), + govid = govid, + years = years, + category = category, + per_capita = per_capita, + adjust_to_year = adjust_to_year, + basis = basis, + recipe = recipe, + expenditure_concept = expenditure_concept + ) +} +``` + +In `.verb_spendrev()`, add `expenditure_concept = c("direct","total")` to the +signature, resolve it with `expenditure_concept <- match.arg(expenditure_concept)` +next to the existing `basis` handling, and pass an IG view into the SQL builder +only for the non-recipe path: + +```r + view <- .select_view(view_base, resolved$basis) + ig_view <- if (identical(expenditure_concept, "total")) { + .select_ig_view(resolved$basis) + } else { + NULL + } + sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view) +``` + +Add the helper beside `.select_view()`: + +```r +#' @noRd +.select_ig_view <- function(basis) { + if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated" +} +``` + +Change `.build_verb_sql()` to take `ig_view = NULL` and union it in. Replace the +`FROM %2$s` clause with a source expression built once: + +```r +.build_verb_sql <- function(view, subtype_col, govid, years, category, + ig_view = NULL) { + govid_lit <- .sql_lit_chr(govid) + years_lit <- paste(as.integer(years), collapse = ",") + category_pred <- if (is.null(category)) { + "" + } else { + sprintf("AND category IN (%s)", .sql_lit_chr(category)) + } + + # expenditure_concept = "total" adds the intergovernmental leg. UNION ALL, + # never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L), + # so de-duplication would be pure cost, and a silent row-drop if two + # governments ever reported identical values. + source_expr <- if (is.null(ig_view)) { + view + } else { + sprintf("(SELECT * FROM %s UNION ALL SELECT * FROM %s)", view, ig_view) + } + ... +``` + +and substitute `source_expr` where the format string previously took `view`. + +**Recipe interaction:** a `recipe` query bypasses the basis views entirely, so it +also bypasses the IG leg. If both `recipe` and `expenditure_concept = "total"` +are supplied, abort — the recipe already defines its own component set: + +```r + if (!is.null(recipe) && identical(expenditure_concept, "total")) { + cli::cli_abort(c( + "`recipe` and `expenditure_concept = \"total\"` are mutually exclusive.", + i = "A recipe defines its own component codes; pass one or the other.", + i = "For a recipe's intergovernmental counterpart, use the matching IG recipe (e.g. `corrections_ig_local_combined`)." + ), class = "uscogdata_recipe_concept_conflict") + } +``` + +Add a test for that abort alongside the others. + +- [ ] **Step 6: Run to green** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); r <- as.data.frame(testthat::test_local(reporter="silent")); cat(sprintf("PASS %d FAIL %d SKIP %d\n", sum(r$passed), sum(r$failed), sum(r$skipped)))' +``` + +Expected: FAIL 0. Every pre-existing test must still pass unchanged — `"direct"` +is the default and must not move a single number. + +- [ ] **Step 7: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add inst/sql R/spending.R R/views.R tests/testthat/test-expenditure-concept.R && \ +git -c commit.gpgsign=false commit -m "feat: expenditure_concept = direct|total in cog_spending() + +total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL +over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate -- +legacy IG lives almost entirely on aggregate rows, and the aggregate codes are +year-disjoint from their modern leaf components, so nothing double-counts. +L-- (the IG-to-state family total) is excluded. direct is the default and is +numerically unchanged." +``` + +--- + +### Task 4: Hard error in the cross-government verbs + +Owner ruling R1: `expenditure_concept = "total"` must **error** in +`cog_geographic_rollup()` and `cog_peer_compare()`, and the message must tell the +user to use `direct` for cross-government work **and explain why**. + +Note the honest justification, which differs per verb and should not be +overstated in the message: `cog_geographic_rollup()` does not itself sum — it +returns per-government rows tagged `state`/`county`/`city` and the user sums +them, which is exactly where hierarchical double-counting bites. +`cog_peer_compare()` emits quantile summary rows over same-type peers. The owner +ruled error for both (2026-07-27) to keep one rule across both repos. + +**Files:** +- Modify: `R/rollup.R` (`cog_geographic_rollup()`), `R/peers.R` (`cog_peer_compare()`) +- Test: `tests/testthat/test-expenditure-concept.R` + +**Interfaces:** +- Consumes: Task 3's `cog_spending()`. +- Produces: both verbs accept `expenditure_concept = c("direct","total")` and + abort with condition class `uscogdata_concept_not_aggregatable` when `"total"`. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("cog_geographic_rollup refuses expenditure_concept = 'total'", { + expect_error( + cog_geographic_rollup( + govids = list(state = "010000226085"), + category = "Police", years = 2019, + expenditure_concept = "total" + ), + class = "uscogdata_concept_not_aggregatable" + ) +}) + +test_that("cog_peer_compare refuses expenditure_concept = 'total'", { + expect_error( + cog_peer_compare( + target_govid = "010000226085", peers = "010000226085", + category = "Police", years = 2019, + expenditure_concept = "total" + ), + class = "uscogdata_concept_not_aggregatable" + ) +}) + +test_that("the refusal message names the fix and the reason", { + err <- tryCatch( + cog_geographic_rollup(govids = list(state = "010000226085"), + category = "Police", years = 2019, + expenditure_concept = "total"), + condition = function(e) e + ) + msg <- paste(conditionMessage(err), collapse = " ") + expect_match(msg, "direct") + expect_match(msg, "double-count|double count") +}) + +test_that("both cross-government verbs still accept the direct default", { + expect_no_error( + cog_geographic_rollup(govids = list(state = "010000226085"), + category = "Police", years = 2019) + ) +}) +``` + +- [ ] **Step 2: Run to verify they fail** + +Expected: `unused argument (expenditure_concept = "total")` from both verbs. + +- [ ] **Step 3: Implement the guard** + +Add a shared helper in `R/spending.R`: + +```r +#' @noRd +.abort_concept_not_aggregatable <- function(verb) { + cli::cli_abort(c( + "{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.", + "*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\ + comparison or sum that spans more than one government.", + "i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\ + money it hands to other governments. The receiving government reports \\ + that same dollar again as its own Direct when it actually spends it, \\ + so combining Total across governments counts intergovernmental \\ + transfers twice.", + "i" = "For one government's own Total, use \\ + {.code cog_spending(expenditure_concept = \"total\")}." + ), class = "uscogdata_concept_not_aggregatable") +} +``` + +In `cog_geographic_rollup()`, add `expenditure_concept = c("direct", "total")` to +the signature and, as the first statement after `call <- match.call()`: + +```r + expenditure_concept <- match.arg(expenditure_concept) + if (identical(expenditure_concept, "total")) { + .abort_concept_not_aggregatable("cog_geographic_rollup") + } +``` + +Do the same in `cog_peer_compare()` with its own verb name. Add roxygen +`@param expenditure_concept` to both, stating that only `"direct"` is accepted +and why. + +- [ ] **Step 4: Run to green** + +Expected: FAIL 0 across the suite. + +- [ ] **Step 5: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add R/rollup.R R/peers.R R/spending.R tests/testthat/test-expenditure-concept.R && \ +git -c commit.gpgsign=false commit -m "feat: refuse expenditure_concept = total in the cross-government verbs + +Owner ruling R1. Combining Census Total across governments counts +intergovernmental transfers twice, and these results land in Tableau where a +warning would be invisible -- so this is a hard error whose message names the +fix and the reason." +``` + +--- + +### Task 5: Provenance records the concept + +**Files:** +- Modify: `R/provenance.R` (`.build_provenance()`), `R/spending.R` (pass it through), `inst/schemas/provenance-v1.json` +- Test: `tests/testthat/test-expenditure-concept.R` + +**Interfaces:** +- Produces: every `cog_spending()` result's provenance carries + `expenditure_concept` (always populated, never implicit), and an + `expenditure_concept_note` naming the year-scoped IG assembly when `"total"`. + +- [ ] **Step 1: Write the failing test** + +```r +test_that("provenance always records the expenditure concept", { + d <- cog_spending("010000226085", years = 2019, category = "Police") + t <- cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total") + expect_equal(attr(d, "provenance")$expenditure_concept, "direct") + expect_equal(attr(t, "provenance")$expenditure_concept, "total") + # The note explains the non-obvious part: how legacy IG was assembled. + expect_true(nzchar(attr(t, "provenance")$expenditure_concept_note)) + expect_true(is.na(attr(d, "provenance")$expenditure_concept_note) || + !nzchar(attr(d, "provenance")$expenditure_concept_note)) +}) + +test_that("the provenance schema documents expenditure_concept", { + sch <- jsonlite::fromJSON( + system.file("schemas", "provenance-v1.json", package = "uscogdata"), + simplifyVector = FALSE + ) + expect_true("expenditure_concept" %in% names(sch$properties)) +}) +``` + +- [ ] **Step 2: Run to verify it fails** (fields are `NULL`; schema key absent). + +- [ ] **Step 3: Implement** + +Add `expenditure_concept = "direct"` and `expenditure_concept_note = NA_character_` +parameters to `.build_provenance()`, place them in the returned list next to +`basis`/`basis_note`, and pass them from `.verb_spendrev()`. The note when +`"total"`: + +``` +"Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded." +``` + +Add to `inst/schemas/provenance-v1.json`'s `properties`: + +```json + "expenditure_concept": { + "type": "string", + "enum": ["direct", "total"], + "description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments." + }, + "expenditure_concept_note": { + "type": ["string", "null"], + "description": "How the intergovernmental leg was assembled; null for 'direct'." + } +``` + +- [ ] **Step 4: Run to green.** Expected FAIL 0. + +- [ ] **Step 5: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add R/provenance.R R/spending.R inst/schemas/provenance-v1.json tests/testthat/test-expenditure-concept.R && \ +git -c commit.gpgsign=false commit -m "feat: record expenditure_concept in provenance and its JSON schema + +Always populated, never implicit, so a downstream artifact says which concept +produced it. cog-api passes provenance through verbatim." +``` + +--- + +### Task 6: Signposting — surface the IG recipe counterpart + +uscogdata #6 item 4: *"when a user filters aggregate codes, the coverage/recipe +machinery should mention the concept choice where relevant."* + +The concrete case: `.build_suggestions()` already fires when a category returns +no rows for some requested years and names the recipe that fills the gap (e.g. +`corrections_combined`). Several of those families have an **intergovernmental +counterpart recipe** — `corrections_ig_local_combined`, `ige_local_m47_wide`, +`ige_local_m89_wide`, `ige_state_l47_wide`, `ige_state_l89_wide`. A user who +took the Direct suggestion has no way to discover the IG one. + +Scope this narrowly: only extend an **already-firing** suggestion. Do NOT add a +new trigger that fires on healthy queries — a message on every call is noise, +and the vignette is where the concept choice is taught. + +**Files:** +- Modify: `R/suggestions.R` (`.build_suggestions()` / its message formatter) +- Test: `tests/testthat/test-expenditure-concept.R` + +**Interfaces:** +- Consumes: Tasks 3–5. +- Produces: when a suggestion fires for a recipe that has an IG counterpart in + `harmonization_recipes`, the suggestion entry gains an `ig_recipe_id` field and + the emitted message names it. + +- [ ] **Step 1: Write the failing test** + +```r +test_that("a firing suggestion names the intergovernmental counterpart recipe", { + # Corrections has no legacy leaf rows, so the coverage-gap suggestion fires; + # corrections_ig_local_combined is its IG counterpart. + r <- suppressMessages( + cog_spending("010000226085", years = c(2005, 2011), category = "Corrections") + ) + sugg <- attr(r, "provenance")$suggestions + expect_gt(length(sugg), 0L) + ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1)) + expect_true("corrections_combined" %in% ids) + ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id)) + expect_true("corrections_ig_local_combined" %in% ig) +}) + +test_that("no suggestion fires for a healthy query", { + r <- cog_spending("010000226085", years = 2019, category = "Police") + expect_length(attr(r, "provenance")$suggestions, 0L) +}) +``` + +- [ ] **Step 2: Run to verify the first fails** (no `ig_recipe_id` field) **and the +second passes** (pinning that we do not add a new trigger). + +- [ ] **Step 3: Implement** + +In `R/suggestions.R`, after the existing suggestion list is built, look up an IG +counterpart per suggested recipe. Derive it by matching on the recipe's function +suffix within `harmonization_recipes`, restricted to component codes whose first +letter is `M` or `L`. Attach as `ig_recipe_id` (`NULL` when there is none) and +append one clause to the emitted message, e.g.: + +``` +• corrections_combined (1967-2023): re-run with recipe = 'corrections_combined' + intergovernmental counterpart: recipe = 'corrections_ig_local_combined' +``` + +- [ ] **Step 4: Run to green.** Expected FAIL 0 across the suite. + +- [ ] **Step 5: Commit** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add R/suggestions.R tests/testthat/test-expenditure-concept.R && \ +git -c commit.gpgsign=false commit -m "feat: name the intergovernmental counterpart in firing recipe suggestions + +Closes uscogdata #6 item 4. Only extends suggestions that already fire -- a +concept hint on every healthy call would be noise." +``` + +--- + +### Task 7: README and the two-archetype vignette + +**Files:** +- Create: `vignettes/total-spending.Rmd` +- Modify: `README.md`, `_pkgdown.yml` (if it enumerates vignettes) +- Test: none (prose), but the vignette must knit + +**Interfaces:** +- Consumes: Tasks 3–5. +- Produces: docs anchored on the two archetype questions with worked code. + +- [ ] **Step 1: Write the vignette** + +Create `vignettes/total-spending.Rmd` with standard front matter +(`%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}`, +`%\VignetteEngine{knitr::rmarkdown}`). It must lead with the two archetype +questions and answer each with runnable code: + +1. **Single-government trend** — *"total spending in my county, 2017 vs today"*. + Show `cog_spending(..., expenditure_concept = "total")`, and note that either + concept is valid here as long as it is applied consistently across years. +2. **Cross-government rollup** — *"all the counties in my state, ten years ago + vs today, vs the neighbouring state"*. Show `cog_geographic_rollup()` with + the default `direct`, then show the error from passing `"total"` and explain + why it exists. + +Include the mechanism in plain language — a state gives a county $10M for roads; +it is in the state's `M44` and again in the county's `E44`/`F44`; summing Total +counts it twice, so the rollup would report $20M of road spending for $10M of +road work. + +State the composition rules explicitly: `expenditure_concept` (whose spending +counts) is orthogonal to `basis` (which vintage of code space); `"total"` is +mutually exclusive with `recipe`. + +- [ ] **Step 2: Knit it** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); rmarkdown::render("vignettes/total-spending.Rmd", quiet = TRUE)' && echo KNIT_OK +``` + +Expected: `KNIT_OK`. Delete the rendered `.html` afterwards if it is not +gitignored — do not commit build output. + +- [ ] **Step 3: Update the README** + +Add a short "Direct vs Total spending" section pointing at the vignette, with +the one-line rule: **any figure that spans more than one government uses +`direct`.** + +- [ ] **Step 4: Full suite + `R CMD check`-level sanity** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +/usr/bin/Rscript -e 'suppressMessages(pkgload::load_all(".", quiet=TRUE)); r <- as.data.frame(testthat::test_local(reporter="silent")); cat(sprintf("PASS %d FAIL %d SKIP %d\n", sum(r$passed), sum(r$failed), sum(r$skipped)))' && \ +/usr/bin/Rscript -e 'devtools::document()' && git diff --stat man/ NAMESPACE +``` + +Expected: FAIL 0; `man/` regenerated for the changed roxygen blocks. + +- [ ] **Step 5: Commit and open the PR** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata && \ +git add -A && \ +git -c commit.gpgsign=false commit -m "docs: total-spending vignette + README section on Direct vs Total" && \ +git push -u origin feat/expenditure-concept +``` + +Then open the PR against `main` with `tea`, closing uscogdata #6. **Stop for +owner review; do not merge.** + +--- + +## Deliberately out of scope + +- **`J` prefix** (`J19`/`J67`/`J68`/`J85`, direct assistance/subsidies) — needs + pipeline category rows first. Tracked as pipeline #58. +- **Public Welfare understatement / signposting** — uscogdata #9. +- **`cog_revenue()`** — `expenditure_concept` is an expenditure concept; revenue's + IG codes are a different axis. +- **Repo 3 (`cog-api` #3)** — separate plan after this merges: same parameter name + and default, HTTP 400 on aggregating endpoints, provenance passthrough, OpenAPI + docs led by the two archetypes, then publish + `docker restart cog-api`. diff --git a/R/explain.R b/R/explain.R index 2f2d44b..f7dd5b9 100644 --- a/R/explain.R +++ b/R/explain.R @@ -60,6 +60,21 @@ cog_explain <- function(result, format = c("print", "list")) { cli::cli_text("Basis: {prov$basis}{note}") } + if (!is.null(prov$expenditure_concept)) { + concept_note <- if (!is.null(prov$expenditure_concept_note) && + !is.na(prov$expenditure_concept_note)) { + sprintf(" (%s)", prov$expenditure_concept_note) + } else { + "" + } + cli::cli_text("Concept: {prov$expenditure_concept}{concept_note}") + if (isTRUE(prov$expenditure_concept_direct_suppressed)) { + cli::cli_alert_warning( + "Direct leg unavailable for at least one requested (year, category) -- affected rows report intergovernmental dollars alone, not Direct + IG. See each row's notes." + ) + } + } + cli::cli_h2("Codes observed") codes <- prov$codes_summed$observed if (length(codes) == 0L) { diff --git a/R/peers.R b/R/peers.R index 892828f..aa36b8f 100644 --- a/R/peers.R +++ b/R/peers.R @@ -143,6 +143,10 @@ cog_find_peers <- function(target_govid, #' @param per_capita Default `TRUE` — peer compare usually normalizes by #' population. #' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`. +#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only +#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for +#' single-government queries but cannot be used here because combining Total +#' across peer sets counts intergovernmental transfers twice. #' @return Tibble matching [cog_spending()]'s columns, plus a `role` #' column taking values `"target"`, `"peer"`, `"summary_p25"`, #' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank @@ -153,8 +157,13 @@ cog_find_peers <- function(target_govid, #' `cohort_year`, and `cohort_govids`. #' @export cog_peer_compare <- function(target_govid, peers, category, years, - per_capita = TRUE, adjust_to_year = NULL) { + per_capita = TRUE, adjust_to_year = NULL, + expenditure_concept = c("direct", "total")) { call <- match.call() + expenditure_concept <- match.arg(expenditure_concept) + if (identical(expenditure_concept, "total")) { + .abort_concept_not_aggregatable("cog_peer_compare") + } if (!is.character(target_govid) || length(target_govid) != 1L) { cli::cli_abort("`target_govid` must be a length-1 character string.") } diff --git a/R/provenance.R b/R/provenance.R index bfd7a03..e8c6d65 100644 --- a/R/provenance.R +++ b/R/provenance.R @@ -6,6 +6,9 @@ per_capita, adjust_to_year, result, sql, subtype_col, basis = NA_character_, basis_note = NA_character_, + expenditure_concept = "direct", + expenditure_concept_note = NA_character_, + expenditure_concept_direct_suppressed = FALSE, harmonization = NULL, recipe = NULL, suggestions = list()) { manifest <- .uscogdata_env$manifest @@ -51,6 +54,9 @@ category = category, basis = basis, basis_note = basis_note, + expenditure_concept = expenditure_concept, + expenditure_concept_note = expenditure_concept_note, + expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed), harmonization = harmonization %||% list( applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0, note = NA_character_ diff --git a/R/rollup.R b/R/rollup.R index cc2adcc..9c90d53 100644 --- a/R/rollup.R +++ b/R/rollup.R @@ -25,6 +25,12 @@ #' population from `gov_population_yearly`. Govs with missing population #' are excluded from the result. #' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`. +#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only +#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for +#' single-government queries but cannot be used here because combining Total +#' across multiple layers of government double-counts intergovernmental +#' transfers (a state's payment to a school district is the same dollar the +#' district reports as its own Direct spending). #' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` / #' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`, @@ -33,8 +39,13 @@ #' and `rollup$included_govids` / `rollup$excluded_govids`. #' @export cog_geographic_rollup <- function(govids, category, years, - per_capita = FALSE, adjust_to_year = NULL) { + per_capita = FALSE, adjust_to_year = NULL, + expenditure_concept = c("direct", "total")) { call <- match.call() + expenditure_concept <- match.arg(expenditure_concept) + if (identical(expenditure_concept, "total")) { + .abort_concept_not_aggregatable("cog_geographic_rollup") + } .validate_rollup_layers(govids) govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]") diff --git a/R/spending.R b/R/spending.R index bb4f641..80a9064 100644 --- a/R/spending.R +++ b/R/spending.R @@ -42,6 +42,34 @@ #' `basis = "recipe"` with an inert `harmonization` block (`applied = #' FALSE`, pointing at the `recipe` block instead) rather than a #' possibly-misleading `"harmonized"`/`"raw"` value. +#' @param expenditure_concept `"direct"` (default) returns only the +#' government's own direct spending (item codes `E`/`F`/`G`), unchanged +#' from prior releases. `"total"` additionally UNIONs in the +#' intergovernmental leg -- payments to local governments (`M` codes) and +#' to the state government (`L` codes, excluding the `L--` family-total +#' rollup) -- so results gain rows with `spend_subtype == +#' "intergovernmental"`. Requires the active corpus's `summary_categories` +#' to carry M/L rows (added by cog_pipeline PR #59); aborts with class +#' `uscogdata_ig_categories_unsupported` on an older corpus rather than +#' silently under-reporting. Mutually exclusive with `recipe` (a recipe +#' already defines its own component codes). **Do not sum `"total"` +#' results across levels of government** (e.g. state + county + city): +#' a state's `M12` payment to a school district is the same dollar the +#' district reports as its own direct `E12`, so summing both double-counts +#' it. This matters in particular with [cog_geographic_rollup()], which +#' sums across exactly that kind of multi-layer government set. +#' +#' In the legacy wide era (<= FY2011), some functions are published ONLY +#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05` +#' split), which the Direct leg excludes by construction but the IG leg +#' deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"` +#' query, any (year, category) where this leaves intergovernmental rows +#' with NO Direct counterpart is flagged: the affected rows' `notes` +#' name the harmonization recipe that recovers the missing Direct +#' component (when one exists), and +#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the +#' figure in those rows is the intergovernmental leg alone, not Direct + +#' IG. #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`, #' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`, @@ -50,12 +78,13 @@ #' @export cog_spending <- function(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, - basis = c("harmonized", "raw"), recipe = NULL) { + basis = c("harmonized", "raw"), recipe = NULL, + expenditure_concept = c("direct", "total")) { .verb_spendrev( verb = "cog_spending", view_base = "spending_annotated", subtype_col = "spend_subtype", - flow_prefixes = c("E", "F", "G", "K"), + flow_prefixes = c("E", "F", "G"), call = match.call(), govid = govid, years = years, @@ -63,22 +92,79 @@ cog_spending <- function(govid, years, category = NULL, per_capita = per_capita, adjust_to_year = adjust_to_year, basis = basis, - recipe = recipe + recipe = recipe, + expenditure_concept = expenditure_concept ) } +#' @noRd +.abort_concept_not_aggregatable <- function(verb) { + cli::cli_abort(c( + "{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.", + "*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\ + comparison or sum that spans more than one government.", + "i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\ + money it hands to other governments. The receiving government reports \\ + that same dollar again as its own Direct when it actually spends it, \\ + so combining Total across governments double-counts intergovernmental \\ + transfers.", + "i" = "For one government's own Total, use \\ + {.code cog_spending(expenditure_concept = \"total\")}." + ), class = "uscogdata_concept_not_aggregatable") +} + #' @noRd .verb_spendrev <- function(verb, view_base, subtype_col, flow_prefixes, call, govid, years, category, per_capita, adjust_to_year, - basis = c("harmonized", "raw"), recipe = NULL) { + basis = c("harmonized", "raw"), recipe = NULL, + expenditure_concept = c("direct", "total")) { basis_explicit <- length(basis) == 1L basis <- match.arg(basis, c("harmonized", "raw")) + # match.arg() itself throws a base `simpleError`, not an rlang-classed + # condition; wrap it so an invalid expenditure_concept aborts consistently + # with the rest of this package's validation (cli::cli_abort -> rlang_error). + expenditure_concept <- tryCatch( + match.arg(expenditure_concept, c("direct", "total")), + error = function(e) { + cli::cli_abort( + "`expenditure_concept` must be one of {.val direct} or {.val total}.", + class = "uscogdata_invalid_expenditure_concept", + parent = e + ) + } + ) govid <- .coerce_govid_input(govid, arg = "govid") .validate_verb_inputs(govid, years, category, per_capita, adjust_to_year, recipe) + if (!is.null(recipe) && identical(expenditure_concept, "total")) { + cli::cli_abort(c( + "`recipe` and `expenditure_concept = \"total\"` are mutually exclusive.", + i = "A recipe defines its own component codes; pass one or the other.", + i = "For a recipe's intergovernmental counterpart, use the matching IG recipe (e.g. `corrections_ig_local_combined`)." + ), class = "uscogdata_recipe_concept_conflict") + } + + # .verb_spendrev() is shared with cog_revenue(), which never exposes + # expenditure_concept and always resolves it to "direct" -- so nothing on + # the public API can reach this today. But it's a cheap guard against a + # future call (direct or via a modified cog_revenue()) that would UNION + # the IG leg's expenditure M/L rows into a revenue result, which has no + # matching IG view and no sensible meaning. + if (identical(expenditure_concept, "total") && + !identical(view_base, "spending_annotated")) { + cli::cli_abort( + paste0( + "`expenditure_concept = \"total\"` is only supported for spending ", + "(view_base = \"spending_annotated\"); got view_base = ", + "{.val {view_base}}." + ), + class = "uscogdata_expenditure_concept_unsupported" + ) + } + years <- as.integer(years) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) @@ -105,7 +191,13 @@ cog_spending <- function(govid, years, category = NULL, category_for_prov <- recipe_label } else { view <- .select_view(view_base, resolved$basis) - sql <- .build_verb_sql(view, subtype_col, govid, years, category) + ig_view <- if (identical(expenditure_concept, "total")) { + .require_ig_categories(con) + .select_ig_view(resolved$basis) + } else { + NULL + } + sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) } @@ -114,8 +206,6 @@ cog_spending <- function(govid, years, category = NULL, result <- .attach_real_dollars(result, adjust_to_year, per_capita) } - result$notes <- .notes_column(result) - # A recipe result doesn't go through spending_annotated(_harmonized) / # revenue_annotated(_harmonized) at all -- .run_recipe()'s generic join # reads `long` directly -- so `basis` and the `harmonization` exclusion @@ -141,7 +231,64 @@ cog_spending <- function(govid, years, category = NULL, harmonization <- .build_harmonization_block( con, govid, years, resolved, flow_prefixes ) - suggestions <- .build_suggestions(con, govid, years, category, result, resolved$basis) + # C1(a): gap detection must run against the Direct leg alone. `result` + # can also carry UNION'd intergovernmental rows (expenditure_concept = + # "total"), and the wide era (<= FY2011) routinely has legacy IG dollars + # surviving (ig_long deliberately keeps aggregate rows) for a + # (year, category) whose legacy Direct dollars were suppressed (spending_ + # long/spending_long_harmonized both filter NOT is_aggregate). Passing + # the UNION'd result here would let a surviving IG row count as coverage + # and silently cancel the recipe-hint suggestion that should fire. + direct_leg_result <- if (identical(expenditure_concept, "total")) { + result[!(result[[subtype_col]] %in% "intergovernmental"), , drop = FALSE] + } else { + result + } + suggestions <- .build_suggestions(con, govid, years, category, + direct_leg_result, + resolved$basis, flow_prefixes) + } + + # C1(b): when expenditure_concept = "total", flag any row where the IG + # leg has dollars but the Direct leg has none for that same (year, + # canonical_govid, category) AND a harmonization recipe actually recovers + # the missing Direct dollars for that exact triple -- see + # .detect_direct_suppressed() for why bare Direct-row absence alone is NOT + # sufficient (the dominant real cause is a government that simply has no + # direct spending in that category, which is correct, ordinary data). When + # a covering recipe is found, both the row-level notes and the provenance + # say so rather than pass silently as a plausible Total. + direct_suppressed_info <- if (identical(expenditure_concept, "total")) { + .detect_direct_suppressed(con, result, subtype_col) + } else { + list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result))) + } + direct_suppressed <- direct_suppressed_info$flag + direct_suppressed_flag <- isTRUE(any(direct_suppressed)) + + result$notes <- .notes_column(result, direct_suppressed_info$notes) + + # Determine expenditure_concept_note: only non-empty for "total", explains + # how the IG leg was assembled from legacy-era aggregates. When the Direct + # leg is suppressed for at least one requested (year, category), append an + # explicit warning rather than let the base note's "Total = Direct + IG" + # framing stand unqualified for rows where that arithmetic didn't happen. + expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) { + base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded." + if (direct_suppressed_flag) { + paste0( + base_note, + " NOTE: for at least one requested (year, category) the Direct leg ", + "has NO rows in this corpus (a legacy aggregate-only family) -- the ", + "affected result rows report the intergovernmental leg alone, not ", + "Direct + IG. See `expenditure_concept_direct_suppressed` and each ", + "affected row's `notes`." + ) + } else { + base_note + } + } else { + NA_character_ } prov <- .build_provenance( @@ -157,6 +304,9 @@ cog_spending <- function(govid, years, category = NULL, subtype_col = subtype_col, basis = basis_for_prov, basis_note = basis_note_for_prov, + expenditure_concept = expenditure_concept, + expenditure_concept_note = expenditure_concept_note_for_prov, + expenditure_concept_direct_suppressed = direct_suppressed_flag, harmonization = harmonization, recipe = recipe_block, suggestions = suggestions @@ -211,6 +361,43 @@ cog_spending <- function(govid, years, category = NULL, if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base } +#' @noRd +.select_ig_view <- function(basis) { + if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated" +} + +#' Abort unless the active corpus's `summary_categories` actually carries +#' intergovernmental (M/L) rows. +#' +#' The 66 M/L category rows arrived via cog_pipeline PR #59 with NO +#' `schema_version` bump (`DESCRIPTION` still declares `MinCorpusSchema: 4`), +#' so `schema_version` alone cannot gate `expenditure_concept = "total"` -- +#' a pre-#59 corpus can validly report schema_version 4, 5, or 6 and still +#' have zero M/L rows in `summary_categories`. Against such a corpus, +#' `ig_annotated`'s LEFT JOIN to `summary_categories` silently produces NA +#' `category`/`spend_subtype` for every IG row: with a `category` filter +#' this returns 0 rows (reads as "no intergovernmental spending" rather than +#' "can't tell"), and with `category = NULL` every IG dollar collapses into +#' one NA-subtype group that is invisible to the `spend_subtype == +#' "intergovernmental"` filter this package's own tests, roxygen, and +#' vignette all rely on. Checking the data directly (rather than +#' schema_version) is the only reliable gate. +#' @noRd +.require_ig_categories <- function(con, what = "expenditure_concept = \"total\"") { + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')" + )$n + if (identical(as.integer(n), 0L)) { + cli::cli_abort(c( + sprintf("%s requires a corpus with intergovernmental category rows.", what), + x = "The active corpus's `summary_categories` has no M/L (intergovernmental) rows.", + i = "This corpus predates the intergovernmental category rows added by cog_pipeline PR #59.", + i = "Point USCOGDATA_URL at a newer corpus that includes the M/L summary_categories rows." + ), class = "uscogdata_ig_categories_unsupported") + } + invisible(TRUE) +} + #' @noRd .sql_lit_chr <- function(x) { safe <- gsub("'", "''", x, fixed = TRUE) @@ -218,7 +405,8 @@ cog_spending <- function(govid, years, category = NULL, } #' @noRd -.build_verb_sql <- function(view, subtype_col, govid, years, category) { +.build_verb_sql <- function(view, subtype_col, govid, years, category, + ig_view = NULL) { govid_lit <- .sql_lit_chr(govid) years_lit <- paste(as.integer(years), collapse = ",") category_pred <- if (is.null(category)) { @@ -227,6 +415,26 @@ cog_spending <- function(govid, years, category = NULL, sprintf("AND category IN (%s)", .sql_lit_chr(category)) } + # expenditure_concept = "total" adds the intergovernmental leg. UNION ALL, + # never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L), + # so de-duplication would be pure cost, and a silent row-drop if two + # governments ever reported identical values. + source_expr <- if (is.null(ig_view)) { + view + } else { + sprintf("(SELECT * FROM %s UNION ALL SELECT * FROM %s)", view, ig_view) + } + + # bool_or(), not bool_and(): a no-op for the Direct/revenue legs (those + # views filter NOT is_aggregate, so no row in any group is ever aggregate), + # but load-bearing for the IG leg, which deliberately keeps aggregate rows + # (see inst/sql/24-ig_long.sql). The wide era is dense -- every government + # has a row for every code in a family, most of them $0 -- so a $0 leaf + # commonly lands in the same (year, gov, subtype, category) group as the + # real aggregate row. bool_and() would then read FALSE for that group even + # though its dollars came entirely from an aggregate row, silently + # suppressing the "Aggregate fallback applied" note on exactly the rows + # this feature exists to surface. sprintf( "SELECT year, @@ -236,14 +444,14 @@ cog_spending <- function(govid, years, category = NULL, category, SUM(amt) * 1000.0 AS amt_nominal, string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included, - bool_and(is_aggregate) AS aggregate_fallback + bool_or(is_aggregate) AS aggregate_fallback FROM %2$s WHERE canonical_govid IN (%3$s) AND year IN (%4$s) %5$s GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category ORDER BY year, canonical_govid, %1$s, category", - subtype_col, view, govid_lit, years_lit, category_pred + subtype_col, source_expr, govid_lit, years_lit, category_pred ) } @@ -296,11 +504,131 @@ cog_spending <- function(govid, years, category = NULL, result } +#' Detect rows where expenditure_concept = "total" is reporting the +#' intergovernmental leg with NO Direct counterpart in the same (year, +#' canonical_govid, category) group AND a harmonization recipe actually +#' recovers the missing Direct dollars for that exact (year, canonical_govid, +#' category) triple. +#' +#' Bare Direct-row absence is deliberately NOT sufficient on its own: the +#' dominant real cause of "no Direct sibling row" is a government that simply +#' has no direct spending in that category (e.g. a state that funds K-12 +#' entirely through school districts), which is correct, ordinary data, not +#' suppression. Genuine suppression -- a legacy aggregate-only family whose +#' Direct-leg basis query excludes it by construction (spending_long/ +#' spending_long_harmonized both filter NOT is_aggregate) -- always has a +#' covering harmonization recipe, because that is exactly what the recipe +#' catalog exists to recover (see R/suggestions.R and `cog_recipes()`). So +#' checking "does a recipe actually cover this triple" cleanly separates the +#' two cases instead of conflating them. +#' +#' Returns `list(flag, notes)`, both the same length as `result`: `flag` is +#' `TRUE` only for the `spend_subtype == "intergovernmental"` row(s) in a +#' suppressed group, and `notes` names the recovering recipe(s) for those +#' rows (`NA` everywhere else). #' @noRd -.notes_column <- function(result) { +.detect_direct_suppressed <- function(con, result, subtype_col) { + n <- nrow(result) + empty_notes <- rep(NA_character_, n) + if (n == 0L) return(list(flag = logical(0), notes = character(0))) + is_ig <- result[[subtype_col]] %in% "intergovernmental" + if (!any(is_ig)) return(list(flag = rep(FALSE, n), notes = empty_notes)) + + key <- paste(result$year, result$canonical_govid, result$category, sep = "\r") + has_direct <- key %in% unique(key[!is_ig]) + candidate <- is_ig & !has_direct + + flag <- rep(FALSE, n) + notes <- empty_notes + if (!any(candidate)) return(list(flag = flag, notes = notes)) + + idx <- which(candidate) + rows <- unique(result[idx, c("year", "canonical_govid", "category")]) + covering <- .covering_recipes(con, rows) + cov_key <- paste(covering$year, covering$canonical_govid, covering$category, + sep = "\r") + + for (i in idx) { + k <- paste(result$year[i], result$canonical_govid[i], result$category[i], + sep = "\r") + m <- match(k, cov_key) + if (is.na(m)) next + ids <- covering$recipe_ids[[m]] + if (length(ids) == 0L) next + flag[i] <- TRUE + notes[i] <- sprintf( + "Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).", + paste(sort(unique(ids)), collapse = "', '") + ) + } + list(flag = flag, notes = notes) +} + +#' For each (year, canonical_govid, category) triple potentially affected by +#' a suppressed Direct leg, find the harmonization recipe(s) that (a) cover +#' this `category` (share a component item_code via `summary_categories`, +#' excluding any recipe that is itself entirely intergovernmental M/L -- the +#' same exclusion `.build_suggestions()` applies, see I2) and (b) actually +#' produce a `long` row for this exact (canonical_govid, year) via the same +#' generic join `.run_recipe()` uses (component year_min/year_max + +#' gov_type_scope, no is_aggregate filter -- a recipe's whole point is to +#' recover data that's aggregate-only). Adds a list-column `recipe_ids` +#' (possibly length-0) to `rows`. +#' @noRd +.covering_recipes <- function(con, rows) { + rows$recipe_ids <- vector("list", nrow(rows)) + cats <- unique(rows$category[!is.na(rows$category)]) + if (length(cats) == 0L) return(rows) + + cand <- DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT sc.category, r.recipe_id + FROM harmonization_recipes r + JOIN summary_categories sc ON sc.item_code = r.component_code + WHERE sc.category IN (%s) + AND r.recipe_id NOT IN ( + SELECT DISTINCT recipe_id FROM harmonization_recipes + WHERE LEFT(component_code, 1) IN ('M', 'L') + )", + .sql_lit_chr(cats) + )) + if (nrow(cand) == 0L) return(rows) + + recipe_ids_all <- unique(cand$recipe_id) + govids <- unique(rows$canonical_govid) + years <- unique(rows$year) + covered <- DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT r.recipe_id, l.canonical_govid, l.year + FROM long l + JOIN harmonization_recipes r + ON l.item_code = r.component_code + AND l.year BETWEEN r.year_min AND r.year_max + AND (r.gov_type_scope = 'all' + OR (r.gov_type_scope = 'state' AND l.type = 0) + OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) + WHERE r.recipe_id IN (%s) + AND l.canonical_govid IN (%s) + AND l.year IN (%s)", + .sql_lit_chr(recipe_ids_all), .sql_lit_chr(govids), paste(years, collapse = ",") + )) + + for (i in seq_len(nrow(rows))) { + cat_i <- rows$category[i] + if (is.na(cat_i)) next + cat_recipe_ids <- cand$recipe_id[cand$category == cat_i] + if (length(cat_recipe_ids) == 0L) next + sub <- covered[covered$canonical_govid == rows$canonical_govid[i] & + covered$year == rows$year[i] & + covered$recipe_id %in% cat_recipe_ids, ] + rows$recipe_ids[[i]] <- sort(unique(sub$recipe_id)) + } + rows +} + +#' @noRd +.notes_column <- function(result, direct_suppressed_notes = NULL) { n <- nrow(result) if (n == 0L) return(character(0)) - parts <- vector("list", 2L) + parts <- vector("list", 3L) agg <- result[["aggregate_fallback"]] parts[[1]] <- if (!is.null(agg)) { ifelse(agg %in% TRUE, @@ -317,6 +645,11 @@ cog_spending <- function(govid, years, category = NULL, } else { rep(NA_character_, n) } + parts[[3]] <- if (!is.null(direct_suppressed_notes)) { + direct_suppressed_notes + } else { + rep(NA_character_, n) + } out <- character(n) for (i in seq_len(n)) { pieces <- vapply(parts, `[[`, character(1), i) diff --git a/R/suggestions.R b/R/suggestions.R index fc59d21..6e64c72 100644 --- a/R/suggestions.R +++ b/R/suggestions.R @@ -22,6 +22,13 @@ # positives from ordinary reporting variance -- most governments don't use # every sibling code in a multi-code category every year, and that is not # a format-boundary gap worth signposting. +# +# C1(a): for expenditure_concept = "total" callers, `result` here must +# already be the Direct-leg subset (the caller filters out +# spend_subtype == "intergovernmental" rows before calling in). A gap year +# is "the requested year has no Direct rows", never "no rows at all" -- +# an IG row surviving on a legacy aggregate that Direct excludes must not +# read as coverage and cancel the very suggestion that would recover it. #' Build the `prov$suggestions` list for a (non-recipe) basis = "harmonized" #' verb call: recipes whose generic join would fill a real gap in `result`. @@ -33,18 +40,41 @@ #' @param category `category` argument as passed to the verb (character #' vector or `NULL`; suggestions are only computed when non-NULL). #' @param result The verb's already-computed result tibble (post basis -#' query, pre per_capita/adjust_to_year). +#' query, pre per_capita/adjust_to_year), pre-filtered to the Direct leg +#' only when the caller's `expenditure_concept = "total"` (see C1(a)). #' @param basis The *resolved* basis (`"harmonized"` or `"raw"`). -#' @return List of `list(recipe_id, label, available_years, hint)`, possibly -#' empty. +#' @param flow_prefixes The calling verb's own flow-type prefixes (e.g. +#' `c("E", "F", "G")` for `cog_spending()`, `c("T", "A", "U", "B", "C", +#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to +#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart +#' lookup scoped to the calling verb's own flow family. +#' @return List of `list(recipe_id, label, available_years, hint, +#' ig_recipe_id)`, possibly empty. #' @noRd -.build_suggestions <- function(con, govid, years, category, result, basis) { +.build_suggestions <- function(con, govid, years, category, result, basis, + flow_prefixes) { if (!identical(basis, "harmonized") || is.null(category)) return(list()) + # Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe -- + # i.e. every one of its own component codes is M/L-prefixed. Without this, + # a category whose summary_categories rows span both a Direct family + # (e.g. E04/E05, "Corrections") and its M/L counterpart (M04/M05, same + # category since Task 1) makes the M/L recipe itself (e.g. + # `corrections_ig_local_combined`) a raw top-level candidate for a plain + # (Direct) cog_spending() call -- following that hint would silently + # return intergovernmental dollars under `expenditure_concept = "direct"` + # provenance. This is a stronger, unconditional exclusion than the + # flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe + # should never be suggested as a coverage-gap filler for EITHER verb, not + # just kept from being named as the *counterpart* of another suggestion. candidates <- DBI::dbGetQuery(con, sprintf( "SELECT DISTINCT recipe_id FROM harmonization_recipes WHERE component_code IN ( SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s) + ) + AND recipe_id NOT IN ( + SELECT DISTINCT recipe_id FROM harmonization_recipes + WHERE LEFT(component_code, 1) IN ('M', 'L') )", .sql_lit_chr(category) ))$recipe_id @@ -98,19 +128,121 @@ hint = sprintf("re-run with recipe = '%s'", rid) ) } - suggestions + .attach_ig_counterparts(con, suggestions, flow_prefixes) +} + +#' Attach `ig_recipe_id` to each suggestion: the intergovernmental-expenditure +#' recipe (an M-to-local or L-to-state recipe) whose component codes cover +#' exactly the same set of function suffixes as the firing recipe's own +#' components, e.g. `corrections_combined`'s {E04, E05} -> suffixes {"04", +#' "05"} matches `corrections_ig_local_combined`'s {M04, M05} -> the same +#' {"04", "05"}. `NULL` when no such recipe exists, which also covers the +#' case where the firing recipe already IS the IG recipe (self-matches are +#' excluded, so an IG recipe never names itself as its own counterpart). +#' +#' Matching is deliberately an exact set match, not "any suffix in common": +#' the two-digit suffix only means the same "function" across recipes that +#' share the underlying Census functional-classification scheme (E/F/G/L/M +#' all use "04"/"05" for corrections). M/L "combined other" codes (47/89/ +#' 91-94) reuse digits for an unrelated catch-all construct, so e.g. +#' `general_gov_e89_wide`'s {E85, E89} -> {"85", "89"} must NOT match +#' `ige_local_m89_wide`'s {"89", "91", "92", "93"} on the shared "89" alone. +#' Checked by hand against the full harmonization_recipes catalog: only the +#' corrections family (E/F/G/M, suffixes 04/05) has an exact-set match in +#' this corpus. +#' +#' Exact-set suffix matching is NOT enough on its own, though: the same +#' reused-digit problem exists ACROSS the revenue-side IG families too. +#' `ig_local_d47_wide` (D47/D94, suffixes {"47","94"}) is an exact-set match +#' for `ige_local_m47_wide` (M47/M94, same suffixes) even though one is +#' intergovernmental REVENUE received from local governments and the other is +#' intergovernmental EXPENDITURE paid to local governments -- unrelated flows +#' that happen to reuse "47"/"94" for their own "transit/utilities" and +#' "other/combined" catch-alls. `ig_federal_b47_wide`, `ig_state_c47_wide`, +#' and their `*_89` siblings all collide the same way. None of this is +#' reachable via `cog_revenue()` in the bundled fixture today (its B/C/D +#' recipes never happen to have a covered gap year for any fixture govid), +#' but it IS reachable via a mis-scoped `cog_spending()` call on a +#' revenue-only category, e.g. `cog_spending(gov, category = "IG Federal")` +#' fires `ig_federal_b47_wide`/`ig_federal_b89_wide` for real in the fixture +#' -- so this is a live, not merely theoretical, gap. +#' +#' Two flow-family checks close this, both required (see +#' `tests/testthat/test-expenditure-concept.R`, "revenue-flavored ... never +#' receives an M/L counterpart" tests, for the pairwise verification): +#' 1. `own_prefix %in% flow_prefixes`: the firing recipe's own component +#' codes must belong to the calling verb's own flow family (the same +#' `flow_prefixes` `.build_harmonization_block()` uses, see +#' `R/basis.R`). This blocks a recipe surfaced through a mis-scoped +#' category from ever reaching the M/L search, e.g. `cog_spending()`'s +#' flow_prefixes are `c("E","F","G")`, which `ig_federal_b47_wide`'s own +#' `"B"` is not part of. +#' 2. `own_prefix %in% c("E","F","G")`: M/L only ever pairs with the +#' DIRECT-expenditure family, never with revenue (`cog_revenue()`'s +#' flow_prefixes already fold B/C/D in as ordinary revenue -- there is +#' no separate "Total" bolt-on for revenue the way `expenditure_concept` +#' adds one for spending) and never with ANOTHER M/L recipe (without +#' this check, `ige_local_m47_wide` would wrongly match sibling +#' `ige_state_l47_wide` on their shared {"47","94"} suffix set). +#' Condition 1 alone does not catch this: under `cog_revenue()`, +#' `ig_federal_b47_wide`'s own `"B"` IS inside revenue's own +#' `flow_prefixes`, so only this second, family-specific check blocks +#' the search. +#' @noRd +.attach_ig_counterparts <- function(con, suggestions, flow_prefixes) { + if (length(suggestions) == 0L) return(suggestions) + + comp <- DBI::dbGetQuery(con, + "SELECT recipe_id, component_code FROM harmonization_recipes") + comp$prefix <- substr(comp$component_code, 1L, 1L) + comp$suffix <- substr(comp$component_code, 2L, nchar(comp$component_code)) + suffix_sets <- lapply(split(comp$suffix, comp$recipe_id), function(x) sort(unique(x))) + prefix_sets <- lapply(split(comp$prefix, comp$recipe_id), function(x) sort(unique(x))) + + ig_recipe_ids <- unique(comp$recipe_id[comp$prefix %in% c("M", "L")]) + + find_counterpart <- function(rid) { + own_prefix <- prefix_sets[[rid]] + own_suffix <- suffix_sets[[rid]] + if (is.null(own_prefix) || is.null(own_suffix)) return(NULL) + if (!all(own_prefix %in% flow_prefixes)) return(NULL) + if (!all(own_prefix %in% c("E", "F", "G"))) return(NULL) + for (cand in ig_recipe_ids) { + if (identical(cand, rid)) next + if (setequal(suffix_sets[[cand]], own_suffix)) return(cand) + } + NULL + } + + lapply(suggestions, function(s) { + # `s$ig_recipe_id <- NULL` would DELETE the element rather than set it + # (standard R list-assignment gotcha), leaving no-match entries missing + # the key entirely instead of carrying it as NULL. Single-bracket + # assignment with a wrapped list preserves a NULL-valued element so the + # field is always present, per the brief's "NULL when there is none". + s["ig_recipe_id"] <- list(find_counterpart(s$recipe_id)) + s + }) } #' Emit the single cli::cli_inform() message summarizing all suggestions #' for a verb call (the brief's "one message", not one per suggestion). #' Bullet text is pre-formatted plain text (no cli/glue `{}` markup) since #' recipe ids/labels are untrusted-ish data values, not literal call-site -#' expressions. +#' expressions. When a suggestion has an `ig_recipe_id`, one indented +#' continuation line is appended naming the intergovernmental counterpart +#' recipe (embedded `\n` renders as a hanging-indent continuation of the +#' same bullet under cli, not a new bullet). #' @noRd .inform_suggestions <- function(suggestions) { bullets <- vapply(suggestions, function(s) { - sprintf("%s (%d-%d): %s", s$recipe_id, + bullet <- sprintf("%s (%d-%d): %s", s$recipe_id, s$available_years[1], s$available_years[2], s$hint) + if (!is.null(s$ig_recipe_id)) { + bullet <- paste0(bullet, sprintf( + "\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id)) + } + bullet }, character(1)) cli::cli_inform(c( i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:", diff --git a/R/views.R b/R/views.R index a5139aa..c5febdd 100644 --- a/R/views.R +++ b/R/views.R @@ -1,23 +1,35 @@ # R/views.R -# SQL files whose view definitions read schema-v5-only parquet tables -# (harmonization_map.parquet, harmonization_recipes.parquet, -# series_breaks.parquet) or select from views built on top of them. DuckDB's -# read_parquet() resolves the file at CREATE VIEW time (even for a view, it -# still needs the source schema) and errors immediately -- "IO Error: No -# files found" -- if the path doesn't exist, so these cannot be registered -# unconditionally against a v4 corpus the way the rest of inst/sql/ is. -# Registration is therefore gated on manifest$schema_version >= 5; verb-level -# *usage* of the resulting views is separately gated by .resolve_basis() / -# .require_schema_v5(). +# SQL files that cannot be registered unconditionally against a v4 corpus, +# for one of two distinct reasons -- both fail at CREATE VIEW time (DuckDB +# resolves a view's source schema eagerly, even though it defers execution), +# so a v4 corpus can't tolerate either unconditionally: +# +# (a) Missing FILE. 33-/34-/35- read_parquet() a v5-only parquet table +# (harmonization_map.parquet, harmonization_recipes.parquet, +# series_breaks.parquet) that doesn't exist at all on a v4 corpus -- +# "IO Error: No files found". +# +# (b) Missing COLUMN. 22-/23-/25- reference `long.harmonized_code`, a +# column that does not exist on a v4 corpus's `long` table (harmonized +# space was introduced in schema v5) -- "Binder Error: Referenced +# column harmonized_code not found". 42-/43-/45- are on this list only +# because they SELECT s.* FROM the (a)/(b) views above, so they'd fail +# to resolve their own source view if it weren't already skipped. +# +# Registration is therefore gated on manifest$schema_version >= 5 for all of +# them; verb-level *usage* of the resulting views is separately gated by +# .resolve_basis() / .require_schema_v5(). .harmonization_view_files <- c( "22-spending_long_harmonized.sql", "23-revenue_long_harmonized.sql", + "25-ig_long_harmonized.sql", "33-harmonization_map.sql", "34-harmonization_recipes.sql", "35-series_breaks_pq.sql", "42-spending_annotated_harmonized.sql", - "43-revenue_annotated_harmonized.sql" + "43-revenue_annotated_harmonized.sql", + "45-ig_annotated_harmonized.sql" ) #' Register DuckDB views from inst/sql/ SQL files diff --git a/README.md b/README.md index d209f11..f700270 100644 --- a/README.md +++ b/README.md @@ -25,14 +25,32 @@ package implements. - `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory - `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600) +## Direct vs Total spending + +`cog_spending(..., expenditure_concept = c("direct", "total"))` controls +whose spending a result counts. `"direct"` (the default) is a government's +own current operations, capital outlay, and other direct spending. `"total"` +additionally adds in the intergovernmental legs — money it hands to other +governments to spend on its behalf — which is meaningful for describing one +government's own budget over time, but double-counts when summed across +governments (a state's payment to a county is the same dollar the county +reports as its own direct spending). + +**Rule of thumb: any figure that spans more than one government uses +`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this +by refusing `expenditure_concept = "total"`. See +`vignette("total-spending", package = "uscogdata")` for the full +explanation with worked examples. + ## Developer notes ### Testing The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` — -a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50 -states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` at this -fixture, so the full test suite runs offline with no network dependency: +a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering +all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` +at this fixture, so the full test suite runs offline with no network +dependency: ```r devtools::test() # uses bundled fixture, no credentials required diff --git a/inst/extdata/fixture_corpus/data/summary_categories.parquet b/inst/extdata/fixture_corpus/data/summary_categories.parquet index 054efa2..d87566b 100644 Binary files a/inst/extdata/fixture_corpus/data/summary_categories.parquet and b/inst/extdata/fixture_corpus/data/summary_categories.parquet differ diff --git a/inst/extdata/fixture_corpus/manifest.json b/inst/extdata/fixture_corpus/manifest.json index a28403e..d86af01 100644 --- a/inst/extdata/fixture_corpus/manifest.json +++ b/inst/extdata/fixture_corpus/manifest.json @@ -1,7 +1,7 @@ { "schema_version": 6, - "built_at": "2026-07-23T16:14:30Z", - "pipeline_commit": "4f992a0", + "built_at": "2026-07-27T13:04:05Z", + "pipeline_commit": "6098baf", "fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.", "data_vintage": { "source_vintages": { @@ -75,7 +75,7 @@ }, { "path": "data/summary_categories.parquet", - "sha256": "8e6fcd4dd9bb4723841a67233b19388c9762dfc23b4479501183cebf7ea3c1b5", + "sha256": "0985b607f3f35a8dff62c0561261ab6922423b81d11c07b03bcb3e3461f85e33", "description": "summary_categories.parquet" }, { diff --git a/inst/schemas/provenance-v1.json b/inst/schemas/provenance-v1.json index 4c040b7..d83a57b 100644 --- a/inst/schemas/provenance-v1.json +++ b/inst/schemas/provenance-v1.json @@ -12,6 +12,19 @@ "category": { "type": ["string", "array", "null"] }, "basis": { "type": ["string", "null"] }, "basis_note": { "type": ["string", "null"] }, + "expenditure_concept": { + "type": "string", + "enum": ["direct", "total"], + "description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments." + }, + "expenditure_concept_note": { + "type": ["string", "null"], + "description": "How the intergovernmental leg was assembled; null for 'direct'." + }, + "expenditure_concept_direct_suppressed": { + "type": "boolean", + "description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any." + }, "harmonization": { "type": "object" }, "recipe": { "type": ["object", "null"] }, "suggestions": { "type": "array" }, diff --git a/inst/sql/20-spending_long.sql b/inst/sql/20-spending_long.sql index 71eeec2..23c4d82 100644 --- a/inst/sql/20-spending_long.sql +++ b/inst/sql/20-spending_long.sql @@ -1,5 +1,5 @@ CREATE OR REPLACE VIEW spending_long AS SELECT * FROM long -WHERE LEFT(item_code, 1) IN ('E', 'F', 'G', 'K') +WHERE LEFT(item_code, 1) IN ('E', 'F', 'G') AND NOT is_aggregate; diff --git a/inst/sql/22-spending_long_harmonized.sql b/inst/sql/22-spending_long_harmonized.sql index ef70ac8..96baa58 100644 --- a/inst/sql/22-spending_long_harmonized.sql +++ b/inst/sql/22-spending_long_harmonized.sql @@ -3,4 +3,4 @@ SELECT * REPLACE (harmonized_code AS item_code) FROM long WHERE NOT is_aggregate AND harmonized_code IS NOT NULL - AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G', 'K'); + AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G'); diff --git a/inst/sql/24-ig_long.sql b/inst/sql/24-ig_long.sql new file mode 100644 index 0000000..ca29324 --- /dev/null +++ b/inst/sql/24-ig_long.sql @@ -0,0 +1,18 @@ +-- Intergovernmental expenditure rows (M = to local govts, L = to state govts). +-- +-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the +-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published +-- ONLY as aggregate-flagged rows -- filtering them would hide ~70% of legacy IG +-- dollars and make Total silently collapse to Direct. This is safe because the +-- aggregate codes and their modern leaf components are strictly year-disjoint +-- (M47 ends 2011 / M94 starts 2012; M89 is aggregate only <= 2011 and a leaf +-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument +-- the pipeline's recipe joins use. +-- +-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up +-- the L-NN codes, so including it would double-count. +CREATE OR REPLACE VIEW ig_long AS +SELECT * +FROM long +WHERE LEFT(item_code, 1) IN ('M', 'L') + AND item_code NOT LIKE '%--'; diff --git a/inst/sql/25-ig_long_harmonized.sql b/inst/sql/25-ig_long_harmonized.sql new file mode 100644 index 0000000..5187265 --- /dev/null +++ b/inst/sql/25-ig_long_harmonized.sql @@ -0,0 +1,15 @@ +-- Harmonized-basis IG rows. Uses COALESCE(harmonized_code, item_code) rather +-- than harmonized_code alone: aggregate rows carry NO harmonized_code by +-- construction (harmonized space is leaf-only), so a plain +-- `harmonized_code IS NOT NULL` filter would drop every legacy IG aggregate -- +-- in the bundled fixture corpus (year 2011; 2012+ all carry a harmonized_code) +-- that is $379,016,063k across 25,688 M rows and $2,277,458k across 19,266 L +-- rows (`SELECT year, LEFT(item_code,1), SUM(amt), COUNT(*) FROM ig_long +-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real +-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+) +-- while never dropping a row. +CREATE OR REPLACE VIEW ig_long_harmonized AS +SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code) +FROM long +WHERE LEFT(item_code, 1) IN ('M', 'L') + AND item_code NOT LIKE '%--'; diff --git a/inst/sql/44-ig_annotated.sql b/inst/sql/44-ig_annotated.sql new file mode 100644 index 0000000..ead46bd --- /dev/null +++ b/inst/sql/44-ig_annotated.sql @@ -0,0 +1,16 @@ +CREATE OR REPLACE VIEW ig_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.spend_subtype +FROM ig_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); diff --git a/inst/sql/45-ig_annotated_harmonized.sql b/inst/sql/45-ig_annotated_harmonized.sql new file mode 100644 index 0000000..68df811 --- /dev/null +++ b/inst/sql/45-ig_annotated_harmonized.sql @@ -0,0 +1,16 @@ +CREATE OR REPLACE VIEW ig_annotated_harmonized AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.spend_subtype +FROM ig_long_harmonized s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); diff --git a/man/cog_geographic_rollup.Rd b/man/cog_geographic_rollup.Rd index ffb2310..bafdc90 100644 --- a/man/cog_geographic_rollup.Rd +++ b/man/cog_geographic_rollup.Rd @@ -9,7 +9,8 @@ cog_geographic_rollup( category, years, per_capita = FALSE, - adjust_to_year = NULL + adjust_to_year = NULL, + expenditure_concept = c("direct", "total") ) } \arguments{ @@ -27,6 +28,13 @@ population from `gov_population_yearly`. Govs with missing population are excluded from the result.} \item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.} + +\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only +`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for +single-government queries but cannot be used here because combining Total +across multiple layers of government double-counts intergovernmental +transfers (a state's payment to a school district is the same dollar the +district reports as its own Direct spending).} } \value{ Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, diff --git a/man/cog_peer_compare.Rd b/man/cog_peer_compare.Rd index a4d9f96..d1a42fc 100644 --- a/man/cog_peer_compare.Rd +++ b/man/cog_peer_compare.Rd @@ -10,7 +10,8 @@ cog_peer_compare( category, years, per_capita = TRUE, - adjust_to_year = NULL + adjust_to_year = NULL, + expenditure_concept = c("direct", "total") ) } \arguments{ @@ -27,6 +28,11 @@ cog_peer_compare( population.} \item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.} + +\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only +`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for +single-government queries but cannot be used here because combining Total +across peer sets counts intergovernmental transfers twice.} } \value{ Tibble matching [cog_spending()]'s columns, plus a `role` diff --git a/man/cog_spending.Rd b/man/cog_spending.Rd index f1016ec..45f8d11 100644 --- a/man/cog_spending.Rd +++ b/man/cog_spending.Rd @@ -11,7 +11,8 @@ cog_spending( per_capita = FALSE, adjust_to_year = NULL, basis = c("harmonized", "raw"), - recipe = NULL + recipe = NULL, + expenditure_concept = c("direct", "total") ) } \arguments{ @@ -55,6 +56,35 @@ argument is ignored and the result's provenance reports `basis = "recipe"` with an inert `harmonization` block (`applied = FALSE`, pointing at the `recipe` block instead) rather than a possibly-misleading `"harmonized"`/`"raw"` value.} + +\item{expenditure_concept}{`"direct"` (default) returns only the +government's own direct spending (item codes `E`/`F`/`G`), unchanged +from prior releases. `"total"` additionally UNIONs in the +intergovernmental leg -- payments to local governments (`M` codes) and +to the state government (`L` codes, excluding the `L--` family-total +rollup) -- so results gain rows with `spend_subtype == +"intergovernmental"`. Requires the active corpus's `summary_categories` +to carry M/L rows (added by cog_pipeline PR #59); aborts with class +`uscogdata_ig_categories_unsupported` on an older corpus rather than +silently under-reporting. Mutually exclusive with `recipe` (a recipe +already defines its own component codes). **Do not sum `"total"` +results across levels of government** (e.g. state + county + city): +a state's `M12` payment to a school district is the same dollar the +district reports as its own direct `E12`, so summing both double-counts +it. This matters in particular with [cog_geographic_rollup()], which +sums across exactly that kind of multi-layer government set. + +In the legacy wide era (<= FY2011), some functions are published ONLY +as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05` +split), which the Direct leg excludes by construction but the IG leg +deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"` +query, any (year, category) where this leaves intergovernmental rows +with NO Direct counterpart is flagged: the affected rows' `notes` +name the harmonization recipe that recovers the missing Direct +component (when one exists), and +`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the +figure in those rows is the intergovernmental leg alone, not Direct + +IG.} } \value{ Tibble with columns `year`, `canonical_govid`, `gov_name`, diff --git a/tests/testthat/helper-fixture.R b/tests/testthat/helper-fixture.R index 0e650a6..30c7225 100644 --- a/tests/testthat/helper-fixture.R +++ b/tests/testthat/helper-fixture.R @@ -57,3 +57,37 @@ with_doctored_schema_version <- function(version, code) { }, add = TRUE) force(code) } + +# Copy the bundled fixture to a temp dir with summary_categories.parquet +# rewritten to drop every M/L (intergovernmental) row, then run `code` +# against it with a clean session (mirrors with_fixture_corpus()/ +# with_doctored_schema_version()). Models a real pre-cog_pipeline-PR#59 +# corpus: the 66 M/L category rows shipped with NO schema_version bump (see +# C2 in the expenditure-concept review), so schema_version is left +# untouched here -- only the category data itself is rolled back. +with_corpus_missing_ig_categories <- function(code) { + src <- fixture_corpus_path() + tmp <- withr::local_tempdir(.local_envir = parent.frame()) + file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf( + "COPY (SELECT * FROM read_parquet(%s) WHERE LEFT(item_code, 1) NOT IN ('M', 'L')) + TO %s (FORMAT PARQUET)", + uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) + )) + file.remove(cats_path) + file.rename(filtered_path, cats_path) + + old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) + uscogdata:::cog_close() + Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) + on.exit({ + uscogdata:::cog_close() + if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) + }, add = TRUE) + force(code) +} diff --git a/tests/testthat/test-categories.R b/tests/testthat/test-categories.R index af802f8..74b6aa1 100644 --- a/tests/testthat/test-categories.R +++ b/tests/testthat/test-categories.R @@ -15,7 +15,18 @@ test_that("cog_categories(type = 'spending') returns only expenditure rows", { skip_if_no_corpus() r <- cog_categories(type = "spending") expect_true(all(r$category_type == "expenditure")) - expect_true(all(r$subtype %in% c("operations", "capital"))) + expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental"))) +}) + +test_that("cog_categories surfaces the intergovernmental spending subtype", { + skip_if_no_corpus() + r <- cog_categories(type = "spending") + expect_true("intergovernmental" %in% r$subtype) + # IG rows reuse the existing functional categories -- they add a subtype, + # not new category values. + ig_cats <- sort(unique(r$category[r$subtype == "intergovernmental"])) + direct_cats <- sort(unique(r$category[r$subtype != "intergovernmental"])) + expect_true(all(ig_cats %in% c(direct_cats, "Other Education"))) }) test_that("cog_categories(type = 'revenue') returns only revenue rows", { diff --git a/tests/testthat/test-expenditure-concept.R b/tests/testthat/test-expenditure-concept.R new file mode 100644 index 0000000..95954ff --- /dev/null +++ b/tests/testthat/test-expenditure-concept.R @@ -0,0 +1,536 @@ +test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", { + con <- .ensure_session() + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM long WHERE LEFT(item_code, 1) = 'K'")$n + expect_equal(n, 0) + + sql_files <- c("20-spending_long.sql", "22-spending_long_harmonized.sql") + for (f in sql_files) { + txt <- paste(readLines(system.file("sql", f, package = "uscogdata")), + collapse = " ") + expect_false(grepl("'K'", txt, fixed = TRUE), + label = paste(f, "must not reference the inert K prefix")) + } +}) + +test_that("expenditure_concept defaults to direct and preserves today's numbers", { + gov <- "010000226085" # Alabama state government + base <- cog_spending(gov, years = 2019, category = "Police") + expl <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "direct") + expect_equal(base$amt_nominal, expl$amt_nominal) + expect_false("intergovernmental" %in% base$spend_subtype) +}) + +test_that("expenditure_concept = 'total' adds an intergovernmental subtype", { + gov <- "010000226085" + d <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "direct") + t <- cog_spending(gov, years = 2019, category = "Police", + expenditure_concept = "total") + expect_true("intergovernmental" %in% t$spend_subtype) + # Direct rows are untouched; Total only ever ADDS. Use %in% rather than + # != : a category = NULL result can contain a NULL-subtype group (codes + # with no summary_categories row, e.g. E16/E21/E85/F16/F85/G16/G21/G85), + # and `NA != "intergovernmental"` is NA, not TRUE, which would silently + # smuggle an all-NA phantom row into dt. + dt <- t[!(t$spend_subtype %in% "intergovernmental"), ] + expect_equal(sort(dt$amt_nominal), sort(d$amt_nominal)) + expect_gt(sum(t$amt_nominal), sum(d$amt_nominal)) +}) + +test_that("legacy-era Total does not collapse to Direct (the is_aggregate trap)", { + # In the wide era the IG dollars live almost entirely on aggregate-flagged + # rows. A Total leg that inherited the Direct leg's NOT is_aggregate filter + # would silently return Total == Direct here. + gov <- "010000226085" + d <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "direct") + t <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total") + expect_true("intergovernmental" %in% t$spend_subtype) + ig <- sum(t$amt_nominal[t$spend_subtype == "intergovernmental"]) + expect_gt(ig, 0) + expect_gt(sum(t$amt_nominal), sum(d$amt_nominal)) +}) + +test_that("the IG leg never includes the L-- family total", { + con <- .ensure_session() + codes <- DBI::dbGetQuery(con, + "SELECT DISTINCT item_code FROM ig_long")$item_code + expect_false(any(grepl("--$", codes))) + expect_true(all(substr(codes, 1, 1) %in% c("M", "L"))) +}) + +test_that("expenditure_concept rejects unknown values", { + expect_error( + cog_spending("010000226085", years = 2019, expenditure_concept = "gross"), + class = "rlang_error" + ) +}) + +test_that("total composes with basis = 'raw' and basis = 'harmonized'", { + gov <- "010000226085" + h <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total", basis = "harmonized") + r <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total", basis = "raw") + ig_h <- sum(h$amt_nominal[h$spend_subtype == "intergovernmental"]) + ig_r <- sum(r$amt_nominal[r$spend_subtype == "intergovernmental"]) + # The only IG harmonization rule is M38 -> M36 (year-disjoint), so the IG + # total must agree between bases even though the code labels may differ. + expect_equal(ig_h, ig_r) +}) + +test_that("recipe = and expenditure_concept = 'total' together aborts", { + expect_error( + cog_spending("121011212191", 2020L, recipe = "corrections_combined", + expenditure_concept = "total"), + class = "uscogdata_recipe_concept_conflict" + ) +}) + +test_that("aggregate-sourced IG dollars are flagged aggregate_fallback = TRUE (bool_or, not bool_and)", { + # Regression test: .build_verb_sql() originally used bool_and(is_aggregate) + # for aggregate_fallback, which is correct for the Direct leg (a group can + # never mix aggregate and non-aggregate rows there -- spending_long filters + # NOT is_aggregate) but wrong for the IG leg. The wide era is dense -- every + # government has a $0 row for every code in a family -- so a $0 leaf sits in + # the same (year, gov, subtype, category) group as the real aggregate row + # and flips bool_and() to FALSE. Measured: AL state 2011 had $5,740,775,000 + # of aggregate-sourced IG dollars (Corrections $31,358,000 + Education K-12 + # $5,152,385,000 + General Government $557,032,000) reporting + # aggregate_fallback = FALSE under bool_and(), with the only TRUE row being + # Transit Utilities at $0. bool_or() reports all of them correctly. + gov <- "010000226085" + t <- cog_spending(gov, years = 2011, category = "Education K-12", + expenditure_concept = "total") + ig <- t[t$spend_subtype == "intergovernmental", ] + expect_equal(nrow(ig), 1L) + expect_true(ig$aggregate_fallback) + expect_true(nzchar(ig$notes)) + expect_match(ig$notes, "Aggregate fallback applied", fixed = TRUE) +}) + +test_that("legacy aggregate IG codes are year-disjoint from their modern leaf components", { + # The safety of ig_long's deliberate omission of `NOT is_aggregate` (see + # inst/sql/24-ig_long.sql) rests entirely on each legacy code's AGGREGATE + # instance being year-disjoint from the modern leaf codes it rolls up -- + # if a future corpus rebuild ever back-filled a leaf into a year where the + # code is still flagged aggregate, `total` would silently double-count and + # this suite would still pass. This test fails loudly if that ever + # happens. + # + # Note the invariant is scoped to the AGGREGATE flag, not bare code + # presence: M89/L89 do NOT disappear after the wide era the way M47/L47 + # do -- they continue past 2011 as their OWN independent leaf line item + # (is_aggregate = FALSE) alongside M91-93/L91-93, which is fine because a + # non-aggregate M89/L89 no longer represents a rollup of those codes. + # (Verified in the fixture: M89/L89 are is_aggregate = TRUE only in 2011, + # when M91-93/L91-93 don't exist yet; from 2012 on M89/L89 are + # is_aggregate = FALSE leaves coexisting with M91-93/L91-93.) + # + # Pairs are the M/L-prefixed components (this package's ig_long only + # covers M/L; other prefixes in the same rollup, e.g. N/O/P/Q/R, fall + # outside its domain and are irrelevant here) enumerated in + # cog_pipeline's data/wide_to_long_xwalk.csv `full_desc` column (read + # once at authoring time, not at test time -- this test stays offline): + # M47 "To local governments, total (includes N47, O47, P47, R47, and M94)" + # M89 "To local governments, total (incl N89, O89, P89, R89, M91, M92, and M93)" + # L47 "To state government (includes L94)" + # L89 "To state government (includes L91, L92, and L93)" + con <- .ensure_session() + pairs <- list( + list(aggregate = "M47", components = "M94"), + list(aggregate = "M89", components = c("M91", "M92", "M93")), + list(aggregate = "L47", components = "L94"), + list(aggregate = "L89", components = c("L91", "L92", "L93")) + ) + agg_years_by_code <- DBI::dbGetQuery(con, + "SELECT DISTINCT year, item_code FROM ig_long WHERE is_aggregate") + codes_by_year <- DBI::dbGetQuery(con, "SELECT DISTINCT year, item_code FROM ig_long") + + for (p in pairs) { + agg_years <- agg_years_by_code$year[agg_years_by_code$item_code == p$aggregate] + for (yr in agg_years) { + codes_yr <- codes_by_year$item_code[codes_by_year$year == yr] + has_component <- any(p$components %in% codes_yr) + expect_false( + has_component, + label = sprintf( + "year %s has aggregate-flagged %s co-occurring with a modern component (%s)", + yr, p$aggregate, paste(p$components, collapse = ",") + ) + ) + } + } +}) + +test_that(".verb_spendrev rejects expenditure_concept = 'total' for a non-spending view_base", { + # cog_revenue() never exposes expenditure_concept and always resolves it + # to the "direct" default, so there is no revenue codepath that reaches + # this today -- but .verb_spendrev() is shared, and nothing else stops a + # future caller from passing expenditure_concept = "total" alongside + # view_base = "revenue_annotated", which would UNION expenditure M/L rows + # into a revenue result. Exercise the internal helper directly. + expect_error( + uscogdata:::.verb_spendrev( + verb = "cog_revenue_test", view_base = "revenue_annotated", + subtype_col = "revenue_subtype", + flow_prefixes = c("T", "A", "U", "B", "C", "D"), + call = quote(cog_revenue_test()), + govid = "010000226085", years = 2019L, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, basis = "raw", + recipe = NULL, expenditure_concept = "total" + ), + class = "uscogdata_expenditure_concept_unsupported" + ) +}) + +test_that("cog_geographic_rollup refuses expenditure_concept = 'total'", { + expect_error( + cog_geographic_rollup( + govids = list(state = "010000226085"), + category = "Police", years = 2019, + expenditure_concept = "total" + ), + class = "uscogdata_concept_not_aggregatable" + ) +}) + +test_that("cog_peer_compare refuses expenditure_concept = 'total'", { + expect_error( + cog_peer_compare( + target_govid = "010000226085", peers = "010000226085", + category = "Police", years = 2019, + expenditure_concept = "total" + ), + class = "uscogdata_concept_not_aggregatable" + ) +}) + +test_that("the refusal message names the fix and the reason", { + err <- tryCatch( + cog_geographic_rollup(govids = list(state = "010000226085"), + category = "Police", years = 2019, + expenditure_concept = "total"), + condition = function(e) e + ) + msg <- paste(conditionMessage(err), collapse = " ") + expect_match(msg, "direct") + expect_match(msg, "double-count|double count") + expect_match(msg, "cog_geographic_rollup") + + # Test that cog_peer_compare's message names its own function + err2 <- tryCatch( + cog_peer_compare(target_govid = "010000226085", peers = "010000226085", + category = "Police", years = 2019, + expenditure_concept = "total"), + condition = function(e) e + ) + msg2 <- paste(conditionMessage(err2), collapse = " ") + expect_match(msg2, "direct") + expect_match(msg2, "double-count|double count") + expect_match(msg2, "cog_peer_compare") +}) + +test_that("both cross-government verbs still accept the direct default", { + expect_no_error( + cog_geographic_rollup(govids = list(state = "010000226085"), + category = "Police", years = 2019) + ) + expect_no_error( + cog_peer_compare(target_govid = "010000226085", peers = "010000226085", + category = "Police", years = 2019) + ) +}) + +test_that("provenance always records the expenditure concept", { + d <- cog_spending("010000226085", years = 2019, category = "Police") + t <- cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total") + expect_equal(attr(d, "provenance")$expenditure_concept, "direct") + expect_equal(attr(t, "provenance")$expenditure_concept, "total") + # The note explains the non-obvious part: how legacy IG was assembled. + expect_true(nzchar(attr(t, "provenance")$expenditure_concept_note)) + expect_true(is.na(attr(d, "provenance")$expenditure_concept_note) || + !nzchar(attr(d, "provenance")$expenditure_concept_note)) +}) + +test_that("the provenance schema documents expenditure_concept", { + sch <- jsonlite::fromJSON( + system.file("schemas", "provenance-v1.json", package = "uscogdata"), + simplifyVector = FALSE + ) + expect_true("expenditure_concept" %in% names(sch$properties)) +}) + +test_that("a firing suggestion names the intergovernmental counterpart recipe", { + # Corrections has no legacy leaf rows, so the coverage-gap suggestion fires; + # corrections_ig_local_combined is its IG counterpart. + r <- suppressMessages( + cog_spending("010000226085", years = c(2005, 2011), category = "Corrections") + ) + sugg <- attr(r, "provenance")$suggestions + expect_gt(length(sugg), 0L) + ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1)) + expect_true("corrections_combined" %in% ids) + ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id)) + expect_true("corrections_ig_local_combined" %in% ig) +}) + +test_that("no suggestion fires for a healthy query", { + r <- cog_spending("010000226085", years = 2019, category = "Police") + expect_length(attr(r, "provenance")$suggestions, 0L) +}) + +test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to a revenue-flavored recipe", { + # "IG Federal" is a revenue-only category (summary_categories maps it to + # B-prefixed component codes only; its recipes are ig_federal_b47_wide / + # ig_federal_b89_wide). A cog_spending() call scoped to it returns zero + # spending rows for every requested year -- there is no spending + # component in this category at all -- so the coverage-gap machinery + # fires for real (not hypothetically) even though this isn't the kind of + # format-boundary gap the recipe catalog is meant to signpost. This is + # exactly the live-corpus risk flagged in review: ig_federal_b47_wide's + # own component codes (B47/B94, suffixes {"47","94"}) are an EXACT + # suffix-set match for the expenditure recipe ige_local_m47_wide + # (M47/M94, same suffixes) -- a coincidence of reused digits, not a real + # Direct/Total pairing. The flow-family gate in + # .attach_ig_counterparts() must keep ig_recipe_id NULL here. + r <- suppressMessages( + cog_spending("010000226085", years = c(2005, 2011), category = "IG Federal") + ) + sugg <- attr(r, "provenance")$suggestions + expect_gt(length(sugg), 0L) + ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1)) + expect_true("ig_federal_b47_wide" %in% ids) + ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id)) + expect_length(ig, 0L) +}) + +test_that("C1: 'total' on a legacy aggregate-only family reports the IG-only figure honestly, not as Direct + IG", { + # AL state government, Corrections, 2011. Measured pre-fix: 'total' + # returned $31,358,000 (the IG leg alone, on an aggregate-flagged M04/M05 + # row) with 0 suggestions (the surviving IG row made the gap-detection + # machinery think the Direct leg was covered) and a note asserting + # "Total = Direct + intergovernmental" with no caveat. True Direct (via + # recipe = "corrections_combined") is $521,651,000 -- the IG-only figure + # is ~6% of it. + gov <- "010000226085" + + d <- cog_spending(gov, years = 2011, category = "Corrections", + expenditure_concept = "direct") + expect_equal(nrow(d), 0L) + + t <- suppressMessages(cog_spending( + gov, years = 2011, category = "Corrections", expenditure_concept = "total" + )) + expect_equal(nrow(t), 1L) + expect_equal(t$spend_subtype, "intergovernmental") + expect_equal(t$amt_nominal, 31358000) + + r <- cog_spending(gov, years = 2011, recipe = "corrections_combined") + expect_equal(r$amt_nominal, 521651000) + + # C1(a): the recipe hints must fire for "total" exactly as they do for + # "direct" -- the surviving IG row must not be mistaken for Direct + # coverage. + prov <- attr(t, "provenance") + expect_gt(length(prov$suggestions), 0L) + ids <- vapply(prov$suggestions, function(s) s$recipe_id %||% "", character(1)) + expect_true("corrections_combined" %in% ids) + + # C1(b): the affected row's notes name a recovering recipe rather than + # staying silent, and the provenance carries a flag a downstream consumer + # (e.g. cog-api, which passes provenance through verbatim) can test. + expect_true(nzchar(t$notes)) + expect_match(t$notes, "unavailable", fixed = TRUE) + expect_match(t$notes, "corrections_combined", fixed = TRUE) + expect_true(prov$expenditure_concept_direct_suppressed) + + # The base "Total = Direct + IG" note must NOT stand unqualified when that + # arithmetic didn't actually happen for this row. + expect_match(prov$expenditure_concept_note, "NOTE", fixed = TRUE) + expect_match(prov$expenditure_concept_note, + "expenditure_concept_direct_suppressed", fixed = TRUE) +}) + +test_that("C1(b): expenditure_concept_direct_suppressed is FALSE when the Direct leg is present", { + d <- cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "direct") + t <- cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total") + expect_false(isTRUE(attr(d, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(isTRUE(attr(t, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(any(nzchar(t$notes[t$spend_subtype == "intergovernmental"]) & + grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"]))) +}) + +# M/I fix: .detect_direct_suppressed() was equating "no Direct sibling row" +# with "Direct was suppressed", but the dominant real cause is a government +# that simply has no direct spending in that category -- correct, ordinary +# data. The fix gates the flag (and its row note) on a harmonization recipe +# ACTUALLY covering that exact (year, canonical_govid, category) triple. + +test_that("M/I: true positive, category supplied explicitly (unchanged behavior)", { + al <- "010000226085" + t_cat <- suppressMessages(cog_spending( + al, years = 2011, category = "Corrections", expenditure_concept = "total" + )) + expect_true(attr(t_cat, "provenance")$expenditure_concept_direct_suppressed) + expect_match(t_cat$notes, "corrections_combined", fixed = TRUE) + expect_match(t_cat$notes, "unavailable", fixed = TRUE) +}) + +test_that("M/I: true positive, category = NULL now also names the recipe (was the fallback bug)", { + # Root bug: .build_suggestions() short-circuits to list() when category is + # NULL, so the note previously always hit its "no covering recipe found" + # fallback here even though corrections_combined genuinely covers this row. + al <- "010000226085" + t_null <- suppressMessages(cog_spending( + al, years = 2011, category = NULL, expenditure_concept = "total" + )) + corr_row <- t_null[t_null$category %in% "Corrections", ] + expect_equal(nrow(corr_row), 1L) + expect_true(attr(t_null, "provenance")$expenditure_concept_direct_suppressed) + expect_match(corr_row$notes, "corrections_combined", fixed = TRUE) + expect_match(corr_row$notes, "unavailable", fixed = TRUE) + expect_false(grepl("no covering recipe found", corr_row$notes, fixed = TRUE)) +}) + +test_that("M/I: false positive -- Virginia Education K-12 FY2019 total is NOT flagged", { + # States fund K-12 through school districts, so the Direct leg (E12/F12/ + # G12) is genuinely, correctly zero -- not suppressed. Must not be flagged + # and must carry no suppression note. + va <- "510000227542" + t_va <- suppressMessages(cog_spending( + va, years = 2019, category = "Education K-12", expenditure_concept = "total" + )) + expect_equal(nrow(t_va), 1L) + expect_equal(t_va$spend_subtype, "intergovernmental") + expect_equal(t_va$amt_nominal, 8028179000) + expect_false(isTRUE(attr(t_va, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(nzchar(t_va$notes) && grepl("unavailable", t_va$notes)) +}) + +test_that("M/I: false positive by construction -- 'Other Education' has no E/F/G code, never flagged", { + # "Other Education" maps only to M21/L21 in summary_categories -- there is + # no E/F/G code for it in this corpus at all, so no Direct-recovering + # recipe can exist and it must never be flagged, in any fixture year. + con <- uscogdata:::.ensure_session() + years_all <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year + states <- DBI::dbGetQuery(con, + "SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid + oe <- suppressMessages(cog_spending( + states, years = years_all, category = "Other Education", + expenditure_concept = "total" + )) + expect_false(isTRUE(attr(oe, "provenance")$expenditure_concept_direct_suppressed)) + expect_false(any(nzchar(oe$notes) & grepl("unavailable", oe$notes))) +}) + +test_that("M/I: a clean FY2019 category = NULL total query flags far fewer than the pre-fix 32/50 states", { + con <- uscogdata:::.ensure_session() + states <- DBI::dbGetQuery(con, + "SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid + r <- suppressMessages(cog_spending( + states, years = 2019, category = NULL, expenditure_concept = "total" + )) + ig <- r[r$spend_subtype == "intergovernmental", ] + flagged <- ig[nzchar(ig$notes) & grepl("unavailable", ig$notes), ] + expect_lt(length(unique(flagged$canonical_govid)), 32L) + # Every remaining flagged row must actually name a covering recipe -- + # never the old no-recipe-found fallback. + expect_true(all(grepl("recipe = '", flagged$notes, fixed = TRUE))) + expect_false(any(grepl("no covering recipe found", flagged$notes, fixed = TRUE))) +}) + +test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", { + with_corpus_missing_ig_categories({ + con <- uscogdata:::.ensure_session() + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')" + )$n + expect_equal(n, 0) + + err <- tryCatch( + cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total"), + condition = function(e) e + ) + expect_s3_class(err, "uscogdata_ig_categories_unsupported") + msg <- conditionMessage(err) + expect_match(msg, "PR #59|predates", perl = TRUE) + }) + + # 'direct' is unaffected on the same corpus -- the guard is scoped to + # expenditure_concept = "total" only. + with_corpus_missing_ig_categories({ + expect_no_error( + cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "direct") + ) + }) +}) + +test_that("C2: expenditure_concept = 'total' still works on a corpus that DOES carry M/L category rows", { + expect_no_error( + cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total") + ) +}) + +test_that("I2: an intergovernmental (M/L) recipe never appears as its own top-level suggestion", { + # Task 1's M04/M05 category rows share the "Corrections" summary_categories + # category with the Direct-flavored E04/E05, so `corrections_ig_local_ + # combined` (entirely M-prefixed) becomes a raw *candidate* in + # .build_suggestions()'s component_code-driven query. Following a + # "re-run with recipe = 'corrections_ig_local_combined'" hint on a plain + # cog_spending() call would silently return intergovernmental dollars + # under provenance$expenditure_concept = "direct". Task 6's gate + # (.attach_ig_counterparts()) already protects the *counterpart* lookup; + # this exercises that the candidate list itself is filtered too. + r <- suppressMessages( + cog_spending("010000226085", years = c(2005, 2011), category = "Corrections") + ) + sugg <- attr(r, "provenance")$suggestions + ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1)) + expect_true("corrections_combined" %in% ids) + expect_false("corrections_ig_local_combined" %in% ids) +}) + +test_that(".attach_ig_counterparts() never pairs a revenue-side recipe with its coincidental M/L suffix twin", { + # Broader version of the case above, run at the matching-helper level + # (the same level code review's pairwise enumeration was done at) rather + # than end-to-end: the fixture has no (govid, year) combination where + # cog_revenue() itself produces a covered gap for any B/C/D recipe, so an + # end-to-end repro for THIS specific set of recipes isn't reachable + # today. Each of these six recipes shares an exact suffix set with an + # M/L expenditure recipe purely by reused-digit coincidence: + # ig_federal_b47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide + # ig_federal_b89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide + # ig_state_c47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide + # ig_state_c89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide + # ig_local_d47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide + # ig_local_d89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide + # None of them may receive an ig_recipe_id under cog_revenue()'s own + # flow_prefixes, since M/L only ever pairs with the direct-expenditure + # (E/F/G) family. + con <- uscogdata:::.ensure_session() + fake_suggestion <- function(rid) { + list(recipe_id = rid, label = "x", available_years = c(1967L, 2023L), + hint = "h") + } + fake_suggestions <- lapply( + c("ig_federal_b47_wide", "ig_federal_b89_wide", + "ig_state_c47_wide", "ig_state_c89_wide", + "ig_local_d47_wide", "ig_local_d89_wide"), + fake_suggestion + ) + out <- uscogdata:::.attach_ig_counterparts( + con, fake_suggestions, c("T", "A", "U", "B", "C", "D") + ) + ig <- unlist(lapply(out, function(s) s$ig_recipe_id)) + expect_length(ig, 0L) +}) diff --git a/tests/testthat/test-explain.R b/tests/testthat/test-explain.R index e824807..fc8f7d3 100644 --- a/tests/testthat/test-explain.R +++ b/tests/testthat/test-explain.R @@ -70,6 +70,37 @@ test_that("cog_explain prints a Suggestions section when the provenance has one" expect_true(grepl("re-run with recipe", txt)) }) +test_that("cog_explain prints the expenditure concept (I1)", { + skip_if_no_corpus() + d <- cog_spending("010000226085", years = 2019, category = "Police") + t <- cog_spending("010000226085", years = 2019, category = "Police", + expenditure_concept = "total") + txt_d <- paste(c( + capture.output(cog_explain(d)), + capture.output(cog_explain(d), type = "message") + ), collapse = "\n") + txt_t <- paste(c( + capture.output(cog_explain(t)), + capture.output(cog_explain(t), type = "message") + ), collapse = "\n") + expect_true(grepl("Concept: direct", txt_d)) + expect_true(grepl("Concept: total", txt_t)) +}) + +test_that("cog_explain surfaces the C1(b) direct-suppressed flag as a warning", { + skip_if_no_corpus() + t <- suppressMessages(cog_spending( + "010000226085", years = 2011, category = "Corrections", + expenditure_concept = "total" + )) + expect_true(attr(t, "provenance")$expenditure_concept_direct_suppressed) + txt <- paste(c( + capture.output(cog_explain(t)), + capture.output(cog_explain(t), type = "message") + ), collapse = "\n") + expect_true(grepl("Direct leg unavailable", txt)) +}) + test_that("cog_explain prints denominator + popyear_range + counts", { skip_if_no_corpus() with_fixture_corpus({ diff --git a/tests/testthat/test-views.R b/tests/testthat/test-views.R index 75eaa5b..3bbfb88 100644 --- a/tests/testthat/test-views.R +++ b/tests/testthat/test-views.R @@ -9,7 +9,9 @@ test_that("all expected views register on session open", { expected <- c( "long", "spending_long", "revenue_long", "canonical_fips_xwalk", "summary_categories", - "spending_annotated", "revenue_annotated" + "spending_annotated", "revenue_annotated", + "ig_long", "ig_annotated", + "ig_long_harmonized", "ig_annotated_harmonized" ) expect_true(all(expected %in% views$table_name)) }) @@ -110,6 +112,70 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate ( expect_equal(rev$amt, 225) }) +test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmonized_code, and exclude the L-- family total (real SQL text, synthetic parquet)", { + # ig_long / ig_long_harmonized have the subtlest predicates in the package: + # a deliberately ABSENT `NOT is_aggregate` (unlike every other *_long view), + # and COALESCE(harmonized_code, item_code) instead of a plain + # `harmonized_code IS NOT NULL` filter. The only end-to-end guard on this + # today is bound to AL state / 2011 / Education K-12, where M12 happens to + # be the sole IG code present -- regenerate the fixture without that one + # row and the guard would die silently while staying green. As with the + # 22-/23- test above, this reads the real inst/sql/24-/25- text off disk + # and executes it against a synthetic hive-partitioned parquet tree, so a + # regression in either predicate changes which rows survive. + skip_if_no_corpus() + + tmp <- withr::local_tempdir() + part_dir <- file.path(tmp, "data", "long", "year=2004") + dir.create(part_dir, recursive = TRUE) + part_path <- file.path(part_dir, "part-0.parquet") + + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is + ('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis + ('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views + ('ig-D', 'L--', 55555, false, 'L--'), -- family total: excluded from BOTH views + ('ig-E', 'T29', 44444, false, 'T29') -- wrong prefix (revenue, not M/L): excluded from BOTH views + ) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(part_path))) + + sql_dir <- system.file("sql", package = "uscogdata") + .read_view_sql <- function(filename) { + txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n") + gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE) + } + + con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(con, .read_view_sql("10-long.sql")) + DBI::dbExecute(con, .read_view_sql("24-ig_long.sql")) + DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql")) + + raw <- DBI::dbGetQuery(con, + "SELECT item_code, SUM(amt) AS amt FROM ig_long + GROUP BY item_code ORDER BY item_code" + ) + # L-- (family total) and T29 (wrong prefix) are gone; the aggregate row + # M47 survives -- proof `NOT is_aggregate` is absent from ig_long. + expect_equal(raw$item_code, c("M04", "M38", "M47")) + expect_equal(raw$amt, c(100, 50, 99999)) + + harmonized <- DBI::dbGetQuery(con, + "SELECT item_code, SUM(amt) AS amt FROM ig_long_harmonized + GROUP BY item_code ORDER BY item_code" + ) + # M38 folds to M36 (real harmonized_code present); M47 keeps its raw code + # via COALESCE(NULL, 'M47') -- proof the aggregate row is NOT dropped by + # a plain `harmonized_code IS NOT NULL` filter. L-- and T29 stay excluded. + expect_equal(harmonized$item_code, c("M04", "M36", "M47")) + expect_equal(harmonized$amt, c(100, 50, 99999)) +}) + test_that(".build_series_break_refs matches fin_code + break_year window", { # No series_breaks_pq row falls inside the bundled fixture's 2011-2020 # window (data-verified; see the "series_break_refs" test in @@ -152,6 +218,108 @@ test_that("schema v5 harmonization views register when the corpus supports them" expect_true(all(expected_v5 %in% views$table_name)) }) +test_that(".harmonization_view_files guard is necessary: registration against a v4-shaped corpus (no harmonized_code column at all) succeeds only because the harmonized views are skipped", { + # with_doctored_schema_version() (used elsewhere in this suite) only + # rewrites manifest.json's schema_version -- the underlying `long` parquet + # is still the bundled v6 fixture, which DOES have a harmonized_code + # column, so it only proves the skip *happens*, not that it is *required*. + # This test builds a genuinely v4-shaped corpus: `long` has no + # harmonized_code column at all, matching a real pre-Phase-R2 publish + # tree, and then shows two things: (1) the real .register_views(), gated + # on manifest$schema_version, registers cleanly against it; (2) the exact + # SQL text of a gated file (25-ig_long_harmonized.sql), executed directly + # against the same corpus without the gate, fails -- proving the gate is + # load-bearing, not incidental. + tmp <- withr::local_tempdir() + part_dir <- file.path(tmp, "data", "long", "year=2004") + dir.create(part_dir, recursive = TRUE) + part_path <- file.path(part_dir, "part-0.parquet") + + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('gov-1', 'E36', 100, false, 500000, 2020) + ) AS t(canonical_govid, item_code, amt, is_aggregate, population, popyear) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(part_path))) + + xwalk_path <- file.path(tmp, "data", "canonical_fips_xwalk.parquet") + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('gov-1', 'Test Gov', 1, 'County', '01', '001', NULL, 500000) + ) AS t(canonical_govid, gov_name, govs_type, type_label, fips_state, + fips_county, fips_place, population_acs) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(xwalk_path))) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('E36', 'Test Category', 'expenditure', 'direct', NULL) + ) AS t(item_code, category, category_type, spend_subtype, revenue_subtype) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(cats_path))) + + # Confirm the synthetic `long` genuinely lacks harmonized_code (not just + # NULL values -- the column itself must be absent) before trusting the + # rest of this test. + cols <- DBI::dbGetQuery(write_con, sprintf( + "DESCRIBE SELECT * FROM read_parquet(%s)", uscogdata:::.sql_lit_chr(part_path) + ))$column_name + expect_false("harmonized_code" %in% cols) + + url <- paste0(tmp, "/") + + # (1) Full .register_views() against this v4-shaped corpus must succeed -- + # this is the behavior the guard exists to protect. + con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE) + expect_no_error( + uscogdata:::.register_views(con, url, manifest = list(schema_version = 4L)) + ) + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'")$table_name + expect_true(all(c("ig_long", "ig_annotated", "spending_annotated") %in% views)) + expect_false(any(c("ig_long_harmonized", "ig_annotated_harmonized", + "spending_long_harmonized") %in% views)) + + # (2) Prove the gate is load-bearing: the exact SQL text of the skipped + # file, executed directly (bypassing .register_views()'s schema_version + # check) against the SAME corpus, fails because it references + # long.harmonized_code, a column this corpus's `long` does not have. + sql_dir <- system.file("sql", package = "uscogdata") + .read_view_sql <- function(filename) { + txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n") + gsub("\\{url\\}", url, txt, fixed = FALSE) + } + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + DBI::dbExecute(con2, .read_view_sql("10-long.sql")) + expect_error(DBI::dbExecute(con2, .read_view_sql("25-ig_long_harmonized.sql"))) + + # Reconciling this test with the C2 guard (expenditure-concept review): + # `ig_annotated`/`spending_annotated` registering cleanly above proves + # only that CREATE VIEW binds against a `summary_categories` with no M/L + # rows at all (this synthetic corpus's own summary_categories has a + # single E36 row, see the COPY above) -- a LEFT JOIN never fails to + # resolve regardless of what the joined-to table contains. It does NOT + # mean querying expenditure_concept = "total" against this shape is safe: + # exactly this corpus (schema_version reported as supported, but + # summary_categories predates the M/L rows cog_pipeline PR #59 added) is + # what .require_ig_categories() exists to catch at the *verb* level, + # since PR #59 shipped those rows with no schema_version bump. Confirm + # the new runtime guard actually fires against this same `con`. + expect_error( + uscogdata:::.require_ig_categories(con), + class = "uscogdata_ig_categories_unsupported" + ) +}) + test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", { skip_if_no_corpus() con <- cog_open() diff --git a/vignettes/total-spending.Rmd b/vignettes/total-spending.Rmd new file mode 100644 index 0000000..0fb8b33 --- /dev/null +++ b/vignettes/total-spending.Rmd @@ -0,0 +1,229 @@ +--- +title: "Total spending: Direct, Total, and when each is right" +output: rmarkdown::html_vignette +vignette: > + %\VignetteIndexEntry{Total spending: Direct, Total, and when each is right} + %\VignetteEngine{knitr::rmarkdown} + %\VignetteEncoding{UTF-8} +--- + +```{r setup, include = FALSE} +knitr::opts_chunk$set(collapse = TRUE, comment = "#>") +``` + +# Two questions that sound the same but aren't + +"Total spending" means two different things depending on whether the question +is about one government or several: + +1. **"What did my county spend in total, a decade ago vs today?"** — one + government, tracked over time. Either `direct` or `total` spending answers + this correctly, as long as the same concept is used for both years. +2. **"How do all the counties in my state compare, a decade ago vs today, + against the neighboring state?"** — several governments, summed together. + Here only `direct` gives the right answer; summing `total` across + governments double-counts money that passes between them. + +`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`) +controls which of these a query answers. This vignette walks through both +questions with code that actually runs against the package's bundled fixture +corpus, then explains why the second question refuses `"total"` outright. + +```{r} +library(uscogdata) + +# Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50 +# states) so this vignette knits without network access. In real use, +# USCOGDATA_URL is instead set to the published corpus URL -- see README.md. +Sys.setenv(USCOGDATA_URL = paste0( + system.file("extdata/fixture_corpus", package = "uscogdata"), "/" +)) +``` + +The fixture doesn't carry 2017 or the present year, so the examples below use +the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs +today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published +corpus and swap in real years; the mechanics are identical. + +# Archetype 1: one government's own trend + +For a single government, `total` is a legitimate way to describe "everything +this government spent, including money it handed to other governments to +spend on its behalf": + +```{r} +al_total <- cog_spending( + "010000226085", # Alabama, the state government + years = c(2012, 2020), + category = "Highways", + expenditure_concept = "total" +) +al_total +``` + +The `intergovernmental` rows are what `"total"` adds on top of `"direct"` +(`capital` + `operations`): Alabama's own payments out to counties and +cities for highway work. Because this query only ever concerns Alabama, +including that piece is safe -- there's no other government's number it +could be double-counted against. + +`"direct"` (the default) answers the same trend question just as validly: + +```{r} +al_direct <- cog_spending( + "010000226085", years = c(2012, 2020), category = "Highways" + # expenditure_concept = "direct" is the default; shown here for contrast +) +al_direct +``` + +Both are internally consistent series. What breaks the comparison is +**switching concepts between the two years being compared** -- e.g. `direct` +for 2012 and `total` for 2020 -- which manufactures a trend that isn't +really there. Pick one concept for a given question and hold it fixed across +every year in the series. + +# Archetype 2: a cross-government rollup + +`cog_geographic_rollup()` sums spending across state/county/city layers for +a place. Its default -- and, as shown below, its *only* accepted value for +`expenditure_concept` -- is `"direct"`: + +```{r} +fl_rollup <- cog_geographic_rollup( + govids = list( + state = "120000226351", # Florida + county = c("121011212191", "121099101897") # Broward + Palm Beach + ), + category = "Highways", + years = c(2012, 2020) +) +fl_rollup +``` + +For the neighboring state, the comparison is a single government, so it's a +plain `cog_spending()` call rather than a rollup: + +```{r} +ga_state <- cog_spending( + "130000226087", years = c(2012, 2020), category = "Highways" # Georgia +) +ga_state +``` + +Now the same rollup, but asking for `expenditure_concept = "total"`: + +```{r, error = TRUE} +cog_geographic_rollup( + govids = list(state = "120000226351", county = "121011212191"), + category = "Highways", + years = 2020, + expenditure_concept = "total" +) +``` + +`cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason) +refuses `"total"` outright rather than silently returning an inflated +number. The next section is why. + +# The mechanism + +Suppose Alabama gives a county $10M toward a highway project. That $10M +shows up **twice** in the underlying corpus: + +- Once on Alabama's own record, coded `M44` ("to local governments, + Highways") -- Alabama's intergovernmental leg. +- Again on the county's record, coded `E44` / `F44` ("Highways, current + operations" / "capital outlay") -- the county's direct spending, because + the county is the government that actually lets the contract and pays the + paving crew. + +`direct` (item codes `E`/`F`/`G`) only ever counts the second of those -- +the government that actually did the spending. `total` (Direct plus the +`M`/`L` intergovernmental legs) counts the first one *as well*, which is +exactly right for describing Alabama's own budget: Alabama's `total` +genuinely includes the $10M it committed to highways, whether it built the +road itself or paid the county to. But sum `total` across Alabama **and** +the county, and that $10M is counted twice -- once as Alabama's payment out, +once as the county's spending in -- reporting $20M of highway work for $10M +actually spent. + +This is exactly the shape of query `cog_geographic_rollup()` exists to run +(summing across layers of government), so it refuses `"total"` rather than +silently overstating every multi-layer figure it produces. + +# How big is the risk in practice + +Intergovernmental transfers aren't evenly distributed by government type. +Measured against the bundled fixture corpus (all 50 states, each of its +four years -- 2011, 2012, 2019, 2020), intergovernmental spending as a +share of a government's own Direct spending is: + +| Government type | Intergovernmental / Direct | +|---|---| +| State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) | +| County | 3.4%-5.1% (varies by year) | +| City | 2.6%-3.1% (varies by year) | + +So the Direct/Total choice matters overwhelmingly for **state** governments +-- a state's Total genuinely differs from its Direct by a meaningful margin, +while for a county or city the two are close. The state range is also far +wider than a single flat figure would suggest: legacy wide-era years (2011: +48.4%) carry proportionally more intergovernmental spending than the modern +era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly +3x larger a decade earlier than it is today. That's also why the mistake +this vignette warns about is easy to make unnoticed at the county/city level +and costly at the state level: rolling up every government in a state using +`total` instead of `direct` overstates the true figure -- measured at 7.6% +for Alabama in FY2019, and 11.6% nationally. + +# Why Total = Direct + M + L, not Direct + M + +It's tempting to assume `total` only needs to add `M`. But `M` and `L` are +both money the queried government itself pays **out** -- they're not two +different accounts of a receiving government's revenue. `M` is what it +pays to other **local** governments (e.g. a county paying a city for a +shared paving contract); `L` is what it pays **up** to its **state** +government (e.g. a county's contribution to a state-administered program). +A local government's Total genuinely includes both legs, because both are +its own spending, just routed to a different kind of recipient. On the +bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for +state governments (a state has no "payments to the state government" leg of +its own) but is 43%-51% the size of `M` for counties (varies by year) and +144%-189% the size of `M` for cities (varies by year; 166% pooled across +all four) -- so a `total` that omitted `L` would silently undercount Total +specifically for local governments, and for cities `L` is often the +*larger* of the two legs. +`cog_spending(expenditure_concept = "total")` includes both legs (excluding +the `L--` family-total rollup row, which would double-count its own +components). + +# Composition rules + +- `expenditure_concept` (whose spending counts -- Direct vs Direct plus + intergovernmental) is **orthogonal** to `basis` (which vintage of the + item-code space a query resolves against -- `"harmonized"` vs `"raw"`). + They combine freely: `expenditure_concept = "total", basis = "raw"` is a + valid, meaningful query, and so is every other pairing. +- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a + recipe already defines its own component codes (some recipes have their + own matching intergovernmental counterpart recipe instead -- see + `cog_recipes()` and the "firing suggestion" notes surfaced in + `cog_spending()`'s provenance), so layering a second, generic `total` + union on top of a recipe query has no well-defined meaning. Passing both + together aborts with an error naming the conflict. +- `expenditure_concept` is a **spending-only** concept: `cog_revenue()` + doesn't expose it (revenue's own intergovernmental codes are a different + axis -- see `?cog_revenue`). + +# Summary + +- Comparing one government to itself over time: `"direct"` or `"total"` + both work -- pick one and hold it fixed across every year compared. +- Comparing or summing across governments -- counties within a state, a + state against its neighbor, cities against counties: use `"direct"`. + `cog_geographic_rollup()` and `cog_peer_compare()` enforce this by + refusing `"total"`. +- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local + governments + `L` to the state government, excluding the `L--` + family-total row).