diff --git a/.superpowers/plans/2026-08-03-cog-balances.md b/.superpowers/plans/2026-08-03-cog-balances.md new file mode 100644 index 0000000..02088d0 --- /dev/null +++ b/.superpowers/plans/2026-08-03-cog-balances.md @@ -0,0 +1,1050 @@ +# cog_balances() Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add `cog_balances()`, a third verb exposing the 14 cash-and-security holding codes (`category_type = 'balance'`), closing `uscogdata#25` requirement 2. + +**Architecture:** Two new DuckDB views (`balance_long`, `balance_annotated`) mirroring the `revenue_long`/`revenue_annotated` pair, registered by the existing `.register_views()` glob behind a new column-presence gate. A dedicated lean verb in `R/balances.R` reuses the shared SQL builder, per-capita join, inflation, recipe runner and provenance builder — it does **not** go through `.verb_spendrev()`, whose concept scoping, intergovernmental leg and `complete=` grid are all flow-specific. + +**Tech Stack:** R, DuckDB via DBI, testthat 3e, roxygen2, cli, tibble, dplyr. + +**Spec:** `specs/2026-08-03-cog-balances-design.md` + +## Global Constraints + +- Run R as `/usr/bin/Rscript`. **Never a conda R** — conda shadows `libPaths` and breaks `arrow`/`duckdb`. +- The test suite runs offline against the bundled fixture; `tests/testthat/setup.R` sets `USCOGDATA_URL` automatically. To run against the live corpus, `USCOGDATA_URL=/home/jared/Nextcloud/Civilytics/SHARE/uscogdata-corpus/` — **trailing slash required**. +- Full suite: `/usr/bin/Rscript -e 'testthat::test_local(".", reporter="silent")'`. Current floor: **716 pass / 0 fail / 0 skip**. Never let it drop. +- Corpus `amt` is in **thousands**; every verb multiplies by `1000.0` in SQL and returns **full US dollars**. Do not apply the conversion twice. +- SQL **view definitions** live in `inst/sql/`; **query construction** is inline `sprintf()` in R. Follow both — `CLAUDE.md`'s "never inline SQL" line is stale and Task 6 fixes it. +- Every verb returns a `tbl_df` with a `provenance` attribute. +- Commits are GPG-signed. If signing times out, retry once — pinentry is non-interactive here. +- **Never verify an absence through the filter that creates it.** Absence assertions read the raw corpus via `arrow::open_dataset()`, never through `cog_balances()`. + +## The test government + +`550000227544` — **WISCONSIN STATE GOVT**. Use this govid and no other. It is +the identifier other agents have standardised on for state governments, it is +stable across corpus vintages, and it is the same id used against the live API +(`/api/v1/governments/550000227544/...`). Verified present in the bundled +fixture: + +| year | codes present | +|---|---| +| 2011 | `X21`, **`X40`**, **`X41`**, `X42`, `X44`, `Y07`, `Y08` | +| 2012 | `W01`, `W31`, `W61`, `X21`, `X30`, `X47`, `Y07`, `Y08`, `Y21`, **`Z77`**, **`Z78`** | +| 2019 | `W01`, `W31`, `W61`, `Y07`, `Y08` | +| 2020 | `W01`, `W31`, `W61`, `Y07`, `Y08`, `Y21` | + +This reaches four of the five subtypes — `general` (W), `employee_retirement` +(X/Z), `unemployment_trust` (Y07/Y08) and `workers_comp_trust` (Y21) — and +carries **both** wide→modern recipe bridges (`X40`→`Z77` and `X41`→`Z78`), so +the whole plan is testable offline. + +`other_insurance_trust` (`Y61`) is not present for Wisconsin in the fixture. No +test below depends on it; do not substitute a different government to reach it. + +Do **not** assert on `gov_name`. The fixture carries both `"WISCONSIN"` (from +`long`) and `"WISCONSIN STATE GOVT"` (from the xwalk), and the verb resolves +`COALESCE(xwalk_gov_name, gov_name)`. + +## File Structure + +| File | Responsibility | +|---|---| +| `inst/sql/26-balance_long.sql` (create) | Restrict `long` to `category_type = 'balance'`, drop aggregates | +| `inst/sql/46-balance_annotated.sql` (create) | Join government xwalk + crosswalk onto `balance_long` | +| `R/views.R` (modify) | Third gate list: skip both views when the corpus lacks `balance_subtype` | +| `R/balances.R` (create) | `cog_balances()`, `.require_balance_support()` | +| `R/balance_caveats.R` (create) | `.balance_caveats()`, `.balance_caveat_once()` | +| `R/session.R` (modify) | Reset the once-per-session caveat log on `cog_close()` | +| `tests/testthat/helper-fixture.R` (modify) | `with_corpus_missing_balance_subtype()` | +| `tests/testthat/test-balances.R` (create) | Verb behaviour, guards, per-capita, recipe, caveats | +| `NAMESPACE`, `man/` | Regenerated by `devtools::document()` | + +--- + +### Task 1: The two views and the registration gate + +**Files:** +- Create: `inst/sql/26-balance_long.sql`, `inst/sql/46-balance_annotated.sql` +- Modify: `R/views.R` +- Modify: `tests/testthat/helper-fixture.R` +- Test: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.register_views(con, url, manifest)`, `.corpus_has_table(manifest, file)` (both existing in `R/views.R`). +- Produces: DuckDB views `balance_long` and `balance_annotated`; `.balance_view_files` (character vector); `.corpus_has_balance_subtype(con)` returning `logical(1)`; test helper `with_corpus_missing_balance_subtype(code)`. + +- [ ] **Step 1: Write the failing test** + +Create `tests/testthat/test-balances.R`: + +```r +test_that("balance views register and carry only balance codes", { + skip_if_no_corpus() + con <- cog_open() + on.exit(cog_close()) + + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + expect_true(all(c("balance_long", "balance_annotated") %in% views)) + + # Every item_code in balance_long is a category_type = 'balance' member. + leak <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long + WHERE item_code NOT IN ( + SELECT item_code FROM summary_categories WHERE category_type = 'balance')" + )$n + expect_identical(as.integer(leak), 0L) + + # And no aggregate row survives, mirroring revenue_long. + agg <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long WHERE is_aggregate" + )$n + expect_identical(as.integer(agg), 0L) + + # balance_annotated exposes the subtype column the verb groups on. + cols <- DBI::dbGetQuery(con, + "SELECT column_name FROM information_schema.columns + WHERE table_name = 'balance_annotated'" + )$column_name + expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) +}) +``` + +- [ ] **Step 2: Run it and watch it fail** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `balance_long` is not among the registered views. + +- [ ] **Step 3: Create the two view files** + +`inst/sql/26-balance_long.sql`: + +```sql +-- Cash and security holdings, classified by crosswalk MEMBERSHIP on +-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot +-- do this job -- the X and Y families each span revenue, expenditure AND +-- balance). +-- +-- These rows are STOCKS: a balance at a point in time, not a flow over a +-- fiscal year. Summing a stock with a flow is meaningless, which is why they +-- live behind a third view rather than as a subtype of either money view, and +-- why neither spending_long nor revenue_long can reach them. +-- +-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era +-- aggregate-only holdings codes (X40/X41) are deliberately outside this view; +-- they are reachable only through the recipe path, which bypasses this filter +-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2). +CREATE OR REPLACE VIEW balance_long AS +SELECT * +FROM long +WHERE item_code IN ( + SELECT item_code FROM summary_categories + WHERE category_type = 'balance' + ) + AND NOT is_aggregate; +``` + +`inst/sql/46-balance_annotated.sql`: + +```sql +CREATE OR REPLACE VIEW balance_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.balance_subtype +FROM balance_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); +``` + +- [ ] **Step 4: Add the gate to `R/views.R`** + +Insert after the `.representation_view_files` block (around line 46): + +```r +# Cash and security holdings (uscogdata#25). 46- selects +# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and +# WITHOUT a schema_version bump -- so neither existing gate applies: +# .harmonization_view_files keys on schema_version, .representation_view_files +# on the presence of a FILE. Here the discriminator is a COLUMN on a table +# that exists either way. CREATE VIEW resolves its source schema eagerly, so +# on an older corpus 46- would fail at registration with "Binder Error: +# Referenced column balance_subtype not found" rather than at query time. +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") + +#' Does the mounted corpus's `summary_categories` carry `balance_subtype`? +#' Probed against the live connection rather than the manifest, because the +#' manifest describes files, not columns. +#' @noRd +.corpus_has_balance_subtype <- function(con) { + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM information_schema.columns + WHERE table_name = 'summary_categories' + AND column_name = 'balance_subtype'" + )$n + isTRUE(as.integer(n) > 0L) +} +``` + +Then inside the `for (f in files)` loop in `.register_views()`, after the +existing two `next` guards: + +```r + if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next +``` + +This is safe ordering: `11-summary_categories.sql` sorts before `26-`, so the +table exists by the time the probe runs. + +- [ ] **Step 5: Run the test — it should pass** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 6: Add the gate helper and its test** + +Append to `tests/testthat/helper-fixture.R`: + +```r +# Copy the bundled fixture to a temp dir with summary_categories.parquet +# rewritten to DROP the balance_subtype column, then run `code` against it. +# Models a corpus published before cog_pipeline #76/#77. schema_version is +# left untouched deliberately: that change shipped without a version bump, so +# column presence is the only honest signal -- this helper is what proves the +# package keys off it. Mirrors with_corpus_missing_ig_categories(). +with_corpus_missing_balance_subtype <- function(code) { + src <- fixture_corpus_path() + tmp <- withr::local_tempdir(.local_envir = parent.frame()) + file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf( + "COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s)) + TO %s (FORMAT PARQUET)", + uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) + )) + file.remove(cats_path) + file.rename(filtered_path, cats_path) + + old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) + uscogdata:::cog_close() + Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) + on.exit({ + uscogdata:::cog_close() + if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) + }, add = TRUE) + force(code) +} +``` + +Add to `tests/testthat/test-balances.R`: + +```r +test_that("balance views are skipped on a corpus without balance_subtype", { + skip_if_no_corpus() + with_corpus_missing_balance_subtype({ + con <- cog_open() + on.exit(cog_close()) + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + # Registration must SKIP them, not error -- an older corpus stays usable. + expect_false(any(c("balance_long", "balance_annotated") %in% views)) + expect_true("revenue_long" %in% views) + }) +}) +``` + +- [ ] **Step 7: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors, and at least 716 passes. + +- [ ] **Step 8: Commit** + +```bash +git add inst/sql/26-balance_long.sql inst/sql/46-balance_annotated.sql R/views.R tests/testthat/helper-fixture.R tests/testthat/test-balances.R +git commit -m "feat: register balance_long / balance_annotated behind a column gate (#25)" +``` + +--- + +### Task 2: `cog_balances()` core + +**Files:** +- Create: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` +- Regenerate: `NAMESPACE`, `man/cog_balances.Rd` + +**Interfaces:** +- Consumes: `.ensure_session()`, `.coerce_govid_input(govid)`, `.check_govids_in_scope(govid)`, `.build_verb_sql(view, subtype_col, govid, years, category, ig_view, subtype_scope)`, `.build_provenance(...)`, `.sql_lit_chr(x)` — all existing internals. +- Produces: exported `cog_balances(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, basis = c("harmonized","raw"), recipe = NULL)` returning a `tbl_df` with columns `year`, `canonical_govid`, `gov_name`, `balance_subtype`, `category`, `amt_nominal`, `codes_included`, `aggregate_fallback` and a `provenance` attribute. Also `.require_balance_support(con)`, which aborts with class `uscogdata_no_balance_support`. + +- [ ] **Step 1: Write the failing tests** + +Add to `tests/testthat/test-balances.R`: + +```r +test_that("cog_balances returns holdings for a government that has them", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + expect_s3_class(r, "tbl_df") + expect_true(nrow(r) > 0L) + expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", + "category", "amt_nominal") %in% names(r))) + expect_identical(sort(unique(r$category)), + c("Fund Balances", "Insurance Trust Balances")) + expect_false(is.null(attr(r, "provenance"))) + expect_identical(attr(r, "provenance")$verb, "cog_balances") + }) +}) + +test_that('category = "Fund Balances" is exactly the general family', { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + expect_identical(codes, c("W01", "W31", "W61")) + }) +}) + +test_that("no flow code can reach cog_balances", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + + # The expected set is read from the RAW corpus, never from the verb -- + # verifying an absence through the filter that creates it proves nothing. + ds <- arrow::open_dataset(file.path(fixture_corpus_path(), "data", "long")) + sc <- arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet")) + sc <- as.data.frame(sc) + balance_codes <- sc$item_code[sc$category_type == "balance"] + + expect_true(all(got %in% balance_codes)) + expect_true(length(setdiff(got, balance_codes)) == 0L) + }) +}) + +test_that("every balance_subtype maps to exactly one category", { + skip_if_no_corpus() + # Dropping the `subtype` argument is only safe while this tree holds. If the + # pipeline ever gives a balance subtype a second category, `category` becomes + # a lossy filter -- fail HERE rather than in a user's analysis. + sc <- as.data.frame(arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) + b <- sc[sc$category_type == "balance", ] + per_subtype <- tapply(b$category, b$balance_subtype, + function(x) length(unique(x))) + expect_true(all(per_subtype == 1L)) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `could not find function "cog_balances"`. + +- [ ] **Step 3: Create `R/balances.R`** + +```r +# R/balances.R +# +# Cash and security holdings. A third verb rather than an argument on a money +# verb because holdings are a STOCK -- a balance at a point in time -- while +# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money +# verbs' whole argument vocabulary (expenditure_concept, revenue_concept, +# complete=) describes flows and is meaningless here, so this deliberately +# does NOT route through .verb_spendrev(). + +#' Cash and security holdings for one or more governments +#' +#' Returns Census cash-and-security holdings (`category_type = "balance"`): +#' fund balances, retirement system holdings and insurance trust balances. +#' +#' @section Holdings are not GAAP fund balance: +#' Census holdings are **gross** -- no liabilities are netted -- so a reserve +#' ratio built from them overstates what is actually available. They are not +#' comparable to a GAAP fund balance from an ACFR. +#' +#' @param govid Canonical govid(s): a character vector, or a data frame with a +#' `canonical_govid` column (e.g. from [cog_gov_search()]). +#' @param years Integer vector of fiscal years. +#' @param category Optional character vector of categories to keep. One of +#' `"Fund Balances"`, `"Insurance Trust Balances"`, +#' `"Retirement System Holdings"`. There is deliberately no `subtype` +#' argument: for holdings, `category` is a strict coarsening of +#' `balance_subtype` (unlike the money verbs, where the two axes cross), so +#' every combination would be either redundant or empty. +#' `category = "Fund Balances"` is exactly the `general` family +#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +#' one `dplyr::filter()` away. +#' @param per_capita Divide holdings by population. Note this is a **stock per +#' resident** (reserves per person), which is *not* comparable to +#' [cog_spending()]'s per-capita figures -- those are a flow per person. +#' @param adjust_to_year Deflate to this year's dollars (CPI-U). +#' @param basis Accepted for uniformity with the money verbs, but currently a +#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized +#' and raw space are identical for holdings. Reported in +#' `provenance$basis_note`. +#' @param recipe Optional harmonization recipe id (see [cog_recipes()]). +#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +#' wide era to the modern one. +#' +#' @return A `tbl_df` with a `provenance` attribute. Amounts are full US +#' dollars. +#' @export +cog_balances <- function(govid, years, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, + basis = c("harmonized", "raw"), recipe = NULL) { + call <- match.call() + basis <- match.arg(basis, c("harmonized", "raw")) + govid <- .coerce_govid_input(govid) + years <- as.integer(years) + if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) + + con <- .ensure_session() + .require_balance_support(con) + .check_govids_in_scope(govid) + + basis_note <- paste0( + "`basis` has no effect on holdings: harmonization_map carries no ", + "balance-code rows, so harmonized and raw space are identical here." + ) + + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + + prov <- .build_provenance( + verb = "cog_balances", call = call, govid = govid, years = years, + category = category, per_capita = per_capita, + adjust_to_year = adjust_to_year, result = result, sql = sql, + subtype_col = "balance_subtype", + basis = basis, basis_note = basis_note, + # Neither concept vocabulary applies to a stock. + expenditure_concept = NA_character_, + revenue_concept = NA_character_ + ) + + attr(result, "provenance") <- prov + result +} + +#' Abort unless the mounted corpus classifies balance codes. +#' +#' `balance_subtype` arrived with cog_pipeline #76/#77 without a +#' schema_version bump, so the check is on the column, not the version. +#' @noRd +.require_balance_support <- function(con) { + if (.corpus_has_balance_subtype(con)) return(invisible(TRUE)) + cli::cli_abort( + c("This corpus does not classify cash and security holdings.", + i = "`summary_categories` has no {.field balance_subtype} column.", + i = "Republish from cog_pipeline at #76/#77 or later."), + class = "uscogdata_no_balance_support" + ) +} +``` + +`codes_included` stays on the returned tibble — `cog_spending()` and +`cog_revenue()` both keep it, and dropping it here would be a gratuitous +asymmetry. It is *also* mirrored into `provenance$codes_summed$observed` by +`.build_provenance()`, which is what the series-break builder reads. + +- [ ] **Step 4: Document and run** + +```bash +/usr/bin/Rscript -e 'devtools::document()' +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS, and `NAMESPACE` gains `export(cog_balances)`. + +- [ ] **Step 5: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors. + +- [ ] **Step 6: Commit** + +```bash +git add R/balances.R NAMESPACE man/ tests/testthat/test-balances.R +git commit -m "feat: cog_balances() core verb (#25)" +``` + +--- + +### Task 3: `per_capita` and `adjust_to_year` + +**Files:** +- Modify: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.attach_per_capita(result, con, govid)` and `.attach_real_dollars(result, adjust_to_year, per_capita)` from `R/spending.R`. The first adds `amt_per_capita_nominal` and `pop_source` and sets the `.popyear_range` attribute; the second adds `amt_real` and, when `per_capita`, `amt_per_capita_real`. +- Produces: no new functions — `cog_balances()` gains the two columns. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("per_capita divides holdings by population", { + skip_if_no_corpus() + with_fixture_corpus({ + plain <- cog_balances("550000227544", 2019, category = "Fund Balances") + pc <- cog_balances("550000227544", 2019, category = "Fund Balances", + per_capita = TRUE) + expect_true("amt_per_capita_nominal" %in% names(pc)) + expect_true("pop_source" %in% names(pc)) + expect_identical(pc$amt_nominal, plain$amt_nominal) + + # Assert against the denominator read from the corpus, NOT against a + # quantity derived from amt_per_capita_nominal itself -- dividing the + # column back out would be tautological and would pass on any value. + pop <- DBI::dbGetQuery(cog_open(), sprintf( + "SELECT population FROM gov_population_yearly + WHERE canonical_govid = %s AND year = 2019", + uscogdata:::.sql_lit_chr("550000227544") + ))$population + expect_length(pop, 1L) + expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, + tolerance = 1e-8) + + prov <- attr(pc, "provenance") + expect_true(prov$transformations$per_capita$applied) + }) +}) + +test_that("adjust_to_year adds real dollars", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r)) + # 2012 dollars inflated to 2020 must exceed nominal. + expect_true(all(r$amt_real > r$amt_nominal)) + prov <- attr(r, "provenance") + expect_true(prov$transformations$inflation$applied) + expect_identical(prov$transformations$inflation$base_year, 2020L) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `amt_per_capita_nominal` not among names. + +- [ ] **Step 3: Wire the two helpers into `cog_balances()`** + +In `R/balances.R`, between the `dbGetQuery()` call and `.build_provenance()`: + +```r + if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid) + if (!is.null(adjust_to_year)) { + result <- .attach_real_dollars(result, adjust_to_year, per_capita) + } +``` + +Order matters and matches `.verb_spendrev()`: per-capita first, so the real +per-capita column is deflated from the nominal per-capita one rather than +recomputed. + +- [ ] **Step 4: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add R/balances.R tests/testthat/test-balances.R +git commit -m "feat: per_capita and adjust_to_year for cog_balances() (#25)" +``` + +--- + +### Task 4: `recipe=` — the wide-era bridge + +**Files:** +- Modify: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.require_schema_v5(con, manifest, what)`, `.validate_recipe_id(con, recipe_id)`, `.recipe_components(con, recipe_id)` (returns a data frame with `label`, `component_code`, `year_min`, `year_max`, `weight`), `.run_recipe(con, recipe_id, govid, years)` (returns a tibble with `sql_query` attribute), `.shape_recipe_result(result, subtype_col, label)`, `.df_to_row_list(df)`. +- Produces: no new functions. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("recipe bridges the wide era into the modern one", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z77_wide") + expect_identical(sort(r$year), c(2011L, 2012L)) + + # The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE + # and therefore invisible to balance_long. If the recipe path ever starts + # filtering aggregates, a 45-year series silently truncates to five -- + # this is the regression guard for phase_r_harmonization_review.md § 0.2. + codes <- attr(r, "provenance")$codes_summed$observed + expect_true("X40" %in% codes) + expect_true("Z77" %in% codes) + expect_true(all(r$amt_nominal > 0)) + + prov <- attr(r, "provenance") + expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide") + }) +}) + +test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z77_wide") + refs <- attr(r, "provenance")$series_break_refs + # SB195 sits on fin_code X40; it can only fire where X40 is observed, + # which is exactly the recipe path. + expect_true("SB195" %in% refs) + }) +}) + +test_that("the second holdings bridge works too", { + skip_if_no_corpus() + with_fixture_corpus({ + # X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011 + # and Z78 in 2012, so both legs are exercised. + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z78_wide") + codes <- attr(r, "provenance")$codes_summed$observed + expect_true(all(c("X41", "Z78") %in% codes)) + expect_identical(sort(r$year), c(2011L, 2012L)) + }) +}) + +test_that("an unknown recipe id is rejected", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_error(cog_balances("550000227544", 2019, recipe = "no_such_recipe")) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `recipe` is accepted but ignored, so 2011 returns no rows. + +- [ ] **Step 3: Add the recipe branch** + +In `R/balances.R`, replace the single `sql <- .build_verb_sql(...)` / +`result <- ...` pair with: + +```r + manifest <- .uscogdata_env$manifest + recipe_block <- NULL + category_for_prov <- category + + if (!is.null(recipe)) { + .require_schema_v5(con, manifest, "recipe =") + .validate_recipe_id(con, recipe) + comps <- .recipe_components(con, recipe) + recipe_label <- comps$label[[1]] + result <- .run_recipe(con, recipe, govid, years) + sql <- attr(result, "sql_query") + result <- .shape_recipe_result(result, "balance_subtype", recipe_label) + recipe_block <- list( + recipe_id = recipe, label = recipe_label, + components = .df_to_row_list(comps) + ) + category_for_prov <- recipe_label + } else { + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + } +``` + +and pass the two new values through to `.build_provenance()`: + +```r + category = category_for_prov, + recipe = recipe_block, +``` + +- [ ] **Step 4: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. If `SB195` is absent, check that `.build_series_break_refs()` +received `X40` — it reads `provenance$codes_summed$observed`, which +`.shape_recipe_result()` populates from `codes_included`. + +- [ ] **Step 5: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +- [ ] **Step 6: Commit** + +```bash +git add R/balances.R tests/testthat/test-balances.R +git commit -m "feat: recipe= bridges the wide-era holdings series (#25)" +``` + +--- + +### Task 5: `balance_caveats` provenance and the once-per-session message + +**Files:** +- Create: `R/balance_caveats.R` +- Modify: `R/balances.R`, `R/session.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.uscogdata_env` (mutable environment from `R/config.R`), `cog_close()` in `R/session.R`. +- Produces: `.balance_caveats(con, codes_observed, years)` returning `list(not_gaap = TRUE, coverage_window = , truncated = )`; `.balance_caveat_once(key)` returning `TRUE` the first time a key is seen in a session and `FALSE` after. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("balance_caveats is always present and flags the GAAP distinction", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + cav <- attr(r, "provenance")$balance_caveats + expect_false(is.null(cav)) + expect_true(cav$not_gaap) + }) +}) + +test_that("coverage_window is computed from the corpus, not hardcoded", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + cav <- attr(r, "provenance")$balance_caveats + + ds <- arrow::open_dataset( + file.path(fixture_corpus_path(), "data", "long")) + sc <- as.data.frame(arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) + gen <- sc$item_code[sc$category_type == "balance" & + sc$balance_subtype == "general"] + obs <- as.data.frame( + dplyr::collect(dplyr::summarise( + dplyr::filter(ds, item_code %in% gen), + y0 = min(year), y1 = max(year)))) + + expect_identical(as.integer(cav$coverage_window$general), + c(as.integer(obs$y0), as.integer(obs$y1))) + }) +}) + +test_that("a request past a family's coverage window is flagged", { + skip_if_no_corpus() + with_fixture_corpus({ + # employee_retirement stops at FY2016; 2019/2020 are past it. + r <- cog_balances("550000227544", c(2012, 2019)) + cav <- attr(r, "provenance")$balance_caveats + expect_true("employee_retirement" %in% cav$truncated) + }) +}) + +test_that("the caveat message fires once per session", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_message(cog_balances("550000227544", 2019), "not.*GAAP") + expect_no_message(cog_balances("550000227544", 2020)) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `balance_caveats` is NULL. + +- [ ] **Step 3: Create `R/balance_caveats.R`** + +```r +# R/balance_caveats.R +# +# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and +# security holdings. Each one silently invalidates an obvious analysis, so +# they travel in provenance (machine-readable, for cog-api#26) rather than +# living only in prose. +# +# Two of the four are already carried by the code-driven series-break +# builders and are deliberately NOT duplicated here: +# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via +# series_break_refs on the recipe path, the only path that observes those +# codes. +# What remains is the GAAP distinction (a constant) and the coverage windows +# (measured, never hardcoded, so they stay correct as the corpus grows). + +#' Per-subtype observed year extents, plus which requested families are +#' truncated relative to the requested span. +#' @noRd +.balance_caveats <- function(con, codes_observed, years) { + windows <- DBI::dbGetQuery(con, + "SELECT c.balance_subtype AS subtype, + MIN(l.year) AS year_min, + MAX(l.year) AS year_max + FROM balance_long l + JOIN summary_categories c USING (item_code) + WHERE c.balance_subtype IS NOT NULL + GROUP BY 1 + ORDER BY 1" + ) + + observed_subtypes <- if (length(codes_observed) == 0L) { + character(0) + } else { + DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT balance_subtype FROM summary_categories + WHERE item_code IN (%s) AND balance_subtype IS NOT NULL", + .sql_lit_chr(codes_observed) + ))$balance_subtype + } + + cw <- stats::setNames( + lapply(seq_len(nrow(windows)), + function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), + windows$subtype + ) + + # A family is "truncated" when the caller asked for years outside the span + # that family actually covers -- the FY2016 employee-retirement termination + # and the FY2021 end of the W family are both this shape. + truncated <- character(0) + if (length(years) > 0L) { + for (s in observed_subtypes) { + w <- cw[[s]] + if (is.null(w)) next + if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s) + } + } + + list( + not_gaap = TRUE, + not_gaap_note = paste0( + "Census holdings are gross -- no liabilities are netted -- and are NOT ", + "GAAP fund balance. A reserve ratio built from them overstates what is ", + "actually available." + ), + coverage_window = cw, + truncated = sort(unique(truncated)) + ) +} + +#' TRUE the first time `key` is seen this session, FALSE thereafter. +#' Reset by cog_close(). +#' @noRd +.balance_caveat_once <- function(key) { + seen <- .uscogdata_env$balance_caveats_shown + if (is.null(seen)) seen <- character(0) + if (key %in% seen) return(FALSE) + .uscogdata_env$balance_caveats_shown <- c(seen, key) + TRUE +} + +#' Emit at most one message per caveat class per session. +#' @noRd +.emit_balance_caveats <- function(caveats) { + if (.balance_caveat_once("not_gaap")) { + cli::cli_inform(c( + "!" = "Census holdings are gross and are {.strong not} GAAP fund balance.", + "i" = "No liabilities are netted; a reserve ratio built from them overstates available funds." + )) + } + if (length(caveats$truncated) > 0L && + .balance_caveat_once("coverage_window")) { + cli::cli_inform(c( + "!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.", + "i" = "See {.code provenance$balance_caveats$coverage_window}." + )) + } + invisible(NULL) +} +``` + +- [ ] **Step 4: Wire it into `cog_balances()`** + +After `prov <- .build_provenance(...)` in `R/balances.R`: + +```r + prov$balance_caveats <- .balance_caveats( + con, prov$codes_summed$observed, years + ) + .emit_balance_caveats(prov$balance_caveats) +``` + +- [ ] **Step 5: Reset the log in `cog_close()`** + +In `R/session.R`, inside `cog_close()`, alongside the existing teardown: + +```r + .uscogdata_env$balance_caveats_shown <- NULL +``` + +This matters for the tests: `with_fixture_corpus()` calls `cog_close()`, so +each test block starts with a clean message log. + +- [ ] **Step 6: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 7: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +- [ ] **Step 8: Commit** + +```bash +git add R/balance_caveats.R R/balances.R R/session.R tests/testthat/test-balances.R +git commit -m "feat: balance_caveats provenance + once-per-session disclosure (#25)" +``` + +--- + +### Task 6: Documentation + +**Files:** +- Modify: `NEWS.md`, `_pkgdown.yml`, `CLAUDE.md`, `README.md` + +**Interfaces:** +- Consumes: the exported `cog_balances()` from Task 2. +- Produces: no code. + +- [ ] **Step 1: Add the NEWS entry** + +At the top of `NEWS.md`, under the development heading: + +```markdown +* New `cog_balances()` exposes the 14 cash-and-security holding codes + (`category_type = "balance"`): fund balances, retirement system holdings and + insurance trust balances (#25). Holdings are a stock, not a flow, so the verb + has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and + no `subtype` argument either -- for holdings, `category` is a strict + coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly + the `general` family (`W01`/`W31`/`W61`). +* `cog_balances()` results carry `provenance$balance_caveats`, recording that + Census holdings are gross rather than GAAP fund balance, and the measured + coverage window of each subtype family. +``` + +- [ ] **Step 2: Add the verb to `_pkgdown.yml`** + +Add `cog_balances` to the same reference section that lists `cog_spending` and +`cog_revenue`. + +- [ ] **Step 3: Correct the stale claims in `CLAUDE.md`** + +Four statements are wrong. Replace the "SQL lives in `inst/sql/` — never +inline SQL strings in R files" bullet with: + +```markdown +- SQL has two layers. **View definitions** live in `inst/sql/` and are + registered by `.register_views()`, which globs the directory in sorted order + and substitutes `{url}`. **Query construction** is inline `sprintf()` in R + (`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as + a numbered `.sql` file; build a query in R. +``` + +and update, in the same file: +- the `inst/sql/` bullet: 7 view definitions → **23** +- the test count: 181 PASS → the current figure from Step 5 +- the fixture description: "years 2019+2020" → **years 2011, 2012, 2019, 2020** + +Add `cog_balances` to the exported-verbs list. + +- [ ] **Step 4: Verify the doc claims are true** + +```bash +/usr/bin/Rscript -e 'cat("sql views:", length(list.files("inst/sql", pattern="[.]sql$")), "\n")' +/usr/bin/Rscript -e 'suppressMessages(library(arrow)); cat("fixture years:", paste(sort(unique(as.data.frame(open_dataset("inst/extdata/fixture_corpus/data/long") |> dplyr::distinct(year) |> dplyr::collect())$year)), collapse=", "), "\n")' +``` + +Put the actual output in `CLAUDE.md` — do not copy the numbers above on faith. + +- [ ] **Step 5: Run the full suite one last time** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors, pass count above 716. + +- [ ] **Step 6: Commit** + +```bash +git add NEWS.md _pkgdown.yml CLAUDE.md README.md +git commit -m "docs: document cog_balances() and correct stale CLAUDE.md claims (#25)" +``` + +--- + +## Out of scope + +- **`cog-api#26`** — the `/balances` endpoint. Needs all three of: handler, `param_contract`, **and** the `plumber.R` route signature. Missing the third makes the endpoint return 200 while silently ignoring the parameter, with every handler test still passing. +- **cog_pipeline series-break entry** for the FY2016 termination of the seven holdings codes. No row exists at 2016/2017 for `Z77`/`Z78`/`X30`, though `docs/phase_r_harmonization_review.md` § 2 recommended exactly that; `SB197`–`SB202` set the precedent (`coverage_restricted` + `with_caution`). Non-blocking — `coverage_window` covers it reader-side meanwhile. Verify corpus-wide and census-to-census before writing the rows. diff --git a/CLAUDE.md b/CLAUDE.md index d739dd8..5fdb610 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -28,8 +28,23 @@ USCOGDATA_URL (local path or https://) - `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()` - `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache) - `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`) -- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated` +- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order + (`10-` through `46-`): the `*_long` layer (`long`, `spending_long`, + `revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of + `spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer + (`spending_annotated`, `revenue_annotated`, `ig_annotated`, + `balance_annotated`, plus `_harmonized` variants of `spending_annotated`/ + `revenue_annotated`/`ig_annotated`), and metadata views + (`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`, + `harmonization_map`, `harmonization_recipes`, `series_breaks_pq`, + `representation`, `code_set`) - `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()` +- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a + **stock** (a balance at a point in time) rather than a **flow** (activity + over a fiscal year), so it does NOT route through `.verb_spendrev()` and has + no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments. + `R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross + disclosure + measured per-subtype coverage windows). - `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer) - `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()` - `R/search.R` — `cog_gov_search()` (name pattern, state, type filters) @@ -45,28 +60,31 @@ USCOGDATA_URL (local path or https://) Any value without `://` is treated as a local path by `.is_local_path()` and reads `manifest.json` directly from disk (no HTTP, no TTL cache). -## Current State (2026-04-27) +## Current State (2026-08-03) **Version:** 0.1.0 (pre-release) -**Branch:** `main`, commit `d65e9fe` -**Tests:** 181 PASS / 0 FAIL / 0 SKIP +**Branch:** `feat/cog-balances-25`, commit `fde62eb` +**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave) **CI:** Gitea Actions green (`.gitea/workflows/ci.yml`) ### Completed (Tasks 2.1–2.7) -All 8 exported verbs implemented and tested: -`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`, -`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`, -plus `cog_categories`. +All **14** exports implemented and tested (measured from `NAMESPACE`): +`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`, +`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`, +`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`, +`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`. -Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years -2019+2020, all 50 states). Tests run fully offline — no credentials needed. +Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years +2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`, +2026-08-03; all 50 states). Tests run fully offline — no credentials needed. ### Remaining to v0.1 release -1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports; - full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`. - Vignettes can be stubbed for v0.1. +1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`, + `README.md` and `_pkgdown.yml` exist, and `vignettes/` carries + `total-spending.Rmd` + `population-denominators.Rmd`. Outstanding: + `pkgdown::build_site()` has never been run (no `docs/`). 2. **Phase 3 — cog_explorer bridge:** create `cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs @@ -99,6 +117,10 @@ devtools::test() - All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()` - Return value is always a `tbl_df` with a `provenance` attribute - govid inputs always go through `.coerce_govid_input()` (accepts character or data frame) -- SQL lives in `inst/sql/` — never inline SQL strings in R files +- SQL has two layers. **View definitions** live in `inst/sql/` and are + registered by `.register_views()`, which globs the directory in sorted order + and substitutes `{url}`. **Query construction** is inline `sprintf()` in R + (`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as + a numbered `.sql` file; build a query in R. - No arrow dependency — DuckDB reads parquet natively - `withr` is a Suggests-only dep; only used in tests diff --git a/NAMESPACE b/NAMESPACE index 1c67f0b..5d956a0 100644 --- a/NAMESPACE +++ b/NAMESPACE @@ -1,5 +1,6 @@ # Generated by roxygen2: do not edit by hand +export(cog_balances) export(cog_basket_resolution) export(cog_basket_unresolved) export(cog_categories) diff --git a/NEWS.md b/NEWS.md index 97c7430..88e8de2 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,5 +1,18 @@ # uscogdata 0.1.0 (development) +## New: `cog_balances()` for cash-and-security holdings + +* New `cog_balances()` exposes the 14 cash-and-security holding codes + (`category_type = "balance"`): fund balances, retirement system holdings and + insurance trust balances (#25). Holdings are a stock, not a flow, so the verb + has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and + no `subtype` argument either -- for holdings, `category` is a strict + coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly + the `general` family (`W01`/`W31`/`W61`). +* `cog_balances()` results carry `provenance$balance_caveats`, recording that + Census holdings are gross rather than GAAP fund balance, and the measured + coverage window of each subtype family. + ## Multi-government aggregates now disclose their reporting coverage * The Census of Governments is a **complete census only in years ending in 2 diff --git a/R/balance_caveats.R b/R/balance_caveats.R new file mode 100644 index 0000000..dc25401 --- /dev/null +++ b/R/balance_caveats.R @@ -0,0 +1,122 @@ +# R/balance_caveats.R +# +# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and +# security holdings. Each one silently invalidates an obvious analysis, so +# they travel in provenance (machine-readable, for cog-api#26) rather than +# living only in prose. +# +# Two of the four are already carried by the code-driven series-break +# builders and are deliberately NOT duplicated here: +# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via +# series_break_refs on the recipe path, the only path that observes those +# codes. +# What remains is the GAAP distinction (a constant) and the coverage windows +# (measured, never hardcoded, so they stay correct as the corpus grows). + +#' Per-subtype observed year extents, plus which requested families are +#' truncated relative to the requested span. +#' @noRd +.balance_caveats <- function(con, codes_observed, years) { + cw <- .balance_coverage_windows(con) + + observed_subtypes <- if (length(codes_observed) == 0L) { + character(0) + } else { + DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT balance_subtype FROM summary_categories + WHERE item_code IN (%s) AND balance_subtype IS NOT NULL", + .sql_lit_chr(codes_observed) + ))$balance_subtype + } + + # A family is "truncated" when the caller asked for years outside the span + # that family actually covers -- the FY2016 employee-retirement termination + # and the FY2021 end of the W family are both this shape. + truncated <- character(0) + if (length(years) > 0L) { + for (s in observed_subtypes) { + w <- cw[[s]] + if (is.null(w)) next + if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s) + } + } + + list( + not_gaap = TRUE, + not_gaap_note = paste0( + "Census holdings are gross -- no liabilities are netted -- and are NOT ", + "GAAP fund balance. A reserve ratio built from them overstates what is ", + "actually available." + ), + coverage_window = cw, + truncated = sort(unique(truncated)) + ) +} + +#' Per-subtype [min year, max year] extents for EVERY balance subtype in the +#' mounted corpus, memoised for the session. +#' +#' The query carries no govid and no year predicate -- its answer is a property +#' of the mounted corpus alone and cannot change between calls -- but it scans +#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime +#' on the bundled fixture and would be a per-request throughput ceiling once +#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and +#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`. +#' +#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking +#' "is there a family I missed?" needs every window. The observed-scoped field +#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json. +#' @noRd +.balance_coverage_windows <- function(con) { + cached <- .uscogdata_env$balance_coverage_windows + if (!is.null(cached)) return(cached) + + windows <- DBI::dbGetQuery(con, + "SELECT c.balance_subtype AS subtype, + MIN(l.year) AS year_min, + MAX(l.year) AS year_max + FROM balance_long l + JOIN summary_categories c USING (item_code) + WHERE c.balance_subtype IS NOT NULL + GROUP BY 1 + ORDER BY 1" + ) + + cw <- stats::setNames( + lapply(seq_len(nrow(windows)), + function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), + windows$subtype + ) + .uscogdata_env$balance_coverage_windows <- cw + cw +} + +#' TRUE the first time `key` is seen this session, FALSE thereafter. +#' Reset by cog_close(). +#' @noRd +.balance_caveat_once <- function(key) { + seen <- .uscogdata_env$balance_caveats_shown + if (is.null(seen)) seen <- character(0) + if (key %in% seen) return(FALSE) + .uscogdata_env$balance_caveats_shown <- c(seen, key) + TRUE +} + +#' Emit at most one message per caveat class per session. +#' @noRd +.emit_balance_caveats <- function(caveats) { + if (.balance_caveat_once("not_gaap")) { + cli::cli_inform(c( + "!" = "Census holdings are gross and are {.strong not} GAAP fund balance.", + "i" = "No liabilities are netted; a reserve ratio built from them overstates available funds." + )) + } + if (length(caveats$truncated) > 0L && + .balance_caveat_once("coverage_window")) { + cli::cli_inform(c( + "!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.", + "i" = "See {.code provenance$balance_caveats$coverage_window}." + )) + } + invisible(NULL) +} diff --git a/R/balances.R b/R/balances.R new file mode 100644 index 0000000..4a9567d --- /dev/null +++ b/R/balances.R @@ -0,0 +1,159 @@ +# R/balances.R +# +# Cash and security holdings. A third verb rather than an argument on a money +# verb because holdings are a STOCK -- a balance at a point in time -- while +# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money +# verbs' whole argument vocabulary (expenditure_concept, revenue_concept, +# complete=) describes flows and is meaningless here, so this deliberately +# does NOT route through .verb_spendrev(). + +#' Cash and security holdings for one or more governments +#' +#' Returns Census cash-and-security holdings (`category_type = "balance"`): +#' fund balances, retirement system holdings and insurance trust balances. +#' +#' @section Holdings are not GAAP fund balance: +#' Census holdings are **gross** -- no liabilities are netted -- so a reserve +#' ratio built from them overstates what is actually available. They are not +#' comparable to a GAAP fund balance from an ACFR. +#' +#' @param govid Canonical govid(s): a character vector, or a data frame with a +#' `canonical_govid` column (e.g. from [cog_gov_search()]). +#' @param years Integer vector of fiscal years. +#' @param category Optional character vector of categories to keep. One of +#' `"Fund Balances"`, `"Insurance Trust Balances"`, +#' `"Retirement System Holdings"`. There is deliberately no `subtype` +#' argument: for holdings, `category` is a strict coarsening of +#' `balance_subtype` (unlike the money verbs, where the two axes cross), so +#' every combination would be either redundant or empty. +#' `category = "Fund Balances"` is exactly the `general` family +#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +#' one `dplyr::filter()` away. +#' @param per_capita Divide holdings by population. Note this is a **stock per +#' resident** (reserves per person), which is *not* comparable to +#' [cog_spending()]'s per-capita figures -- those are a flow per person. +#' @param adjust_to_year Deflate to this year's dollars (CPI-U). +#' @param basis Accepted for uniformity with the money verbs, but currently a +#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized +#' and raw space are identical for holdings. Reported in +#' `provenance$basis_note`. +#' @param recipe Optional harmonization recipe id (see [cog_recipes()]). +#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +#' wide era to the modern one. +#' +#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, +#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`, +#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and +#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when +#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when +#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no +#' nominal per-capita column to deflate otherwise). Amounts are full US +#' dollars. +#' +#' Carries a `provenance` attribute matching +#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports +#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for +#' every balance subtype in the mounted corpus, not only the observed ones) +#' and `truncated` (the observed subtypes whose coverage falls short of the +#' requested years). `expenditure_concept`/`revenue_concept` are `NA` -- +#' holdings are a stock, not a flow, so neither concept vocabulary applies. +#' @export +cog_balances <- function(govid, years, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, + basis = c("harmonized", "raw"), recipe = NULL) { + call <- match.call() + basis <- match.arg(basis, c("harmonized", "raw")) + # Coerce FIRST, validate second: .validate_verb_inputs() asserts + # is.character(govid), and a data-frame govid (cog_gov_search() output) has + # not been unwrapped yet at this point. + govid <- .coerce_govid_input(govid) + # The money verbs' validator, reused rather than re-implemented (R/spending.R). + # It covers the exact superset cog_balances() needs -- including the + # recipe/category mutual-exclusivity guard -- so a second local copy would + # only be a place for the two to drift apart. This is the same kind of + # helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT + # route the verb through .verb_spendrev(), which stays deliberately unused + # here because its flow vocabulary is meaningless for a stock. + .validate_verb_inputs(govid, years, category, per_capita, adjust_to_year, + recipe) + years <- as.integer(years) + if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) + + con <- .ensure_session() + .require_balance_support(con) + scope <- .check_govids_in_scope(govid) + + basis_note <- paste0( + "`basis` has no effect on holdings: harmonization_map carries no ", + "balance-code rows, so harmonized and raw space are identical here." + ) + + manifest <- .uscogdata_env$manifest + recipe_block <- NULL + category_for_prov <- category + + if (!is.null(recipe)) { + .require_schema_v5(con, manifest, "recipe =") + .validate_recipe_id(con, recipe) + comps <- .recipe_components(con, recipe) + recipe_label <- comps$label[[1]] + result <- .run_recipe(con, recipe, govid, years) + sql <- attr(result, "sql_query") + result <- .shape_recipe_result(result, "balance_subtype", recipe_label) + recipe_block <- list( + recipe_id = recipe, label = recipe_label, + components = .df_to_row_list(comps) + ) + category_for_prov <- recipe_label + } else { + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + } + + # Order matters (matches .verb_spendrev()): per-capita first, so + # .attach_real_dollars() deflates the nominal per-capita column into + # amt_per_capita_real rather than needing amt_per_capita_nominal recomputed. + if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid) + if (!is.null(adjust_to_year)) { + result <- .attach_real_dollars(result, adjust_to_year, per_capita) + } + + prov <- .build_provenance( + verb = "cog_balances", call = call, govid = govid, years = years, + category = category_for_prov, per_capita = per_capita, + adjust_to_year = adjust_to_year, result = result, sql = sql, + subtype_col = "balance_subtype", + basis = basis, basis_note = basis_note, + # Neither concept vocabulary applies to a stock. + expenditure_concept = NA_character_, + revenue_concept = NA_character_, + recipe = recipe_block + ) + prov$scope$govids_found <- scope$found + prov$scope$govids_missing <- scope$missing + + prov$balance_caveats <- .balance_caveats( + con, prov$codes_summed$observed, years + ) + .emit_balance_caveats(prov$balance_caveats) + + attr(result, "provenance") <- prov + result +} + +#' Abort unless the mounted corpus classifies balance codes. +#' +#' `balance_subtype` arrived with cog_pipeline #76/#77 without a +#' schema_version bump, so the check is on the column, not the version. +#' @noRd +.require_balance_support <- function(con) { + if (.corpus_has_balance_subtype(con)) return(invisible(TRUE)) + cli::cli_abort( + c("This corpus does not classify cash and security holdings.", + i = "`summary_categories` has no {.field balance_subtype} column.", + i = "Republish from cog_pipeline at #76/#77 or later."), + class = "uscogdata_no_balance_support" + ) +} diff --git a/R/explain.R b/R/explain.R index 5309e97..37380c4 100644 --- a/R/explain.R +++ b/R/explain.R @@ -68,6 +68,11 @@ cog_explain <- function(result, format = c("print", "list")) { if (!is.null(prov$revenue_concept)) { cli::cli_text("Concept: {prov$revenue_concept} revenue") } + } else if (identical(prov$verb, "cog_balances")) { + # Both concept fields are deliberately NA here (a stock has no flow + # concept). Printing the raw NA reads as a missing value rather than an + # intentional one, so say what it means instead. + cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)") } else if (!is.null(prov$expenditure_concept)) { concept_note <- if (!is.null(prov$expenditure_concept_note) && !is.na(prov$expenditure_concept_note)) { @@ -174,6 +179,31 @@ cog_explain <- function(result, format = c("print", "list")) { cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs)) } + # Balance results only (NULL on money-verb provenance, so they are + # unaffected). This is the ONLY on-demand surface for the GAAP disclosure: + # .emit_balance_caveats() fires at most once per session, and is routinely + # consumed by a suppressMessages() call or by a knitted chunk nobody reads, + # so a caller who deliberately audits a result with cog_explain() must still + # be told. + bc <- prov$balance_caveats + if (!is.null(bc)) { + cli::cli_h2("Holdings caveats") + if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note) + if (length(bc$truncated) > 0L) { + cli::cli_text( + "Requested years extend beyond what these families actually cover:" + ) + cli::cli_ul(vapply(bc$truncated, function(s) { + w <- bc$coverage_window[[s]] + if (length(w) == 2L) { + sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2]) + } else { + s + } + }, character(1))) + } + } + cli::cli_h2("Transformations") uc <- prov$transformations$units_conversion if (isTRUE(uc$applied)) { diff --git a/R/session.R b/R/session.R index 2a1d7d0..ef1c28c 100644 --- a/R/session.R +++ b/R/session.R @@ -95,4 +95,7 @@ cog_close <- function() { } .uscogdata_env$con <- NULL .uscogdata_env$manifest <- NULL + .uscogdata_env$balance_caveats_shown <- NULL + # Memoised corpus-constant; a different corpus may be mounted next. + .uscogdata_env$balance_coverage_windows <- NULL } diff --git a/R/views.R b/R/views.R index 172d6e9..28bf42c 100644 --- a/R/views.R +++ b/R/views.R @@ -45,6 +45,29 @@ "37-code_set.sql" = "code_set.parquet" ) +# Cash and security holdings (uscogdata#25). 46- selects +# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and +# WITHOUT a schema_version bump -- so neither existing gate applies: +# .harmonization_view_files keys on schema_version, .representation_view_files +# on the presence of a FILE. Here the discriminator is a COLUMN on a table +# that exists either way. CREATE VIEW resolves its source schema eagerly, so +# on an older corpus 46- would fail at registration with "Binder Error: +# Referenced column balance_subtype not found" rather than at query time. +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") + +#' Does the mounted corpus's `summary_categories` carry `balance_subtype`? +#' Probed against the live connection rather than the manifest, because the +#' manifest describes files, not columns. +#' @noRd +.corpus_has_balance_subtype <- function(con) { + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM information_schema.columns + WHERE table_name = 'summary_categories' + AND column_name = 'balance_subtype'" + )$n + isTRUE(as.integer(n) > 0L) +} + #' Does the mounted corpus publish `file` (e.g. "code_set.parquet")? #' Reads the manifest's metadata list rather than stat-ing the URL, so it #' works identically for a local fixture and a remote share. @@ -66,6 +89,7 @@ if (base %in% .harmonization_view_files && schema_version < 5L) next if (base %in% names(.representation_view_files) && !.corpus_has_table(manifest, .representation_view_files[[base]])) next + if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next sql <- paste(readLines(f, warn = FALSE), collapse = "\n") sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) DBI::dbExecute(con, sql) diff --git a/_pkgdown.yml b/_pkgdown.yml index 2022548..1541211 100644 --- a/_pkgdown.yml +++ b/_pkgdown.yml @@ -3,6 +3,12 @@ template: bootstrap: 5 reference: + - title: Financial data + desc: Spending, revenue and balance-sheet holdings for one or more governments. + contents: + - cog_spending + - cog_revenue + - cog_balances - title: Search & basket desc: Resolve place names into canonical govids. contents: diff --git a/inst/schemas/provenance-v1.json b/inst/schemas/provenance-v1.json index 779fad6..b7e942b 100644 --- a/inst/schemas/provenance-v1.json +++ b/inst/schemas/provenance-v1.json @@ -53,6 +53,16 @@ "items": { "type": "string" }, "description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series." }, + "balance_caveats": { + "type": ["object", "null"], + "description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.", + "properties": { + "not_gaap": { "type": "boolean" }, + "not_gaap_note": { "type": "string" }, + "coverage_window": { "type": "object" }, + "truncated": { "type": "array", "items": { "type": "string" } } + } + }, "manifest": { "type": "object" }, "sql_query": { "type": "string" } } diff --git a/inst/sql/26-balance_long.sql b/inst/sql/26-balance_long.sql new file mode 100644 index 0000000..22095e8 --- /dev/null +++ b/inst/sql/26-balance_long.sql @@ -0,0 +1,22 @@ +-- Cash and security holdings, classified by crosswalk MEMBERSHIP on +-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot +-- do this job -- the X and Y families each span revenue, expenditure AND +-- balance). +-- +-- These rows are STOCKS: a balance at a point in time, not a flow over a +-- fiscal year. Summing a stock with a flow is meaningless, which is why they +-- live behind a third view rather than as a subtype of either money view, and +-- why neither spending_long nor revenue_long can reach them. +-- +-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era +-- aggregate-only holdings codes (X40/X41) are deliberately outside this view; +-- they are reachable only through the recipe path, which bypasses this filter +-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2). +CREATE OR REPLACE VIEW balance_long AS +SELECT * +FROM long +WHERE item_code IN ( + SELECT item_code FROM summary_categories + WHERE category_type = 'balance' + ) + AND NOT is_aggregate; diff --git a/inst/sql/46-balance_annotated.sql b/inst/sql/46-balance_annotated.sql new file mode 100644 index 0000000..8a40a91 --- /dev/null +++ b/inst/sql/46-balance_annotated.sql @@ -0,0 +1,16 @@ +CREATE OR REPLACE VIEW balance_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.balance_subtype +FROM balance_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); diff --git a/man/cog_balances.Rd b/man/cog_balances.Rd new file mode 100644 index 0000000..f594eba --- /dev/null +++ b/man/cog_balances.Rd @@ -0,0 +1,76 @@ +% Generated by roxygen2: do not edit by hand +% Please edit documentation in R/balances.R +\name{cog_balances} +\alias{cog_balances} +\title{Cash and security holdings for one or more governments} +\usage{ +cog_balances( + govid, + years, + category = NULL, + per_capita = FALSE, + adjust_to_year = NULL, + basis = c("harmonized", "raw"), + recipe = NULL +) +} +\arguments{ +\item{govid}{Canonical govid(s): a character vector, or a data frame with a +`canonical_govid` column (e.g. from [cog_gov_search()]).} + +\item{years}{Integer vector of fiscal years.} + +\item{category}{Optional character vector of categories to keep. One of +`"Fund Balances"`, `"Insurance Trust Balances"`, +`"Retirement System Holdings"`. There is deliberately no `subtype` +argument: for holdings, `category` is a strict coarsening of +`balance_subtype` (unlike the money verbs, where the two axes cross), so +every combination would be either redundant or empty. +`category = "Fund Balances"` is exactly the `general` family +(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +one `dplyr::filter()` away.} + +\item{per_capita}{Divide holdings by population. Note this is a **stock per +resident** (reserves per person), which is *not* comparable to +[cog_spending()]'s per-capita figures -- those are a flow per person.} + +\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).} + +\item{basis}{Accepted for uniformity with the money verbs, but currently a +**no-op**: `harmonization_map` carries no balance-code rows, so harmonized +and raw space are identical for holdings. Reported in +`provenance$basis_note`.} + +\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]). +`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +wide era to the modern one.} +} +\value{ +Tibble with columns `year`, `canonical_govid`, `gov_name`, + `balance_subtype`, `category`, `amt_nominal`, `codes_included`, + `aggregate_fallback`, plus optional `amt_per_capita_nominal` and + `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when + `adjust_to_year` is set), and optional `amt_per_capita_real` (only when + **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no + nominal per-capita column to deflate otherwise). Amounts are full US + dollars. + + Carries a `provenance` attribute matching + `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports + `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for + every balance subtype in the mounted corpus, not only the observed ones) + and `truncated` (the observed subtypes whose coverage falls short of the + requested years). `expenditure_concept`/`revenue_concept` are `NA` -- + holdings are a stock, not a flow, so neither concept vocabulary applies. +} +\description{ +Returns Census cash-and-security holdings (`category_type = "balance"`): +fund balances, retirement system holdings and insurance trust balances. +} +\section{Holdings are not GAAP fund balance}{ + +Census holdings are **gross** -- no liabilities are netted -- so a reserve +ratio built from them overstates what is actually available. They are not +comparable to a GAAP fund balance from an ACFR. +} + diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md new file mode 100644 index 0000000..2f97945 --- /dev/null +++ b/specs/2026-08-03-cog-balances-design.md @@ -0,0 +1,281 @@ +# `cog_balances()` — a reader surface for cash and security holdings + +**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26` +**Date:** 2026-08-03 · **Status:** design, awaiting approval + +Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped +with `#11`/`#12` and is asserted at both view and verb level. This spec covers +requirement 2 only: a way to query holdings. + +## Decision: a verb, not an argument + +`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`. + +Holdings are a **stock** — a balance at a point in time — while the money verbs +return **flows** over a fiscal year. The flow verbs' whole argument vocabulary +is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe +which flows Census aggregates into a published total, and `complete=` fills a +grid of fiscal-year cells. Overloading a money verb would put a stock behind +arguments that all assume a flow. + +## The 14 codes + +Measured against the published corpus 2026-08-03, not transcribed from the +issue. `year_min`/`year_max` are observed row extents. + +| `balance_subtype` | `category` | codes | observed years | +|---|---|---|---| +| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 | +| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 | +| | | `X47` | 1988–2016 | +| | | `X30`, `Z77`, `Z78` | 2012–2016 | +| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 | +| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 | +| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 | + +## Architecture + +### Two new views + +Mirroring the `revenue_long` / `revenue_annotated` pair exactly: + +- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate` +- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and + `summary_categories`, exposing `category`, `category_type`, `balance_subtype` + +`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register +with no new registration code. + +### A third gate list in `R/views.R` + +`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails +at registration time, not at query time. `46-balance_annotated.sql` selects +`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`. +That arrived without a `schema_version` bump, so neither existing gate applies: +`.harmonization_view_files` keys on `schema_version`, `.representation_view_files` +on the presence of a *file*. The discriminator here is a **column on an existing +table**. + +```r +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") +``` + +gated by probing `summary_categories` for `balance_subtype`, with +`cog_balances()` erroring cleanly via `.require_balance_support()` on an older +corpus — mirroring how `.require_schema_v5()` gates the harmonized views. + +### `R/balances.R` — a dedicated path, not `.verb_spendrev()` + +`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and +`complete=` grid are all flow-specific, and four verbs depend on it. Threading a +third mode through it adds branching to shared code for no reuse benefit. + +Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`, +`.build_corpus_break_refs()`, the population join, `.inflate()`, and +`.coerce_govid_input()`. + +Following the package's real two-layer convention: **view definitions** live in +`inst/sql/`; **query construction** is inline `sprintf()` in R, as in +`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R +files", which the verb layer has never obeyed. Corrected in a separate commit — +see Out of scope.) + +## Signature + +```r +cog_balances(govid, years, + category = NULL, # Fund Balances | Insurance Trust Balances | + # Retirement System Holdings + per_capita = FALSE, + adjust_to_year = NULL, + basis = c("harmonized", "raw"), + recipe = NULL) +``` + +Returns a `tbl_df` with a `provenance` attribute, like every other verb. + +**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`, +and `subtype` — see below. + +**`per_capita` is offered.** Holdings per resident is a real measure (pension +assets per capita, fund balance per resident). The roxygen `@param` states +plainly that this is a *stock per resident* and is **not** comparable to +`cog_spending()`'s per-capita figures. + +**`basis` is currently a no-op** — `harmonization_map` has zero balance-code +rows, so harmonized and raw are identical for holdings. Kept for uniformity +with the money verbs (the API would otherwise special-case), and +`provenance$basis_note` says so outright rather than letting it look meaningful. + +**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era +to the modern one: + +``` +cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023) +cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023) +``` + +`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they +are invisible to `balance_long`, which filters `NOT is_aggregate` like every +other basis view. That is by design, not a defect: +`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide +era exposes these split families *only* as aggregates, and that the recipe join +must therefore **not** filter `is_aggregate` — safe by construction, because +wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and +every component is year-scoped, so no double-count is possible. § 1 records the +matching decision that the planned `X40→Z77` harmonization *map* rows were +dropped and the continuity ships as recipes instead, which is why +`harmonization_map` has no balance-code rows. + +The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is +verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose +wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the +live corpus. So a recipe query reaches rows the verb's own view cannot, exactly +as intended. + +### No `subtype` argument: `category` is a strict coarsening + +`balance` is the only `category_type` in which `category` and the subtype column +are **not** orthogonal. Measured against the published crosswalk: + +| `category_type` | subtypes spanning more than one category | +|---|---| +| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) | +| revenue | 1 of 7 (`own_source`) | +| **balance** | **0 of 5** | + +For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire) +× *economic character* (operations, capital) — so both earn their place. For +balance the relation is a strict tree: + +``` +Fund Balances = {general} W01 W31 W61 +Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78 +Insurance Trust Balances = {unemployment_trust, + workers_comp_trust, + other_insurance_trust} Y07 Y08 Y21 Y61 +``` + +Exposing both would therefore admit no useful combination. Of the 15 possible +pairs, 3 are redundant (the subtype already implies its category) and **12 are +guaranteed empty for every government in every year** — and an impossible query +would fail by returning an empty tibble, which reads as "this government holds +none" rather than "you asked a contradiction." + +Dropping `subtype` also keeps the verb aligned with the rest of the package: no +uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in +`.verb_spendrev()`, and the API layers its own `subtype` row filter on top +(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for +`/balances`. + +`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the +`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is +isolating one of the three insurance funds in a single argument; +`balance_subtype` remains a returned column, so that is one `dplyr::filter()` +away. + +## Caveat surfacing + +`provenance$balance_caveats`, always present, plus one `cli_inform()` per +session per caveat class when a query actually touches an affected family or +year. Structured so `cog-api#26` can forward the fields verbatim. + +Verified against `series_breaks.csv`, not assumed: + +| # | Caveat | Covered by existing machinery? | +|---|---|---| +| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` | +| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus | +| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) | +| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed | + +On caveat 4's second condition: `.build_series_break_refs()` matches +`break_year BETWEEN min(years) AND max(years)`, so a request spanning only +2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a +series sits entirely after the change, on one consistent basis, and flagging a +break it never crosses would be noise. The same rule is applied deliberately in +`.build_corpus_break_refs()`. An earlier draft of this row omitted the span +condition and overclaimed. + +`coverage_window` is derived per observed subtype family from the corpus, never +hardcoded, so it stays correct as the corpus grows. + +`series_break_refs` and `corpus_break_refs` are otherwise populated by the +existing code-driven builders and need no change. + +## Testing + +New `tests/testthat/test-balances.R`. The bundled fixture covers all four +fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y` +throughout — so every test below runs offline. + +- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing + the already-asserted forward guard. Absence is verified against the raw corpus + via `read_parquet` on `data/long`, never through the verb that creates it. +- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019; + `coverage_window` reports the termination and the console message fires once. +- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the + measured table above; the FY2002 valuation caveat fires only when the year + range crosses 2002 *and* touches `employee_retirement`. +- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`. +- **`category = "Fund Balances"` is the `general` family.** Returns exactly + `W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted + rather than assumed. +- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to + exactly one `category`. Asserted against the crosswalk so that an upstream + change breaking the tree — which would silently make `category` lossy — + fails here rather than in a user's analysis. +- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the + `X40` leg for a pre-2012 year, proving the aggregate-only wide rows are + reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A + regression here would silently truncate a 45-year series to five. +- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning + FY2002 carries both in `provenance$series_break_refs`, so the book → market + basis change is disclosed wherever `X40`/`X41` are actually observed. +- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose + `summary_categories` lacks `balance_subtype`. + +## Out of scope, tracked separately + +1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of + the seven holdings codes in `series_breaks.csv`. There is currently **no** + entry at 2016/2017 for `Z77`/`Z78`/`X30`, although + `docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended + exactly this — *"candidate new `series_breaks.csv` entries (recommend + `with_caution` documentation rows, no map action)"*. The follow-through never + happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow + codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also + what keeps this out of the `joinable = "no"` identity-change rule, which + would otherwise oblige a harmonization-map row. + + Verify the break corpus-wide and census-to-census before writing the rows. + `cog_balances()` does not wait on this — caveat 3 is covered reader-side by + `coverage_window` meanwhile, and the entry simply adds a second, catalogued + signpost when it lands. + + **Superseded:** an earlier draft of this spec proposed adding + `summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked. + Both were wrong. `X40`/`X41` are deliberately aggregate-only per + `phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows + are the documented § 1 decision, and the recipe path reaches them by design. +2. **`cog-api#26`.** Adds `/balances` in all three required places — handler, + `param_contract`, and the `plumber.R` route signature. Lands after this. + + **Two contract facts the API must carry forward**, both settled during + implementation and easy to get wrong from the outside: + + - `provenance$balance_caveats$coverage_window` is **corpus-scoped, not + result-scoped**. It reports the observed year extent of *every* balance + subtype in the corpus, not only the subtypes a given query returned — so a + `category = "Fund Balances"` query still returns all five windows. That is + deliberate: the windows describe what the corpus holds, which is what a + consumer needs in order to know what it did *not* ask for. The sibling + field `truncated` is the result-scoped one. Documented in + `inst/schemas/provenance-v1.json` and mutation-guarded against silent + inversion. + - `balance_caveats` appears **only** on `cog_balances()` results. It is + absent from `cog_spending()`/`cog_revenue()` provenance, and the schema + says so — an API layer that assumes it is universal will read `NULL`. +3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7 + SQL views (there are 21), 181 tests (716), a two-year fixture (four years), + and a "never inline SQL" rule the verb layer does not follow. diff --git a/tests/testthat/helper-fixture.R b/tests/testthat/helper-fixture.R index 3ea3744..9d23512 100644 --- a/tests/testthat/helper-fixture.R +++ b/tests/testthat/helper-fixture.R @@ -154,3 +154,36 @@ with_corpus_missing_ig_categories <- function(code) { }, add = TRUE) force(code) } + +# Copy the bundled fixture to a temp dir with summary_categories.parquet +# rewritten to DROP the balance_subtype column, then run `code` against it. +# Models a corpus published before cog_pipeline #76/#77. schema_version is +# left untouched deliberately: that change shipped without a version bump, so +# column presence is the only honest signal -- this helper is what proves the +# package keys off it. Mirrors with_corpus_missing_ig_categories(). +with_corpus_missing_balance_subtype <- function(code) { + src <- fixture_corpus_path() + tmp <- withr::local_tempdir(.local_envir = parent.frame()) + file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf( + "COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s)) + TO %s (FORMAT PARQUET)", + uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) + )) + file.remove(cats_path) + file.rename(filtered_path, cats_path) + + old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) + uscogdata:::cog_close() + Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) + on.exit({ + uscogdata:::cog_close() + if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) + }, add = TRUE) + force(code) +} diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R new file mode 100644 index 0000000..86d5ba4 --- /dev/null +++ b/tests/testthat/test-balances.R @@ -0,0 +1,539 @@ +test_that("balance views register and carry only balance codes", { + skip_if_no_corpus() + con <- cog_open() + on.exit(cog_close()) + + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + expect_true(all(c("balance_long", "balance_annotated") %in% views)) + + # Every item_code in balance_long is a category_type = 'balance' member. + leak <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long + WHERE item_code NOT IN ( + SELECT item_code FROM summary_categories WHERE category_type = 'balance')" + )$n + expect_identical(as.integer(leak), 0L) + + # balance_annotated exposes the subtype column the verb groups on. + cols <- DBI::dbGetQuery(con, + "SELECT column_name FROM information_schema.columns + WHERE table_name = 'balance_annotated'" + )$column_name + expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) +}) + +test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", { + # Every category_type = 'balance' item_code in the bundled fixture has + # is_aggregate = FALSE for every row of every year -- there is no real row + # that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An + # assertion against the live fixture (`WHERE is_aggregate` returns 0) is + # therefore vacuous: it passes identically whether or not the view's + # predicate is present. As with the 22-/23- and 24-/25- tests above, this + # reads the real inst/sql/26-balance_long.sql text off disk and executes it + # -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against + # a synthetic hive-partitioned parquet tree that DOES contain an aggregate + # row under a real balance item_code (W01), so a regression that drops the + # predicate changes which rows survive. + skip_if_no_corpus() + + tmp <- withr::local_tempdir() + part_dir <- file.path(tmp, "data", "long", "year=2004") + dir.create(part_dir, recursive = TRUE) + part_path <- file.path(part_dir, "part-0.parquet") + + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives + ('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate` + ) AS t(canonical_govid, item_code, amt, is_aggregate) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(part_path))) + + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('W01', 'Fund Balances', 'balance', NULL, NULL, 'general') + ) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet")))) + + sql_dir <- system.file("sql", package = "uscogdata") + .read_view_sql <- function(filename) { + txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n") + gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE) + } + + con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(con, .read_view_sql("10-long.sql")) + DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql")) + DBI::dbExecute(con, .read_view_sql("26-balance_long.sql")) + + rows <- DBI::dbGetQuery(con, + "SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid" + ) + expect_equal(nrow(rows), 1L) + expect_equal(rows$canonical_govid, "bal-A") + expect_equal(rows$amt, 100) +}) + +test_that("balance views are skipped on a corpus without balance_subtype", { + skip_if_no_corpus() + with_corpus_missing_balance_subtype({ + con <- cog_open() + on.exit(cog_close()) + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + # Registration must SKIP them, not error -- an older corpus stays usable. + expect_false(any(c("balance_long", "balance_annotated") %in% views)) + expect_true("revenue_long" %in% views) + + # ...and calling the verb on such a corpus must hit + # .require_balance_support()'s curated abort (spec § Testing: "Gating"), + # not a DuckDB binder error naming a view that was never registered. + # Asserted on the CLASS: removing the guard still errors, so a bare + # expect_error() would pass on the regression. + expect_error( + cog_balances("550000227544", 2019), + class = "uscogdata_no_balance_support" + ) + }) +}) + +test_that("cog_balances returns holdings for a government that has them", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + expect_s3_class(r, "tbl_df") + expect_true(nrow(r) > 0L) + expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", + "category", "amt_nominal") %in% names(r))) + expect_identical(sort(unique(r$category)), + c("Fund Balances", "Insurance Trust Balances")) + expect_false(is.null(attr(r, "provenance"))) + expect_identical(attr(r, "provenance")$verb, "cog_balances") + }) +}) + +test_that('category = "Fund Balances" is exactly the general family', { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + expect_identical(codes, c("W01", "W31", "W61")) + }) +}) + +test_that("no flow code can reach cog_balances", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + + # The expected set is read from the RAW corpus, never from the verb -- + # verifying an absence through the filter that creates it proves nothing. + # A fresh, direct DuckDB connection against the raw parquet files (never + # cog_open()'s session, never balance_long/balance_annotated) reads + # parquet natively -- no arrow dependency needed (see CLAUDE.md). + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + balance_codes <- DBI::dbGetQuery(con2, sprintf( + "SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'", + uscogdata:::.sql_lit_chr(cats_path) + ))$item_code + + expect_true(length(got) > 0L) + expect_true(all(got %in% balance_codes)) + }) +}) + +test_that("every balance_subtype maps to exactly one category", { + skip_if_no_corpus() + # Dropping the `subtype` argument is only safe while this tree holds. If the + # pipeline ever gives a balance subtype a second category, `category` becomes + # a lossy filter -- fail HERE rather than in a user's analysis. Read via a + # fresh direct DuckDB connection against the raw parquet file, not through + # any registered view. + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + b <- DBI::dbGetQuery(con2, sprintf( + "SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'", + uscogdata:::.sql_lit_chr(cats_path) + )) + per_subtype <- tapply(b$category, b$balance_subtype, + function(x) length(unique(x))) + expect_true(all(per_subtype == 1L)) +}) + +test_that("cog_balances records found + missing govids in provenance", { + skip_if_no_corpus() + with_fixture_corpus({ + suppressMessages( + r <- cog_balances(c("550000227544", "XXXINVALID"), 2019) + ) + prov <- attr(r, "provenance") + expect_equal(sort(prov$scope$govids_found), "550000227544") + expect_equal(sort(prov$scope$govids_missing), "XXXINVALID") + }) +}) + +test_that("per_capita divides holdings by population", { + skip_if_no_corpus() + with_fixture_corpus({ + plain <- cog_balances("550000227544", 2019, category = "Fund Balances") + pc <- cog_balances("550000227544", 2019, category = "Fund Balances", + per_capita = TRUE) + expect_true("amt_per_capita_nominal" %in% names(pc)) + expect_true("pop_source" %in% names(pc)) + expect_identical(pc$amt_nominal, plain$amt_nominal) + + # Assert against the denominator read from the corpus, NOT against a + # quantity derived from amt_per_capita_nominal itself -- dividing the + # column back out would be tautological and would pass on any value. + pop <- DBI::dbGetQuery(cog_open(), sprintf( + "SELECT population FROM gov_population_yearly + WHERE canonical_govid = %s AND year = 2019", + uscogdata:::.sql_lit_chr("550000227544") + ))$population + expect_length(pop, 1L) + expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, + tolerance = 1e-8) + + prov <- attr(pc, "provenance") + expect_true(prov$transformations$per_capita$applied) + }) +}) + +test_that("adjust_to_year adds real dollars", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r)) + # 2012 dollars inflated to 2020 must exceed nominal. + expect_true(all(r$amt_real > r$amt_nominal)) + prov <- attr(r, "provenance") + expect_true(prov$transformations$inflation$applied) + expect_identical(prov$transformations$inflation$base_year, 2020L) + }) +}) + +test_that("per_capita and adjust_to_year compose", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2012, category = "Fund Balances", + per_capita = TRUE, adjust_to_year = 2020) + expect_true("amt_per_capita_real" %in% names(r)) + # The per-capita column must be deflated by the SAME factor as the level + # column -- this is what the ordering at R/balances.R:101-103 guarantees. + # .attach_real_dollars() silently no-ops on the per-capita leg when + # amt_per_capita_nominal does not exist yet (R/spending.R:664), so + # reversing those two calls drops this column with no error at all. + expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal, + r$amt_real / r$amt_nominal, tolerance = 1e-8) + + # And the documented condition is a conjunction: adjust_to_year ALONE + # must not produce amt_per_capita_real (pins the @return wording). + r2 <- cog_balances("550000227544", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r2)) + expect_false("amt_per_capita_real" %in% names(r2)) + }) +}) + +# --- input validation ------------------------------------------------------ + +test_that("cog_balances validates its inputs like the money verbs", { + skip_if_no_corpus() + with_fixture_corpus({ + G <- "550000227544" + # Pinned to the message, not bare expect_error(): every one of these + # already produces *some* error or *some* quiet wrong answer today -- + # years = integer(0) leaks `Parser Error ... AND year IN ()` with the + # generated SQL, recipe = c("a","b") throws "the condition has length > 1", + # and the govid/category cases return 0 rows with no error at all. + expect_error(cog_balances(G, integer(0)), "non-empty integer vector") + expect_error(cog_balances(character(0), 2019), "non-empty character vector") + expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL") + expect_error(cog_balances(G, 2019, recipe = c("a", "b")), + "length-1 character string") + }) +}) + +test_that("validation runs after govid coercion, so a data frame still works", { + skip_if_no_corpus() + with_fixture_corpus({ + # .validate_verb_inputs() asserts is.character(govid); it must therefore + # run AFTER .coerce_govid_input(), never before, or the documented + # data-frame input (cog_gov_search() output) would abort. + df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE) + r <- suppressMessages(cog_balances(df, 2019)) + expect_true(nrow(r) > 0L) + expect_identical(unique(r$canonical_govid), "550000227544") + }) +}) + +test_that("recipe and category are mutually exclusive", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_error( + cog_balances("550000227544", c(2011, 2012), + category = "Fund Balances", + recipe = "cash_securities_z77_wide"), + class = "uscogdata_recipe_category_conflict" + ) + }) +}) + +# --- recipe = : the wide-era holdings bridge ------------------------------- + +test_that("recipe bridges the wide era into the modern one", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z77_wide") + # .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait, + # not specific to this recipe -- see the money-verb recipe tests, which + # only ever assert on it with expect_equal), so compare numerically rather + # than with expect_identical()'s type-strict comparison. + expect_equal(sort(r$year), c(2011, 2012)) + + # The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE + # and therefore invisible to balance_long. If the recipe path ever starts + # filtering aggregates, a 45-year series silently truncates to five -- + # this is the regression guard for phase_r_harmonization_review.md § 0.2. + codes <- attr(r, "provenance")$codes_summed$observed + expect_true("X40" %in% codes) + expect_true("Z77" %in% codes) + expect_true(all(r$amt_nominal > 0)) + + prov <- attr(r, "provenance") + expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide") + }) +}) + +test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { + skip_if_no_corpus() + with_fixture_corpus({ + # 2002 is in the year vector deliberately, and must stay -- do not + # "simplify" this back to c(2011, 2012). + # + # .build_series_break_refs() (R/series_breaks.R, shared with every verb) + # matches breaks with `break_year BETWEEN min(years) AND max(years)`, and + # SB195's break_year is 2002. A c(2011, 2012) span never crosses the + # FY2002 book -> market change -- that whole span sits after it, on one + # consistent basis -- so NOT disclosing SB195 there is correct behaviour, + # not a gap (same reasoning as the "a request that never crosses the + # boundary is not affected by it" comment on .build_corpus_break_refs()). + # + # The property actually worth testing is: a recipe query that observes + # X40 AND spans FY2002 discloses SB195. This fixture has no 2002 + # partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020 + # partitions exist), so including 2002 in `years` widens the + # break-matching window without changing which rows the recipe join + # returns -- verified empirically: r$year below is exactly {2011, 2012} + # whether or not 2002 is in the request (see task-4-report.md). + # Removing 2002 would silently turn this back into the non-crossing case + # above and destroy the test's purpose. + r <- cog_balances("550000227544", c(2002, 2011, 2012), + recipe = "cash_securities_z77_wide") + expect_equal(sort(r$year), c(2011, 2012)) + refs <- attr(r, "provenance")$series_break_refs + # SB195 sits on fin_code X40; it can only fire where X40 is observed, + # which is exactly the recipe path. + expect_true("SB195" %in% refs) + }) +}) + +test_that("the second holdings bridge works too", { + skip_if_no_corpus() + with_fixture_corpus({ + # X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011 + # and Z78 in 2012, so both legs are exercised. + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z78_wide") + codes <- attr(r, "provenance")$codes_summed$observed + expect_true(all(c("X41", "Z78") %in% codes)) + expect_equal(sort(r$year), c(2011, 2012)) + }) +}) + +test_that("an unknown recipe id is rejected", { + skip_if_no_corpus() + with_fixture_corpus({ + # Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84). + # Without it the test is non-discriminating: deleting the validation call + # leaves .recipe_components() returning 0 rows and comps$label[[1]] + # throwing "subscript out of bounds", which a bare expect_error() accepts + # while the user loses the curated "valid recipe ids are ..." message. + expect_error( + cog_balances("550000227544", 2019, recipe = "no_such_recipe"), + class = "uscogdata_unknown_recipe" + ) + }) +}) + +# --- balance_caveats: GAAP disclosure + measured coverage windows ---------- + +test_that("balance_caveats is always present and flags the GAAP distinction", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + cav <- attr(r, "provenance")$balance_caveats + expect_false(is.null(cav)) + expect_true(cav$not_gaap) + }) +}) + +test_that("coverage_window is computed from the corpus, not hardcoded", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + cav <- attr(r, "provenance")$balance_caveats + + # Read the "general" family's true year extent independently, via a + # fresh DuckDB connection against the raw parquet files (never through + # balance_long/.balance_caveats() itself, and never via arrow -- this + # package reads parquet through DuckDB only, see CLAUDE.md). Replicates + # the same predicates 26-balance_long.sql applies (category_type = + # 'balance', NOT is_aggregate) so this is a faithful, independent + # measurement rather than a re-statement of the view under test. + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet") + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + obs <- DBI::dbGetQuery(con2, sprintf( + "SELECT MIN(l.year) AS y0, MAX(l.year) AS y1 + FROM read_parquet(%s, hive_partitioning = true) l + JOIN read_parquet(%s) c USING (item_code) + WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate", + uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path) + )) + + expect_identical(as.integer(cav$coverage_window$general), + c(as.integer(obs$y0), as.integer(obs$y1))) + }) +}) + +test_that("coverage_window covers every corpus subtype, not just observed ones", { + skip_if_no_corpus() + with_fixture_corpus({ + # Deliberate contract (provenance-v1.json): the window block is corpus- + # scoped so a caller can ask "is there a family I missed?", while + # `truncated` is the observed-scoped field. A single-category query must + # therefore still report every balance family in the mounted corpus. + r <- cog_balances("550000227544", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + all_subtypes <- DBI::dbGetQuery(con2, sprintf( + "SELECT DISTINCT balance_subtype FROM read_parquet(%s) + WHERE balance_subtype IS NOT NULL", + uscogdata:::.sql_lit_chr(cats_path) + ))$balance_subtype + + cav <- attr(r, "provenance")$balance_caveats + expect_setequal(names(cav$coverage_window), all_subtypes) + expect_true(length(all_subtypes) > 1L) + # ...while `truncated` stays scoped to what this query actually observed. + expect_true(all(cav$truncated %in% unique(r$balance_subtype))) + }) +}) + +test_that("the corpus-constant coverage windows are memoised per session", { + skip_if_no_corpus() + with_fixture_corpus({ + # The windows query has no govid/year predicate: its answer depends only + # on which corpus is mounted, so re-running the full balance_long scan on + # every call is pure waste (35% of verb runtime on the fixture). Same + # memoise-and-invalidate pattern as .uscogdata_env$manifest. + expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows) + suppressMessages(cog_balances("550000227544", 2019)) + memo <- uscogdata:::.uscogdata_env$balance_coverage_windows + expect_false(is.null(memo)) + expect_true("general" %in% names(memo)) + + uscogdata:::cog_close() + expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows) + }) +}) + +test_that("a request past a family's coverage window is flagged", { + skip_if_no_corpus() + with_fixture_corpus({ + # employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in + # the LIVE corpus -- Census moved employee retirement reporting to the + # Annual Survey of Public Pensions after that year. This bundled FIXTURE + # doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present), + # so the family's *observed* max here is 2012, not 2016. Either way the + # requested span (2012, 2019) reaches past what the family covers in + # THIS corpus, which is what makes .balance_caveats() flag it -- the + # assertion below is about the fixture's measured window, not the FY2016 + # live-corpus cutoff. + r <- cog_balances("550000227544", c(2012, 2019)) + cav <- attr(r, "provenance")$balance_caveats + expect_true("employee_retirement" %in% cav$truncated) + }) +}) + +test_that("the provenance schema documents balance_caveats", { + sch <- jsonlite::fromJSON( + system.file("schemas", "provenance-v1.json", package = "uscogdata"), + simplifyVector = FALSE + ) + expect_true("balance_caveats" %in% names(sch$properties)) +}) + +test_that("cog_explain surfaces the balance caveats", { + skip_if_no_corpus() + with_fixture_corpus({ + # Asserted on the RENDERED text, not on prov$balance_caveats: the field + # is already covered above, and the once-per-session cli_inform() means + # cog_explain() is the only surface a caller who missed (or suppressed) + # the first message can still audit. + r <- suppressMessages(cog_balances("550000227544", c(2012, 2019))) + # Both streams: cli routes most of its output through conditions that + # land on stderr, so a stdout-only capture would be empty (the pattern + # used throughout test-explain.R). + out <- paste(c(capture.output(cog_explain(r)), + capture.output(cog_explain(r), type = "message")), + collapse = "\n") + expect_match(out, "GAAP") + expect_match(out, "employee_retirement") + }) +}) + +test_that("cog_explain on a money-verb result has no balance caveat section", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- suppressMessages(cog_spending("550000227544", 2019)) + out <- paste(c(capture.output(cog_explain(r)), + capture.output(cog_explain(r), type = "message")), + collapse = "\n") + # Guard against the capture itself being vacuous: the section must be + # absent from output that demonstrably contains the rest of the report. + expect_match(out, "Data vintage") + expect_false(grepl("GAAP", out)) + }) +}) + +test_that("the caveat message fires once per session", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_message(cog_balances("550000227544", 2019), "not.*GAAP") + expect_no_message(cog_balances("550000227544", 2020)) + }) +})