# cog_balances() Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Add `cog_balances()`, a third verb exposing the 14 cash-and-security holding codes (`category_type = 'balance'`), closing `uscogdata#25` requirement 2. **Architecture:** Two new DuckDB views (`balance_long`, `balance_annotated`) mirroring the `revenue_long`/`revenue_annotated` pair, registered by the existing `.register_views()` glob behind a new column-presence gate. A dedicated lean verb in `R/balances.R` reuses the shared SQL builder, per-capita join, inflation, recipe runner and provenance builder — it does **not** go through `.verb_spendrev()`, whose concept scoping, intergovernmental leg and `complete=` grid are all flow-specific. **Tech Stack:** R, DuckDB via DBI, testthat 3e, roxygen2, cli, tibble, dplyr. **Spec:** `specs/2026-08-03-cog-balances-design.md` ## Global Constraints - Run R as `/usr/bin/Rscript`. **Never a conda R** — conda shadows `libPaths` and breaks `arrow`/`duckdb`. - The test suite runs offline against the bundled fixture; `tests/testthat/setup.R` sets `USCOGDATA_URL` automatically. To run against the live corpus, `USCOGDATA_URL=/home/jared/Nextcloud/Civilytics/SHARE/uscogdata-corpus/` — **trailing slash required**. - Full suite: `/usr/bin/Rscript -e 'testthat::test_local(".", reporter="silent")'`. Current floor: **716 pass / 0 fail / 0 skip**. Never let it drop. - Corpus `amt` is in **thousands**; every verb multiplies by `1000.0` in SQL and returns **full US dollars**. Do not apply the conversion twice. - SQL **view definitions** live in `inst/sql/`; **query construction** is inline `sprintf()` in R. Follow both — `CLAUDE.md`'s "never inline SQL" line is stale and Task 6 fixes it. - Every verb returns a `tbl_df` with a `provenance` attribute. - Commits are GPG-signed. If signing times out, retry once — pinentry is non-interactive here. - **Never verify an absence through the filter that creates it.** Absence assertions read the raw corpus via `arrow::open_dataset()`, never through `cog_balances()`. ## The test government `010000226085` — ALABAMA STATE GOVT. One government covers every case: | year | codes present | |---|---| | 2011 | `X21`, **`X40`**, `Y07`, `Y08` | | 2012 | `W01`, `W31`, `W61`, `X21`, `Y07`, `Y08`, `Z77` | | 2019 | `W01`, `W31`, `W61`, `Y07`, `Y08` | | 2020 | `W01`, `W31`, `W61`, `Y07`, `Y08` | `X40` in 2011 + `Z77` in 2012 is what makes the recipe bridge testable in the fixture. ## File Structure | File | Responsibility | |---|---| | `inst/sql/26-balance_long.sql` (create) | Restrict `long` to `category_type = 'balance'`, drop aggregates | | `inst/sql/46-balance_annotated.sql` (create) | Join government xwalk + crosswalk onto `balance_long` | | `R/views.R` (modify) | Third gate list: skip both views when the corpus lacks `balance_subtype` | | `R/balances.R` (create) | `cog_balances()`, `.require_balance_support()` | | `R/balance_caveats.R` (create) | `.balance_caveats()`, `.balance_caveat_once()` | | `R/session.R` (modify) | Reset the once-per-session caveat log on `cog_close()` | | `tests/testthat/helper-fixture.R` (modify) | `with_corpus_missing_balance_subtype()` | | `tests/testthat/test-balances.R` (create) | Verb behaviour, guards, per-capita, recipe, caveats | | `NAMESPACE`, `man/` | Regenerated by `devtools::document()` | --- ### Task 1: The two views and the registration gate **Files:** - Create: `inst/sql/26-balance_long.sql`, `inst/sql/46-balance_annotated.sql` - Modify: `R/views.R` - Modify: `tests/testthat/helper-fixture.R` - Test: `tests/testthat/test-balances.R` **Interfaces:** - Consumes: `.register_views(con, url, manifest)`, `.corpus_has_table(manifest, file)` (both existing in `R/views.R`). - Produces: DuckDB views `balance_long` and `balance_annotated`; `.balance_view_files` (character vector); `.corpus_has_balance_subtype(con)` returning `logical(1)`; test helper `with_corpus_missing_balance_subtype(code)`. - [ ] **Step 1: Write the failing test** Create `tests/testthat/test-balances.R`: ```r test_that("balance views register and carry only balance codes", { skip_if_no_corpus() con <- cog_open() on.exit(cog_close()) views <- DBI::dbGetQuery(con, "SELECT table_name FROM information_schema.tables WHERE table_schema = 'main' AND table_type = 'VIEW'" )$table_name expect_true(all(c("balance_long", "balance_annotated") %in% views)) # Every item_code in balance_long is a category_type = 'balance' member. leak <- DBI::dbGetQuery(con, "SELECT COUNT(*) AS n FROM balance_long WHERE item_code NOT IN ( SELECT item_code FROM summary_categories WHERE category_type = 'balance')" )$n expect_identical(as.integer(leak), 0L) # And no aggregate row survives, mirroring revenue_long. agg <- DBI::dbGetQuery(con, "SELECT COUNT(*) AS n FROM balance_long WHERE is_aggregate" )$n expect_identical(as.integer(agg), 0L) # balance_annotated exposes the subtype column the verb groups on. cols <- DBI::dbGetQuery(con, "SELECT column_name FROM information_schema.columns WHERE table_name = 'balance_annotated'" )$column_name expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) }) ``` - [ ] **Step 2: Run it and watch it fail** ```bash cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: FAIL — `balance_long` is not among the registered views. - [ ] **Step 3: Create the two view files** `inst/sql/26-balance_long.sql`: ```sql -- Cash and security holdings, classified by crosswalk MEMBERSHIP on -- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot -- do this job -- the X and Y families each span revenue, expenditure AND -- balance). -- -- These rows are STOCKS: a balance at a point in time, not a flow over a -- fiscal year. Summing a stock with a flow is meaningless, which is why they -- live behind a third view rather than as a subtype of either money view, and -- why neither spending_long nor revenue_long can reach them. -- -- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era -- aggregate-only holdings codes (X40/X41) are deliberately outside this view; -- they are reachable only through the recipe path, which bypasses this filter -- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2). CREATE OR REPLACE VIEW balance_long AS SELECT * FROM long WHERE item_code IN ( SELECT item_code FROM summary_categories WHERE category_type = 'balance' ) AND NOT is_aggregate; ``` `inst/sql/46-balance_annotated.sql`: ```sql CREATE OR REPLACE VIEW balance_annotated AS SELECT s.*, x.gov_name AS xwalk_gov_name, x.govs_type, x.type_label, x.fips_state AS xwalk_fips_state, x.fips_county AS xwalk_fips_county, x.fips_place, x.population_acs, c.category, c.category_type, c.balance_subtype FROM balance_long s LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) LEFT JOIN summary_categories c USING (item_code); ``` - [ ] **Step 4: Add the gate to `R/views.R`** Insert after the `.representation_view_files` block (around line 46): ```r # Cash and security holdings (uscogdata#25). 46- selects # `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and # WITHOUT a schema_version bump -- so neither existing gate applies: # .harmonization_view_files keys on schema_version, .representation_view_files # on the presence of a FILE. Here the discriminator is a COLUMN on a table # that exists either way. CREATE VIEW resolves its source schema eagerly, so # on an older corpus 46- would fail at registration with "Binder Error: # Referenced column balance_subtype not found" rather than at query time. .balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") #' Does the mounted corpus's `summary_categories` carry `balance_subtype`? #' Probed against the live connection rather than the manifest, because the #' manifest describes files, not columns. #' @noRd .corpus_has_balance_subtype <- function(con) { n <- DBI::dbGetQuery(con, "SELECT COUNT(*) AS n FROM information_schema.columns WHERE table_name = 'summary_categories' AND column_name = 'balance_subtype'" )$n isTRUE(as.integer(n) > 0L) } ``` Then inside the `for (f in files)` loop in `.register_views()`, after the existing two `next` guards: ```r if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next ``` This is safe ordering: `11-summary_categories.sql` sorts before `26-`, so the table exists by the time the probe runs. - [ ] **Step 5: Run the test — it should pass** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: PASS. - [ ] **Step 6: Add the gate helper and its test** Append to `tests/testthat/helper-fixture.R`: ```r # Copy the bundled fixture to a temp dir with summary_categories.parquet # rewritten to DROP the balance_subtype column, then run `code` against it. # Models a corpus published before cog_pipeline #76/#77. schema_version is # left untouched deliberately: that change shipped without a version bump, so # column presence is the only honest signal -- this helper is what proves the # package keys off it. Mirrors with_corpus_missing_ig_categories(). with_corpus_missing_balance_subtype <- function(code) { src <- fixture_corpus_path() tmp <- withr::local_tempdir(.local_envir = parent.frame()) file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) cats_path <- file.path(tmp, "data", "summary_categories.parquet") filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") write_con <- DBI::dbConnect(duckdb::duckdb()) on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) DBI::dbExecute(write_con, sprintf( "COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s)) TO %s (FORMAT PARQUET)", uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) )) file.remove(cats_path) file.rename(filtered_path, cats_path) old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) uscogdata:::cog_close() Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) on.exit({ uscogdata:::cog_close() if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) }, add = TRUE) force(code) } ``` Add to `tests/testthat/test-balances.R`: ```r test_that("balance views are skipped on a corpus without balance_subtype", { skip_if_no_corpus() with_corpus_missing_balance_subtype({ con <- cog_open() on.exit(cog_close()) views <- DBI::dbGetQuery(con, "SELECT table_name FROM information_schema.tables WHERE table_schema = 'main' AND table_type = 'VIEW'" )$table_name # Registration must SKIP them, not error -- an older corpus stays usable. expect_false(any(c("balance_long", "balance_annotated") %in% views)) expect_true("revenue_long" %in% views) }) }) ``` - [ ] **Step 7: Run the full suite** ```bash /usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' ``` Expected: 0 failures, 0 errors, and at least 716 passes. - [ ] **Step 8: Commit** ```bash git add inst/sql/26-balance_long.sql inst/sql/46-balance_annotated.sql R/views.R tests/testthat/helper-fixture.R tests/testthat/test-balances.R git commit -m "feat: register balance_long / balance_annotated behind a column gate (#25)" ``` --- ### Task 2: `cog_balances()` core **Files:** - Create: `R/balances.R` - Modify: `tests/testthat/test-balances.R` - Regenerate: `NAMESPACE`, `man/cog_balances.Rd` **Interfaces:** - Consumes: `.ensure_session()`, `.coerce_govid_input(govid)`, `.check_govids_in_scope(govid)`, `.build_verb_sql(view, subtype_col, govid, years, category, ig_view, subtype_scope)`, `.build_provenance(...)`, `.sql_lit_chr(x)` — all existing internals. - Produces: exported `cog_balances(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, basis = c("harmonized","raw"), recipe = NULL)` returning a `tbl_df` with columns `year`, `canonical_govid`, `gov_name`, `balance_subtype`, `category`, `amt_nominal`, `codes_included`, `aggregate_fallback` and a `provenance` attribute. Also `.require_balance_support(con)`, which aborts with class `uscogdata_no_balance_support`. - [ ] **Step 1: Write the failing tests** Add to `tests/testthat/test-balances.R`: ```r test_that("cog_balances returns holdings for a government that has them", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", 2019) expect_s3_class(r, "tbl_df") expect_true(nrow(r) > 0L) expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", "category", "amt_nominal") %in% names(r))) expect_identical(sort(unique(r$category)), c("Fund Balances", "Insurance Trust Balances")) expect_false(is.null(attr(r, "provenance"))) expect_identical(attr(r, "provenance")$verb, "cog_balances") }) }) test_that('category = "Fund Balances" is exactly the general family', { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", 2019, category = "Fund Balances") expect_identical(unique(r$balance_subtype), "general") codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) expect_identical(codes, c("W01", "W31", "W61")) }) }) test_that("no flow code can reach cog_balances", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) # The expected set is read from the RAW corpus, never from the verb -- # verifying an absence through the filter that creates it proves nothing. ds <- arrow::open_dataset(file.path(fixture_corpus_path(), "data", "long")) sc <- arrow::read_parquet( file.path(fixture_corpus_path(), "data", "summary_categories.parquet")) sc <- as.data.frame(sc) balance_codes <- sc$item_code[sc$category_type == "balance"] expect_true(all(got %in% balance_codes)) expect_true(length(setdiff(got, balance_codes)) == 0L) }) }) test_that("every balance_subtype maps to exactly one category", { skip_if_no_corpus() # Dropping the `subtype` argument is only safe while this tree holds. If the # pipeline ever gives a balance subtype a second category, `category` becomes # a lossy filter -- fail HERE rather than in a user's analysis. sc <- as.data.frame(arrow::read_parquet( file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) b <- sc[sc$category_type == "balance", ] per_subtype <- tapply(b$category, b$balance_subtype, function(x) length(unique(x))) expect_true(all(per_subtype == 1L)) }) ``` - [ ] **Step 2: Run and watch it fail** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: FAIL — `could not find function "cog_balances"`. - [ ] **Step 3: Create `R/balances.R`** ```r # R/balances.R # # Cash and security holdings. A third verb rather than an argument on a money # verb because holdings are a STOCK -- a balance at a point in time -- while # cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money # verbs' whole argument vocabulary (expenditure_concept, revenue_concept, # complete=) describes flows and is meaningless here, so this deliberately # does NOT route through .verb_spendrev(). #' Cash and security holdings for one or more governments #' #' Returns Census cash-and-security holdings (`category_type = "balance"`): #' fund balances, retirement system holdings and insurance trust balances. #' #' @section Holdings are not GAAP fund balance: #' Census holdings are **gross** -- no liabilities are netted -- so a reserve #' ratio built from them overstates what is actually available. They are not #' comparable to a GAAP fund balance from an ACFR. #' #' @param govid Canonical govid(s): a character vector, or a data frame with a #' `canonical_govid` column (e.g. from [cog_gov_search()]). #' @param years Integer vector of fiscal years. #' @param category Optional character vector of categories to keep. One of #' `"Fund Balances"`, `"Insurance Trust Balances"`, #' `"Retirement System Holdings"`. There is deliberately no `subtype` #' argument: for holdings, `category` is a strict coarsening of #' `balance_subtype` (unlike the money verbs, where the two axes cross), so #' every combination would be either redundant or empty. #' `category = "Fund Balances"` is exactly the `general` family #' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is #' one `dplyr::filter()` away. #' @param per_capita Divide holdings by population. Note this is a **stock per #' resident** (reserves per person), which is *not* comparable to #' [cog_spending()]'s per-capita figures -- those are a flow per person. #' @param adjust_to_year Deflate to this year's dollars (CPI-U). #' @param basis Accepted for uniformity with the money verbs, but currently a #' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized #' and raw space are identical for holdings. Reported in #' `provenance$basis_note`. #' @param recipe Optional harmonization recipe id (see [cog_recipes()]). #' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the #' wide era to the modern one. #' #' @return A `tbl_df` with a `provenance` attribute. Amounts are full US #' dollars. #' @export cog_balances <- function(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, basis = c("harmonized", "raw"), recipe = NULL) { call <- match.call() basis <- match.arg(basis, c("harmonized", "raw")) govid <- .coerce_govid_input(govid) years <- as.integer(years) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) con <- .ensure_session() .require_balance_support(con) .check_govids_in_scope(govid) basis_note <- paste0( "`basis` has no effect on holdings: harmonization_map carries no ", "balance-code rows, so harmonized and raw space are identical here." ) sql <- .build_verb_sql("balance_annotated", "balance_subtype", govid, years, category, ig_view = NULL, subtype_scope = NULL) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) prov <- .build_provenance( verb = "cog_balances", call = call, govid = govid, years = years, category = category, per_capita = per_capita, adjust_to_year = adjust_to_year, result = result, sql = sql, subtype_col = "balance_subtype", basis = basis, basis_note = basis_note, # Neither concept vocabulary applies to a stock. expenditure_concept = NA_character_, revenue_concept = NA_character_ ) attr(result, "provenance") <- prov result } #' Abort unless the mounted corpus classifies balance codes. #' #' `balance_subtype` arrived with cog_pipeline #76/#77 without a #' schema_version bump, so the check is on the column, not the version. #' @noRd .require_balance_support <- function(con) { if (.corpus_has_balance_subtype(con)) return(invisible(TRUE)) cli::cli_abort( c("This corpus does not classify cash and security holdings.", i = "`summary_categories` has no {.field balance_subtype} column.", i = "Republish from cog_pipeline at #76/#77 or later."), class = "uscogdata_no_balance_support" ) } ``` `codes_included` stays on the returned tibble — `cog_spending()` and `cog_revenue()` both keep it, and dropping it here would be a gratuitous asymmetry. It is *also* mirrored into `provenance$codes_summed$observed` by `.build_provenance()`, which is what the series-break builder reads. - [ ] **Step 4: Document and run** ```bash /usr/bin/Rscript -e 'devtools::document()' /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: PASS, and `NAMESPACE` gains `export(cog_balances)`. - [ ] **Step 5: Run the full suite** ```bash /usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' ``` Expected: 0 failures, 0 errors. - [ ] **Step 6: Commit** ```bash git add R/balances.R NAMESPACE man/ tests/testthat/test-balances.R git commit -m "feat: cog_balances() core verb (#25)" ``` --- ### Task 3: `per_capita` and `adjust_to_year` **Files:** - Modify: `R/balances.R` - Modify: `tests/testthat/test-balances.R` **Interfaces:** - Consumes: `.attach_per_capita(result, con, govid)` and `.attach_real_dollars(result, adjust_to_year, per_capita)` from `R/spending.R`. The first adds `amt_per_capita_nominal` and `pop_source` and sets the `.popyear_range` attribute; the second adds `amt_real` and, when `per_capita`, `amt_per_capita_real`. - Produces: no new functions — `cog_balances()` gains the two columns. - [ ] **Step 1: Write the failing tests** ```r test_that("per_capita divides holdings by population", { skip_if_no_corpus() with_fixture_corpus({ plain <- cog_balances("010000226085", 2019, category = "Fund Balances") pc <- cog_balances("010000226085", 2019, category = "Fund Balances", per_capita = TRUE) expect_true("amt_per_capita_nominal" %in% names(pc)) expect_true("pop_source" %in% names(pc)) expect_identical(pc$amt_nominal, plain$amt_nominal) # Assert against the denominator read from the corpus, NOT against a # quantity derived from amt_per_capita_nominal itself -- dividing the # column back out would be tautological and would pass on any value. pop <- DBI::dbGetQuery(cog_open(), sprintf( "SELECT population FROM gov_population_yearly WHERE canonical_govid = %s AND year = 2019", uscogdata:::.sql_lit_chr("010000226085") ))$population expect_length(pop, 1L) expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, tolerance = 1e-8) prov <- attr(pc, "provenance") expect_true(prov$transformations$per_capita$applied) }) }) test_that("adjust_to_year adds real dollars", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", 2012, category = "Fund Balances", adjust_to_year = 2020) expect_true("amt_real" %in% names(r)) # 2012 dollars inflated to 2020 must exceed nominal. expect_true(all(r$amt_real > r$amt_nominal)) prov <- attr(r, "provenance") expect_true(prov$transformations$inflation$applied) expect_identical(prov$transformations$inflation$base_year, 2020L) }) }) ``` - [ ] **Step 2: Run and watch it fail** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: FAIL — `amt_per_capita_nominal` not among names. - [ ] **Step 3: Wire the two helpers into `cog_balances()`** In `R/balances.R`, between the `dbGetQuery()` call and `.build_provenance()`: ```r if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid) if (!is.null(adjust_to_year)) { result <- .attach_real_dollars(result, adjust_to_year, per_capita) } ``` Order matters and matches `.verb_spendrev()`: per-capita first, so the real per-capita column is deflated from the nominal per-capita one rather than recomputed. - [ ] **Step 4: Run the tests** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: PASS. - [ ] **Step 5: Commit** ```bash git add R/balances.R tests/testthat/test-balances.R git commit -m "feat: per_capita and adjust_to_year for cog_balances() (#25)" ``` --- ### Task 4: `recipe=` — the wide-era bridge **Files:** - Modify: `R/balances.R` - Modify: `tests/testthat/test-balances.R` **Interfaces:** - Consumes: `.require_schema_v5(con, manifest, what)`, `.validate_recipe_id(con, recipe_id)`, `.recipe_components(con, recipe_id)` (returns a data frame with `label`, `component_code`, `year_min`, `year_max`, `weight`), `.run_recipe(con, recipe_id, govid, years)` (returns a tibble with `sql_query` attribute), `.shape_recipe_result(result, subtype_col, label)`, `.df_to_row_list(df)`. - Produces: no new functions. - [ ] **Step 1: Write the failing tests** ```r test_that("recipe bridges the wide era into the modern one", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", c(2011, 2012), recipe = "cash_securities_z77_wide") expect_identical(sort(r$year), c(2011L, 2012L)) # The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE # and therefore invisible to balance_long. If the recipe path ever starts # filtering aggregates, a 45-year series silently truncates to five -- # this is the regression guard for phase_r_harmonization_review.md § 0.2. codes <- attr(r, "provenance")$codes_summed$observed expect_true("X40" %in% codes) expect_true("Z77" %in% codes) expect_true(all(r$amt_nominal > 0)) prov <- attr(r, "provenance") expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide") }) }) test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", c(2011, 2012), recipe = "cash_securities_z77_wide") refs <- attr(r, "provenance")$series_break_refs # SB195 sits on fin_code X40; it can only fire where X40 is observed, # which is exactly the recipe path. expect_true("SB195" %in% refs) }) }) test_that("an unknown recipe id is rejected", { skip_if_no_corpus() with_fixture_corpus({ expect_error(cog_balances("010000226085", 2019, recipe = "no_such_recipe")) }) }) ``` - [ ] **Step 2: Run and watch it fail** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: FAIL — `recipe` is accepted but ignored, so 2011 returns no rows. - [ ] **Step 3: Add the recipe branch** In `R/balances.R`, replace the single `sql <- .build_verb_sql(...)` / `result <- ...` pair with: ```r manifest <- .uscogdata_env$manifest recipe_block <- NULL category_for_prov <- category if (!is.null(recipe)) { .require_schema_v5(con, manifest, "recipe =") .validate_recipe_id(con, recipe) comps <- .recipe_components(con, recipe) recipe_label <- comps$label[[1]] result <- .run_recipe(con, recipe, govid, years) sql <- attr(result, "sql_query") result <- .shape_recipe_result(result, "balance_subtype", recipe_label) recipe_block <- list( recipe_id = recipe, label = recipe_label, components = .df_to_row_list(comps) ) category_for_prov <- recipe_label } else { sql <- .build_verb_sql("balance_annotated", "balance_subtype", govid, years, category, ig_view = NULL, subtype_scope = NULL) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) } ``` and pass the two new values through to `.build_provenance()`: ```r category = category_for_prov, recipe = recipe_block, ``` - [ ] **Step 4: Run the tests** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: PASS. If `SB195` is absent, check that `.build_series_break_refs()` received `X40` — it reads `provenance$codes_summed$observed`, which `.shape_recipe_result()` populates from `codes_included`. - [ ] **Step 5: Run the full suite** ```bash /usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' ``` - [ ] **Step 6: Commit** ```bash git add R/balances.R tests/testthat/test-balances.R git commit -m "feat: recipe= bridges the wide-era holdings series (#25)" ``` --- ### Task 5: `balance_caveats` provenance and the once-per-session message **Files:** - Create: `R/balance_caveats.R` - Modify: `R/balances.R`, `R/session.R` - Modify: `tests/testthat/test-balances.R` **Interfaces:** - Consumes: `.uscogdata_env` (mutable environment from `R/config.R`), `cog_close()` in `R/session.R`. - Produces: `.balance_caveats(con, codes_observed, years)` returning `list(not_gaap = TRUE, coverage_window = , truncated = )`; `.balance_caveat_once(key)` returning `TRUE` the first time a key is seen in a session and `FALSE` after. - [ ] **Step 1: Write the failing tests** ```r test_that("balance_caveats is always present and flags the GAAP distinction", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", 2019) cav <- attr(r, "provenance")$balance_caveats expect_false(is.null(cav)) expect_true(cav$not_gaap) }) }) test_that("coverage_window is computed from the corpus, not hardcoded", { skip_if_no_corpus() with_fixture_corpus({ r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) cav <- attr(r, "provenance")$balance_caveats ds <- arrow::open_dataset( file.path(fixture_corpus_path(), "data", "long")) sc <- as.data.frame(arrow::read_parquet( file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) gen <- sc$item_code[sc$category_type == "balance" & sc$balance_subtype == "general"] obs <- as.data.frame( dplyr::collect(dplyr::summarise( dplyr::filter(ds, item_code %in% gen), y0 = min(year), y1 = max(year)))) expect_identical(as.integer(cav$coverage_window$general), c(as.integer(obs$y0), as.integer(obs$y1))) }) }) test_that("a request past a family's coverage window is flagged", { skip_if_no_corpus() with_fixture_corpus({ # employee_retirement stops at FY2016; 2019/2020 are past it. r <- cog_balances("010000226085", c(2012, 2019)) cav <- attr(r, "provenance")$balance_caveats expect_true("employee_retirement" %in% cav$truncated) }) }) test_that("the caveat message fires once per session", { skip_if_no_corpus() with_fixture_corpus({ expect_message(cog_balances("010000226085", 2019), "not.*GAAP") expect_no_message(cog_balances("010000226085", 2020)) }) }) ``` - [ ] **Step 2: Run and watch it fail** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: FAIL — `balance_caveats` is NULL. - [ ] **Step 3: Create `R/balance_caveats.R`** ```r # R/balance_caveats.R # # The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and # security holdings. Each one silently invalidates an obvious analysis, so # they travel in provenance (machine-readable, for cog-api#26) rather than # living only in prose. # # Two of the four are already carried by the code-driven series-break # builders and are deliberately NOT duplicated here: # * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via # series_break_refs on the recipe path, the only path that observes those # codes. # What remains is the GAAP distinction (a constant) and the coverage windows # (measured, never hardcoded, so they stay correct as the corpus grows). #' Per-subtype observed year extents, plus which requested families are #' truncated relative to the requested span. #' @noRd .balance_caveats <- function(con, codes_observed, years) { windows <- DBI::dbGetQuery(con, "SELECT c.balance_subtype AS subtype, MIN(l.year) AS year_min, MAX(l.year) AS year_max FROM balance_long l JOIN summary_categories c USING (item_code) WHERE c.balance_subtype IS NOT NULL GROUP BY 1 ORDER BY 1" ) observed_subtypes <- if (length(codes_observed) == 0L) { character(0) } else { DBI::dbGetQuery(con, sprintf( "SELECT DISTINCT balance_subtype FROM summary_categories WHERE item_code IN (%s) AND balance_subtype IS NOT NULL", .sql_lit_chr(codes_observed) ))$balance_subtype } cw <- stats::setNames( lapply(seq_len(nrow(windows)), function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), windows$subtype ) # A family is "truncated" when the caller asked for years outside the span # that family actually covers -- the FY2016 employee-retirement termination # and the FY2021 end of the W family are both this shape. truncated <- character(0) if (length(years) > 0L) { for (s in observed_subtypes) { w <- cw[[s]] if (is.null(w)) next if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s) } } list( not_gaap = TRUE, not_gaap_note = paste0( "Census holdings are gross -- no liabilities are netted -- and are NOT ", "GAAP fund balance. A reserve ratio built from them overstates what is ", "actually available." ), coverage_window = cw, truncated = sort(unique(truncated)) ) } #' TRUE the first time `key` is seen this session, FALSE thereafter. #' Reset by cog_close(). #' @noRd .balance_caveat_once <- function(key) { seen <- .uscogdata_env$balance_caveats_shown if (is.null(seen)) seen <- character(0) if (key %in% seen) return(FALSE) .uscogdata_env$balance_caveats_shown <- c(seen, key) TRUE } #' Emit at most one message per caveat class per session. #' @noRd .emit_balance_caveats <- function(caveats) { if (.balance_caveat_once("not_gaap")) { cli::cli_inform(c( "!" = "Census holdings are gross and are {.strong not} GAAP fund balance.", "i" = "No liabilities are netted; a reserve ratio built from them overstates available funds." )) } if (length(caveats$truncated) > 0L && .balance_caveat_once("coverage_window")) { cli::cli_inform(c( "!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.", "i" = "See {.code provenance$balance_caveats$coverage_window}." )) } invisible(NULL) } ``` - [ ] **Step 4: Wire it into `cog_balances()`** After `prov <- .build_provenance(...)` in `R/balances.R`: ```r prov$balance_caveats <- .balance_caveats( con, prov$codes_summed$observed, years ) .emit_balance_caveats(prov$balance_caveats) ``` - [ ] **Step 5: Reset the log in `cog_close()`** In `R/session.R`, inside `cog_close()`, alongside the existing teardown: ```r .uscogdata_env$balance_caveats_shown <- NULL ``` This matters for the tests: `with_fixture_corpus()` calls `cog_close()`, so each test block starts with a clean message log. - [ ] **Step 6: Run the tests** ```bash /usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' ``` Expected: PASS. - [ ] **Step 7: Run the full suite** ```bash /usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' ``` - [ ] **Step 8: Commit** ```bash git add R/balance_caveats.R R/balances.R R/session.R tests/testthat/test-balances.R git commit -m "feat: balance_caveats provenance + once-per-session disclosure (#25)" ``` --- ### Task 6: Documentation **Files:** - Modify: `NEWS.md`, `_pkgdown.yml`, `CLAUDE.md`, `README.md` **Interfaces:** - Consumes: the exported `cog_balances()` from Task 2. - Produces: no code. - [ ] **Step 1: Add the NEWS entry** At the top of `NEWS.md`, under the development heading: ```markdown * New `cog_balances()` exposes the 14 cash-and-security holding codes (`category_type = "balance"`): fund balances, retirement system holdings and insurance trust balances (#25). Holdings are a stock, not a flow, so the verb has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and no `subtype` argument either -- for holdings, `category` is a strict coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly the `general` family (`W01`/`W31`/`W61`). * `cog_balances()` results carry `provenance$balance_caveats`, recording that Census holdings are gross rather than GAAP fund balance, and the measured coverage window of each subtype family. ``` - [ ] **Step 2: Add the verb to `_pkgdown.yml`** Add `cog_balances` to the same reference section that lists `cog_spending` and `cog_revenue`. - [ ] **Step 3: Correct the stale claims in `CLAUDE.md`** Four statements are wrong. Replace the "SQL lives in `inst/sql/` — never inline SQL strings in R files" bullet with: ```markdown - SQL has two layers. **View definitions** live in `inst/sql/` and are registered by `.register_views()`, which globs the directory in sorted order and substitutes `{url}`. **Query construction** is inline `sprintf()` in R (`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as a numbered `.sql` file; build a query in R. ``` and update, in the same file: - the `inst/sql/` bullet: 7 view definitions → **23** - the test count: 181 PASS → the current figure from Step 5 - the fixture description: "years 2019+2020" → **years 2011, 2012, 2019, 2020** Add `cog_balances` to the exported-verbs list. - [ ] **Step 4: Verify the doc claims are true** ```bash /usr/bin/Rscript -e 'cat("sql views:", length(list.files("inst/sql", pattern="[.]sql$")), "\n")' /usr/bin/Rscript -e 'suppressMessages(library(arrow)); cat("fixture years:", paste(sort(unique(as.data.frame(open_dataset("inst/extdata/fixture_corpus/data/long") |> dplyr::distinct(year) |> dplyr::collect())$year)), collapse=", "), "\n")' ``` Put the actual output in `CLAUDE.md` — do not copy the numbers above on faith. - [ ] **Step 5: Run the full suite one last time** ```bash /usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' ``` Expected: 0 failures, 0 errors, pass count above 716. - [ ] **Step 6: Commit** ```bash git add NEWS.md _pkgdown.yml CLAUDE.md README.md git commit -m "docs: document cog_balances() and correct stale CLAUDE.md claims (#25)" ``` --- ## Out of scope - **`cog-api#26`** — the `/balances` endpoint. Needs all three of: handler, `param_contract`, **and** the `plumber.R` route signature. Missing the third makes the endpoint return 200 while silently ignoring the parameter, with every handler test still passing. - **cog_pipeline series-break entry** for the FY2016 termination of the seven holdings codes. No row exists at 2016/2017 for `Z77`/`Z78`/`X30`, though `docs/phase_r_harmonization_review.md` § 2 recommended exactly that; `SB197`–`SB202` set the precedent (`coverage_restricted` + `with_caution`). Non-blocking — `coverage_window` covers it reader-side meanwhile. Verify corpus-wide and census-to-census before writing the rows.