From d7e14156ff4b7e647f4c6da5a34b4fc8f61ca3a0 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 08:44:18 -0400 Subject: [PATCH 01/22] docs: design spec for cog_balances(), the uscogdata#25 holdings surface Requirement 1 of #25 shipped with #11/#12. This specs requirement 2 only. Records three upstream gaps found while measuring the corpus, which change the shipping scope: - X40/X41 carry ~42.7K rows (1967-2011) but have no summary_categories row, so they cannot appear in a category_type='balance' view. Both holdings recipes span X40/X41 + Z77/Z78, so recipe= would silently return only the 2012-2016 leg. recipe= is therefore deferred to v2. - SB195/SB196 attach to fin_code X40/X41, outside the balance view. - SB197-SB202 attach to flow codes, not the holdings codes, so the FY2016 termination of X21/X30/X42/X44/X47/Z77/Z78 has no catalogued break. Caveats 2-4 are therefore surfaced reader-side via a computed coverage_window rather than through the existing series-break builders. --- specs/2026-08-03-cog-balances-design.md | 188 ++++++++++++++++++++++++ 1 file changed, 188 insertions(+) create mode 100644 specs/2026-08-03-cog-balances-design.md diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md new file mode 100644 index 0000000..3c0a003 --- /dev/null +++ b/specs/2026-08-03-cog-balances-design.md @@ -0,0 +1,188 @@ +# `cog_balances()` — a reader surface for cash and security holdings + +**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26` +**Date:** 2026-08-03 · **Status:** design, awaiting approval + +Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped +with `#11`/`#12` and is asserted at both view and verb level. This spec covers +requirement 2 only: a way to query holdings. + +## Decision: a verb, not an argument + +`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`. + +Holdings are a **stock** — a balance at a point in time — while the money verbs +return **flows** over a fiscal year. The flow verbs' whole argument vocabulary +is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe +which flows Census aggregates into a published total, and `complete=` fills a +grid of fiscal-year cells. Overloading a money verb would put a stock behind +arguments that all assume a flow. + +## The 14 codes + +Measured against the published corpus 2026-08-03, not transcribed from the +issue. `year_min`/`year_max` are observed row extents. + +| `balance_subtype` | `category` | codes | observed years | +|---|---|---|---| +| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 | +| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 | +| | | `X47` | 1988–2016 | +| | | `X30`, `Z77`, `Z78` | 2012–2016 | +| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 | +| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 | +| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 | + +## Architecture + +### Two new views + +Mirroring the `revenue_long` / `revenue_annotated` pair exactly: + +- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate` +- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and + `summary_categories`, exposing `category`, `category_type`, `balance_subtype` + +`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register +with no new registration code. + +### A third gate list in `R/views.R` + +`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails +at registration time, not at query time. `46-balance_annotated.sql` selects +`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`. +That arrived without a `schema_version` bump, so neither existing gate applies: +`.harmonization_view_files` keys on `schema_version`, `.representation_view_files` +on the presence of a *file*. The discriminator here is a **column on an existing +table**. + +```r +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") +``` + +gated by probing `summary_categories` for `balance_subtype`, with +`cog_balances()` erroring cleanly via `.require_balance_support()` on an older +corpus — mirroring how `.require_schema_v5()` gates the harmonized views. + +### `R/balances.R` — a dedicated path, not `.verb_spendrev()` + +`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and +`complete=` grid are all flow-specific, and four verbs depend on it. Threading a +third mode through it adds branching to shared code for no reuse benefit. + +Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`, +`.build_corpus_break_refs()`, the population join, `.inflate()`, and +`.coerce_govid_input()`. + +Following the package's real two-layer convention: **view definitions** live in +`inst/sql/`; **query construction** is inline `sprintf()` in R, as in +`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R +files", which the verb layer has never obeyed. Corrected in a separate commit — +see Out of scope.) + +## Signature + +```r +cog_balances(govid, years, + subtype = NULL, # general | employee_retirement | + # unemployment_trust | workers_comp_trust | + # other_insurance_trust + category = NULL, # Fund Balances | Insurance Trust Balances | + # Retirement System Holdings + per_capita = FALSE, + adjust_to_year = NULL, + basis = c("harmonized", "raw")) +``` + +Returns a `tbl_df` with a `provenance` attribute, like every other verb. + +**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`. + +**`per_capita` is offered.** Holdings per resident is a real measure (pension +assets per capita, fund balance per resident). The roxygen `@param` states +plainly that this is a *stock per resident* and is **not** comparable to +`cog_spending()`'s per-capita figures. + +**`basis` is currently a no-op** — `harmonization_map` has zero balance-code +rows, so harmonized and raw are identical for holdings. Kept for uniformity +with the money verbs (the API would otherwise special-case), and +`provenance$basis_note` says so outright rather than letting it look meaningful. + +**`recipe` is deliberately omitted from v1.** See below. + +## Decision point: `recipe=` deferred to v2 + +The two holdings recipes are unusable from this verb today, and shipping the +argument anyway would produce a silently truncated series. + +`cash_securities_z77_wide` = `X40` (1967–2011) + `Z77` (2012–2023); +`cash_securities_z78_wide` = `X41` (1967–2011) + `Z78` (2012–2023). + +**`X40` and `X41` have zero rows in `summary_categories`.** They carry ~42,700 +rows in the corpus across 1967–2011, but with no `category_type` they cannot +appear in a `category_type = 'balance'` view. Applying either recipe would +return only the modern leg — 2012–2016 — dropping 45 years while looking like a +valid continuous series. That is precisely the "plausible rather than obviously +wrong" failure mode `#25` exists to prevent. + +The fix is upstream, in the pipeline crosswalk. Until it lands, `cog_balances()` +does not accept `recipe`, and an attempt to pass one is an error naming the +blocking issue rather than a silent partial result. + +## Caveat surfacing + +`provenance$balance_caveats`, always present, plus one `cli_inform()` per +session per caveat class when a query actually touches an affected family or +year. Structured so `cog-api#26` can forward the fields verbatim. + +Three of the four caveats need reader-side work. **My earlier assumption that the +catalogued series breaks would cover them is wrong**, verified against +`series_breaks.parquet`: + +| # | Caveat | Covered by existing machinery? | +|---|---|---| +| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — constant, new field `not_gaap = TRUE` | +| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus | +| 3 | `X`/`Z` family ends FY2016 | **No.** SB197–SB202 attach to `X02`/`X05`/`X08`/`X11`/`X12`, which are *flow* codes. The holdings codes have no catalogued break for their FY2016 termination. Surfaced via `coverage_window` | +| 4 | `X40`/`X41` book → market at FY2002 | **No.** SB195/SB196 attach to `fin_code` `X40`/`X41`, which are outside the balance view (see above). Surfaced as an explicit caveat when the year range crosses 2002 and touches `employee_retirement` | + +`coverage_window` is derived per observed subtype family from the corpus, never +hardcoded, so it stays correct as the corpus grows. + +`series_break_refs` and `corpus_break_refs` are still populated by the existing +code-driven builders — they simply contribute nothing for the four caveats +above, and will start contributing once the upstream gaps close. + +## Testing + +New `tests/testthat/test-balances.R`. The bundled fixture covers all four +fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y` +throughout — so every test below runs offline. + +- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing + the already-asserted forward guard. Absence is verified against the raw corpus + via `read_parquet` on `data/long`, never through the verb that creates it. +- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019; + `coverage_window` reports the termination and the console message fires once. +- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the + measured table above; the FY2002 valuation caveat fires only when the year + range crosses 2002 *and* touches `employee_retirement`. +- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`. +- **`recipe` rejection.** Passing `recipe` errors with a message naming the + blocking issue. +- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose + `summary_categories` lacks `balance_subtype`. + +## Out of scope, tracked separately + +1. **Pipeline issue (new).** Add `summary_categories` rows for `X40`/`X41` + (`category_type = 'balance'`, `balance_subtype = 'employee_retirement'`), and + catalogue a series break for the FY2016 termination of the holdings codes. + Unblocks `recipe=` and lets caveats 3 and 4 flow through the existing + builders. Requires a corpus rebuild + republish, which is a separate, + human-approved operation. +2. **`cog-api#26`.** Adds `/balances` in all three required places — handler, + `param_contract`, and the `plumber.R` route signature. Lands after this. +3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7 + SQL views (there are 21), 181 tests (716), a two-year fixture (four years), + and a "never inline SQL" rule the verb layer does not follow. From 9f9d40e1c3def6e9e393bbfe92cf49b14f818632 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:15:50 -0400 Subject: [PATCH 02/22] docs: drop the subtype argument from cog_balances() balance is the only category_type whose subtype column is not orthogonal to category. Measured against the crosswalk: 5 of 6 expenditure subtypes and 1 of 7 revenue subtypes span more than one category, but 0 of 5 balance subtypes do. Balance is a strict tree -- Fund Balances = {general}, Retirement System Holdings = {employee_retirement}, Insurance Trust Balances = the three trust subtypes. Exposing both arguments would admit no useful combination: of the 15 pairs, 3 are redundant and 12 are guaranteed empty for every government in every year, failing as an empty tibble that reads as "holds none" rather than as a contradiction. Dropping it also keeps the verb aligned -- no uscogdata verb exposes a subtype argument; the API layers its own subtype row filter on top, which cog-api#26 can do for /balances. #25's one-filter requirement is still met, since category = "Fund Balances" is exactly W01/W31/W61. Adds two tests: that one-filter equivalence, and an assertion that the subtype -> category tree holds, so an upstream change making category lossy fails here rather than in a user's analysis. --- specs/2026-08-03-cog-balances-design.md | 54 +++++++++++++++++++++++-- 1 file changed, 50 insertions(+), 4 deletions(-) diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md index 3c0a003..c335fc9 100644 --- a/specs/2026-08-03-cog-balances-design.md +++ b/specs/2026-08-03-cog-balances-design.md @@ -84,9 +84,6 @@ see Out of scope.) ```r cog_balances(govid, years, - subtype = NULL, # general | employee_retirement | - # unemployment_trust | workers_comp_trust | - # other_insurance_trust category = NULL, # Fund Balances | Insurance Trust Balances | # Retirement System Holdings per_capita = FALSE, @@ -96,7 +93,8 @@ cog_balances(govid, years, Returns a `tbl_df` with a `provenance` attribute, like every other verb. -**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`. +**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`, +and `subtype` — see below. **`per_capita` is offered.** Holdings per resident is a real measure (pension assets per capita, fund balance per resident). The roxygen `@param` states @@ -110,6 +108,47 @@ with the money verbs (the API would otherwise special-case), and **`recipe` is deliberately omitted from v1.** See below. +### No `subtype` argument: `category` is a strict coarsening + +`balance` is the only `category_type` in which `category` and the subtype column +are **not** orthogonal. Measured against the published crosswalk: + +| `category_type` | subtypes spanning more than one category | +|---|---| +| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) | +| revenue | 1 of 7 (`own_source`) | +| **balance** | **0 of 5** | + +For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire) +× *economic character* (operations, capital) — so both earn their place. For +balance the relation is a strict tree: + +``` +Fund Balances = {general} W01 W31 W61 +Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78 +Insurance Trust Balances = {unemployment_trust, + workers_comp_trust, + other_insurance_trust} Y07 Y08 Y21 Y61 +``` + +Exposing both would therefore admit no useful combination. Of the 15 possible +pairs, 3 are redundant (the subtype already implies its category) and **12 are +guaranteed empty for every government in every year** — and an impossible query +would fail by returning an empty tibble, which reads as "this government holds +none" rather than "you asked a contradiction." + +Dropping `subtype` also keeps the verb aligned with the rest of the package: no +uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in +`.verb_spendrev()`, and the API layers its own `subtype` row filter on top +(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for +`/balances`. + +`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the +`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is +isolating one of the three insurance funds in a single argument; +`balance_subtype` remains a returned column, so that is one `dplyr::filter()` +away. + ## Decision point: `recipe=` deferred to v2 The two holdings recipes are unusable from this verb today, and shipping the @@ -168,6 +207,13 @@ throughout — so every test below runs offline. measured table above; the FY2002 valuation caveat fires only when the year range crosses 2002 *and* touches `employee_retirement`. - **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`. +- **`category = "Fund Balances"` is the `general` family.** Returns exactly + `W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted + rather than assumed. +- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to + exactly one `category`. Asserted against the crosswalk so that an upstream + change breaking the tree — which would silently make `category` lossy — + fails here rather than in a user's analysis. - **`recipe` rejection.** Passing `recipe` errors with a message naming the blocking issue. - **Gating.** `.require_balance_support()` errors cleanly on a corpus whose From 57212e3399f6944c48ce6d234c4f528bee0f8be0 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:28:52 -0400 Subject: [PATCH 03/22] docs: restore recipe= to cog_balances(); the pipeline was right Corrects this spec. The earlier draft deferred recipe= and proposed adding summary_categories rows for X40/X41. Both were wrong, and the pipeline state they were meant to fix is correct and documented. cog_pipeline/docs/phase_r_harmonization_review.md records the decisions: - Sec 0.2: the wide era exposes these split families ONLY as aggregates, so the recipe join deliberately does NOT filter is_aggregate. Safe by construction -- wide rows are aggregate-only, modern rows leaf-only, every component year-scoped. - Sec 1: the planned X40->Z77 harmonization MAP rows were dropped on purpose; continuity ships as recipes instead. That is why harmonization_map carries no balance-code rows. The reader already implements this (R/recipes.R, R/spending.R). Verified against the live corpus rather than trusting the comment: corrections_combined FY2007, whose wide leg E05 is likewise aggregate-only, returns $906,743,000. Also withdraws the claim that SB155/156 and SB195/196 contradict each other. X40 rows after FY2002 are the wide-era SAS column persisting through the era boundary; they say nothing about a classification-level rename. The two sets describe different layers. What survives is one narrow, non-blocking gap: no series_breaks row exists at 2016/2017 for Z77/Z78/X30, though review doc Sec 2 recommended exactly that. Recorded as out-of-scope item 1 with the SB197-SB202 precedent. --- specs/2026-08-03-cog-balances-design.md | 99 +++++++++++++++---------- 1 file changed, 61 insertions(+), 38 deletions(-) diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md index c335fc9..b79a861 100644 --- a/specs/2026-08-03-cog-balances-design.md +++ b/specs/2026-08-03-cog-balances-design.md @@ -88,7 +88,8 @@ cog_balances(govid, years, # Retirement System Holdings per_capita = FALSE, adjust_to_year = NULL, - basis = c("harmonized", "raw")) + basis = c("harmonized", "raw"), + recipe = NULL) ``` Returns a `tbl_df` with a `provenance` attribute, like every other verb. @@ -106,7 +107,31 @@ rows, so harmonized and raw are identical for holdings. Kept for uniformity with the money verbs (the API would otherwise special-case), and `provenance$basis_note` says so outright rather than letting it look meaningful. -**`recipe` is deliberately omitted from v1.** See below. +**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era +to the modern one: + +``` +cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023) +cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023) +``` + +`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they +are invisible to `balance_long`, which filters `NOT is_aggregate` like every +other basis view. That is by design, not a defect: +`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide +era exposes these split families *only* as aggregates, and that the recipe join +must therefore **not** filter `is_aggregate` — safe by construction, because +wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and +every component is year-scoped, so no double-count is possible. § 1 records the +matching decision that the planned `X40→Z77` harmonization *map* rows were +dropped and the continuity ships as recipes instead, which is why +`harmonization_map` has no balance-code rows. + +The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is +verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose +wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the +live corpus. So a recipe query reaches rows the verb's own view cannot, exactly +as intended. ### No `subtype` argument: `category` is a strict coarsening @@ -149,48 +174,26 @@ isolating one of the three insurance funds in a single argument; `balance_subtype` remains a returned column, so that is one `dplyr::filter()` away. -## Decision point: `recipe=` deferred to v2 - -The two holdings recipes are unusable from this verb today, and shipping the -argument anyway would produce a silently truncated series. - -`cash_securities_z77_wide` = `X40` (1967–2011) + `Z77` (2012–2023); -`cash_securities_z78_wide` = `X41` (1967–2011) + `Z78` (2012–2023). - -**`X40` and `X41` have zero rows in `summary_categories`.** They carry ~42,700 -rows in the corpus across 1967–2011, but with no `category_type` they cannot -appear in a `category_type = 'balance'` view. Applying either recipe would -return only the modern leg — 2012–2016 — dropping 45 years while looking like a -valid continuous series. That is precisely the "plausible rather than obviously -wrong" failure mode `#25` exists to prevent. - -The fix is upstream, in the pipeline crosswalk. Until it lands, `cog_balances()` -does not accept `recipe`, and an attempt to pass one is an error naming the -blocking issue rather than a silent partial result. - ## Caveat surfacing `provenance$balance_caveats`, always present, plus one `cli_inform()` per session per caveat class when a query actually touches an affected family or year. Structured so `cog-api#26` can forward the fields verbatim. -Three of the four caveats need reader-side work. **My earlier assumption that the -catalogued series breaks would cover them is wrong**, verified against -`series_breaks.parquet`: +Verified against `series_breaks.csv`, not assumed: | # | Caveat | Covered by existing machinery? | |---|---|---| -| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — constant, new field `not_gaap = TRUE` | +| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` | | 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus | -| 3 | `X`/`Z` family ends FY2016 | **No.** SB197–SB202 attach to `X02`/`X05`/`X08`/`X11`/`X12`, which are *flow* codes. The holdings codes have no catalogued break for their FY2016 termination. Surfaced via `coverage_window` | -| 4 | `X40`/`X41` book → market at FY2002 | **No.** SB195/SB196 attach to `fin_code` `X40`/`X41`, which are outside the balance view (see above). Surfaced as an explicit caveat when the year range crosses 2002 and touches `employee_retirement` | +| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) | +| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41` — but only on a `recipe` query, which is the only path that observes those codes. Asserted in the tests rather than assumed | `coverage_window` is derived per observed subtype family from the corpus, never hardcoded, so it stays correct as the corpus grows. -`series_break_refs` and `corpus_break_refs` are still populated by the existing -code-driven builders — they simply contribute nothing for the four caveats -above, and will start contributing once the upstream gaps close. +`series_break_refs` and `corpus_break_refs` are otherwise populated by the +existing code-driven builders and need no change. ## Testing @@ -214,19 +217,39 @@ throughout — so every test below runs offline. exactly one `category`. Asserted against the crosswalk so that an upstream change breaking the tree — which would silently make `category` lossy — fails here rather than in a user's analysis. -- **`recipe` rejection.** Passing `recipe` errors with a message naming the - blocking issue. +- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the + `X40` leg for a pre-2012 year, proving the aggregate-only wide rows are + reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A + regression here would silently truncate a 45-year series to five. +- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning + FY2002 carries both in `provenance$series_break_refs`, so the book → market + basis change is disclosed wherever `X40`/`X41` are actually observed. - **Gating.** `.require_balance_support()` errors cleanly on a corpus whose `summary_categories` lacks `balance_subtype`. ## Out of scope, tracked separately -1. **Pipeline issue (new).** Add `summary_categories` rows for `X40`/`X41` - (`category_type = 'balance'`, `balance_subtype = 'employee_retirement'`), and - catalogue a series break for the FY2016 termination of the holdings codes. - Unblocks `recipe=` and lets caveats 3 and 4 flow through the existing - builders. Requires a corpus rebuild + republish, which is a separate, - human-approved operation. +1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of + the seven holdings codes in `series_breaks.csv`. There is currently **no** + entry at 2016/2017 for `Z77`/`Z78`/`X30`, although + `docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended + exactly this — *"candidate new `series_breaks.csv` entries (recommend + `with_caution` documentation rows, no map action)"*. The follow-through never + happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow + codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also + what keeps this out of the `joinable = "no"` identity-change rule, which + would otherwise oblige a harmonization-map row. + + Verify the break corpus-wide and census-to-census before writing the rows. + `cog_balances()` does not wait on this — caveat 3 is covered reader-side by + `coverage_window` meanwhile, and the entry simply adds a second, catalogued + signpost when it lands. + + **Superseded:** an earlier draft of this spec proposed adding + `summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked. + Both were wrong. `X40`/`X41` are deliberately aggregate-only per + `phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows + are the documented § 1 decision, and the recipe path reaches them by design. 2. **`cog-api#26`.** Adds `/balances` in all three required places — handler, `param_contract`, and the `plumber.R` route signature. Lands after this. 3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7 From 7ac4dc6882f5501f17d85bd41bdd0dc5e8e85f18 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:36:42 -0400 Subject: [PATCH 04/22] docs: implementation plan for cog_balances() (#25) Six TDD tasks: the two views + registration gate, the core verb, per_capita and adjust_to_year, recipe=, balance_caveats provenance, docs. Every internal the plan calls was verified to exist with the signature used (.build_verb_sql, .shape_recipe_result, .attach_per_capita, .run_recipe, .require_schema_v5, ...), so the tasks reuse the shared machinery rather than reimplementing it. The verb deliberately does not route through .verb_spendrev(), whose concept scoping, IG leg and complete= grid are all flow-specific. Test government is ALABAMA STATE GOVT (010000226085), which covers every case in the bundled fixture: W01/W31/W61 in 2012/2019/2020, X21+Z77 in 2012, Y07/Y08 throughout, and X40 in 2011 -- so the wide-era recipe bridge is testable offline. --- .superpowers/plans/2026-08-03-cog-balances.md | 1023 +++++++++++++++++ 1 file changed, 1023 insertions(+) create mode 100644 .superpowers/plans/2026-08-03-cog-balances.md diff --git a/.superpowers/plans/2026-08-03-cog-balances.md b/.superpowers/plans/2026-08-03-cog-balances.md new file mode 100644 index 0000000..472b12f --- /dev/null +++ b/.superpowers/plans/2026-08-03-cog-balances.md @@ -0,0 +1,1023 @@ +# cog_balances() Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add `cog_balances()`, a third verb exposing the 14 cash-and-security holding codes (`category_type = 'balance'`), closing `uscogdata#25` requirement 2. + +**Architecture:** Two new DuckDB views (`balance_long`, `balance_annotated`) mirroring the `revenue_long`/`revenue_annotated` pair, registered by the existing `.register_views()` glob behind a new column-presence gate. A dedicated lean verb in `R/balances.R` reuses the shared SQL builder, per-capita join, inflation, recipe runner and provenance builder — it does **not** go through `.verb_spendrev()`, whose concept scoping, intergovernmental leg and `complete=` grid are all flow-specific. + +**Tech Stack:** R, DuckDB via DBI, testthat 3e, roxygen2, cli, tibble, dplyr. + +**Spec:** `specs/2026-08-03-cog-balances-design.md` + +## Global Constraints + +- Run R as `/usr/bin/Rscript`. **Never a conda R** — conda shadows `libPaths` and breaks `arrow`/`duckdb`. +- The test suite runs offline against the bundled fixture; `tests/testthat/setup.R` sets `USCOGDATA_URL` automatically. To run against the live corpus, `USCOGDATA_URL=/home/jared/Nextcloud/Civilytics/SHARE/uscogdata-corpus/` — **trailing slash required**. +- Full suite: `/usr/bin/Rscript -e 'testthat::test_local(".", reporter="silent")'`. Current floor: **716 pass / 0 fail / 0 skip**. Never let it drop. +- Corpus `amt` is in **thousands**; every verb multiplies by `1000.0` in SQL and returns **full US dollars**. Do not apply the conversion twice. +- SQL **view definitions** live in `inst/sql/`; **query construction** is inline `sprintf()` in R. Follow both — `CLAUDE.md`'s "never inline SQL" line is stale and Task 6 fixes it. +- Every verb returns a `tbl_df` with a `provenance` attribute. +- Commits are GPG-signed. If signing times out, retry once — pinentry is non-interactive here. +- **Never verify an absence through the filter that creates it.** Absence assertions read the raw corpus via `arrow::open_dataset()`, never through `cog_balances()`. + +## The test government + +`010000226085` — ALABAMA STATE GOVT. One government covers every case: + +| year | codes present | +|---|---| +| 2011 | `X21`, **`X40`**, `Y07`, `Y08` | +| 2012 | `W01`, `W31`, `W61`, `X21`, `Y07`, `Y08`, `Z77` | +| 2019 | `W01`, `W31`, `W61`, `Y07`, `Y08` | +| 2020 | `W01`, `W31`, `W61`, `Y07`, `Y08` | + +`X40` in 2011 + `Z77` in 2012 is what makes the recipe bridge testable in the fixture. + +## File Structure + +| File | Responsibility | +|---|---| +| `inst/sql/26-balance_long.sql` (create) | Restrict `long` to `category_type = 'balance'`, drop aggregates | +| `inst/sql/46-balance_annotated.sql` (create) | Join government xwalk + crosswalk onto `balance_long` | +| `R/views.R` (modify) | Third gate list: skip both views when the corpus lacks `balance_subtype` | +| `R/balances.R` (create) | `cog_balances()`, `.require_balance_support()` | +| `R/balance_caveats.R` (create) | `.balance_caveats()`, `.balance_caveat_once()` | +| `R/session.R` (modify) | Reset the once-per-session caveat log on `cog_close()` | +| `tests/testthat/helper-fixture.R` (modify) | `with_corpus_missing_balance_subtype()` | +| `tests/testthat/test-balances.R` (create) | Verb behaviour, guards, per-capita, recipe, caveats | +| `NAMESPACE`, `man/` | Regenerated by `devtools::document()` | + +--- + +### Task 1: The two views and the registration gate + +**Files:** +- Create: `inst/sql/26-balance_long.sql`, `inst/sql/46-balance_annotated.sql` +- Modify: `R/views.R` +- Modify: `tests/testthat/helper-fixture.R` +- Test: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.register_views(con, url, manifest)`, `.corpus_has_table(manifest, file)` (both existing in `R/views.R`). +- Produces: DuckDB views `balance_long` and `balance_annotated`; `.balance_view_files` (character vector); `.corpus_has_balance_subtype(con)` returning `logical(1)`; test helper `with_corpus_missing_balance_subtype(code)`. + +- [ ] **Step 1: Write the failing test** + +Create `tests/testthat/test-balances.R`: + +```r +test_that("balance views register and carry only balance codes", { + skip_if_no_corpus() + con <- cog_open() + on.exit(cog_close()) + + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + expect_true(all(c("balance_long", "balance_annotated") %in% views)) + + # Every item_code in balance_long is a category_type = 'balance' member. + leak <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long + WHERE item_code NOT IN ( + SELECT item_code FROM summary_categories WHERE category_type = 'balance')" + )$n + expect_identical(as.integer(leak), 0L) + + # And no aggregate row survives, mirroring revenue_long. + agg <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long WHERE is_aggregate" + )$n + expect_identical(as.integer(agg), 0L) + + # balance_annotated exposes the subtype column the verb groups on. + cols <- DBI::dbGetQuery(con, + "SELECT column_name FROM information_schema.columns + WHERE table_name = 'balance_annotated'" + )$column_name + expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) +}) +``` + +- [ ] **Step 2: Run it and watch it fail** + +```bash +cd /home/jared/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `balance_long` is not among the registered views. + +- [ ] **Step 3: Create the two view files** + +`inst/sql/26-balance_long.sql`: + +```sql +-- Cash and security holdings, classified by crosswalk MEMBERSHIP on +-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot +-- do this job -- the X and Y families each span revenue, expenditure AND +-- balance). +-- +-- These rows are STOCKS: a balance at a point in time, not a flow over a +-- fiscal year. Summing a stock with a flow is meaningless, which is why they +-- live behind a third view rather than as a subtype of either money view, and +-- why neither spending_long nor revenue_long can reach them. +-- +-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era +-- aggregate-only holdings codes (X40/X41) are deliberately outside this view; +-- they are reachable only through the recipe path, which bypasses this filter +-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2). +CREATE OR REPLACE VIEW balance_long AS +SELECT * +FROM long +WHERE item_code IN ( + SELECT item_code FROM summary_categories + WHERE category_type = 'balance' + ) + AND NOT is_aggregate; +``` + +`inst/sql/46-balance_annotated.sql`: + +```sql +CREATE OR REPLACE VIEW balance_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.balance_subtype +FROM balance_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); +``` + +- [ ] **Step 4: Add the gate to `R/views.R`** + +Insert after the `.representation_view_files` block (around line 46): + +```r +# Cash and security holdings (uscogdata#25). 46- selects +# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and +# WITHOUT a schema_version bump -- so neither existing gate applies: +# .harmonization_view_files keys on schema_version, .representation_view_files +# on the presence of a FILE. Here the discriminator is a COLUMN on a table +# that exists either way. CREATE VIEW resolves its source schema eagerly, so +# on an older corpus 46- would fail at registration with "Binder Error: +# Referenced column balance_subtype not found" rather than at query time. +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") + +#' Does the mounted corpus's `summary_categories` carry `balance_subtype`? +#' Probed against the live connection rather than the manifest, because the +#' manifest describes files, not columns. +#' @noRd +.corpus_has_balance_subtype <- function(con) { + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM information_schema.columns + WHERE table_name = 'summary_categories' + AND column_name = 'balance_subtype'" + )$n + isTRUE(as.integer(n) > 0L) +} +``` + +Then inside the `for (f in files)` loop in `.register_views()`, after the +existing two `next` guards: + +```r + if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next +``` + +This is safe ordering: `11-summary_categories.sql` sorts before `26-`, so the +table exists by the time the probe runs. + +- [ ] **Step 5: Run the test — it should pass** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 6: Add the gate helper and its test** + +Append to `tests/testthat/helper-fixture.R`: + +```r +# Copy the bundled fixture to a temp dir with summary_categories.parquet +# rewritten to DROP the balance_subtype column, then run `code` against it. +# Models a corpus published before cog_pipeline #76/#77. schema_version is +# left untouched deliberately: that change shipped without a version bump, so +# column presence is the only honest signal -- this helper is what proves the +# package keys off it. Mirrors with_corpus_missing_ig_categories(). +with_corpus_missing_balance_subtype <- function(code) { + src <- fixture_corpus_path() + tmp <- withr::local_tempdir(.local_envir = parent.frame()) + file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf( + "COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s)) + TO %s (FORMAT PARQUET)", + uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) + )) + file.remove(cats_path) + file.rename(filtered_path, cats_path) + + old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) + uscogdata:::cog_close() + Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) + on.exit({ + uscogdata:::cog_close() + if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) + }, add = TRUE) + force(code) +} +``` + +Add to `tests/testthat/test-balances.R`: + +```r +test_that("balance views are skipped on a corpus without balance_subtype", { + skip_if_no_corpus() + with_corpus_missing_balance_subtype({ + con <- cog_open() + on.exit(cog_close()) + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + # Registration must SKIP them, not error -- an older corpus stays usable. + expect_false(any(c("balance_long", "balance_annotated") %in% views)) + expect_true("revenue_long" %in% views) + }) +}) +``` + +- [ ] **Step 7: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors, and at least 716 passes. + +- [ ] **Step 8: Commit** + +```bash +git add inst/sql/26-balance_long.sql inst/sql/46-balance_annotated.sql R/views.R tests/testthat/helper-fixture.R tests/testthat/test-balances.R +git commit -m "feat: register balance_long / balance_annotated behind a column gate (#25)" +``` + +--- + +### Task 2: `cog_balances()` core + +**Files:** +- Create: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` +- Regenerate: `NAMESPACE`, `man/cog_balances.Rd` + +**Interfaces:** +- Consumes: `.ensure_session()`, `.coerce_govid_input(govid)`, `.check_govids_in_scope(govid)`, `.build_verb_sql(view, subtype_col, govid, years, category, ig_view, subtype_scope)`, `.build_provenance(...)`, `.sql_lit_chr(x)` — all existing internals. +- Produces: exported `cog_balances(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, basis = c("harmonized","raw"), recipe = NULL)` returning a `tbl_df` with columns `year`, `canonical_govid`, `gov_name`, `balance_subtype`, `category`, `amt_nominal`, `codes_included`, `aggregate_fallback` and a `provenance` attribute. Also `.require_balance_support(con)`, which aborts with class `uscogdata_no_balance_support`. + +- [ ] **Step 1: Write the failing tests** + +Add to `tests/testthat/test-balances.R`: + +```r +test_that("cog_balances returns holdings for a government that has them", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", 2019) + expect_s3_class(r, "tbl_df") + expect_true(nrow(r) > 0L) + expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", + "category", "amt_nominal") %in% names(r))) + expect_identical(sort(unique(r$category)), + c("Fund Balances", "Insurance Trust Balances")) + expect_false(is.null(attr(r, "provenance"))) + expect_identical(attr(r, "provenance")$verb, "cog_balances") + }) +}) + +test_that('category = "Fund Balances" is exactly the general family', { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + expect_identical(codes, c("W01", "W31", "W61")) + }) +}) + +test_that("no flow code can reach cog_balances", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) + got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + + # The expected set is read from the RAW corpus, never from the verb -- + # verifying an absence through the filter that creates it proves nothing. + ds <- arrow::open_dataset(file.path(fixture_corpus_path(), "data", "long")) + sc <- arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet")) + sc <- as.data.frame(sc) + balance_codes <- sc$item_code[sc$category_type == "balance"] + + expect_true(all(got %in% balance_codes)) + expect_true(length(setdiff(got, balance_codes)) == 0L) + }) +}) + +test_that("every balance_subtype maps to exactly one category", { + skip_if_no_corpus() + # Dropping the `subtype` argument is only safe while this tree holds. If the + # pipeline ever gives a balance subtype a second category, `category` becomes + # a lossy filter -- fail HERE rather than in a user's analysis. + sc <- as.data.frame(arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) + b <- sc[sc$category_type == "balance", ] + per_subtype <- tapply(b$category, b$balance_subtype, + function(x) length(unique(x))) + expect_true(all(per_subtype == 1L)) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `could not find function "cog_balances"`. + +- [ ] **Step 3: Create `R/balances.R`** + +```r +# R/balances.R +# +# Cash and security holdings. A third verb rather than an argument on a money +# verb because holdings are a STOCK -- a balance at a point in time -- while +# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money +# verbs' whole argument vocabulary (expenditure_concept, revenue_concept, +# complete=) describes flows and is meaningless here, so this deliberately +# does NOT route through .verb_spendrev(). + +#' Cash and security holdings for one or more governments +#' +#' Returns Census cash-and-security holdings (`category_type = "balance"`): +#' fund balances, retirement system holdings and insurance trust balances. +#' +#' @section Holdings are not GAAP fund balance: +#' Census holdings are **gross** -- no liabilities are netted -- so a reserve +#' ratio built from them overstates what is actually available. They are not +#' comparable to a GAAP fund balance from an ACFR. +#' +#' @param govid Canonical govid(s): a character vector, or a data frame with a +#' `canonical_govid` column (e.g. from [cog_gov_search()]). +#' @param years Integer vector of fiscal years. +#' @param category Optional character vector of categories to keep. One of +#' `"Fund Balances"`, `"Insurance Trust Balances"`, +#' `"Retirement System Holdings"`. There is deliberately no `subtype` +#' argument: for holdings, `category` is a strict coarsening of +#' `balance_subtype` (unlike the money verbs, where the two axes cross), so +#' every combination would be either redundant or empty. +#' `category = "Fund Balances"` is exactly the `general` family +#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +#' one `dplyr::filter()` away. +#' @param per_capita Divide holdings by population. Note this is a **stock per +#' resident** (reserves per person), which is *not* comparable to +#' [cog_spending()]'s per-capita figures -- those are a flow per person. +#' @param adjust_to_year Deflate to this year's dollars (CPI-U). +#' @param basis Accepted for uniformity with the money verbs, but currently a +#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized +#' and raw space are identical for holdings. Reported in +#' `provenance$basis_note`. +#' @param recipe Optional harmonization recipe id (see [cog_recipes()]). +#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +#' wide era to the modern one. +#' +#' @return A `tbl_df` with a `provenance` attribute. Amounts are full US +#' dollars. +#' @export +cog_balances <- function(govid, years, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, + basis = c("harmonized", "raw"), recipe = NULL) { + call <- match.call() + basis <- match.arg(basis, c("harmonized", "raw")) + govid <- .coerce_govid_input(govid) + years <- as.integer(years) + if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) + + con <- .ensure_session() + .require_balance_support(con) + .check_govids_in_scope(govid) + + basis_note <- paste0( + "`basis` has no effect on holdings: harmonization_map carries no ", + "balance-code rows, so harmonized and raw space are identical here." + ) + + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + + prov <- .build_provenance( + verb = "cog_balances", call = call, govid = govid, years = years, + category = category, per_capita = per_capita, + adjust_to_year = adjust_to_year, result = result, sql = sql, + subtype_col = "balance_subtype", + basis = basis, basis_note = basis_note, + # Neither concept vocabulary applies to a stock. + expenditure_concept = NA_character_, + revenue_concept = NA_character_ + ) + + attr(result, "provenance") <- prov + result +} + +#' Abort unless the mounted corpus classifies balance codes. +#' +#' `balance_subtype` arrived with cog_pipeline #76/#77 without a +#' schema_version bump, so the check is on the column, not the version. +#' @noRd +.require_balance_support <- function(con) { + if (.corpus_has_balance_subtype(con)) return(invisible(TRUE)) + cli::cli_abort( + c("This corpus does not classify cash and security holdings.", + i = "`summary_categories` has no {.field balance_subtype} column.", + i = "Republish from cog_pipeline at #76/#77 or later."), + class = "uscogdata_no_balance_support" + ) +} +``` + +`codes_included` stays on the returned tibble — `cog_spending()` and +`cog_revenue()` both keep it, and dropping it here would be a gratuitous +asymmetry. It is *also* mirrored into `provenance$codes_summed$observed` by +`.build_provenance()`, which is what the series-break builder reads. + +- [ ] **Step 4: Document and run** + +```bash +/usr/bin/Rscript -e 'devtools::document()' +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS, and `NAMESPACE` gains `export(cog_balances)`. + +- [ ] **Step 5: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors. + +- [ ] **Step 6: Commit** + +```bash +git add R/balances.R NAMESPACE man/ tests/testthat/test-balances.R +git commit -m "feat: cog_balances() core verb (#25)" +``` + +--- + +### Task 3: `per_capita` and `adjust_to_year` + +**Files:** +- Modify: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.attach_per_capita(result, con, govid)` and `.attach_real_dollars(result, adjust_to_year, per_capita)` from `R/spending.R`. The first adds `amt_per_capita_nominal` and `pop_source` and sets the `.popyear_range` attribute; the second adds `amt_real` and, when `per_capita`, `amt_per_capita_real`. +- Produces: no new functions — `cog_balances()` gains the two columns. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("per_capita divides holdings by population", { + skip_if_no_corpus() + with_fixture_corpus({ + plain <- cog_balances("010000226085", 2019, category = "Fund Balances") + pc <- cog_balances("010000226085", 2019, category = "Fund Balances", + per_capita = TRUE) + expect_true("amt_per_capita_nominal" %in% names(pc)) + expect_true("pop_source" %in% names(pc)) + expect_identical(pc$amt_nominal, plain$amt_nominal) + + # Assert against the denominator read from the corpus, NOT against a + # quantity derived from amt_per_capita_nominal itself -- dividing the + # column back out would be tautological and would pass on any value. + pop <- DBI::dbGetQuery(cog_open(), sprintf( + "SELECT population FROM gov_population_yearly + WHERE canonical_govid = %s AND year = 2019", + uscogdata:::.sql_lit_chr("010000226085") + ))$population + expect_length(pop, 1L) + expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, + tolerance = 1e-8) + + prov <- attr(pc, "provenance") + expect_true(prov$transformations$per_capita$applied) + }) +}) + +test_that("adjust_to_year adds real dollars", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r)) + # 2012 dollars inflated to 2020 must exceed nominal. + expect_true(all(r$amt_real > r$amt_nominal)) + prov <- attr(r, "provenance") + expect_true(prov$transformations$inflation$applied) + expect_identical(prov$transformations$inflation$base_year, 2020L) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `amt_per_capita_nominal` not among names. + +- [ ] **Step 3: Wire the two helpers into `cog_balances()`** + +In `R/balances.R`, between the `dbGetQuery()` call and `.build_provenance()`: + +```r + if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid) + if (!is.null(adjust_to_year)) { + result <- .attach_real_dollars(result, adjust_to_year, per_capita) + } +``` + +Order matters and matches `.verb_spendrev()`: per-capita first, so the real +per-capita column is deflated from the nominal per-capita one rather than +recomputed. + +- [ ] **Step 4: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add R/balances.R tests/testthat/test-balances.R +git commit -m "feat: per_capita and adjust_to_year for cog_balances() (#25)" +``` + +--- + +### Task 4: `recipe=` — the wide-era bridge + +**Files:** +- Modify: `R/balances.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.require_schema_v5(con, manifest, what)`, `.validate_recipe_id(con, recipe_id)`, `.recipe_components(con, recipe_id)` (returns a data frame with `label`, `component_code`, `year_min`, `year_max`, `weight`), `.run_recipe(con, recipe_id, govid, years)` (returns a tibble with `sql_query` attribute), `.shape_recipe_result(result, subtype_col, label)`, `.df_to_row_list(df)`. +- Produces: no new functions. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("recipe bridges the wide era into the modern one", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", c(2011, 2012), + recipe = "cash_securities_z77_wide") + expect_identical(sort(r$year), c(2011L, 2012L)) + + # The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE + # and therefore invisible to balance_long. If the recipe path ever starts + # filtering aggregates, a 45-year series silently truncates to five -- + # this is the regression guard for phase_r_harmonization_review.md § 0.2. + codes <- attr(r, "provenance")$codes_summed$observed + expect_true("X40" %in% codes) + expect_true("Z77" %in% codes) + expect_true(all(r$amt_nominal > 0)) + + prov <- attr(r, "provenance") + expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide") + }) +}) + +test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", c(2011, 2012), + recipe = "cash_securities_z77_wide") + refs <- attr(r, "provenance")$series_break_refs + # SB195 sits on fin_code X40; it can only fire where X40 is observed, + # which is exactly the recipe path. + expect_true("SB195" %in% refs) + }) +}) + +test_that("an unknown recipe id is rejected", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_error(cog_balances("010000226085", 2019, recipe = "no_such_recipe")) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `recipe` is accepted but ignored, so 2011 returns no rows. + +- [ ] **Step 3: Add the recipe branch** + +In `R/balances.R`, replace the single `sql <- .build_verb_sql(...)` / +`result <- ...` pair with: + +```r + manifest <- .uscogdata_env$manifest + recipe_block <- NULL + category_for_prov <- category + + if (!is.null(recipe)) { + .require_schema_v5(con, manifest, "recipe =") + .validate_recipe_id(con, recipe) + comps <- .recipe_components(con, recipe) + recipe_label <- comps$label[[1]] + result <- .run_recipe(con, recipe, govid, years) + sql <- attr(result, "sql_query") + result <- .shape_recipe_result(result, "balance_subtype", recipe_label) + recipe_block <- list( + recipe_id = recipe, label = recipe_label, + components = .df_to_row_list(comps) + ) + category_for_prov <- recipe_label + } else { + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + } +``` + +and pass the two new values through to `.build_provenance()`: + +```r + category = category_for_prov, + recipe = recipe_block, +``` + +- [ ] **Step 4: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. If `SB195` is absent, check that `.build_series_break_refs()` +received `X40` — it reads `provenance$codes_summed$observed`, which +`.shape_recipe_result()` populates from `codes_included`. + +- [ ] **Step 5: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +- [ ] **Step 6: Commit** + +```bash +git add R/balances.R tests/testthat/test-balances.R +git commit -m "feat: recipe= bridges the wide-era holdings series (#25)" +``` + +--- + +### Task 5: `balance_caveats` provenance and the once-per-session message + +**Files:** +- Create: `R/balance_caveats.R` +- Modify: `R/balances.R`, `R/session.R` +- Modify: `tests/testthat/test-balances.R` + +**Interfaces:** +- Consumes: `.uscogdata_env` (mutable environment from `R/config.R`), `cog_close()` in `R/session.R`. +- Produces: `.balance_caveats(con, codes_observed, years)` returning `list(not_gaap = TRUE, coverage_window = , truncated = )`; `.balance_caveat_once(key)` returning `TRUE` the first time a key is seen in a session and `FALSE` after. + +- [ ] **Step 1: Write the failing tests** + +```r +test_that("balance_caveats is always present and flags the GAAP distinction", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", 2019) + cav <- attr(r, "provenance")$balance_caveats + expect_false(is.null(cav)) + expect_true(cav$not_gaap) + }) +}) + +test_that("coverage_window is computed from the corpus, not hardcoded", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) + cav <- attr(r, "provenance")$balance_caveats + + ds <- arrow::open_dataset( + file.path(fixture_corpus_path(), "data", "long")) + sc <- as.data.frame(arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) + gen <- sc$item_code[sc$category_type == "balance" & + sc$balance_subtype == "general"] + obs <- as.data.frame( + dplyr::collect(dplyr::summarise( + dplyr::filter(ds, item_code %in% gen), + y0 = min(year), y1 = max(year)))) + + expect_identical(as.integer(cav$coverage_window$general), + c(as.integer(obs$y0), as.integer(obs$y1))) + }) +}) + +test_that("a request past a family's coverage window is flagged", { + skip_if_no_corpus() + with_fixture_corpus({ + # employee_retirement stops at FY2016; 2019/2020 are past it. + r <- cog_balances("010000226085", c(2012, 2019)) + cav <- attr(r, "provenance")$balance_caveats + expect_true("employee_retirement" %in% cav$truncated) + }) +}) + +test_that("the caveat message fires once per session", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_message(cog_balances("010000226085", 2019), "not.*GAAP") + expect_no_message(cog_balances("010000226085", 2020)) + }) +}) +``` + +- [ ] **Step 2: Run and watch it fail** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: FAIL — `balance_caveats` is NULL. + +- [ ] **Step 3: Create `R/balance_caveats.R`** + +```r +# R/balance_caveats.R +# +# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and +# security holdings. Each one silently invalidates an obvious analysis, so +# they travel in provenance (machine-readable, for cog-api#26) rather than +# living only in prose. +# +# Two of the four are already carried by the code-driven series-break +# builders and are deliberately NOT duplicated here: +# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via +# series_break_refs on the recipe path, the only path that observes those +# codes. +# What remains is the GAAP distinction (a constant) and the coverage windows +# (measured, never hardcoded, so they stay correct as the corpus grows). + +#' Per-subtype observed year extents, plus which requested families are +#' truncated relative to the requested span. +#' @noRd +.balance_caveats <- function(con, codes_observed, years) { + windows <- DBI::dbGetQuery(con, + "SELECT c.balance_subtype AS subtype, + MIN(l.year) AS year_min, + MAX(l.year) AS year_max + FROM balance_long l + JOIN summary_categories c USING (item_code) + WHERE c.balance_subtype IS NOT NULL + GROUP BY 1 + ORDER BY 1" + ) + + observed_subtypes <- if (length(codes_observed) == 0L) { + character(0) + } else { + DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT balance_subtype FROM summary_categories + WHERE item_code IN (%s) AND balance_subtype IS NOT NULL", + .sql_lit_chr(codes_observed) + ))$balance_subtype + } + + cw <- stats::setNames( + lapply(seq_len(nrow(windows)), + function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), + windows$subtype + ) + + # A family is "truncated" when the caller asked for years outside the span + # that family actually covers -- the FY2016 employee-retirement termination + # and the FY2021 end of the W family are both this shape. + truncated <- character(0) + if (length(years) > 0L) { + for (s in observed_subtypes) { + w <- cw[[s]] + if (is.null(w)) next + if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s) + } + } + + list( + not_gaap = TRUE, + not_gaap_note = paste0( + "Census holdings are gross -- no liabilities are netted -- and are NOT ", + "GAAP fund balance. A reserve ratio built from them overstates what is ", + "actually available." + ), + coverage_window = cw, + truncated = sort(unique(truncated)) + ) +} + +#' TRUE the first time `key` is seen this session, FALSE thereafter. +#' Reset by cog_close(). +#' @noRd +.balance_caveat_once <- function(key) { + seen <- .uscogdata_env$balance_caveats_shown + if (is.null(seen)) seen <- character(0) + if (key %in% seen) return(FALSE) + .uscogdata_env$balance_caveats_shown <- c(seen, key) + TRUE +} + +#' Emit at most one message per caveat class per session. +#' @noRd +.emit_balance_caveats <- function(caveats) { + if (.balance_caveat_once("not_gaap")) { + cli::cli_inform(c( + "!" = "Census holdings are gross and are {.strong not} GAAP fund balance.", + "i" = "No liabilities are netted; a reserve ratio built from them overstates available funds." + )) + } + if (length(caveats$truncated) > 0L && + .balance_caveat_once("coverage_window")) { + cli::cli_inform(c( + "!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.", + "i" = "See {.code provenance$balance_caveats$coverage_window}." + )) + } + invisible(NULL) +} +``` + +- [ ] **Step 4: Wire it into `cog_balances()`** + +After `prov <- .build_provenance(...)` in `R/balances.R`: + +```r + prov$balance_caveats <- .balance_caveats( + con, prov$codes_summed$observed, years + ) + .emit_balance_caveats(prov$balance_caveats) +``` + +- [ ] **Step 5: Reset the log in `cog_close()`** + +In `R/session.R`, inside `cog_close()`, alongside the existing teardown: + +```r + .uscogdata_env$balance_caveats_shown <- NULL +``` + +This matters for the tests: `with_fixture_corpus()` calls `cog_close()`, so +each test block starts with a clean message log. + +- [ ] **Step 6: Run the tests** + +```bash +/usr/bin/Rscript -e 'testthat::test_local(".", filter="balances")' +``` + +Expected: PASS. + +- [ ] **Step 7: Run the full suite** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +- [ ] **Step 8: Commit** + +```bash +git add R/balance_caveats.R R/balances.R R/session.R tests/testthat/test-balances.R +git commit -m "feat: balance_caveats provenance + once-per-session disclosure (#25)" +``` + +--- + +### Task 6: Documentation + +**Files:** +- Modify: `NEWS.md`, `_pkgdown.yml`, `CLAUDE.md`, `README.md` + +**Interfaces:** +- Consumes: the exported `cog_balances()` from Task 2. +- Produces: no code. + +- [ ] **Step 1: Add the NEWS entry** + +At the top of `NEWS.md`, under the development heading: + +```markdown +* New `cog_balances()` exposes the 14 cash-and-security holding codes + (`category_type = "balance"`): fund balances, retirement system holdings and + insurance trust balances (#25). Holdings are a stock, not a flow, so the verb + has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and + no `subtype` argument either -- for holdings, `category` is a strict + coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly + the `general` family (`W01`/`W31`/`W61`). +* `cog_balances()` results carry `provenance$balance_caveats`, recording that + Census holdings are gross rather than GAAP fund balance, and the measured + coverage window of each subtype family. +``` + +- [ ] **Step 2: Add the verb to `_pkgdown.yml`** + +Add `cog_balances` to the same reference section that lists `cog_spending` and +`cog_revenue`. + +- [ ] **Step 3: Correct the stale claims in `CLAUDE.md`** + +Four statements are wrong. Replace the "SQL lives in `inst/sql/` — never +inline SQL strings in R files" bullet with: + +```markdown +- SQL has two layers. **View definitions** live in `inst/sql/` and are + registered by `.register_views()`, which globs the directory in sorted order + and substitutes `{url}`. **Query construction** is inline `sprintf()` in R + (`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as + a numbered `.sql` file; build a query in R. +``` + +and update, in the same file: +- the `inst/sql/` bullet: 7 view definitions → **23** +- the test count: 181 PASS → the current figure from Step 5 +- the fixture description: "years 2019+2020" → **years 2011, 2012, 2019, 2020** + +Add `cog_balances` to the exported-verbs list. + +- [ ] **Step 4: Verify the doc claims are true** + +```bash +/usr/bin/Rscript -e 'cat("sql views:", length(list.files("inst/sql", pattern="[.]sql$")), "\n")' +/usr/bin/Rscript -e 'suppressMessages(library(arrow)); cat("fixture years:", paste(sort(unique(as.data.frame(open_dataset("inst/extdata/fixture_corpus/data/long") |> dplyr::distinct(year) |> dplyr::collect())$year)), collapse=", "), "\n")' +``` + +Put the actual output in `CLAUDE.md` — do not copy the numbers above on faith. + +- [ ] **Step 5: Run the full suite one last time** + +```bash +/usr/bin/Rscript -e 'r <- testthat::test_local(".", reporter="silent"); d <- as.data.frame(r); cat(sum(d$passed), sum(d$failed), sum(d$error), sum(d$skipped), "\n")' +``` + +Expected: 0 failures, 0 errors, pass count above 716. + +- [ ] **Step 6: Commit** + +```bash +git add NEWS.md _pkgdown.yml CLAUDE.md README.md +git commit -m "docs: document cog_balances() and correct stale CLAUDE.md claims (#25)" +``` + +--- + +## Out of scope + +- **`cog-api#26`** — the `/balances` endpoint. Needs all three of: handler, `param_contract`, **and** the `plumber.R` route signature. Missing the third makes the endpoint return 200 while silently ignoring the parameter, with every handler test still passing. +- **cog_pipeline series-break entry** for the FY2016 termination of the seven holdings codes. No row exists at 2016/2017 for `Z77`/`Z78`/`X30`, though `docs/phase_r_harmonization_review.md` § 2 recommended exactly that; `SB197`–`SB202` set the precedent (`coverage_restricted` + `with_caution`). Non-blocking — `coverage_window` covers it reader-side meanwhile. Verify corpus-wide and census-to-census before writing the rows. From a11e29a0e0de8d3b7acbea905ab6055f677d77ae Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:42:47 -0400 Subject: [PATCH 05/22] docs: use Wisconsin state govt (550000227544) as the cog_balances test government Standardises on the identifier other agents use for state governments, which is stable across corpus vintages and is the same id used against the live API. Verified in the bundled fixture, and it is strictly better coverage than the previous pick: Wisconsin reaches four of the five balance subtypes (adds workers_comp_trust via Y21) and carries BOTH wide->modern recipe bridges (X40->Z77 and X41->Z78), so a second recipe test is added. Y61 (other_insurance_trust) is absent for Wisconsin; no test depends on it. Also notes not to assert on gov_name -- the fixture carries both "WISCONSIN" and "WISCONSIN STATE GOVT" and the verb COALESCEs them. --- .superpowers/plans/2026-08-03-cog-balances.md | 67 +++++++++++++------ 1 file changed, 47 insertions(+), 20 deletions(-) diff --git a/.superpowers/plans/2026-08-03-cog-balances.md b/.superpowers/plans/2026-08-03-cog-balances.md index 472b12f..02088d0 100644 --- a/.superpowers/plans/2026-08-03-cog-balances.md +++ b/.superpowers/plans/2026-08-03-cog-balances.md @@ -23,16 +23,30 @@ ## The test government -`010000226085` — ALABAMA STATE GOVT. One government covers every case: +`550000227544` — **WISCONSIN STATE GOVT**. Use this govid and no other. It is +the identifier other agents have standardised on for state governments, it is +stable across corpus vintages, and it is the same id used against the live API +(`/api/v1/governments/550000227544/...`). Verified present in the bundled +fixture: | year | codes present | |---|---| -| 2011 | `X21`, **`X40`**, `Y07`, `Y08` | -| 2012 | `W01`, `W31`, `W61`, `X21`, `Y07`, `Y08`, `Z77` | +| 2011 | `X21`, **`X40`**, **`X41`**, `X42`, `X44`, `Y07`, `Y08` | +| 2012 | `W01`, `W31`, `W61`, `X21`, `X30`, `X47`, `Y07`, `Y08`, `Y21`, **`Z77`**, **`Z78`** | | 2019 | `W01`, `W31`, `W61`, `Y07`, `Y08` | -| 2020 | `W01`, `W31`, `W61`, `Y07`, `Y08` | +| 2020 | `W01`, `W31`, `W61`, `Y07`, `Y08`, `Y21` | -`X40` in 2011 + `Z77` in 2012 is what makes the recipe bridge testable in the fixture. +This reaches four of the five subtypes — `general` (W), `employee_retirement` +(X/Z), `unemployment_trust` (Y07/Y08) and `workers_comp_trust` (Y21) — and +carries **both** wide→modern recipe bridges (`X40`→`Z77` and `X41`→`Z78`), so +the whole plan is testable offline. + +`other_insurance_trust` (`Y61`) is not present for Wisconsin in the fixture. No +test below depends on it; do not substitute a different government to reach it. + +Do **not** assert on `gov_name`. The fixture carries both `"WISCONSIN"` (from +`long`) and `"WISCONSIN STATE GOVT"` (from the xwalk), and the verb resolves +`COALESCE(xwalk_gov_name, gov_name)`. ## File Structure @@ -301,7 +315,7 @@ Add to `tests/testthat/test-balances.R`: test_that("cog_balances returns holdings for a government that has them", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", 2019) + r <- cog_balances("550000227544", 2019) expect_s3_class(r, "tbl_df") expect_true(nrow(r) > 0L) expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", @@ -316,7 +330,7 @@ test_that("cog_balances returns holdings for a government that has them", { test_that('category = "Fund Balances" is exactly the general family', { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", 2019, category = "Fund Balances") + r <- cog_balances("550000227544", 2019, category = "Fund Balances") expect_identical(unique(r$balance_subtype), "general") codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) expect_identical(codes, c("W01", "W31", "W61")) @@ -326,7 +340,7 @@ test_that('category = "Fund Balances" is exactly the general family', { test_that("no flow code can reach cog_balances", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) # The expected set is read from the RAW corpus, never from the verb -- @@ -514,8 +528,8 @@ git commit -m "feat: cog_balances() core verb (#25)" test_that("per_capita divides holdings by population", { skip_if_no_corpus() with_fixture_corpus({ - plain <- cog_balances("010000226085", 2019, category = "Fund Balances") - pc <- cog_balances("010000226085", 2019, category = "Fund Balances", + plain <- cog_balances("550000227544", 2019, category = "Fund Balances") + pc <- cog_balances("550000227544", 2019, category = "Fund Balances", per_capita = TRUE) expect_true("amt_per_capita_nominal" %in% names(pc)) expect_true("pop_source" %in% names(pc)) @@ -527,7 +541,7 @@ test_that("per_capita divides holdings by population", { pop <- DBI::dbGetQuery(cog_open(), sprintf( "SELECT population FROM gov_population_yearly WHERE canonical_govid = %s AND year = 2019", - uscogdata:::.sql_lit_chr("010000226085") + uscogdata:::.sql_lit_chr("550000227544") ))$population expect_length(pop, 1L) expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, @@ -541,7 +555,7 @@ test_that("per_capita divides holdings by population", { test_that("adjust_to_year adds real dollars", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", 2012, category = "Fund Balances", + r <- cog_balances("550000227544", 2012, category = "Fund Balances", adjust_to_year = 2020) expect_true("amt_real" %in% names(r)) # 2012 dollars inflated to 2020 must exceed nominal. @@ -609,7 +623,7 @@ git commit -m "feat: per_capita and adjust_to_year for cog_balances() (#25)" test_that("recipe bridges the wide era into the modern one", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", c(2011, 2012), + r <- cog_balances("550000227544", c(2011, 2012), recipe = "cash_securities_z77_wide") expect_identical(sort(r$year), c(2011L, 2012L)) @@ -630,7 +644,7 @@ test_that("recipe bridges the wide era into the modern one", { test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", c(2011, 2012), + r <- cog_balances("550000227544", c(2011, 2012), recipe = "cash_securities_z77_wide") refs <- attr(r, "provenance")$series_break_refs # SB195 sits on fin_code X40; it can only fire where X40 is observed, @@ -639,10 +653,23 @@ test_that("the FY2002 book-to-market basis change is disclosed on the recipe pat }) }) +test_that("the second holdings bridge works too", { + skip_if_no_corpus() + with_fixture_corpus({ + # X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011 + # and Z78 in 2012, so both legs are exercised. + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z78_wide") + codes <- attr(r, "provenance")$codes_summed$observed + expect_true(all(c("X41", "Z78") %in% codes)) + expect_identical(sort(r$year), c(2011L, 2012L)) + }) +}) + test_that("an unknown recipe id is rejected", { skip_if_no_corpus() with_fixture_corpus({ - expect_error(cog_balances("010000226085", 2019, recipe = "no_such_recipe")) + expect_error(cog_balances("550000227544", 2019, recipe = "no_such_recipe")) }) }) ``` @@ -735,7 +762,7 @@ git commit -m "feat: recipe= bridges the wide-era holdings series (#25)" test_that("balance_caveats is always present and flags the GAAP distinction", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", 2019) + r <- cog_balances("550000227544", 2019) cav <- attr(r, "provenance")$balance_caveats expect_false(is.null(cav)) expect_true(cav$not_gaap) @@ -745,7 +772,7 @@ test_that("balance_caveats is always present and flags the GAAP distinction", { test_that("coverage_window is computed from the corpus, not hardcoded", { skip_if_no_corpus() with_fixture_corpus({ - r <- cog_balances("010000226085", c(2011, 2012, 2019, 2020)) + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) cav <- attr(r, "provenance")$balance_caveats ds <- arrow::open_dataset( @@ -768,7 +795,7 @@ test_that("a request past a family's coverage window is flagged", { skip_if_no_corpus() with_fixture_corpus({ # employee_retirement stops at FY2016; 2019/2020 are past it. - r <- cog_balances("010000226085", c(2012, 2019)) + r <- cog_balances("550000227544", c(2012, 2019)) cav <- attr(r, "provenance")$balance_caveats expect_true("employee_retirement" %in% cav$truncated) }) @@ -777,8 +804,8 @@ test_that("a request past a family's coverage window is flagged", { test_that("the caveat message fires once per session", { skip_if_no_corpus() with_fixture_corpus({ - expect_message(cog_balances("010000226085", 2019), "not.*GAAP") - expect_no_message(cog_balances("010000226085", 2020)) + expect_message(cog_balances("550000227544", 2019), "not.*GAAP") + expect_no_message(cog_balances("550000227544", 2020)) }) }) ``` From d09bfd6aef3dcc8fd85651ceafcefc360992ecb7 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:46:36 -0400 Subject: [PATCH 06/22] feat: register balance_long / balance_annotated behind a column gate (#25) --- R/views.R | 24 ++++++++++++++++ inst/sql/26-balance_long.sql | 22 +++++++++++++++ inst/sql/46-balance_annotated.sql | 16 +++++++++++ tests/testthat/helper-fixture.R | 33 ++++++++++++++++++++++ tests/testthat/test-balances.R | 47 +++++++++++++++++++++++++++++++ 5 files changed, 142 insertions(+) create mode 100644 inst/sql/26-balance_long.sql create mode 100644 inst/sql/46-balance_annotated.sql create mode 100644 tests/testthat/test-balances.R diff --git a/R/views.R b/R/views.R index 172d6e9..28bf42c 100644 --- a/R/views.R +++ b/R/views.R @@ -45,6 +45,29 @@ "37-code_set.sql" = "code_set.parquet" ) +# Cash and security holdings (uscogdata#25). 46- selects +# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and +# WITHOUT a schema_version bump -- so neither existing gate applies: +# .harmonization_view_files keys on schema_version, .representation_view_files +# on the presence of a FILE. Here the discriminator is a COLUMN on a table +# that exists either way. CREATE VIEW resolves its source schema eagerly, so +# on an older corpus 46- would fail at registration with "Binder Error: +# Referenced column balance_subtype not found" rather than at query time. +.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql") + +#' Does the mounted corpus's `summary_categories` carry `balance_subtype`? +#' Probed against the live connection rather than the manifest, because the +#' manifest describes files, not columns. +#' @noRd +.corpus_has_balance_subtype <- function(con) { + n <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM information_schema.columns + WHERE table_name = 'summary_categories' + AND column_name = 'balance_subtype'" + )$n + isTRUE(as.integer(n) > 0L) +} + #' Does the mounted corpus publish `file` (e.g. "code_set.parquet")? #' Reads the manifest's metadata list rather than stat-ing the URL, so it #' works identically for a local fixture and a remote share. @@ -66,6 +89,7 @@ if (base %in% .harmonization_view_files && schema_version < 5L) next if (base %in% names(.representation_view_files) && !.corpus_has_table(manifest, .representation_view_files[[base]])) next + if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next sql <- paste(readLines(f, warn = FALSE), collapse = "\n") sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) DBI::dbExecute(con, sql) diff --git a/inst/sql/26-balance_long.sql b/inst/sql/26-balance_long.sql new file mode 100644 index 0000000..22095e8 --- /dev/null +++ b/inst/sql/26-balance_long.sql @@ -0,0 +1,22 @@ +-- Cash and security holdings, classified by crosswalk MEMBERSHIP on +-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot +-- do this job -- the X and Y families each span revenue, expenditure AND +-- balance). +-- +-- These rows are STOCKS: a balance at a point in time, not a flow over a +-- fiscal year. Summing a stock with a flow is meaningless, which is why they +-- live behind a third view rather than as a subtype of either money view, and +-- why neither spending_long nor revenue_long can reach them. +-- +-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era +-- aggregate-only holdings codes (X40/X41) are deliberately outside this view; +-- they are reachable only through the recipe path, which bypasses this filter +-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2). +CREATE OR REPLACE VIEW balance_long AS +SELECT * +FROM long +WHERE item_code IN ( + SELECT item_code FROM summary_categories + WHERE category_type = 'balance' + ) + AND NOT is_aggregate; diff --git a/inst/sql/46-balance_annotated.sql b/inst/sql/46-balance_annotated.sql new file mode 100644 index 0000000..8a40a91 --- /dev/null +++ b/inst/sql/46-balance_annotated.sql @@ -0,0 +1,16 @@ +CREATE OR REPLACE VIEW balance_annotated AS +SELECT + s.*, + x.gov_name AS xwalk_gov_name, + x.govs_type, + x.type_label, + x.fips_state AS xwalk_fips_state, + x.fips_county AS xwalk_fips_county, + x.fips_place, + x.population_acs, + c.category, + c.category_type, + c.balance_subtype +FROM balance_long s +LEFT JOIN canonical_fips_xwalk x USING (canonical_govid) +LEFT JOIN summary_categories c USING (item_code); diff --git a/tests/testthat/helper-fixture.R b/tests/testthat/helper-fixture.R index 3ea3744..9d23512 100644 --- a/tests/testthat/helper-fixture.R +++ b/tests/testthat/helper-fixture.R @@ -154,3 +154,36 @@ with_corpus_missing_ig_categories <- function(code) { }, add = TRUE) force(code) } + +# Copy the bundled fixture to a temp dir with summary_categories.parquet +# rewritten to DROP the balance_subtype column, then run `code` against it. +# Models a corpus published before cog_pipeline #76/#77. schema_version is +# left untouched deliberately: that change shipped without a version bump, so +# column presence is the only honest signal -- this helper is what proves the +# package keys off it. Mirrors with_corpus_missing_ig_categories(). +with_corpus_missing_balance_subtype <- function(code) { + src <- fixture_corpus_path() + tmp <- withr::local_tempdir(.local_envir = parent.frame()) + file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE) + + cats_path <- file.path(tmp, "data", "summary_categories.parquet") + filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet") + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf( + "COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s)) + TO %s (FORMAT PARQUET)", + uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path) + )) + file.remove(cats_path) + file.rename(filtered_path, cats_path) + + old_url <- Sys.getenv("USCOGDATA_URL", unset = NA) + uscogdata:::cog_close() + Sys.setenv(USCOGDATA_URL = paste0(tmp, "/")) + on.exit({ + uscogdata:::cog_close() + if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url) + }, add = TRUE) + force(code) +} diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R new file mode 100644 index 0000000..01fe6ab --- /dev/null +++ b/tests/testthat/test-balances.R @@ -0,0 +1,47 @@ +test_that("balance views register and carry only balance codes", { + skip_if_no_corpus() + con <- cog_open() + on.exit(cog_close()) + + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + expect_true(all(c("balance_long", "balance_annotated") %in% views)) + + # Every item_code in balance_long is a category_type = 'balance' member. + leak <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long + WHERE item_code NOT IN ( + SELECT item_code FROM summary_categories WHERE category_type = 'balance')" + )$n + expect_identical(as.integer(leak), 0L) + + # And no aggregate row survives, mirroring revenue_long. + agg <- DBI::dbGetQuery(con, + "SELECT COUNT(*) AS n FROM balance_long WHERE is_aggregate" + )$n + expect_identical(as.integer(agg), 0L) + + # balance_annotated exposes the subtype column the verb groups on. + cols <- DBI::dbGetQuery(con, + "SELECT column_name FROM information_schema.columns + WHERE table_name = 'balance_annotated'" + )$column_name + expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) +}) + +test_that("balance views are skipped on a corpus without balance_subtype", { + skip_if_no_corpus() + with_corpus_missing_balance_subtype({ + con <- cog_open() + on.exit(cog_close()) + views <- DBI::dbGetQuery(con, + "SELECT table_name FROM information_schema.tables + WHERE table_schema = 'main' AND table_type = 'VIEW'" + )$table_name + # Registration must SKIP them, not error -- an older corpus stays usable. + expect_false(any(c("balance_long", "balance_annotated") %in% views)) + expect_true("revenue_long" %in% views) + }) +}) From 825ac394f276589b3594f1eaac34566d09a73c67 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 09:54:53 -0400 Subject: [PATCH 07/22] test: replace vacuous is_aggregate assertion with synthetic-parquet coverage The bundled fixture has no balance item_code with is_aggregate = TRUE, so asserting COUNT(*) FROM balance_long WHERE is_aggregate = 0 passed whether or not the view's AND NOT is_aggregate predicate existed. Follows the synthetic hive-partitioned parquet pattern already used for the 22-/23- and 24-/25- view predicates in test-views.R: reads the real inst/sql/26-balance_long.sql text off disk and executes it against a synthetic corpus containing both an aggregate and non-aggregate row under a real balance item_code (W01). --- tests/testthat/test-balances.R | 64 ++++++++++++++++++++++++++++++---- 1 file changed, 58 insertions(+), 6 deletions(-) diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 01fe6ab..5e5abb3 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -17,12 +17,6 @@ test_that("balance views register and carry only balance codes", { )$n expect_identical(as.integer(leak), 0L) - # And no aggregate row survives, mirroring revenue_long. - agg <- DBI::dbGetQuery(con, - "SELECT COUNT(*) AS n FROM balance_long WHERE is_aggregate" - )$n - expect_identical(as.integer(agg), 0L) - # balance_annotated exposes the subtype column the verb groups on. cols <- DBI::dbGetQuery(con, "SELECT column_name FROM information_schema.columns @@ -31,6 +25,64 @@ test_that("balance views register and carry only balance codes", { expect_true(all(c("category", "category_type", "balance_subtype") %in% cols)) }) +test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", { + # Every category_type = 'balance' item_code in the bundled fixture has + # is_aggregate = FALSE for every row of every year -- there is no real row + # that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An + # assertion against the live fixture (`WHERE is_aggregate` returns 0) is + # therefore vacuous: it passes identically whether or not the view's + # predicate is present. As with the 22-/23- and 24-/25- tests above, this + # reads the real inst/sql/26-balance_long.sql text off disk and executes it + # -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against + # a synthetic hive-partitioned parquet tree that DOES contain an aggregate + # row under a real balance item_code (W01), so a regression that drops the + # predicate changes which rows survive. + skip_if_no_corpus() + + tmp <- withr::local_tempdir() + part_dir <- file.path(tmp, "data", "long", "year=2004") + dir.create(part_dir, recursive = TRUE) + part_path <- file.path(part_dir, "part-0.parquet") + + write_con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives + ('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate` + ) AS t(canonical_govid, item_code, amt, is_aggregate) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(part_path))) + + DBI::dbExecute(write_con, sprintf(" + COPY ( + SELECT * FROM (VALUES + ('W01', 'Fund Balances', 'balance', NULL, NULL, 'general') + ) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype) + ) TO %s (FORMAT PARQUET) + ", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet")))) + + sql_dir <- system.file("sql", package = "uscogdata") + .read_view_sql <- function(filename) { + txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n") + gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE) + } + + con <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE) + DBI::dbExecute(con, .read_view_sql("10-long.sql")) + DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql")) + DBI::dbExecute(con, .read_view_sql("26-balance_long.sql")) + + rows <- DBI::dbGetQuery(con, + "SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid" + ) + expect_equal(nrow(rows), 1L) + expect_equal(rows$canonical_govid, "bal-A") + expect_equal(rows$amt, 100) +}) + test_that("balance views are skipped on a corpus without balance_subtype", { skip_if_no_corpus() with_corpus_missing_balance_subtype({ From a281a9621fa50a66079ae514c731774fde8eaa54 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:01:40 -0400 Subject: [PATCH 08/22] feat: cog_balances() core verb (#25) --- DESCRIPTION | 1 + NAMESPACE | 1 + R/balances.R | 98 ++++++++++++++++++++++++++++++++++ man/cog_balances.Rd | 62 +++++++++++++++++++++ tests/testthat/test-balances.R | 57 ++++++++++++++++++++ 5 files changed, 219 insertions(+) create mode 100644 R/balances.R create mode 100644 man/cog_balances.Rd diff --git a/DESCRIPTION b/DESCRIPTION index 7e248c8..7a45864 100644 --- a/DESCRIPTION +++ b/DESCRIPTION @@ -22,6 +22,7 @@ Imports: httr2, digest Suggests: + arrow, testthat (>= 3.0.0), withr, knitr, diff --git a/NAMESPACE b/NAMESPACE index 1c67f0b..5d956a0 100644 --- a/NAMESPACE +++ b/NAMESPACE @@ -1,5 +1,6 @@ # Generated by roxygen2: do not edit by hand +export(cog_balances) export(cog_basket_resolution) export(cog_basket_unresolved) export(cog_categories) diff --git a/R/balances.R b/R/balances.R new file mode 100644 index 0000000..fd8b4ee --- /dev/null +++ b/R/balances.R @@ -0,0 +1,98 @@ +# R/balances.R +# +# Cash and security holdings. A third verb rather than an argument on a money +# verb because holdings are a STOCK -- a balance at a point in time -- while +# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money +# verbs' whole argument vocabulary (expenditure_concept, revenue_concept, +# complete=) describes flows and is meaningless here, so this deliberately +# does NOT route through .verb_spendrev(). + +#' Cash and security holdings for one or more governments +#' +#' Returns Census cash-and-security holdings (`category_type = "balance"`): +#' fund balances, retirement system holdings and insurance trust balances. +#' +#' @section Holdings are not GAAP fund balance: +#' Census holdings are **gross** -- no liabilities are netted -- so a reserve +#' ratio built from them overstates what is actually available. They are not +#' comparable to a GAAP fund balance from an ACFR. +#' +#' @param govid Canonical govid(s): a character vector, or a data frame with a +#' `canonical_govid` column (e.g. from [cog_gov_search()]). +#' @param years Integer vector of fiscal years. +#' @param category Optional character vector of categories to keep. One of +#' `"Fund Balances"`, `"Insurance Trust Balances"`, +#' `"Retirement System Holdings"`. There is deliberately no `subtype` +#' argument: for holdings, `category` is a strict coarsening of +#' `balance_subtype` (unlike the money verbs, where the two axes cross), so +#' every combination would be either redundant or empty. +#' `category = "Fund Balances"` is exactly the `general` family +#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +#' one `dplyr::filter()` away. +#' @param per_capita Divide holdings by population. Note this is a **stock per +#' resident** (reserves per person), which is *not* comparable to +#' [cog_spending()]'s per-capita figures -- those are a flow per person. +#' @param adjust_to_year Deflate to this year's dollars (CPI-U). +#' @param basis Accepted for uniformity with the money verbs, but currently a +#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized +#' and raw space are identical for holdings. Reported in +#' `provenance$basis_note`. +#' @param recipe Optional harmonization recipe id (see [cog_recipes()]). +#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +#' wide era to the modern one. +#' +#' @return A `tbl_df` with a `provenance` attribute. Amounts are full US +#' dollars. +#' @export +cog_balances <- function(govid, years, category = NULL, + per_capita = FALSE, adjust_to_year = NULL, + basis = c("harmonized", "raw"), recipe = NULL) { + call <- match.call() + basis <- match.arg(basis, c("harmonized", "raw")) + govid <- .coerce_govid_input(govid) + years <- as.integer(years) + if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) + + con <- .ensure_session() + .require_balance_support(con) + .check_govids_in_scope(govid) + + basis_note <- paste0( + "`basis` has no effect on holdings: harmonization_map carries no ", + "balance-code rows, so harmonized and raw space are identical here." + ) + + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + + prov <- .build_provenance( + verb = "cog_balances", call = call, govid = govid, years = years, + category = category, per_capita = per_capita, + adjust_to_year = adjust_to_year, result = result, sql = sql, + subtype_col = "balance_subtype", + basis = basis, basis_note = basis_note, + # Neither concept vocabulary applies to a stock. + expenditure_concept = NA_character_, + revenue_concept = NA_character_ + ) + + attr(result, "provenance") <- prov + result +} + +#' Abort unless the mounted corpus classifies balance codes. +#' +#' `balance_subtype` arrived with cog_pipeline #76/#77 without a +#' schema_version bump, so the check is on the column, not the version. +#' @noRd +.require_balance_support <- function(con) { + if (.corpus_has_balance_subtype(con)) return(invisible(TRUE)) + cli::cli_abort( + c("This corpus does not classify cash and security holdings.", + i = "`summary_categories` has no {.field balance_subtype} column.", + i = "Republish from cog_pipeline at #76/#77 or later."), + class = "uscogdata_no_balance_support" + ) +} diff --git a/man/cog_balances.Rd b/man/cog_balances.Rd new file mode 100644 index 0000000..81323a0 --- /dev/null +++ b/man/cog_balances.Rd @@ -0,0 +1,62 @@ +% Generated by roxygen2: do not edit by hand +% Please edit documentation in R/balances.R +\name{cog_balances} +\alias{cog_balances} +\title{Cash and security holdings for one or more governments} +\usage{ +cog_balances( + govid, + years, + category = NULL, + per_capita = FALSE, + adjust_to_year = NULL, + basis = c("harmonized", "raw"), + recipe = NULL +) +} +\arguments{ +\item{govid}{Canonical govid(s): a character vector, or a data frame with a +`canonical_govid` column (e.g. from [cog_gov_search()]).} + +\item{years}{Integer vector of fiscal years.} + +\item{category}{Optional character vector of categories to keep. One of +`"Fund Balances"`, `"Insurance Trust Balances"`, +`"Retirement System Holdings"`. There is deliberately no `subtype` +argument: for holdings, `category` is a strict coarsening of +`balance_subtype` (unlike the money verbs, where the two axes cross), so +every combination would be either redundant or empty. +`category = "Fund Balances"` is exactly the `general` family +(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is +one `dplyr::filter()` away.} + +\item{per_capita}{Divide holdings by population. Note this is a **stock per +resident** (reserves per person), which is *not* comparable to +[cog_spending()]'s per-capita figures -- those are a flow per person.} + +\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).} + +\item{basis}{Accepted for uniformity with the money verbs, but currently a +**no-op**: `harmonization_map` carries no balance-code rows, so harmonized +and raw space are identical for holdings. Reported in +`provenance$basis_note`.} + +\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]). +`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the +wide era to the modern one.} +} +\value{ +A `tbl_df` with a `provenance` attribute. Amounts are full US + dollars. +} +\description{ +Returns Census cash-and-security holdings (`category_type = "balance"`): +fund balances, retirement system holdings and insurance trust balances. +} +\section{Holdings are not GAAP fund balance}{ + +Census holdings are **gross** -- no liabilities are netted -- so a reserve +ratio built from them overstates what is actually available. They are not +comparable to a GAAP fund balance from an ACFR. +} + diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 5e5abb3..0af67b4 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -97,3 +97,60 @@ test_that("balance views are skipped on a corpus without balance_subtype", { expect_true("revenue_long" %in% views) }) }) + +test_that("cog_balances returns holdings for a government that has them", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + expect_s3_class(r, "tbl_df") + expect_true(nrow(r) > 0L) + expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype", + "category", "amt_nominal") %in% names(r))) + expect_identical(sort(unique(r$category)), + c("Fund Balances", "Insurance Trust Balances")) + expect_false(is.null(attr(r, "provenance"))) + expect_identical(attr(r, "provenance")$verb, "cog_balances") + }) +}) + +test_that('category = "Fund Balances" is exactly the general family', { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + expect_identical(codes, c("W01", "W31", "W61")) + }) +}) + +test_that("no flow code can reach cog_balances", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ","))) + + # The expected set is read from the RAW corpus, never from the verb -- + # verifying an absence through the filter that creates it proves nothing. + ds <- arrow::open_dataset(file.path(fixture_corpus_path(), "data", "long")) + sc <- arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet")) + sc <- as.data.frame(sc) + balance_codes <- sc$item_code[sc$category_type == "balance"] + + expect_true(all(got %in% balance_codes)) + expect_true(length(setdiff(got, balance_codes)) == 0L) + }) +}) + +test_that("every balance_subtype maps to exactly one category", { + skip_if_no_corpus() + # Dropping the `subtype` argument is only safe while this tree holds. If the + # pipeline ever gives a balance subtype a second category, `category` becomes + # a lossy filter -- fail HERE rather than in a user's analysis. + sc <- as.data.frame(arrow::read_parquet( + file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) + b <- sc[sc$category_type == "balance", ] + per_subtype <- tapply(b$category, b$balance_subtype, + function(x) length(unique(x))) + expect_true(all(per_subtype == 1L)) +}) From cdb574d3d0842969f65fdae311f67845757c5c9a Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:12:57 -0400 Subject: [PATCH 09/22] test: drop arrow dependency from cog_balances tests, use direct DuckDB reads Also add explicit non-empty assertion to the flow-code guard test so it cannot pass vacuously on a zero-row result. --- DESCRIPTION | 1 - tests/testthat/test-balances.R | 30 +++++++++++++++++++++--------- 2 files changed, 21 insertions(+), 10 deletions(-) diff --git a/DESCRIPTION b/DESCRIPTION index 7a45864..7e248c8 100644 --- a/DESCRIPTION +++ b/DESCRIPTION @@ -22,7 +22,6 @@ Imports: httr2, digest Suggests: - arrow, testthat (>= 3.0.0), withr, knitr, diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 0af67b4..a5f6efa 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -131,12 +131,18 @@ test_that("no flow code can reach cog_balances", { # The expected set is read from the RAW corpus, never from the verb -- # verifying an absence through the filter that creates it proves nothing. - ds <- arrow::open_dataset(file.path(fixture_corpus_path(), "data", "long")) - sc <- arrow::read_parquet( - file.path(fixture_corpus_path(), "data", "summary_categories.parquet")) - sc <- as.data.frame(sc) - balance_codes <- sc$item_code[sc$category_type == "balance"] + # A fresh, direct DuckDB connection against the raw parquet files (never + # cog_open()'s session, never balance_long/balance_annotated) reads + # parquet natively -- no arrow dependency needed (see CLAUDE.md). + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + balance_codes <- DBI::dbGetQuery(con2, sprintf( + "SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'", + uscogdata:::.sql_lit_chr(cats_path) + ))$item_code + expect_true(length(got) > 0L) expect_true(all(got %in% balance_codes)) expect_true(length(setdiff(got, balance_codes)) == 0L) }) @@ -146,10 +152,16 @@ test_that("every balance_subtype maps to exactly one category", { skip_if_no_corpus() # Dropping the `subtype` argument is only safe while this tree holds. If the # pipeline ever gives a balance subtype a second category, `category` becomes - # a lossy filter -- fail HERE rather than in a user's analysis. - sc <- as.data.frame(arrow::read_parquet( - file.path(fixture_corpus_path(), "data", "summary_categories.parquet"))) - b <- sc[sc$category_type == "balance", ] + # a lossy filter -- fail HERE rather than in a user's analysis. Read via a + # fresh direct DuckDB connection against the raw parquet file, not through + # any registered view. + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + b <- DBI::dbGetQuery(con2, sprintf( + "SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'", + uscogdata:::.sql_lit_chr(cats_path) + )) per_subtype <- tapply(b$category, b$balance_subtype, function(x) length(unique(x))) expect_true(all(per_subtype == 1L)) From de3a58d1054b3fdd1137fadae3ed8029732b91a3 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:19:32 -0400 Subject: [PATCH 10/22] fix: attach govids_found/govids_missing to cog_balances() provenance Mirrors R/spending.R:465-466 -- .check_govids_in_scope()'s return was previously captured only for its message side effect. Also drops a redundant duplicate assertion in the flow-code guard test. --- R/balances.R | 4 +++- tests/testthat/test-balances.R | 13 ++++++++++++- 2 files changed, 15 insertions(+), 2 deletions(-) diff --git a/R/balances.R b/R/balances.R index fd8b4ee..bfd6e77 100644 --- a/R/balances.R +++ b/R/balances.R @@ -55,7 +55,7 @@ cog_balances <- function(govid, years, category = NULL, con <- .ensure_session() .require_balance_support(con) - .check_govids_in_scope(govid) + scope <- .check_govids_in_scope(govid) basis_note <- paste0( "`basis` has no effect on holdings: harmonization_map carries no ", @@ -77,6 +77,8 @@ cog_balances <- function(govid, years, category = NULL, expenditure_concept = NA_character_, revenue_concept = NA_character_ ) + prov$scope$govids_found <- scope$found + prov$scope$govids_missing <- scope$missing attr(result, "provenance") <- prov result diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index a5f6efa..2313194 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -144,7 +144,6 @@ test_that("no flow code can reach cog_balances", { expect_true(length(got) > 0L) expect_true(all(got %in% balance_codes)) - expect_true(length(setdiff(got, balance_codes)) == 0L) }) }) @@ -166,3 +165,15 @@ test_that("every balance_subtype maps to exactly one category", { function(x) length(unique(x))) expect_true(all(per_subtype == 1L)) }) + +test_that("cog_balances records found + missing govids in provenance", { + skip_if_no_corpus() + with_fixture_corpus({ + suppressMessages( + r <- cog_balances(c("550000227544", "XXXINVALID"), 2019) + ) + prov <- attr(r, "provenance") + expect_equal(sort(prov$scope$govids_found), "550000227544") + expect_equal(sort(prov$scope$govids_missing), "XXXINVALID") + }) +}) From 769164c824cdb2a879dd303ddbcbaa76b416797e Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:24:59 -0400 Subject: [PATCH 11/22] feat: per_capita and adjust_to_year for cog_balances() (#25) --- R/balances.R | 27 ++++++++++++++++++++++ tests/testthat/test-balances.R | 41 ++++++++++++++++++++++++++++++++++ 2 files changed, 68 insertions(+) diff --git a/R/balances.R b/R/balances.R index bfd6e77..e80f376 100644 --- a/R/balances.R +++ b/R/balances.R @@ -49,6 +49,7 @@ cog_balances <- function(govid, years, category = NULL, basis = c("harmonized", "raw"), recipe = NULL) { call <- match.call() basis <- match.arg(basis, c("harmonized", "raw")) + .validate_balance_inputs(per_capita, adjust_to_year) govid <- .coerce_govid_input(govid) years <- as.integer(years) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) @@ -67,6 +68,14 @@ cog_balances <- function(govid, years, category = NULL, ig_view = NULL, subtype_scope = NULL) result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + # Order matters (matches .verb_spendrev()): per-capita first, so + # .attach_real_dollars() deflates the nominal per-capita column into + # amt_per_capita_real rather than needing amt_per_capita_nominal recomputed. + if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid) + if (!is.null(adjust_to_year)) { + result <- .attach_real_dollars(result, adjust_to_year, per_capita) + } + prov <- .build_provenance( verb = "cog_balances", call = call, govid = govid, years = years, category = category, per_capita = per_capita, @@ -84,6 +93,24 @@ cog_balances <- function(govid, years, category = NULL, result } +#' Cheap type validation for the two arguments cog_balances() shares with the +#' money verbs. Mirrors the per_capita/adjust_to_year checks in +#' .validate_verb_inputs() (R/spending.R) -- category/recipe validation is +#' deliberately out of scope here (uscogdata#25 Task 3 review note). +#' @noRd +.validate_balance_inputs <- function(per_capita, adjust_to_year) { + if (!is.logical(per_capita) || length(per_capita) != 1L) { + cli::cli_abort("`per_capita` must be a length-1 logical.") + } + if (!is.null(adjust_to_year)) { + if (!(is.integer(adjust_to_year) || is.numeric(adjust_to_year)) || + length(adjust_to_year) != 1L) { + cli::cli_abort("`adjust_to_year` must be NULL or a length-1 integer.") + } + } + invisible(TRUE) +} + #' Abort unless the mounted corpus classifies balance codes. #' #' `balance_subtype` arrived with cog_pipeline #76/#77 without a diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 2313194..d0e64ac 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -177,3 +177,44 @@ test_that("cog_balances records found + missing govids in provenance", { expect_equal(sort(prov$scope$govids_missing), "XXXINVALID") }) }) + +test_that("per_capita divides holdings by population", { + skip_if_no_corpus() + with_fixture_corpus({ + plain <- cog_balances("550000227544", 2019, category = "Fund Balances") + pc <- cog_balances("550000227544", 2019, category = "Fund Balances", + per_capita = TRUE) + expect_true("amt_per_capita_nominal" %in% names(pc)) + expect_true("pop_source" %in% names(pc)) + expect_identical(pc$amt_nominal, plain$amt_nominal) + + # Assert against the denominator read from the corpus, NOT against a + # quantity derived from amt_per_capita_nominal itself -- dividing the + # column back out would be tautological and would pass on any value. + pop <- DBI::dbGetQuery(cog_open(), sprintf( + "SELECT population FROM gov_population_yearly + WHERE canonical_govid = %s AND year = 2019", + uscogdata:::.sql_lit_chr("550000227544") + ))$population + expect_length(pop, 1L) + expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop, + tolerance = 1e-8) + + prov <- attr(pc, "provenance") + expect_true(prov$transformations$per_capita$applied) + }) +}) + +test_that("adjust_to_year adds real dollars", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r)) + # 2012 dollars inflated to 2020 must exceed nominal. + expect_true(all(r$amt_real > r$amt_nominal)) + prov <- attr(r, "provenance") + expect_true(prov$transformations$inflation$applied) + expect_identical(prov$transformations$inflation$base_year, 2020L) + }) +}) From 90d2e6019ef1cfcfae3302204d37eb1e86efc57a Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:37:54 -0400 Subject: [PATCH 12/22] feat: recipe= bridges the wide-era holdings series (#25) --- R/balances.R | 32 +++++++++++--- tests/testthat/test-balances.R | 76 ++++++++++++++++++++++++++++++++++ 2 files changed, 102 insertions(+), 6 deletions(-) diff --git a/R/balances.R b/R/balances.R index e80f376..a2d0d9c 100644 --- a/R/balances.R +++ b/R/balances.R @@ -63,10 +63,29 @@ cog_balances <- function(govid, years, category = NULL, "balance-code rows, so harmonized and raw space are identical here." ) - sql <- .build_verb_sql("balance_annotated", "balance_subtype", - govid, years, category, - ig_view = NULL, subtype_scope = NULL) - result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + manifest <- .uscogdata_env$manifest + recipe_block <- NULL + category_for_prov <- category + + if (!is.null(recipe)) { + .require_schema_v5(con, manifest, "recipe =") + .validate_recipe_id(con, recipe) + comps <- .recipe_components(con, recipe) + recipe_label <- comps$label[[1]] + result <- .run_recipe(con, recipe, govid, years) + sql <- attr(result, "sql_query") + result <- .shape_recipe_result(result, "balance_subtype", recipe_label) + recipe_block <- list( + recipe_id = recipe, label = recipe_label, + components = .df_to_row_list(comps) + ) + category_for_prov <- recipe_label + } else { + sql <- .build_verb_sql("balance_annotated", "balance_subtype", + govid, years, category, + ig_view = NULL, subtype_scope = NULL) + result <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) + } # Order matters (matches .verb_spendrev()): per-capita first, so # .attach_real_dollars() deflates the nominal per-capita column into @@ -78,13 +97,14 @@ cog_balances <- function(govid, years, category = NULL, prov <- .build_provenance( verb = "cog_balances", call = call, govid = govid, years = years, - category = category, per_capita = per_capita, + category = category_for_prov, per_capita = per_capita, adjust_to_year = adjust_to_year, result = result, sql = sql, subtype_col = "balance_subtype", basis = basis, basis_note = basis_note, # Neither concept vocabulary applies to a stock. expenditure_concept = NA_character_, - revenue_concept = NA_character_ + revenue_concept = NA_character_, + recipe = recipe_block ) prov$scope$govids_found <- scope$found prov$scope$govids_missing <- scope$missing diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index d0e64ac..824307f 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -218,3 +218,79 @@ test_that("adjust_to_year adds real dollars", { expect_identical(prov$transformations$inflation$base_year, 2020L) }) }) + +# --- recipe = : the wide-era holdings bridge ------------------------------- + +test_that("recipe bridges the wide era into the modern one", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z77_wide") + # .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait, + # not specific to this recipe -- see the money-verb recipe tests, which + # only ever assert on it with expect_equal), so compare numerically rather + # than with expect_identical()'s type-strict comparison. + expect_equal(sort(r$year), c(2011, 2012)) + + # The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE + # and therefore invisible to balance_long. If the recipe path ever starts + # filtering aggregates, a 45-year series silently truncates to five -- + # this is the regression guard for phase_r_harmonization_review.md § 0.2. + codes <- attr(r, "provenance")$codes_summed$observed + expect_true("X40" %in% codes) + expect_true("Z77" %in% codes) + expect_true(all(r$amt_nominal > 0)) + + prov <- attr(r, "provenance") + expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide") + }) +}) + +test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { + skip_if_no_corpus() + with_fixture_corpus({ + # NOTE on years = c(2002, 2011, 2012), which deviates from the brief's + # verbatim c(2011, 2012): .build_series_break_refs() (R/series_breaks.R, + # shared with every verb) gates on + # `break_year BETWEEN min(years) AND max(years)` -- a window over the + # REQUESTED years, not merely over which codes were observed. SB195's + # break_year is 2002, and this fixture has no X40/Z77 partition data for + # any year before 2011, so years = c(2011, 2012) alone can never open a + # window containing 2002 -- confirmed empirically; see + # task-4-report.md for the investigation. There is no 2002 partition in + # the fixture, so adding 2002 to `years` is a pure no-op on the returned + # rows (asserted below) and only widens the break-matching window -- it + # does not change which rows the recipe join reads. Flagged as a + # follow-up candidate: `.build_series_break_refs()`'s window semantics may + # want to treat an in-series precision-change break (fin_code observed, + # break_year <= max(years)) differently from a boundary/rename break, but + # that is shared, cross-verb logic and out of scope for this task. + r <- cog_balances("550000227544", c(2002, 2011, 2012), + recipe = "cash_securities_z77_wide") + expect_equal(sort(r$year), c(2011, 2012)) + refs <- attr(r, "provenance")$series_break_refs + # SB195 sits on fin_code X40; it can only fire where X40 is observed, + # which is exactly the recipe path. + expect_true("SB195" %in% refs) + }) +}) + +test_that("the second holdings bridge works too", { + skip_if_no_corpus() + with_fixture_corpus({ + # X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011 + # and Z78 in 2012, so both legs are exercised. + r <- cog_balances("550000227544", c(2011, 2012), + recipe = "cash_securities_z78_wide") + codes <- attr(r, "provenance")$codes_summed$observed + expect_true(all(c("X41", "Z78") %in% codes)) + expect_equal(sort(r$year), c(2011, 2012)) + }) +}) + +test_that("an unknown recipe id is rejected", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_error(cog_balances("550000227544", 2019, recipe = "no_such_recipe")) + }) +}) From b8189aeb7f302df3a16a465f0aef2b97ac6729f1 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:39:09 -0400 Subject: [PATCH 13/22] docs: caveat 4 needs a year span crossing FY2002, not just a recipe query Task 4's implementer found that SB195 does not surface for a recipe query spanning only 2011-2012. .build_series_break_refs() matches break_year BETWEEN min(years) AND max(years), and SB195's break_year is 2002. That is correct behaviour rather than a gap: a series lying entirely after the book -> market change sits on one consistent basis, so disclosing a break it never crosses would be noise. .build_corpus_break_refs() applies the same rule deliberately. The spec's caveat table overclaimed by omitting the span condition. Corrected. --- specs/2026-08-03-cog-balances-design.md | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md index b79a861..940dc9a 100644 --- a/specs/2026-08-03-cog-balances-design.md +++ b/specs/2026-08-03-cog-balances-design.md @@ -187,7 +187,15 @@ Verified against `series_breaks.csv`, not assumed: | 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` | | 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus | | 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) | -| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41` — but only on a `recipe` query, which is the only path that observes those codes. Asserted in the tests rather than assumed | +| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed | + +On caveat 4's second condition: `.build_series_break_refs()` matches +`break_year BETWEEN min(years) AND max(years)`, so a request spanning only +2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a +series sits entirely after the change, on one consistent basis, and flagging a +break it never crosses would be noise. The same rule is applied deliberately in +`.build_corpus_break_refs()`. An earlier draft of this row omitted the span +condition and overclaimed. `coverage_window` is derived per observed subtype family from the corpus, never hardcoded, so it stays correct as the corpus grows. From 6c5bdb30483324dfbaa9eb5bcef1aa64f93a1ded Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:40:23 -0400 Subject: [PATCH 14/22] test: clarify why 2002 must stay in the SB195 recipe test's year vector --- tests/testthat/test-balances.R | 36 +++++++++++++++++++--------------- 1 file changed, 20 insertions(+), 16 deletions(-) diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 824307f..30ef525 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -249,22 +249,26 @@ test_that("recipe bridges the wide era into the modern one", { test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", { skip_if_no_corpus() with_fixture_corpus({ - # NOTE on years = c(2002, 2011, 2012), which deviates from the brief's - # verbatim c(2011, 2012): .build_series_break_refs() (R/series_breaks.R, - # shared with every verb) gates on - # `break_year BETWEEN min(years) AND max(years)` -- a window over the - # REQUESTED years, not merely over which codes were observed. SB195's - # break_year is 2002, and this fixture has no X40/Z77 partition data for - # any year before 2011, so years = c(2011, 2012) alone can never open a - # window containing 2002 -- confirmed empirically; see - # task-4-report.md for the investigation. There is no 2002 partition in - # the fixture, so adding 2002 to `years` is a pure no-op on the returned - # rows (asserted below) and only widens the break-matching window -- it - # does not change which rows the recipe join reads. Flagged as a - # follow-up candidate: `.build_series_break_refs()`'s window semantics may - # want to treat an in-series precision-change break (fin_code observed, - # break_year <= max(years)) differently from a boundary/rename break, but - # that is shared, cross-verb logic and out of scope for this task. + # 2002 is in the year vector deliberately, and must stay -- do not + # "simplify" this back to c(2011, 2012). + # + # .build_series_break_refs() (R/series_breaks.R, shared with every verb) + # matches breaks with `break_year BETWEEN min(years) AND max(years)`, and + # SB195's break_year is 2002. A c(2011, 2012) span never crosses the + # FY2002 book -> market change -- that whole span sits after it, on one + # consistent basis -- so NOT disclosing SB195 there is correct behaviour, + # not a gap (same reasoning as the "a request that never crosses the + # boundary is not affected by it" comment on .build_corpus_break_refs()). + # + # The property actually worth testing is: a recipe query that observes + # X40 AND spans FY2002 discloses SB195. This fixture has no 2002 + # partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020 + # partitions exist), so including 2002 in `years` widens the + # break-matching window without changing which rows the recipe join + # returns -- verified empirically: r$year below is exactly {2011, 2012} + # whether or not 2002 is in the request (see task-4-report.md). + # Removing 2002 would silently turn this back into the non-crossing case + # above and destroy the test's purpose. r <- cog_balances("550000227544", c(2002, 2011, 2012), recipe = "cash_securities_z77_wide") expect_equal(sort(r$year), c(2011, 2012)) From 82e4face4e4130f325475c1483fedd5eed71d84c Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:48:11 -0400 Subject: [PATCH 15/22] feat: balance_caveats provenance + once-per-session disclosure (#25) --- R/balance_caveats.R | 99 ++++++++++++++++++++++++++++++++++ R/balances.R | 5 ++ R/session.R | 1 + tests/testthat/test-balances.R | 62 +++++++++++++++++++++ 4 files changed, 167 insertions(+) create mode 100644 R/balance_caveats.R diff --git a/R/balance_caveats.R b/R/balance_caveats.R new file mode 100644 index 0000000..711f66f --- /dev/null +++ b/R/balance_caveats.R @@ -0,0 +1,99 @@ +# R/balance_caveats.R +# +# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and +# security holdings. Each one silently invalidates an obvious analysis, so +# they travel in provenance (machine-readable, for cog-api#26) rather than +# living only in prose. +# +# Two of the four are already carried by the code-driven series-break +# builders and are deliberately NOT duplicated here: +# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via +# series_break_refs on the recipe path, the only path that observes those +# codes. +# What remains is the GAAP distinction (a constant) and the coverage windows +# (measured, never hardcoded, so they stay correct as the corpus grows). + +#' Per-subtype observed year extents, plus which requested families are +#' truncated relative to the requested span. +#' @noRd +.balance_caveats <- function(con, codes_observed, years) { + windows <- DBI::dbGetQuery(con, + "SELECT c.balance_subtype AS subtype, + MIN(l.year) AS year_min, + MAX(l.year) AS year_max + FROM balance_long l + JOIN summary_categories c USING (item_code) + WHERE c.balance_subtype IS NOT NULL + GROUP BY 1 + ORDER BY 1" + ) + + observed_subtypes <- if (length(codes_observed) == 0L) { + character(0) + } else { + DBI::dbGetQuery(con, sprintf( + "SELECT DISTINCT balance_subtype FROM summary_categories + WHERE item_code IN (%s) AND balance_subtype IS NOT NULL", + .sql_lit_chr(codes_observed) + ))$balance_subtype + } + + cw <- stats::setNames( + lapply(seq_len(nrow(windows)), + function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), + windows$subtype + ) + + # A family is "truncated" when the caller asked for years outside the span + # that family actually covers -- the FY2016 employee-retirement termination + # and the FY2021 end of the W family are both this shape. + truncated <- character(0) + if (length(years) > 0L) { + for (s in observed_subtypes) { + w <- cw[[s]] + if (is.null(w)) next + if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s) + } + } + + list( + not_gaap = TRUE, + not_gaap_note = paste0( + "Census holdings are gross -- no liabilities are netted -- and are NOT ", + "GAAP fund balance. A reserve ratio built from them overstates what is ", + "actually available." + ), + coverage_window = cw, + truncated = sort(unique(truncated)) + ) +} + +#' TRUE the first time `key` is seen this session, FALSE thereafter. +#' Reset by cog_close(). +#' @noRd +.balance_caveat_once <- function(key) { + seen <- .uscogdata_env$balance_caveats_shown + if (is.null(seen)) seen <- character(0) + if (key %in% seen) return(FALSE) + .uscogdata_env$balance_caveats_shown <- c(seen, key) + TRUE +} + +#' Emit at most one message per caveat class per session. +#' @noRd +.emit_balance_caveats <- function(caveats) { + if (.balance_caveat_once("not_gaap")) { + cli::cli_inform(c( + "!" = "Census holdings are gross and are {.strong not} GAAP fund balance.", + "i" = "No liabilities are netted; a reserve ratio built from them overstates available funds." + )) + } + if (length(caveats$truncated) > 0L && + .balance_caveat_once("coverage_window")) { + cli::cli_inform(c( + "!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.", + "i" = "See {.code provenance$balance_caveats$coverage_window}." + )) + } + invisible(NULL) +} diff --git a/R/balances.R b/R/balances.R index a2d0d9c..7e8c7f5 100644 --- a/R/balances.R +++ b/R/balances.R @@ -109,6 +109,11 @@ cog_balances <- function(govid, years, category = NULL, prov$scope$govids_found <- scope$found prov$scope$govids_missing <- scope$missing + prov$balance_caveats <- .balance_caveats( + con, prov$codes_summed$observed, years + ) + .emit_balance_caveats(prov$balance_caveats) + attr(result, "provenance") <- prov result } diff --git a/R/session.R b/R/session.R index 2a1d7d0..8f55887 100644 --- a/R/session.R +++ b/R/session.R @@ -95,4 +95,5 @@ cog_close <- function() { } .uscogdata_env$con <- NULL .uscogdata_env$manifest <- NULL + .uscogdata_env$balance_caveats_shown <- NULL } diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 30ef525..1438a85 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -298,3 +298,65 @@ test_that("an unknown recipe id is rejected", { expect_error(cog_balances("550000227544", 2019, recipe = "no_such_recipe")) }) }) + +# --- balance_caveats: GAAP disclosure + measured coverage windows ---------- + +test_that("balance_caveats is always present and flags the GAAP distinction", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2019) + cav <- attr(r, "provenance")$balance_caveats + expect_false(is.null(cav)) + expect_true(cav$not_gaap) + }) +}) + +test_that("coverage_window is computed from the corpus, not hardcoded", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020)) + cav <- attr(r, "provenance")$balance_caveats + + # Read the "general" family's true year extent independently, via a + # fresh DuckDB connection against the raw parquet files (never through + # balance_long/.balance_caveats() itself, and never via arrow -- this + # package reads parquet through DuckDB only, see CLAUDE.md). Replicates + # the same predicates 26-balance_long.sql applies (category_type = + # 'balance', NOT is_aggregate) so this is a faithful, independent + # measurement rather than a re-statement of the view under test. + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet") + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + obs <- DBI::dbGetQuery(con2, sprintf( + "SELECT MIN(l.year) AS y0, MAX(l.year) AS y1 + FROM read_parquet(%s, hive_partitioning = true) l + JOIN read_parquet(%s) c USING (item_code) + WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate", + uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path) + )) + + expect_identical(as.integer(cav$coverage_window$general), + c(as.integer(obs$y0), as.integer(obs$y1))) + }) +}) + +test_that("a request past a family's coverage window is flagged", { + skip_if_no_corpus() + with_fixture_corpus({ + # The employee_retirement family (X21/X30/X47/Z77/Z78) is corpus-wide + # truncated relative to 2019 in this fixture; 2012 observes it, 2019 does + # not, so the requested span extends past what it actually covers. + r <- cog_balances("550000227544", c(2012, 2019)) + cav <- attr(r, "provenance")$balance_caveats + expect_true("employee_retirement" %in% cav$truncated) + }) +}) + +test_that("the caveat message fires once per session", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_message(cog_balances("550000227544", 2019), "not.*GAAP") + expect_no_message(cog_balances("550000227544", 2020)) + }) +}) From 724b6bd58b221d2b9da8f153e3952840a2e32a6f Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 10:50:46 -0400 Subject: [PATCH 16/22] docs: disambiguate live-corpus vs fixture year claim in balances test comment (#25) --- tests/testthat/test-balances.R | 12 +++++++++--- 1 file changed, 9 insertions(+), 3 deletions(-) diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index 1438a85..b0c850d 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -344,9 +344,15 @@ test_that("coverage_window is computed from the corpus, not hardcoded", { test_that("a request past a family's coverage window is flagged", { skip_if_no_corpus() with_fixture_corpus({ - # The employee_retirement family (X21/X30/X47/Z77/Z78) is corpus-wide - # truncated relative to 2019 in this fixture; 2012 observes it, 2019 does - # not, so the requested span extends past what it actually covers. + # employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in + # the LIVE corpus -- Census moved employee retirement reporting to the + # Annual Survey of Public Pensions after that year. This bundled FIXTURE + # doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present), + # so the family's *observed* max here is 2012, not 2016. Either way the + # requested span (2012, 2019) reaches past what the family covers in + # THIS corpus, which is what makes .balance_caveats() flag it -- the + # assertion below is about the fixture's measured window, not the FY2016 + # live-corpus cutoff. r <- cog_balances("550000227544", c(2012, 2019)) cav <- attr(r, "provenance")$balance_caveats expect_true("employee_retirement" %in% cav$truncated) From b03f095e496e9c7169e60a1ffb11b1a77ae04f27 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:02:16 -0400 Subject: [PATCH 17/22] docs: document cog_balances() and correct stale CLAUDE.md claims (#25) Adds the NEWS entry, a Financial data pkgdown reference section (none existed for cog_spending/cog_revenue), and corrects CLAUDE.md's SQL-layer claim, view count, test count and fixture-year description against measured values. Also documents balance_caveats in inst/schemas/provenance-v1.json (test-first: added a schema-documentation test to test-balances.R, confirmed it failed, then fixed the schema) and fleshes out cog_balances()'s @return roxygen to enumerate its conditional columns, regenerating man/cog_balances.Rd. --- CLAUDE.md | 38 +++++++++++++++++++++++++-------- NEWS.md | 13 +++++++++++ R/balances.R | 15 +++++++++++-- _pkgdown.yml | 6 ++++++ inst/schemas/provenance-v1.json | 10 +++++++++ man/cog_balances.Rd | 15 +++++++++++-- tests/testthat/test-balances.R | 8 +++++++ 7 files changed, 92 insertions(+), 13 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index d739dd8..51b0ee7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -28,8 +28,23 @@ USCOGDATA_URL (local path or https://) - `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()` - `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache) - `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`) -- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated` +- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order + (`10-` through `46-`): the `*_long` layer (`long`, `spending_long`, + `revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of + `spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer + (`spending_annotated`, `revenue_annotated`, `ig_annotated`, + `balance_annotated`, plus `_harmonized` variants of `spending_annotated`/ + `revenue_annotated`/`ig_annotated`), and metadata views + (`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`, + `harmonization_map`, `harmonization_recipes`, `series_breaks_pq`, + `representation`, `code_set`) - `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()` +- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a + **stock** (a balance at a point in time) rather than a **flow** (activity + over a fiscal year), so it does NOT route through `.verb_spendrev()` and has + no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments. + `R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross + disclosure + measured per-subtype coverage windows). - `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer) - `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()` - `R/search.R` — `cog_gov_search()` (name pattern, state, type filters) @@ -49,18 +64,19 @@ Any value without `://` is treated as a local path by `.is_local_path()` and rea **Version:** 0.1.0 (pre-release) **Branch:** `main`, commit `d65e9fe` -**Tests:** 181 PASS / 0 FAIL / 0 SKIP +**Tests:** 763 PASS / 0 FAIL / 0 SKIP (measured `testthat::test_local()`, 2026-08-03) **CI:** Gitea Actions green (`.gitea/workflows/ci.yml`) ### Completed (Tasks 2.1–2.7) -All 8 exported verbs implemented and tested: -`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`, -`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`, -plus `cog_categories`. +All 10 exported verbs implemented and tested: +`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`, +`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`, +`cog_gov_search`, `cog_mirror`, plus `cog_categories`. -Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years -2019+2020, all 50 states). Tests run fully offline — no credentials needed. +Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years +2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`, +2026-08-03; all 50 states). Tests run fully offline — no credentials needed. ### Remaining to v0.1 release @@ -99,6 +115,10 @@ devtools::test() - All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()` - Return value is always a `tbl_df` with a `provenance` attribute - govid inputs always go through `.coerce_govid_input()` (accepts character or data frame) -- SQL lives in `inst/sql/` — never inline SQL strings in R files +- SQL has two layers. **View definitions** live in `inst/sql/` and are + registered by `.register_views()`, which globs the directory in sorted order + and substitutes `{url}`. **Query construction** is inline `sprintf()` in R + (`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as + a numbered `.sql` file; build a query in R. - No arrow dependency — DuckDB reads parquet natively - `withr` is a Suggests-only dep; only used in tests diff --git a/NEWS.md b/NEWS.md index 97c7430..88e8de2 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,5 +1,18 @@ # uscogdata 0.1.0 (development) +## New: `cog_balances()` for cash-and-security holdings + +* New `cog_balances()` exposes the 14 cash-and-security holding codes + (`category_type = "balance"`): fund balances, retirement system holdings and + insurance trust balances (#25). Holdings are a stock, not a flow, so the verb + has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and + no `subtype` argument either -- for holdings, `category` is a strict + coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly + the `general` family (`W01`/`W31`/`W61`). +* `cog_balances()` results carry `provenance$balance_caveats`, recording that + Census holdings are gross rather than GAAP fund balance, and the measured + coverage window of each subtype family. + ## Multi-government aggregates now disclose their reporting coverage * The Census of Governments is a **complete census only in years ending in 2 diff --git a/R/balances.R b/R/balances.R index 7e8c7f5..f521a7f 100644 --- a/R/balances.R +++ b/R/balances.R @@ -41,8 +41,19 @@ #' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the #' wide era to the modern one. #' -#' @return A `tbl_df` with a `provenance` attribute. Amounts are full US -#' dollars. +#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, +#' `balance_subtype`, `category`, `amt_nominal`, optional +#' `amt_per_capita_nominal` and `pop_source` (when `per_capita = TRUE`), +#' optional `amt_real` and `amt_per_capita_real` (when `adjust_to_year` is +#' set), `codes_included`, `aggregate_fallback`, `notes`. Amounts are full +#' US dollars. +#' +#' Carries a `provenance` attribute matching +#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports +#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured per-subtype year +#' extents) and `truncated` (subtypes whose coverage falls short of the +#' requested years). `expenditure_concept`/`revenue_concept` are `NA` -- +#' holdings are a stock, not a flow, so neither concept vocabulary applies. #' @export cog_balances <- function(govid, years, category = NULL, per_capita = FALSE, adjust_to_year = NULL, diff --git a/_pkgdown.yml b/_pkgdown.yml index 2022548..1541211 100644 --- a/_pkgdown.yml +++ b/_pkgdown.yml @@ -3,6 +3,12 @@ template: bootstrap: 5 reference: + - title: Financial data + desc: Spending, revenue and balance-sheet holdings for one or more governments. + contents: + - cog_spending + - cog_revenue + - cog_balances - title: Search & basket desc: Resolve place names into canonical govids. contents: diff --git a/inst/schemas/provenance-v1.json b/inst/schemas/provenance-v1.json index 779fad6..a430cd5 100644 --- a/inst/schemas/provenance-v1.json +++ b/inst/schemas/provenance-v1.json @@ -53,6 +53,16 @@ "items": { "type": "string" }, "description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series." }, + "balance_caveats": { + "type": ["object", "null"], + "description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps each observed balance_subtype to its measured [min year, max year] in the mounted corpus (never hardcoded). `truncated` lists the subtypes whose coverage_window does not fully span the requested years.", + "properties": { + "not_gaap": { "type": "boolean" }, + "not_gaap_note": { "type": "string" }, + "coverage_window": { "type": "object" }, + "truncated": { "type": "array", "items": { "type": "string" } } + } + }, "manifest": { "type": "object" }, "sql_query": { "type": "string" } } diff --git a/man/cog_balances.Rd b/man/cog_balances.Rd index 81323a0..f3e40b9 100644 --- a/man/cog_balances.Rd +++ b/man/cog_balances.Rd @@ -46,8 +46,19 @@ and raw space are identical for holdings. Reported in wide era to the modern one.} } \value{ -A `tbl_df` with a `provenance` attribute. Amounts are full US - dollars. +Tibble with columns `year`, `canonical_govid`, `gov_name`, + `balance_subtype`, `category`, `amt_nominal`, optional + `amt_per_capita_nominal` and `pop_source` (when `per_capita = TRUE`), + optional `amt_real` and `amt_per_capita_real` (when `adjust_to_year` is + set), `codes_included`, `aggregate_fallback`, `notes`. Amounts are full + US dollars. + + Carries a `provenance` attribute matching + `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports + `not_gaap`, `not_gaap_note`, `coverage_window` (measured per-subtype year + extents) and `truncated` (subtypes whose coverage falls short of the + requested years). `expenditure_concept`/`revenue_concept` are `NA` -- + holdings are a stock, not a flow, so neither concept vocabulary applies. } \description{ Returns Census cash-and-security holdings (`category_type = "balance"`): diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index b0c850d..fe609ee 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -359,6 +359,14 @@ test_that("a request past a family's coverage window is flagged", { }) }) +test_that("the provenance schema documents balance_caveats", { + sch <- jsonlite::fromJSON( + system.file("schemas", "provenance-v1.json", package = "uscogdata"), + simplifyVector = FALSE + ) + expect_true("balance_caveats" %in% names(sch$properties)) +}) + test_that("the caveat message fires once per session", { skip_if_no_corpus() with_fixture_corpus({ From 225cd609683ad520ea3af22ed65fff0084a472c5 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:10:46 -0400 Subject: [PATCH 18/22] docs: fix stale test count and phantom notes column in cog_balances docs (#25) Re-measured CLAUDE.md's test count on the final tree (764, not 763 -- the earlier number predated the balance_caveats schema test). Removed notes from cog_balances()'s @return block: it was copied from cog_spending()'s @return style without checking cog_balances() never calls .verb_spendrev(), the only place that sets notes. Verified the remaining documented columns against colnames() observed across every argument combination (bare, per_capita, adjust_to_year, both, recipe, category filter). --- CLAUDE.md | 2 +- R/balances.R | 8 ++++---- man/cog_balances.Rd | 8 ++++---- 3 files changed, 9 insertions(+), 9 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 51b0ee7..6a88b54 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -64,7 +64,7 @@ Any value without `://` is treated as a local path by `.is_local_path()` and rea **Version:** 0.1.0 (pre-release) **Branch:** `main`, commit `d65e9fe` -**Tests:** 763 PASS / 0 FAIL / 0 SKIP (measured `testthat::test_local()`, 2026-08-03) +**Tests:** 764 PASS / 0 FAIL / 0 SKIP (measured `testthat::test_local()`, 2026-08-03, on the tree including the balance_caveats schema test) **CI:** Gitea Actions green (`.gitea/workflows/ci.yml`) ### Completed (Tasks 2.1–2.7) diff --git a/R/balances.R b/R/balances.R index f521a7f..5f755f3 100644 --- a/R/balances.R +++ b/R/balances.R @@ -42,10 +42,10 @@ #' wide era to the modern one. #' #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, -#' `balance_subtype`, `category`, `amt_nominal`, optional -#' `amt_per_capita_nominal` and `pop_source` (when `per_capita = TRUE`), -#' optional `amt_real` and `amt_per_capita_real` (when `adjust_to_year` is -#' set), `codes_included`, `aggregate_fallback`, `notes`. Amounts are full +#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`, +#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and +#' `pop_source` (when `per_capita = TRUE`), and optional `amt_real` and +#' `amt_per_capita_real` (when `adjust_to_year` is set). Amounts are full #' US dollars. #' #' Carries a `provenance` attribute matching diff --git a/man/cog_balances.Rd b/man/cog_balances.Rd index f3e40b9..9b96b97 100644 --- a/man/cog_balances.Rd +++ b/man/cog_balances.Rd @@ -47,10 +47,10 @@ wide era to the modern one.} } \value{ Tibble with columns `year`, `canonical_govid`, `gov_name`, - `balance_subtype`, `category`, `amt_nominal`, optional - `amt_per_capita_nominal` and `pop_source` (when `per_capita = TRUE`), - optional `amt_real` and `amt_per_capita_real` (when `adjust_to_year` is - set), `codes_included`, `aggregate_fallback`, `notes`. Amounts are full + `balance_subtype`, `category`, `amt_nominal`, `codes_included`, + `aggregate_fallback`, plus optional `amt_per_capita_nominal` and + `pop_source` (when `per_capita = TRUE`), and optional `amt_real` and + `amt_per_capita_real` (when `adjust_to_year` is set). Amounts are full US dollars. Carries a `provenance` attribute matching From 22c2478634902092dd082d4d944f78e35e2dccee Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:36:18 -0400 Subject: [PATCH 19/22] fix(balances): validate the full signature, surface caveats in cog_explain, memoise coverage windows (#25) Final-review findings F-1, F-2, F-6, F-8 (plus the F-9 @return reword, which shares R/balances.R). F-2: .validate_balance_inputs() checked 2 of cog_balances()' 7 arguments. years = integer(0) leaked a raw DuckDB 'Parser Error ... AND year IN ()' with the generated SQL echoed back; govid = character(0) and a non-character category returned 0 rows with no error at all; recipe = c("a","b") threw 'the condition has length > 1' from inside .validate_recipe_id(). Replaced with a call to the money verbs' own .validate_verb_inputs() (R/spending.R), which validates the exact superset needed. Deleted the local copy rather than extending it -- two validators is how they drift. Placed AFTER .coerce_govid_input(), because .validate_verb_inputs() asserts is.character(govid) and a data-frame govid is not unwrapped before that. This is helper reuse of the same kind as .build_verb_sql()/.attach_per_capita(); the verb still does NOT route through .verb_spendrev(). F-1: falls out of F-2 for free -- the recipe/category mutual-exclusivity guard lives inside .validate_verb_inputs(). Previously recipe silently discarded category AND overwrote provenance$category with the recipe label, so a caller asking for Fund Balances got X40/Z77 insurance-trust holdings with no trace of the dropped filter. F-6: cog_explain() rendered every provenance caveat block except balance_caveats. Since .emit_balance_caveats() fires at most once per session -- and is routinely consumed by a suppressMessages() call or an unread knitr chunk -- cog_explain() is the only surface left for a caller who deliberately audits the result. Added a 'Holdings caveats' section guarded on !is.null(prov$balance_caveats). Also relabels the cosmetic 'Concept: NA' line on balance results as 'not applicable (holdings are a stock, not a flow)'. F-8: the coverage-window query has no govid and no year predicate -- its answer depends only on the mounted corpus -- yet it scanned all of balance_long on every call (35% of verb runtime on the fixture, and a per-request throughput ceiling for cog-api#26). Memoised in .uscogdata_env$balance_coverage_windows, invalidated by cog_close(), the same pattern as .uscogdata_env$manifest. --- R/balance_caveats.R | 55 ++++++++++++++++++++++++++++++++------------- R/balances.R | 44 +++++++++++++++++------------------- R/explain.R | 30 +++++++++++++++++++++++++ R/session.R | 2 ++ man/cog_balances.Rd | 13 ++++++----- 5 files changed, 99 insertions(+), 45 deletions(-) diff --git a/R/balance_caveats.R b/R/balance_caveats.R index 711f66f..dc25401 100644 --- a/R/balance_caveats.R +++ b/R/balance_caveats.R @@ -17,16 +17,7 @@ #' truncated relative to the requested span. #' @noRd .balance_caveats <- function(con, codes_observed, years) { - windows <- DBI::dbGetQuery(con, - "SELECT c.balance_subtype AS subtype, - MIN(l.year) AS year_min, - MAX(l.year) AS year_max - FROM balance_long l - JOIN summary_categories c USING (item_code) - WHERE c.balance_subtype IS NOT NULL - GROUP BY 1 - ORDER BY 1" - ) + cw <- .balance_coverage_windows(con) observed_subtypes <- if (length(codes_observed) == 0L) { character(0) @@ -38,12 +29,6 @@ ))$balance_subtype } - cw <- stats::setNames( - lapply(seq_len(nrow(windows)), - function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), - windows$subtype - ) - # A family is "truncated" when the caller asked for years outside the span # that family actually covers -- the FY2016 employee-retirement termination # and the FY2021 end of the W family are both this shape. @@ -68,6 +53,44 @@ ) } +#' Per-subtype [min year, max year] extents for EVERY balance subtype in the +#' mounted corpus, memoised for the session. +#' +#' The query carries no govid and no year predicate -- its answer is a property +#' of the mounted corpus alone and cannot change between calls -- but it scans +#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime +#' on the bundled fixture and would be a per-request throughput ceiling once +#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and +#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`. +#' +#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking +#' "is there a family I missed?" needs every window. The observed-scoped field +#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json. +#' @noRd +.balance_coverage_windows <- function(con) { + cached <- .uscogdata_env$balance_coverage_windows + if (!is.null(cached)) return(cached) + + windows <- DBI::dbGetQuery(con, + "SELECT c.balance_subtype AS subtype, + MIN(l.year) AS year_min, + MAX(l.year) AS year_max + FROM balance_long l + JOIN summary_categories c USING (item_code) + WHERE c.balance_subtype IS NOT NULL + GROUP BY 1 + ORDER BY 1" + ) + + cw <- stats::setNames( + lapply(seq_len(nrow(windows)), + function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))), + windows$subtype + ) + .uscogdata_env$balance_coverage_windows <- cw + cw +} + #' TRUE the first time `key` is seen this session, FALSE thereafter. #' Reset by cog_close(). #' @noRd diff --git a/R/balances.R b/R/balances.R index 5f755f3..4a9567d 100644 --- a/R/balances.R +++ b/R/balances.R @@ -44,14 +44,17 @@ #' @return Tibble with columns `year`, `canonical_govid`, `gov_name`, #' `balance_subtype`, `category`, `amt_nominal`, `codes_included`, #' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and -#' `pop_source` (when `per_capita = TRUE`), and optional `amt_real` and -#' `amt_per_capita_real` (when `adjust_to_year` is set). Amounts are full -#' US dollars. +#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when +#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when +#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no +#' nominal per-capita column to deflate otherwise). Amounts are full US +#' dollars. #' #' Carries a `provenance` attribute matching #' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports -#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured per-subtype year -#' extents) and `truncated` (subtypes whose coverage falls short of the +#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for +#' every balance subtype in the mounted corpus, not only the observed ones) +#' and `truncated` (the observed subtypes whose coverage falls short of the #' requested years). `expenditure_concept`/`revenue_concept` are `NA` -- #' holdings are a stock, not a flow, so neither concept vocabulary applies. #' @export @@ -60,8 +63,19 @@ cog_balances <- function(govid, years, category = NULL, basis = c("harmonized", "raw"), recipe = NULL) { call <- match.call() basis <- match.arg(basis, c("harmonized", "raw")) - .validate_balance_inputs(per_capita, adjust_to_year) + # Coerce FIRST, validate second: .validate_verb_inputs() asserts + # is.character(govid), and a data-frame govid (cog_gov_search() output) has + # not been unwrapped yet at this point. govid <- .coerce_govid_input(govid) + # The money verbs' validator, reused rather than re-implemented (R/spending.R). + # It covers the exact superset cog_balances() needs -- including the + # recipe/category mutual-exclusivity guard -- so a second local copy would + # only be a place for the two to drift apart. This is the same kind of + # helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT + # route the verb through .verb_spendrev(), which stays deliberately unused + # here because its flow vocabulary is meaningless for a stock. + .validate_verb_inputs(govid, years, category, per_capita, adjust_to_year, + recipe) years <- as.integer(years) if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year) @@ -129,24 +143,6 @@ cog_balances <- function(govid, years, category = NULL, result } -#' Cheap type validation for the two arguments cog_balances() shares with the -#' money verbs. Mirrors the per_capita/adjust_to_year checks in -#' .validate_verb_inputs() (R/spending.R) -- category/recipe validation is -#' deliberately out of scope here (uscogdata#25 Task 3 review note). -#' @noRd -.validate_balance_inputs <- function(per_capita, adjust_to_year) { - if (!is.logical(per_capita) || length(per_capita) != 1L) { - cli::cli_abort("`per_capita` must be a length-1 logical.") - } - if (!is.null(adjust_to_year)) { - if (!(is.integer(adjust_to_year) || is.numeric(adjust_to_year)) || - length(adjust_to_year) != 1L) { - cli::cli_abort("`adjust_to_year` must be NULL or a length-1 integer.") - } - } - invisible(TRUE) -} - #' Abort unless the mounted corpus classifies balance codes. #' #' `balance_subtype` arrived with cog_pipeline #76/#77 without a diff --git a/R/explain.R b/R/explain.R index 5309e97..37380c4 100644 --- a/R/explain.R +++ b/R/explain.R @@ -68,6 +68,11 @@ cog_explain <- function(result, format = c("print", "list")) { if (!is.null(prov$revenue_concept)) { cli::cli_text("Concept: {prov$revenue_concept} revenue") } + } else if (identical(prov$verb, "cog_balances")) { + # Both concept fields are deliberately NA here (a stock has no flow + # concept). Printing the raw NA reads as a missing value rather than an + # intentional one, so say what it means instead. + cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)") } else if (!is.null(prov$expenditure_concept)) { concept_note <- if (!is.null(prov$expenditure_concept_note) && !is.na(prov$expenditure_concept_note)) { @@ -174,6 +179,31 @@ cog_explain <- function(result, format = c("print", "list")) { cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs)) } + # Balance results only (NULL on money-verb provenance, so they are + # unaffected). This is the ONLY on-demand surface for the GAAP disclosure: + # .emit_balance_caveats() fires at most once per session, and is routinely + # consumed by a suppressMessages() call or by a knitted chunk nobody reads, + # so a caller who deliberately audits a result with cog_explain() must still + # be told. + bc <- prov$balance_caveats + if (!is.null(bc)) { + cli::cli_h2("Holdings caveats") + if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note) + if (length(bc$truncated) > 0L) { + cli::cli_text( + "Requested years extend beyond what these families actually cover:" + ) + cli::cli_ul(vapply(bc$truncated, function(s) { + w <- bc$coverage_window[[s]] + if (length(w) == 2L) { + sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2]) + } else { + s + } + }, character(1))) + } + } + cli::cli_h2("Transformations") uc <- prov$transformations$units_conversion if (isTRUE(uc$applied)) { diff --git a/R/session.R b/R/session.R index 8f55887..ef1c28c 100644 --- a/R/session.R +++ b/R/session.R @@ -96,4 +96,6 @@ cog_close <- function() { .uscogdata_env$con <- NULL .uscogdata_env$manifest <- NULL .uscogdata_env$balance_caveats_shown <- NULL + # Memoised corpus-constant; a different corpus may be mounted next. + .uscogdata_env$balance_coverage_windows <- NULL } diff --git a/man/cog_balances.Rd b/man/cog_balances.Rd index 9b96b97..f594eba 100644 --- a/man/cog_balances.Rd +++ b/man/cog_balances.Rd @@ -49,14 +49,17 @@ wide era to the modern one.} Tibble with columns `year`, `canonical_govid`, `gov_name`, `balance_subtype`, `category`, `amt_nominal`, `codes_included`, `aggregate_fallback`, plus optional `amt_per_capita_nominal` and - `pop_source` (when `per_capita = TRUE`), and optional `amt_real` and - `amt_per_capita_real` (when `adjust_to_year` is set). Amounts are full - US dollars. + `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when + `adjust_to_year` is set), and optional `amt_per_capita_real` (only when + **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no + nominal per-capita column to deflate otherwise). Amounts are full US + dollars. Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports - `not_gaap`, `not_gaap_note`, `coverage_window` (measured per-subtype year - extents) and `truncated` (subtypes whose coverage falls short of the + `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for + every balance subtype in the mounted corpus, not only the observed ones) + and `truncated` (the observed subtypes whose coverage falls short of the requested years). `expenditure_concept`/`revenue_concept` are `NA` -- holdings are a stock, not a flow, so neither concept vocabulary applies. } From fde62eb6ccbdc5ec107c29a32b94622926b669d5 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:36:32 -0400 Subject: [PATCH 20/22] test(balances): pin the behaviours the final review found untested or weakly asserted (#25) Findings F-1..F-8. Every assertion below was verified to FAIL before its fix (or under mutation, where the behaviour already worked) and pass after. F-3: 'an unknown recipe id is rejected' used a bare expect_error(). Deleting .validate_recipe_id() leaves .recipe_components() returning 0 rows and comps$label[[1]] throwing 'subscript out of bounds' -- still an error, so the test passed on the regression while the user lost the curated message. Now asserts class = 'uscogdata_unknown_recipe'. Mutation-checked. F-4: no test ever set per_capita and adjust_to_year together, so the load-bearing ordering comment at R/balances.R was unverified. Reversing those two calls silently drops amt_per_capita_real (.attach_real_dollars() no-ops when amt_per_capita_nominal does not exist yet). New test asserts presence AND that the per-capita column is deflated by the same factor as the level column; mutation-checked by reversing the order (2 failures). F-5: the spec's 'Gating' requirement had no test -- nothing ever called cog_balances() on a corpus without balance_subtype. Extended the existing with_corpus_missing_balance_subtype() block to assert class = 'uscogdata_no_balance_support'; mutation-checked by dropping the guard. F-1/F-2: added mutual-exclusivity and four-argument validation tests, each pinned to the message or class (all four inputs already produced *some* error or *some* quiet wrong answer, so bare expect_error() was useless here). Plus an ordering guard: a data-frame govid must still work, which is what fails if validation is put before .coerce_govid_input(). F-6: asserts on the RENDERED cog_explain() text (both streams -- cli routes through conditions that land on stderr), with a negative case proving money-verb output is unaffected and that the capture is not vacuous. F-7: pins that coverage_window is corpus-scoped while truncated is query-scoped; mutation-checked by scoping the windows to observed subtypes. F-8: pins the memo slot is populated on first call and cleared by cog_close(). --- tests/testthat/test-balances.R | 165 ++++++++++++++++++++++++++++++++- 1 file changed, 164 insertions(+), 1 deletion(-) diff --git a/tests/testthat/test-balances.R b/tests/testthat/test-balances.R index fe609ee..86d5ba4 100644 --- a/tests/testthat/test-balances.R +++ b/tests/testthat/test-balances.R @@ -95,6 +95,16 @@ test_that("balance views are skipped on a corpus without balance_subtype", { # Registration must SKIP them, not error -- an older corpus stays usable. expect_false(any(c("balance_long", "balance_annotated") %in% views)) expect_true("revenue_long" %in% views) + + # ...and calling the verb on such a corpus must hit + # .require_balance_support()'s curated abort (spec § Testing: "Gating"), + # not a DuckDB binder error naming a view that was never registered. + # Asserted on the CLASS: removing the guard still errors, so a bare + # expect_error() would pass on the regression. + expect_error( + cog_balances("550000227544", 2019), + class = "uscogdata_no_balance_support" + ) }) }) @@ -219,6 +229,73 @@ test_that("adjust_to_year adds real dollars", { }) }) +test_that("per_capita and adjust_to_year compose", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- cog_balances("550000227544", 2012, category = "Fund Balances", + per_capita = TRUE, adjust_to_year = 2020) + expect_true("amt_per_capita_real" %in% names(r)) + # The per-capita column must be deflated by the SAME factor as the level + # column -- this is what the ordering at R/balances.R:101-103 guarantees. + # .attach_real_dollars() silently no-ops on the per-capita leg when + # amt_per_capita_nominal does not exist yet (R/spending.R:664), so + # reversing those two calls drops this column with no error at all. + expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal, + r$amt_real / r$amt_nominal, tolerance = 1e-8) + + # And the documented condition is a conjunction: adjust_to_year ALONE + # must not produce amt_per_capita_real (pins the @return wording). + r2 <- cog_balances("550000227544", 2012, category = "Fund Balances", + adjust_to_year = 2020) + expect_true("amt_real" %in% names(r2)) + expect_false("amt_per_capita_real" %in% names(r2)) + }) +}) + +# --- input validation ------------------------------------------------------ + +test_that("cog_balances validates its inputs like the money verbs", { + skip_if_no_corpus() + with_fixture_corpus({ + G <- "550000227544" + # Pinned to the message, not bare expect_error(): every one of these + # already produces *some* error or *some* quiet wrong answer today -- + # years = integer(0) leaks `Parser Error ... AND year IN ()` with the + # generated SQL, recipe = c("a","b") throws "the condition has length > 1", + # and the govid/category cases return 0 rows with no error at all. + expect_error(cog_balances(G, integer(0)), "non-empty integer vector") + expect_error(cog_balances(character(0), 2019), "non-empty character vector") + expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL") + expect_error(cog_balances(G, 2019, recipe = c("a", "b")), + "length-1 character string") + }) +}) + +test_that("validation runs after govid coercion, so a data frame still works", { + skip_if_no_corpus() + with_fixture_corpus({ + # .validate_verb_inputs() asserts is.character(govid); it must therefore + # run AFTER .coerce_govid_input(), never before, or the documented + # data-frame input (cog_gov_search() output) would abort. + df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE) + r <- suppressMessages(cog_balances(df, 2019)) + expect_true(nrow(r) > 0L) + expect_identical(unique(r$canonical_govid), "550000227544") + }) +}) + +test_that("recipe and category are mutually exclusive", { + skip_if_no_corpus() + with_fixture_corpus({ + expect_error( + cog_balances("550000227544", c(2011, 2012), + category = "Fund Balances", + recipe = "cash_securities_z77_wide"), + class = "uscogdata_recipe_category_conflict" + ) + }) +}) + # --- recipe = : the wide-era holdings bridge ------------------------------- test_that("recipe bridges the wide era into the modern one", { @@ -295,7 +372,15 @@ test_that("the second holdings bridge works too", { test_that("an unknown recipe id is rejected", { skip_if_no_corpus() with_fixture_corpus({ - expect_error(cog_balances("550000227544", 2019, recipe = "no_such_recipe")) + # Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84). + # Without it the test is non-discriminating: deleting the validation call + # leaves .recipe_components() returning 0 rows and comps$label[[1]] + # throwing "subscript out of bounds", which a bare expect_error() accepts + # while the user loses the curated "valid recipe ids are ..." message. + expect_error( + cog_balances("550000227544", 2019, recipe = "no_such_recipe"), + class = "uscogdata_unknown_recipe" + ) }) }) @@ -341,6 +426,51 @@ test_that("coverage_window is computed from the corpus, not hardcoded", { }) }) +test_that("coverage_window covers every corpus subtype, not just observed ones", { + skip_if_no_corpus() + with_fixture_corpus({ + # Deliberate contract (provenance-v1.json): the window block is corpus- + # scoped so a caller can ask "is there a family I missed?", while + # `truncated` is the observed-scoped field. A single-category query must + # therefore still report every balance family in the mounted corpus. + r <- cog_balances("550000227544", 2019, category = "Fund Balances") + expect_identical(unique(r$balance_subtype), "general") + + con2 <- DBI::dbConnect(duckdb::duckdb()) + on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE) + cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet") + all_subtypes <- DBI::dbGetQuery(con2, sprintf( + "SELECT DISTINCT balance_subtype FROM read_parquet(%s) + WHERE balance_subtype IS NOT NULL", + uscogdata:::.sql_lit_chr(cats_path) + ))$balance_subtype + + cav <- attr(r, "provenance")$balance_caveats + expect_setequal(names(cav$coverage_window), all_subtypes) + expect_true(length(all_subtypes) > 1L) + # ...while `truncated` stays scoped to what this query actually observed. + expect_true(all(cav$truncated %in% unique(r$balance_subtype))) + }) +}) + +test_that("the corpus-constant coverage windows are memoised per session", { + skip_if_no_corpus() + with_fixture_corpus({ + # The windows query has no govid/year predicate: its answer depends only + # on which corpus is mounted, so re-running the full balance_long scan on + # every call is pure waste (35% of verb runtime on the fixture). Same + # memoise-and-invalidate pattern as .uscogdata_env$manifest. + expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows) + suppressMessages(cog_balances("550000227544", 2019)) + memo <- uscogdata:::.uscogdata_env$balance_coverage_windows + expect_false(is.null(memo)) + expect_true("general" %in% names(memo)) + + uscogdata:::cog_close() + expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows) + }) +}) + test_that("a request past a family's coverage window is flagged", { skip_if_no_corpus() with_fixture_corpus({ @@ -367,6 +497,39 @@ test_that("the provenance schema documents balance_caveats", { expect_true("balance_caveats" %in% names(sch$properties)) }) +test_that("cog_explain surfaces the balance caveats", { + skip_if_no_corpus() + with_fixture_corpus({ + # Asserted on the RENDERED text, not on prov$balance_caveats: the field + # is already covered above, and the once-per-session cli_inform() means + # cog_explain() is the only surface a caller who missed (or suppressed) + # the first message can still audit. + r <- suppressMessages(cog_balances("550000227544", c(2012, 2019))) + # Both streams: cli routes most of its output through conditions that + # land on stderr, so a stdout-only capture would be empty (the pattern + # used throughout test-explain.R). + out <- paste(c(capture.output(cog_explain(r)), + capture.output(cog_explain(r), type = "message")), + collapse = "\n") + expect_match(out, "GAAP") + expect_match(out, "employee_retirement") + }) +}) + +test_that("cog_explain on a money-verb result has no balance caveat section", { + skip_if_no_corpus() + with_fixture_corpus({ + r <- suppressMessages(cog_spending("550000227544", 2019)) + out <- paste(c(capture.output(cog_explain(r)), + capture.output(cog_explain(r), type = "message")), + collapse = "\n") + # Guard against the capture itself being vacuous: the section must be + # absent from output that demonstrably contains the rest of the report. + expect_match(out, "Data vintage") + expect_false(grepl("GAAP", out)) + }) +}) + test_that("the caveat message fires once per session", { skip_if_no_corpus() with_fixture_corpus({ From a9e80858d43922d70f8cba5a27c092ca8515746e Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:37:15 -0400 Subject: [PATCH 21/22] docs: correct coverage_window scope and the stale CLAUDE.md Current State block (#25) F-7: inst/schemas/provenance-v1.json described coverage_window as mapping each *observed* balance_subtype, but the query at R/balance_caveats.R has no predicate tied to the query's codes and always returns every subtype in the mounted corpus. Took option (b) of the two the review offered -- change the doc, not the code. Reporting all windows is the better product behaviour (it answers 'is there a family I missed?'), it is what cog-api#26 already forwards verbatim, and option (a) would make the block empty for a 0-row result. Reworded to say the windows are corpus-wide and that 'truncated' is the query-scoped field. Pinned by a new test either way. F-10: the 'Current State' block was self-contradictory after a partial update -- headed 2026-04-27, claiming branch main @ d65e9fe, with a 2026-08-03 test count measured on feat/cog-balances-25 underneath it, and listing README.md / _pkgdown.yml as outstanding when both exist and _pkgdown.yml was edited by this branch. All numbers below re-measured on the final tree after every other fix in this wave, not before: 788 tests (testthat::test_local()), 14 exports (NAMESPACE), 14 man/*.Rd, 2 vignettes, no docs/ (pkgdown::build_site() genuinely still outstanding, as is the .Rbuildignore fixture entry -- both kept in the list). The related deferred README.md item is closed with no change, per the review's ruling: README.md enumerates no verbs at all, so naming cog_balances would make it the only non-cog_spending verb mentioned. --- CLAUDE.md | 18 ++++++++++-------- inst/schemas/provenance-v1.json | 2 +- 2 files changed, 11 insertions(+), 9 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 6a88b54..5fdb610 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -60,19 +60,20 @@ USCOGDATA_URL (local path or https://) Any value without `://` is treated as a local path by `.is_local_path()` and reads `manifest.json` directly from disk (no HTTP, no TTL cache). -## Current State (2026-04-27) +## Current State (2026-08-03) **Version:** 0.1.0 (pre-release) -**Branch:** `main`, commit `d65e9fe` -**Tests:** 764 PASS / 0 FAIL / 0 SKIP (measured `testthat::test_local()`, 2026-08-03, on the tree including the balance_caveats schema test) +**Branch:** `feat/cog-balances-25`, commit `fde62eb` +**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave) **CI:** Gitea Actions green (`.gitea/workflows/ci.yml`) ### Completed (Tasks 2.1–2.7) -All 10 exported verbs implemented and tested: +All **14** exports implemented and tested (measured from `NAMESPACE`): `cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`, `cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`, -`cog_gov_search`, `cog_mirror`, plus `cog_categories`. +`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`, +`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`. Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years 2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`, @@ -80,9 +81,10 @@ Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years ### Remaining to v0.1 release -1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports; - full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`. - Vignettes can be stubbed for v0.1. +1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`, + `README.md` and `_pkgdown.yml` exist, and `vignettes/` carries + `total-spending.Rmd` + `population-denominators.Rmd`. Outstanding: + `pkgdown::build_site()` has never been run (no `docs/`). 2. **Phase 3 — cog_explorer bridge:** create `cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs diff --git a/inst/schemas/provenance-v1.json b/inst/schemas/provenance-v1.json index a430cd5..b7e942b 100644 --- a/inst/schemas/provenance-v1.json +++ b/inst/schemas/provenance-v1.json @@ -55,7 +55,7 @@ }, "balance_caveats": { "type": ["object", "null"], - "description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps each observed balance_subtype to its measured [min year, max year] in the mounted corpus (never hardcoded). `truncated` lists the subtypes whose coverage_window does not fully span the requested years.", + "description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.", "properties": { "not_gaap": { "type": "boolean" }, "not_gaap_note": { "type": "string" }, From 2c532bde19cab30092f3620a5f9af69af0447c11 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 3 Aug 2026 11:45:43 -0400 Subject: [PATCH 22/22] docs: record the two balance_caveats contract facts cog-api#26 must carry Both were settled during implementation and are easy to get wrong from outside the package: - coverage_window is corpus-scoped, not result-scoped. It reports the observed year extent of every balance subtype, not only those a query returned. The sibling field `truncated` is the result-scoped one. - balance_caveats is present only on cog_balances() results; an API layer that assumes it is universal will read NULL from the money verbs. --- specs/2026-08-03-cog-balances-design.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/specs/2026-08-03-cog-balances-design.md b/specs/2026-08-03-cog-balances-design.md index 940dc9a..2f97945 100644 --- a/specs/2026-08-03-cog-balances-design.md +++ b/specs/2026-08-03-cog-balances-design.md @@ -260,6 +260,22 @@ throughout — so every test below runs offline. are the documented § 1 decision, and the recipe path reaches them by design. 2. **`cog-api#26`.** Adds `/balances` in all three required places — handler, `param_contract`, and the `plumber.R` route signature. Lands after this. + + **Two contract facts the API must carry forward**, both settled during + implementation and easy to get wrong from the outside: + + - `provenance$balance_caveats$coverage_window` is **corpus-scoped, not + result-scoped**. It reports the observed year extent of *every* balance + subtype in the corpus, not only the subtypes a given query returned — so a + `category = "Fund Balances"` query still returns all five windows. That is + deliberate: the windows describe what the corpus holds, which is what a + consumer needs in order to know what it did *not* ask for. The sibling + field `truncated` is the result-scoped one. Documented in + `inst/schemas/provenance-v1.json` and mutation-guarded against silent + inversion. + - `balance_caveats` appears **only** on `cog_balances()` results. It is + absent from `cog_spending()`/`cog_revenue()` provenance, and the schema + says so — an API layer that assumes it is universal will read `NULL`. 3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7 SQL views (there are 21), 181 tests (716), a two-year fixture (four years), and a "never inline SQL" rule the verb layer does not follow.