feat: cog_balances(), a reader surface for cash and security holdings (#25) #28

Merged
jared merged 22 commits from feat/cog-balances-25 into main 2026-08-03 11:52:13 -04:00
Showing only changes of commit 57212e3399 - Show all commits
+61 -38
View File
@@ -88,7 +88,8 @@ cog_balances(govid, years,
# Retirement System Holdings # Retirement System Holdings
per_capita = FALSE, per_capita = FALSE,
adjust_to_year = NULL, adjust_to_year = NULL,
basis = c("harmonized", "raw")) basis = c("harmonized", "raw"),
recipe = NULL)
``` ```
Returns a `tbl_df` with a `provenance` attribute, like every other verb. Returns a `tbl_df` with a `provenance` attribute, like every other verb.
@@ -106,7 +107,31 @@ rows, so harmonized and raw are identical for holdings. Kept for uniformity
with the money verbs (the API would otherwise special-case), and with the money verbs (the API would otherwise special-case), and
`provenance$basis_note` says so outright rather than letting it look meaningful. `provenance$basis_note` says so outright rather than letting it look meaningful.
**`recipe` is deliberately omitted from v1.** See below. **`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
to the modern one:
```
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
```
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
other basis view. That is by design, not a defect:
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
era exposes these split families *only* as aggregates, and that the recipe join
must therefore **not** filter `is_aggregate` — safe by construction, because
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
every component is year-scoped, so no double-count is possible. § 1 records the
matching decision that the planned `X40→Z77` harmonization *map* rows were
dropped and the continuity ships as recipes instead, which is why
`harmonization_map` has no balance-code rows.
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
as intended.
### No `subtype` argument: `category` is a strict coarsening ### No `subtype` argument: `category` is a strict coarsening
@@ -149,48 +174,26 @@ isolating one of the three insurance funds in a single argument;
`balance_subtype` remains a returned column, so that is one `dplyr::filter()` `balance_subtype` remains a returned column, so that is one `dplyr::filter()`
away. away.
## Decision point: `recipe=` deferred to v2
The two holdings recipes are unusable from this verb today, and shipping the
argument anyway would produce a silently truncated series.
`cash_securities_z77_wide` = `X40` (1967–2011) + `Z77` (2012–2023);
`cash_securities_z78_wide` = `X41` (1967–2011) + `Z78` (2012–2023).
**`X40` and `X41` have zero rows in `summary_categories`.** They carry ~42,700
rows in the corpus across 1967–2011, but with no `category_type` they cannot
appear in a `category_type = 'balance'` view. Applying either recipe would
return only the modern leg — 2012–2016 — dropping 45 years while looking like a
valid continuous series. That is precisely the "plausible rather than obviously
wrong" failure mode `#25` exists to prevent.
The fix is upstream, in the pipeline crosswalk. Until it lands, `cog_balances()`
does not accept `recipe`, and an attempt to pass one is an error naming the
blocking issue rather than a silent partial result.
## Caveat surfacing ## Caveat surfacing
`provenance$balance_caveats`, always present, plus one `cli_inform()` per `provenance$balance_caveats`, always present, plus one `cli_inform()` per
session per caveat class when a query actually touches an affected family or session per caveat class when a query actually touches an affected family or
year. Structured so `cog-api#26` can forward the fields verbatim. year. Structured so `cog-api#26` can forward the fields verbatim.
Three of the four caveats need reader-side work. **My earlier assumption that the Verified against `series_breaks.csv`, not assumed:
catalogued series breaks would cover them is wrong**, verified against
`series_breaks.parquet`:
| # | Caveat | Covered by existing machinery? | | # | Caveat | Covered by existing machinery? |
|---|---|---| |---|---|---|
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — constant, new field `not_gaap = TRUE` | | 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus | | 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
| 3 | `X`/`Z` family ends FY2016 | **No.** SB197–SB202 attach to `X02`/`X05`/`X08`/`X11`/`X12`, which are *flow* codes. The holdings codes have no catalogued break for their FY2016 termination. Surfaced via `coverage_window` | | 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
| 4 | `X40`/`X41` book → market at FY2002 | **No.** SB195/SB196 attach to `fin_code` `X40`/`X41`, which are outside the balance view (see above). Surfaced as an explicit caveat when the year range crosses 2002 and touches `employee_retirement` | | 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41` — but only on a `recipe` query, which is the only path that observes those codes. Asserted in the tests rather than assumed |
`coverage_window` is derived per observed subtype family from the corpus, never `coverage_window` is derived per observed subtype family from the corpus, never
hardcoded, so it stays correct as the corpus grows. hardcoded, so it stays correct as the corpus grows.
`series_break_refs` and `corpus_break_refs` are still populated by the existing `series_break_refs` and `corpus_break_refs` are otherwise populated by the
code-driven builders — they simply contribute nothing for the four caveats existing code-driven builders and need no change.
above, and will start contributing once the upstream gaps close.
## Testing ## Testing
@@ -214,19 +217,39 @@ throughout — so every test below runs offline.
exactly one `category`. Asserted against the crosswalk so that an upstream exactly one `category`. Asserted against the crosswalk so that an upstream
change breaking the tree — which would silently make `category` lossy — change breaking the tree — which would silently make `category` lossy —
fails here rather than in a user's analysis. fails here rather than in a user's analysis.
- **`recipe` rejection.** Passing `recipe` errors with a message naming the - **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
blocking issue. `X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
regression here would silently truncate a 45-year series to five.
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
FY2002 carries both in `provenance$series_break_refs`, so the book → market
basis change is disclosed wherever `X40`/`X41` are actually observed.
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose - **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
`summary_categories` lacks `balance_subtype`. `summary_categories` lacks `balance_subtype`.
## Out of scope, tracked separately ## Out of scope, tracked separately
1. **Pipeline issue (new).** Add `summary_categories` rows for `X40`/`X41` 1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
(`category_type = 'balance'`, `balance_subtype = 'employee_retirement'`), and the seven holdings codes in `series_breaks.csv`. There is currently **no**
catalogue a series break for the FY2016 termination of the holdings codes. entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
Unblocks `recipe=` and lets caveats 3 and 4 flow through the existing `docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
builders. Requires a corpus rebuild + republish, which is a separate, exactly this — *"candidate new `series_breaks.csv` entries (recommend
human-approved operation. `with_caution` documentation rows, no map action)"*. The follow-through never
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
what keeps this out of the `joinable = "no"` identity-change rule, which
would otherwise oblige a harmonization-map row.
Verify the break corpus-wide and census-to-census before writing the rows.
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
signpost when it lands.
**Superseded:** an earlier draft of this spec proposed adding
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
are the documented § 1 decision, and the recipe path reaches them by design.
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler, 2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
`param_contract`, and the `plumber.R` route signature. Lands after this. `param_contract`, and the `plumber.R` route signature. Lands after this.
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7 3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7