docs: restore recipe= to cog_balances(); the pipeline was right

Corrects this spec. The earlier draft deferred recipe= and proposed adding
summary_categories rows for X40/X41. Both were wrong, and the pipeline state
they were meant to fix is correct and documented.

cog_pipeline/docs/phase_r_harmonization_review.md records the decisions:

- Sec 0.2: the wide era exposes these split families ONLY as aggregates, so
  the recipe join deliberately does NOT filter is_aggregate. Safe by
  construction -- wide rows are aggregate-only, modern rows leaf-only, every
  component year-scoped.
- Sec 1: the planned X40->Z77 harmonization MAP rows were dropped on purpose;
  continuity ships as recipes instead. That is why harmonization_map carries
  no balance-code rows.

The reader already implements this (R/recipes.R, R/spending.R). Verified
against the live corpus rather than trusting the comment: corrections_combined
FY2007, whose wide leg E05 is likewise aggregate-only, returns $906,743,000.

Also withdraws the claim that SB155/156 and SB195/196 contradict each other.
X40 rows after FY2002 are the wide-era SAS column persisting through the era
boundary; they say nothing about a classification-level rename. The two sets
describe different layers.

What survives is one narrow, non-blocking gap: no series_breaks row exists at
2016/2017 for Z77/Z78/X30, though review doc Sec 2 recommended exactly that.
Recorded as out-of-scope item 1 with the SB197-SB202 precedent.
This commit is contained in:
2026-08-03 09:28:52 -04:00
parent 9f9d40e1c3
commit 57212e3399
+61 -38
View File
@@ -88,7 +88,8 @@ cog_balances(govid, years,
# Retirement System Holdings
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"))
basis = c("harmonized", "raw"),
recipe = NULL)
```
Returns a `tbl_df` with a `provenance` attribute, like every other verb.
@@ -106,7 +107,31 @@ rows, so harmonized and raw are identical for holdings. Kept for uniformity
with the money verbs (the API would otherwise special-case), and
`provenance$basis_note` says so outright rather than letting it look meaningful.
**`recipe` is deliberately omitted from v1.** See below.
**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
to the modern one:
```
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
```
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
other basis view. That is by design, not a defect:
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
era exposes these split families *only* as aggregates, and that the recipe join
must therefore **not** filter `is_aggregate` — safe by construction, because
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
every component is year-scoped, so no double-count is possible. § 1 records the
matching decision that the planned `X40→Z77` harmonization *map* rows were
dropped and the continuity ships as recipes instead, which is why
`harmonization_map` has no balance-code rows.
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
as intended.
### No `subtype` argument: `category` is a strict coarsening
@@ -149,48 +174,26 @@ isolating one of the three insurance funds in a single argument;
`balance_subtype` remains a returned column, so that is one `dplyr::filter()`
away.
## Decision point: `recipe=` deferred to v2
The two holdings recipes are unusable from this verb today, and shipping the
argument anyway would produce a silently truncated series.
`cash_securities_z77_wide` = `X40` (1967–2011) + `Z77` (2012–2023);
`cash_securities_z78_wide` = `X41` (1967–2011) + `Z78` (2012–2023).
**`X40` and `X41` have zero rows in `summary_categories`.** They carry ~42,700
rows in the corpus across 1967–2011, but with no `category_type` they cannot
appear in a `category_type = 'balance'` view. Applying either recipe would
return only the modern leg — 2012–2016 — dropping 45 years while looking like a
valid continuous series. That is precisely the "plausible rather than obviously
wrong" failure mode `#25` exists to prevent.
The fix is upstream, in the pipeline crosswalk. Until it lands, `cog_balances()`
does not accept `recipe`, and an attempt to pass one is an error naming the
blocking issue rather than a silent partial result.
## Caveat surfacing
`provenance$balance_caveats`, always present, plus one `cli_inform()` per
session per caveat class when a query actually touches an affected family or
year. Structured so `cog-api#26` can forward the fields verbatim.
Three of the four caveats need reader-side work. **My earlier assumption that the
catalogued series breaks would cover them is wrong**, verified against
`series_breaks.parquet`:
Verified against `series_breaks.csv`, not assumed:
| # | Caveat | Covered by existing machinery? |
|---|---|---|
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — constant, new field `not_gaap = TRUE` |
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
| 3 | `X`/`Z` family ends FY2016 | **No.** SB197–SB202 attach to `X02`/`X05`/`X08`/`X11`/`X12`, which are *flow* codes. The holdings codes have no catalogued break for their FY2016 termination. Surfaced via `coverage_window` |
| 4 | `X40`/`X41` book → market at FY2002 | **No.** SB195/SB196 attach to `fin_code` `X40`/`X41`, which are outside the balance view (see above). Surfaced as an explicit caveat when the year range crosses 2002 and touches `employee_retirement` |
| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41` — but only on a `recipe` query, which is the only path that observes those codes. Asserted in the tests rather than assumed |
`coverage_window` is derived per observed subtype family from the corpus, never
hardcoded, so it stays correct as the corpus grows.
`series_break_refs` and `corpus_break_refs` are still populated by the existing
code-driven builders — they simply contribute nothing for the four caveats
above, and will start contributing once the upstream gaps close.
`series_break_refs` and `corpus_break_refs` are otherwise populated by the
existing code-driven builders and need no change.
## Testing
@@ -214,19 +217,39 @@ throughout — so every test below runs offline.
exactly one `category`. Asserted against the crosswalk so that an upstream
change breaking the tree — which would silently make `category` lossy —
fails here rather than in a user's analysis.
- **`recipe` rejection.** Passing `recipe` errors with a message naming the
blocking issue.
- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
`X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
regression here would silently truncate a 45-year series to five.
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
FY2002 carries both in `provenance$series_break_refs`, so the book → market
basis change is disclosed wherever `X40`/`X41` are actually observed.
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
`summary_categories` lacks `balance_subtype`.
## Out of scope, tracked separately
1. **Pipeline issue (new).** Add `summary_categories` rows for `X40`/`X41`
(`category_type = 'balance'`, `balance_subtype = 'employee_retirement'`), and
catalogue a series break for the FY2016 termination of the holdings codes.
Unblocks `recipe=` and lets caveats 3 and 4 flow through the existing
builders. Requires a corpus rebuild + republish, which is a separate,
human-approved operation.
1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
the seven holdings codes in `series_breaks.csv`. There is currently **no**
entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
`docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
exactly this — *"candidate new `series_breaks.csv` entries (recommend
`with_caution` documentation rows, no map action)"*. The follow-through never
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
what keeps this out of the `joinable = "no"` identity-change rule, which
would otherwise oblige a harmonization-map row.
Verify the break corpus-wide and census-to-census before writing the rows.
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
signpost when it lands.
**Superseded:** an earlier draft of this spec proposed adding
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
are the documented § 1 decision, and the recipe path reaches them by design.
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
`param_contract`, and the `plumber.R` route signature. Lands after this.
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7