Task 4's implementer found that SB195 does not surface for a
recipe query spanning only 2011-2012. .build_series_break_refs() matches
break_year BETWEEN min(years) AND max(years), and SB195's break_year is 2002.
That is correct behaviour rather than a gap: a series lying entirely after the
book -> market change sits on one consistent basis, so disclosing a break it
never crosses would be noise. .build_corpus_break_refs() applies the same rule
deliberately.
The spec's caveat table overclaimed by omitting the span condition. Corrected.
Corrects this spec. The earlier draft deferred recipe= and proposed adding
summary_categories rows for X40/X41. Both were wrong, and the pipeline state
they were meant to fix is correct and documented.
cog_pipeline/docs/phase_r_harmonization_review.md records the decisions:
- Sec 0.2: the wide era exposes these split families ONLY as aggregates, so
the recipe join deliberately does NOT filter is_aggregate. Safe by
construction -- wide rows are aggregate-only, modern rows leaf-only, every
component year-scoped.
- Sec 1: the planned X40->Z77 harmonization MAP rows were dropped on purpose;
continuity ships as recipes instead. That is why harmonization_map carries
no balance-code rows.
The reader already implements this (R/recipes.R, R/spending.R). Verified
against the live corpus rather than trusting the comment: corrections_combined
FY2007, whose wide leg E05 is likewise aggregate-only, returns $906,743,000.
Also withdraws the claim that SB155/156 and SB195/196 contradict each other.
X40 rows after FY2002 are the wide-era SAS column persisting through the era
boundary; they say nothing about a classification-level rename. The two sets
describe different layers.
What survives is one narrow, non-blocking gap: no series_breaks row exists at
2016/2017 for Z77/Z78/X30, though review doc Sec 2 recommended exactly that.
Recorded as out-of-scope item 1 with the SB197-SB202 precedent.
balance is the only category_type whose subtype column is not orthogonal to
category. Measured against the crosswalk: 5 of 6 expenditure subtypes and 1 of
7 revenue subtypes span more than one category, but 0 of 5 balance subtypes do.
Balance is a strict tree -- Fund Balances = {general}, Retirement System
Holdings = {employee_retirement}, Insurance Trust Balances = the three trust
subtypes.
Exposing both arguments would admit no useful combination: of the 15 pairs, 3
are redundant and 12 are guaranteed empty for every government in every year,
failing as an empty tibble that reads as "holds none" rather than as a
contradiction.
Dropping it also keeps the verb aligned -- no uscogdata verb exposes a subtype
argument; the API layers its own subtype row filter on top, which cog-api#26
can do for /balances. #25's one-filter requirement is still met, since
category = "Fund Balances" is exactly W01/W31/W61.
Adds two tests: that one-filter equivalence, and an assertion that the
subtype -> category tree holds, so an upstream change making category lossy
fails here rather than in a user's analysis.
Requirement 1 of #25 shipped with #11/#12. This specs requirement 2 only.
Records three upstream gaps found while measuring the corpus, which change
the shipping scope:
- X40/X41 carry ~42.7K rows (1967-2011) but have no summary_categories row,
so they cannot appear in a category_type='balance' view. Both holdings
recipes span X40/X41 + Z77/Z78, so recipe= would silently return only the
2012-2016 leg. recipe= is therefore deferred to v2.
- SB195/SB196 attach to fin_code X40/X41, outside the balance view.
- SB197-SB202 attach to flow codes, not the holdings codes, so the FY2016
termination of X21/X30/X42/X44/X47/Z77/Z78 has no catalogued break.
Caveats 2-4 are therefore surfaced reader-side via a computed coverage_window
rather than through the existing series-break builders.