A reserved value nobody can discover is a trap, and this is the view
the API's /categories endpoint is built from. Emitted for the two flow
vocabularies only -- cog_balances() returns a stock and has no concept
to sum within.
Also fix test-categories.R to exclude pseudo-category rows from
crosswalk-specific assertions (one row per (category, subtype) pair,
non-empty item_codes, valid subtypes).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The balance work added category_type = "balance" rows to the corpus and
cog_balances() to read them, but left cog_categories() -- the discovery
surface -- unable to describe them:
- subtype COALESCEd only spend_subtype and revenue_subtype, so every balance
row came back with subtype = NA
- type rejected "balance", so there was no way to ask for the holdings
taxonomy at all
Both matter downstream: cog-api derives its subtype vocabulary from
cog_categories(), so an NA subtype becomes an unusable API parameter. Found
while implementing cog-api#26.
Note cog_balances() itself still takes no subtype argument -- for holdings
category is a strict coarsening of balance_subtype -- but the value belongs
in the discovery surface regardless.
Tests read the expected subtype set independently from the crosswalk parquet
rather than from the function under test.
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.
`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:
general = own_source + federal + state + local_aid (the default)
total = general + utility + liquor_store + insurance_trust
The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.
- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
subtype; the concept narrows in R via the existing subtype_scope
machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
printed `expenditure_concept` unconditionally, so a cog_revenue()
caller was told "Concept: primary" -- a spending concept their result
has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).
Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).
Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.
Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.
Closes#12
Tracks the corpus published 2026-07-30, which adds category_type = 'balance'
(pipeline#76) and the I/Q/Y flow codes (pipeline#78) -- the crosswalk
prerequisite for #11's three-concept expenditure model.
Fixture crosswalk goes 291 -> 324 rows and gains balance_subtype. Only three
files change (series_breaks, summary_categories, manifest); no long partition
moves, because the published change was metadata-only.
test-categories.R's vocabulary assertions extended for the new values:
category_type gains 'balance', spending subtypes gain 'interest' and
'insurance_benefits', revenue subtypes gain 'insurance_trust'.
cog_categories() is a catalogue verb so it surfaces every category_type the
corpus carries; the stock/flow guard belongs on the money verbs.
Suite: 0 failures, 2 skips (the #11 and #12 blocks).
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.
Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.
Three test repairs, each a real consequence of sparsification rather than
a number to bump:
test-categories.R "assistance" joined the spending subtype
vocabulary with the J-prefix codes.
test-spending.R The harmonization block counts rows that
exist. Broward's E21/F21/G21 were zero-pads
and are gone, so the anchor moves to FL state,
whose three NA-mapped rows carry $2.83B --
the amount accounting was previously asserted
only against 0 and could not have caught a
bug. Broward keeps a test of its own, now
asserting the zero-pads are absent.
test-expenditure-concept.R Coverage-gap suggestions are presence-based.
AL state's only FY2011 B47 cell was an
explicit zero, so ig_federal_b47_wide stopped
being a candidate there; FL state carries a
real amount, so the counterpart guard is
exercised against a suggestion that fires.
test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.
Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).
cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.
Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().
Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:
type = NULL | "spending" | "revenue"
pattern = regex matched case-insensitively on category
The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.
Also: fix @noRd placement in session.R so devtools::document() stops
warning.
Tests: +16 new / 181 total pass. check 0E/0W/0N.