.build_suggestions()'s recipe-candidate sub-select was keyed on `WHERE category IN (<category>)`. The reserved pseudo-category "All Categories" is never itself a row in summary_categories.category, so in all-categories mode `candidates` always came back empty and coverage signposting (uscogdata#9) was structurally impossible for the one mode whose entire premise is "you cannot sum the wrong scope" -- measured on Los Angeles County FY2011: category = "Public Welfare" reports 2 suggestions (incl. $271,589,000 excluded E68), category = "All Categories" reported 0, silently losing that same signal. Apply the branch's own design principle: the concept boundary is subtype, not category. .build_suggestions() now accepts all_categories/subtype_col/ subtype_scope (all optional, default off, so no other caller's behaviour changes) and, when all-categories mode is active, scopes the candidate sub-select by `<subtype_col> IN (<subtype_scope>)` instead -- symmetric with .build_verb_sql()'s own WHERE predicate. The M/L recipe exclusion and the is.null(category) early return are unchanged. After the fix, LA County FY2011 "All Categories" reports 5 suggestions, including welfare_cash_e68_wide for the exact $271,589,000 gap. Adds two covering tests to test-all-categories.R using the bundled fixture (AL state gov, FY2011, "Corrections"): one end-to-end (per-category and all-categories both signpost the same recipe) and one direct on .build_suggestions() proving the subtype-vs-category branch is what changes the query. Updates the 0.2.0 NEWS entry.
17 KiB
uscogdata 0.2.0
New features
-
cog_spending()andcog_revenue()accept the reserved category"All Categories", returning one summed row per(year, canonical_govid, subtype)across every category inside the requested concept's subtype scope. Filtering the result tospend_subtype == "operations"gives an operating-expenditure total.cog_geographic_rollup()inherits it, which is the efficient way to build a geographic total — previously a caller had to issue one rollup per category and sum the results (cog-api#37)."All Categories"is not the same thing asexpenditure_concept = "total". The concept chooses which subtypes are in scope;"All Categories"chooses whether the rows inside that scope are broken out or summed. -
cog_categories()advertises"All Categories"for the expenditure and revenue vocabularies, so the reserved value is discoverable. -
Coverage signposting (see "Signposting now catches partially-suppressed categories" below) now also works in
category = "All Categories"mode. The recipe-suggestion candidate query used to be scoped bycategory, which is never a match for the reserved"All Categories"value, soprovenance$suggestionsalways came back empty there — the one mode whose whole point is "you cannot sum the wrong scope" was silently unable to signal a wrong scope. The candidate query is now scoped by the concept's subtype allowlist instead, symmetric with how.build_verb_sql()itself scopes the summed total: Los Angeles County FY2011,category = "All Categories"still excludes $271,589,000 of aggregate-published Public Welfare (E68), but now namesrecipe = "welfare_cash_e68_wide"to recover it instead of reporting zero suggestions.
Documentation
cog_geographic_rollup()andcog_peer_compare()now document thatprovenance$coverage'sn_units_reportingis category-conditional and is not a response rate: a government that was surveyed and genuinely spends nothing in the requested category is indistinguishable from one never surveyed (uscogdata#36).
uscogdata 0.1.0 (development)
Signposting now catches partially-suppressed categories
- A coverage suggestion used to fire only when a category returned no rows
at all in a requested year. That missed the more dangerous case: a
category that still returns rows while silently dropping component codes
the wide era publishes only as aggregates (#9).
cog_spending(category = "Public Welfare")for FY2011 returned a plausible figure that omittedE67/E68entirely -- for Los Angeles County, $2,075,461,000 of a true $5,261,404,000, a 39% understatement, withprovenance$suggestionsempty. - Suggestions now also fire on partial coverage, and every suggestion
carries
trigger("empty_year"or"suppressed_component"),suppressed_amount,suppressed_yearsandsuppressed_codes, so a caller can see how much is missing and decide whether to re-run with the recipe. cog_revenue()gets the same fix through the shared verb path. Alaska's FY2011Miscellaneous Revenuereported $943,842,000 while dropping $1,899,995,000 of aggregate-publishedU4-rents and royalties.- The trigger stays recipe-driven, so it only fires where a harmonization
recipe actually exists to name the fix.
higher_ed_e18_wideandgeneral_gov_e89_widestay silent in every year measured on the bundled fixture, because their components are ordinary classified leaves even pre-2012. - The
suppressed_componenttrigger (and anysuppressed_amount/suppressed_codesanempty_yearfire also carries) is scoped to the calling verb's own flow family:cog_spending()only ever measures E/F/G component dollars,cog_revenue()only T/A/U/B/C/D. A component from the OTHER flow family reportssuppressed_amount = 0rather than a fabricated claim. Theempty_yeartrigger itself is not flow-scoped -- a category belonging to the other flow (e.g.cog_spending(category = "IG Local")) still returns zero rows and can still fire, in any year including modern ones, naming the recipe whose own generic join finds real data for this government. That is a mis-scoped query, not a corpus-format gap, so itssuppressed_amountis correctly 0.
New: cog_balances() for cash-and-security holdings
- New
cog_balances()exposes the 14 cash-and-security holding codes (category_type = "balance"): fund balances, retirement system holdings and insurance trust balances (#25). Holdings are a stock, not a flow, so the verb has noexpenditure_concept/revenue_concept/completearguments, and nosubtypeargument either -- for holdings,categoryis a strict coarsening ofbalance_subtype, socategory = "Fund Balances"is exactly thegeneralfamily (W01/W31/W61). cog_balances()results carryprovenance$balance_caveats, recording that Census holdings are gross rather than GAAP fund balance, and the measured coverage window of each subtype family.
Multi-government aggregates now disclose their reporting coverage
-
The Census of Governments is a complete census only in years ending in 2 and 7; every other year is a sample, and the sample varies enormously. On the bundled fixture, Wisconsin's 608-city universe rolls up 597 governments in FY2012 and 112 in FY2019 — an 18%-to-98% swing the return value said nothing about, so a statewide total resting on a fifth of the universe looked exactly like one resting on all of it.
-
cog_geographic_rollup(),cog_peer_compare()andcog_find_peers()gaincoverage:value effect "all"(default)every unit that reported that year — unchanged behaviour "census"census years only; aborts if the range holds none rather than returning nothing "consistent"only units reporting in every requested year — a balanced panel -
Regardless of mode, every result now carries
provenance$coveragewith per-yearn_units_reporting,n_units_expectedandis_census_year, plusprovenance$coverage_mode.cog_explain()prints a "Reporting coverage" section. So the default mode can no longer mislead silently. -
is_census_yearis a statement about the survey calendar, never a claim of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.n_units_reportingis the number that tells the truth. -
On
cog_peer_compare()the target is exempt from"consistent"balancing — it is the subject of the comparison, not a member of the cohort — and thesummary_*quantiles are computed after the filter, so they describe the cohort actually returned.n_units_reportingcounts peers only, against the cohort size. -
On
cog_find_peers(),coveragegoverns the cohort vintage whenyearisNULL:"census"snaps to the most recent census year with an observed population, so a cohort is not built from a sample year in which most of the candidate universe is absent.
complete = TRUE: absent cells, labelled with why they are absent
-
cog_spending()andcog_revenue()gaincomplete, defaulting toFALSE(today's behaviour). Withcomplete = TRUEthe requested grid is filled from the corpus'scode_settable and every row carries a newvalue_sourcecolumn:value_sourcemeaning amt_nominalreportedthe corpus carries this cell as published census_zerodense-source year (≤ FY2011), cell absent — Census published $00not_reportedsparse-source year (≥ FY2012), cell absent — unknown NAThe
NAis deliberate and is the whole point: filling a modern absence with0would invent data, which is precisely the error the corpus's representation contract exists to prevent. -
This restores information the reader lost when the corpus was sparsified (
SB194, cog_pipeline#64) — a wide-era query whose cells were all$0had begun returning nothing at all — and improves on what came before it, since the pre-sparsification corpus could not distinguish a published zero from an unreported cell either. -
The grid is scoped to each government's own type, so a county is never filled with cells only a state can report.
-
Needs a corpus published from 2026-07-29 onward (when
representationandcode_setbegan shipping); aborts with classuscogdata_representation_unavailableotherwise. Gated on the manifest listing those tables rather than onschema_version, which was never bumped for the change. Not available withrecipeorexpenditure_concept = "total"— neither draws its cells fromcode_set. -
provenance$completionreportsapplied,rows_filled, and the per-yearabsence_meansrule;cog_explain()prints a "Completion" section.
Corpus-wide series breaks now reach users (corpus_break_refs)
- Four catalogued series breaks carry
fin_code = "ALL"— caveats about the corpus as a whole rather than about one item code.series_break_refsis built by matchingfin_codeagainst the item codes in the result, and no row'sitem_codeis ever the literal"ALL", so none of them could ever be surfaced:SB085(dollar precision across the 1976/1977 boundary),SB087(imputation exclusion from FY2002),SB194(the dense → sparse representation change at FY2012) andSB086(the government id scheme change at FY2017). - Provenance gains
corpus_break_refs, selected on the break-year window alone and disjoint fromseries_break_refsby construction, so a consumer can tell a whole-result caveat from a break in one series.cog_explain()prints them under their own "Corpus-wide caveats" heading. cog-api passes provenance through verbatim, so the field appears there without an API change. SB194is the one that made this urgent: a query spanning FY2011 → FY2012 crosses the boundary where an absent cell stops meaning "Census published$0" and starts meaning "not reported", and until now nothing said so.
Bundled fixture regenerated against the sparsified corpus
inst/extdata/fixture_corpus/now tracks the corpus published on 2026-07-29 (pipeline_commit 83f9715, schema v6). The wide era no longer stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of which none are$0. Absence now means two different things — in adense_sourceyear (≤ FY2011) an absent cell means Census published$0; in asparse_sourceyear (≥ FY2012) it means not reported. The corpus carries that rule in two new tables the fixture now ships,representation.parquetandcode_set.parquet, alongsidecensus_collection_coverage.parquetandlineage_events.parquet(all ten publish-tree metadata tables, up from six). Catalogued upstream as series breakSB194.cog_categories()gains anassistancespending subtype: the J-prefix aid/benefit codes (J19,J67,J68,J85) are categorised now that the upstream crosswalk covers every flow code carrying dollars.- Two consequences worth knowing about, both visible in provenance rather
than in returned dollars. The harmonization block's
na_rows_excludedcounts only rows that exist, so wide-era codes that were zero-padded no longer appear there. Coverage-gapsuggestionsare presence-based for the same reason, so a recipe whose component codes were all$0for a given government-year is no longer suggested for it. tests/testthat/test-fixture-vintage.Rpins these structural facts, so a fixture left behind by a future publish fails loudly instead of letting the suite pass against a corpus that no longer exists.
Breaking: corpus schema_version 4 (Phase P canonical ids)
- The package now requires corpus
schema_version = 4(MinCorpusSchema/MaxCorpusSchemainDESCRIPTIONare both4); older corpora built against schema 3 are rejected bycog_open()with a clear version-mismatch error.canonical_govidis now uniformly 12 characters across every vintage the corpus covers (previously a mix of 9-char legacy ids and 12-char FIPS ids depending on source year) — every hardcodedcanonical_govidliteral from a pre-Phase-P corpus is now invalid and must be re-resolved viacog_gov_search()or the newcanonical_aliaslookup table.canonical_fips_xwalkgains four columns (legacy_govs_id,census_geoid,id_source;confidenceis renamed topop_confidence) and a companioncanonical_aliastable ships in the corpus for mapping legacy/alternate ids onto the current canonical namespace. The bundled fixture corpus (inst/extdata/fixture_corpus/) has been regenerated against the Phase P publish tree, now ships the fullcanonical_fips_xwalkandcanonical_aliasmaster tables alongside the 2019-2020 long partitions, and is reproducible viadata-raw/regenerate_fixture_corpus.R.
Clearer errors when USCOGDATA_URL is unconfigured or returns non-JSON
cog_open()now aborts with theuscogdata_url_not_configurederror class when the resolved corpus URL still contains the placeholderREPLACE_WITH_SHARE_TOKENsentinel (or is empty). The message lists both remediation paths (Sys.setenv(USCOGDATA_URL = ...)andoptions(uscogdata.url = ...)) and points at the bundled fixture for offline testing. Previously the package proceeded to fetch the placeholder URL, cached the resulting HTML welcome page, and failed downstream with a crypticjsonlitelexical-error..fetch_or_cache_manifest()now parses the HTTP response body before persisting it. Non-JSON responses (login pages, 404 HTML) raiseuscogdata_invalid_manifestwith the URL, Content-Type, and underlying parse error — and never write to the on-disk cache.- Manifest cache writes are now atomic (write to
manifest.json.tmp.<pid>incache_dir, thenfile.renameover the target), so an interrupted fetch cannot replace a previously-good cache. - Existing caches with non-JSON content (poisoned by the prior code path) are silently refetched instead of returning a parse error to the caller.
- Local
USCOGDATA_URLpaths whosemanifest.jsonis not valid JSON now surface the sameuscogdata_invalid_manifestclass with file context.
Per-capita denominators now use per-year Census F-33 population
cog_spending()andcog_revenue()previously divided all years' amounts by a single ACS 2018-2022 estimate (canonical_fips_xwalk.population_acs), producing biased per-capita values for time-series analysis. They now divide by the F-33populationrecorded on each gov-year via the newgov_population_yearlyview. Result tibbles gain apop_sourcecolumn with values"census_f33"or"unavailable".notesis updated to concatenate multiple notes with"; ".
Peer cohorts can be set to a chosen year
cog_find_peers()adds ayearargument (default: most recent year for which the target has an observed population ingov_population_yearly). The returned column previously namedpopulation_acsis nowpopulationand reflects the cohort year's vintage. The cohort year is attached to the returned tibble asattr(x, "cohort_year").cog_peer_compare()now stamps acohort_yearcolumn on its result (read from the peers tibble's attribute) and recordscohort_yearpluscohort_govidsin provenance. When the caller supplies a bare character vector instead of acog_find_peers()result,cohort_yearisNA.
Rollups exclude govs missing population
cog_geographic_rollup(per_capita = TRUE)drops rows whose government haspop_source == "unavailable"and records the dropped govids inprovenance$rollup$excluded_govids. This excludes special districts (type 4) and school districts (type 5) from per-capita rollups by design.
New: vignette and provenance metadata
- New vignette
population-denominatorscovers the four population sources, the type-4/5 coverage gap, the popyear quirk, and how to build moving-window peer cohorts manually. - Provenance gains
transformations$per_capita$popyear_rangeandpop_source_counts.cog_explain()renders both.
New features
cog_gov_search()gains a basket mode: passing vectorname/state/typearguments resolves multiple place names in one call and returns a tibble of canonical rows in input order, ready to pipe intocog_spending()/cog_revenue(). Per-row resolution follows an exact-then-substring matching algorithm with deterministic disambiguation; ambiguous and missing entries are surfaced via a sidecar audit tibble plus a single console summary message.- New exports
cog_basket_resolution()andcog_basket_unresolved()expose the basket sidecar for iterative query refinement.
Breaking changes
- The first formal of
cog_gov_search()was renamed frompatterntoname. All existing call sites incog_explorer/and the package itself use positional first-arg, so this rename is non-breaking in practice. Callers that passpattern = ...by name must update toname = ....