Bumps the minor version because 'All Categories' adds public surface
without breaking any existing call.
The bump is load-bearing, not cosmetic: cog-api installs this package
with install_local(), which no-ops when the version already matches.
Without it, Phase 1 would silently test against the 0.1.0 reader and
pass while proving nothing.
Final whole-branch review fix wave for the partial-coverage signposting
feature:
- I1: .suppressed_components() now filters measured recipe components to
the calling verb's own flow_prefixes. Without this, a candidate recipe
from the OTHER flow family was always absent from the verb's own view by
construction and so was always reported as "suppressed" -- fabricating a
dollar claim across flow families (cog_revenue(category = "Corrections")
claimed $3.63B excluded that cog_spending() actually reports in full).
- I2: reworded provenance-v1.json's trigger/suppressed_amount descriptions
to describe what the code actually measures (the verb's underlying long
view, not "the result"), and to note suppressed_amount can be negative.
Added ig_recipe_id to the suggestions items' required list, matching the
key's always-set/nullable runtime behavior.
- I3(a): restated the year/govid literals inside .suppressed_components()'s
NOT EXISTS subquery so DuckDB can partition-prune that side too (verified
via EXPLAIN: Scanning Files 1/4 instead of an unfiltered full scan;
all.equal(old, new) results confirmed unchanged).
- I3(b): added .needs_suppression_query(), a free, exact pre-check reusing
the verb's own already-computed result$codes_included to skip the anti-
join round trip on the common fully-covered path, without weakening the
"suppression can fire with zero gap years" guarantee.
- M4: corrected the overbroad "confines every fire to 2011" scope claim in
R/suggestions.R and NEWS.md -- the suppressed-dollar measurement is now
flow-scoped (post-I1), but the empty_year trigger itself is not, and can
still fire in modern years for a mis-scoped cross-flow-family category.
Added a regression test for I1 plus direct unit-test coverage for the new
flow-family filter and the I3(b) pre-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds the NEWS entry, a Financial data pkgdown reference section (none
existed for cog_spending/cog_revenue), and corrects CLAUDE.md's SQL-layer
claim, view count, test count and fixture-year description against
measured values. Also documents balance_caveats in
inst/schemas/provenance-v1.json (test-first: added a schema-documentation
test to test-balances.R, confirmed it failed, then fixed the schema) and
fleshes out cog_balances()'s @return roxygen to enumerate its conditional
columns, regenerating man/cog_balances.Rd.
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".
On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.
Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.
Decisions worth stating:
- n_units_expected is the universe the CALLER named, not the national one.
That is what makes the ratio mean something: "597 of the 608 Wisconsin
cities you asked about". For peers it is the cohort size, counted over
peer rows only -- including the target would inflate every count by one
and make a cohort that has entirely stopped reporting look non-empty.
- The coverage table is built from the REQUESTED years, not the years
present in the result, so a year in which nothing reported still appears
with n_units_reporting = 0. A year that vanishes silently is precisely
the disclosure failure at issue.
- "census" filters years BEFORE the query, and aborts when the range holds
no census year rather than returning an empty result for a query the
caller believes they made.
- "consistent" exempts the peer-comparison target: it is the subject of the
comparison, not a member of the cohort being balanced, and dropping it
would leave nothing to compare. The summary_* quantiles are computed
AFTER the filter so they describe the cohort actually returned.
- is_census_year is documented as a statement about the survey CALENDAR,
never a claim of completeness -- FY1967 is a census year in which only 97
of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
that tells the truth.
On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().
One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.
Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:
<= FY2011 dense_source absent => Census published $0
>= FY2012 sparse_source absent => not reported, unknown
A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.
complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.
Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.
Scoping decisions, each of which would invent rows if taken loosely:
- The grid is per government TYPE (code_set.type). Filling against the
union of all types would give a county cells like "state IG transfer to
school districts", indistinguishable from real census zeros.
- NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
grid offers cells those views never return, so each would fill as a
phantom $0.
- Filling happens BEFORE per_capita and inflation, so a census_zero stays
0 through both and a not_reported stays NA rather than becoming 0.
Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.
Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.
Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.
Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.
Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:
SB085 1977 dollar precision across the 1976/1977 boundary
SB087 2002 imputation exclusion FY2002-2006
SB194 2012 dense -> sparse representation change
SB086 2017 government id scheme change
SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.
Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.
On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.
Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.
Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
The fixture predated three shipped corpus changes at once: no J rows in
summary_categories (it was built before the crosswalk completion), no
representation.parquet or code_set.parquet, and a still-dense wide era.
Every test in this package and in cog-api runs against it, so both suites
were green against a corpus that no longer exists. This is #18's stated
prerequisite; it proves nothing about production until it lands.
Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6,
built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of
the old partition was explicit zeros -- and the fixture now ships all ten
publish-tree metadata tables rather than six. The generator's file list is
a single constant now, so the copy step and the manifest step cannot drift.
Three test repairs, each a real consequence of sparsification rather than
a number to bump:
test-categories.R "assistance" joined the spending subtype
vocabulary with the J-prefix codes.
test-spending.R The harmonization block counts rows that
exist. Broward's E21/F21/G21 were zero-pads
and are gone, so the anchor moves to FL state,
whose three NA-mapped rows carry $2.83B --
the amount accounting was previously asserted
only against 0 and could not have caught a
bug. Broward keeps a test of its own, now
asserting the zero-pads are absent.
test-expenditure-concept.R Coverage-gap suggestions are presence-based.
AL state's only FY2011 B47 cell was an
explicit zero, so ig_federal_b47_wide stopped
being a candidate there; FL state carries a
real amount, so the counterpart guard is
exercised against a suggestion that fires.
test-fixture-vintage.R pins the structural facts that separate this vintage
from its predecessor -- the ten metadata tables, the dense/sparse
representation contract, zero explicit zeros in FY2011, code_set coverage,
and J19's category. Checked against the old fixture: FY2011 carried
2,368,208 explicit zeros, so the assertion discriminates rather than
merely passing.
Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6).
cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture,
unchanged from its baseline.
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:
1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
when the resolved URL is empty or still contains the
`REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
`Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
existing caches with non-JSON content are silently refetched instead
of returning a parse error to the caller.
Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.