724b6bd58b221d2b9da8f153e3952840a2e32a6f
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d95c9032c5
|
feat: coverage argument + always-on reporting-coverage metadata (#13)
The Census of Governments is a complete census only in years ending in 2 and
7. Every other year is a sample, and the sample varies enormously. Neither
cog_geographic_rollup() nor cog_peer_compare()/cog_find_peers() had any
concept of "the universe": each summed or labelled whichever govids happened
to have rows and returned that with nothing distinguishing "every government
reported" from "a fifth of them did".
On the bundled fixture, Wisconsin's 608-city universe rolls up 597
governments in FY2012 and 112 in FY2019. The peer side is worse exposure, not
better: a Madison-scale cohort looks stable because Madison is large, while
governments matched to a small target sit in exactly the population band the
sample cycle hits hardest. Chilton's 15-peer cohort reports 15 of 15 in
FY2012 and 3 of 15 in FY2019.
Implements the owner's settled design: coverage = c("all", "census",
"consistent") on all three verbs, defaulting to "all" so nothing currently
calling them changes, PLUS always-on provenance$coverage carrying per-year
n_units_reporting / n_units_expected / is_census_year and
provenance$coverage_mode. cog_explain() prints a "Reporting coverage"
section. The default mode can no longer mislead silently, which is the point
-- using these verbs correctly must not require knowing the survey calendar.
Decisions worth stating:
- n_units_expected is the universe the CALLER named, not the national one.
That is what makes the ratio mean something: "597 of the 608 Wisconsin
cities you asked about". For peers it is the cohort size, counted over
peer rows only -- including the target would inflate every count by one
and make a cohort that has entirely stopped reporting look non-empty.
- The coverage table is built from the REQUESTED years, not the years
present in the result, so a year in which nothing reported still appears
with n_units_reporting = 0. A year that vanishes silently is precisely
the disclosure failure at issue.
- "census" filters years BEFORE the query, and aborts when the range holds
no census year rather than returning an empty result for a query the
caller believes they made.
- "consistent" exempts the peer-comparison target: it is the subject of the
comparison, not a member of the cohort being balanced, and dropping it
would leave nothing to compare. The summary_* quantiles are computed
AFTER the filter so they describe the cohort actually returned.
- is_census_year is documented as a statement about the survey CALENDAR,
never a claim of completeness -- FY1967 is a census year in which only 97
of Wisconsin's 608 cities report (DoD 3). n_units_reporting is the number
that tells the truth.
On cog_find_peers(), where there is no year range, coverage governs the
cohort VINTAGE: "census" snaps to the most recent census year with an
observed population, so a cohort is not built from a sample year in which
most of the candidate universe is absent. "consistent" is a comparison-time
concept and selects like "all" there, carried on the result for
cog_peer_compare().
One fix to the committed test, which was internally inconsistent. It pinned
n_units_reporting == 597 for FY2012 AND asserted that number equals a raw
cross-check that answers 595. Both numbers are right for different questions:
VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 in `long` (their
as-of-year identity, as townships) while the xwalk lists them as govs_type =
2 (their present identity, as villages) -- schema v6 made the long table's
geography present-harmonized but `type` still reads as-of-year. The rollup
counts against the requested govid set, so 597 answers "how many of the
governments I asked about reported". The cross-check now scopes to that same
universe instead of to long.type/long.fips_state; it still reads raw parquet
rather than going through the verb under test.
Suite: 670 pass / 0 fail / 2 skip (was 658/0/3). rcmdcheck clean.
The two remaining skips are #11 and #12.
|
||
|
|
af85a23ea7
|
feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean. |
||
|
|
1d553a788f
|
fix: surface ALL-scoped series breaks in provenance (#19)
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`. No row's item_code is ever the literal "ALL", so the four corpus-wide entries could never match and reached no user: SB085 1977 dollar precision across the 1976/1977 boundary SB087 2002 imputation exclusion FY2002-2006 SB194 2012 dense -> sparse representation change SB086 2017 government id scheme change SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv carries an ALL @ 2012 entry describing the representation change, SO cog_explain() surfaces it". The entry shipped; the reader dropped it. A query spanning FY2011 -> FY2012 crosses the boundary where an absent cell stops meaning "Census published $0" and starts meaning "not reported", and nothing said so. Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs() on the break_year window alone -- which codes a result happens to contain is irrelevant to a caveat about the corpus. A separate field rather than more entries in series_break_refs, because an ALL caveat qualifies the whole result and folding the two together invites reading it as a caveat about one series; .build_series_break_refs() now excludes 'ALL' explicitly so the two stay disjoint by construction. cog_explain() prints them under their own "Corpus-wide caveats" heading, and cog-api passes provenance through verbatim, so the field reaches the API with no change there. On the year rule: all four entries are BOUNDARY caveats -- their own join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not being comparable across FY2012, of pre- vs post-2017 ids -- so the same `break_year BETWEEN min(years) AND max(years)` rule the code-specific path uses is the right one, and matches the issue's DoD 1. The issue's DoD 3 also asks that a FY2011 query surface SB085; that cannot hold under DoD 1 and does not hold under any reading of SB085's text, whose boundary is 1976/1977. Tested with a range that actually spans it, and flagged on the issue. Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's bundled fixture, which predates the break being catalogued. Suite: 606 pass / 0 fail / 6 skip (was 594/0/6). cog-api 357 / 0 / 8, unchanged. |
||
|
|
c375c55da7
|
fix: regenerate the bundled fixture against the sparsified corpus (#18)
The fixture predated three shipped corpus changes at once: no J rows in summary_categories (it was built before the crosswalk completion), no representation.parquet or code_set.parquet, and a still-dense wide era. Every test in this package and in cog-api runs against it, so both suites were green against a corpus that no longer exists. This is #18's stated prerequisite; it proves nothing about production until it lands. Regenerated from the publish tree at pipeline_commit 83f9715 (schema v6, built 2026-07-29). FY2011 goes from 2,864,212 rows to 496,004 -- 82.7% of the old partition was explicit zeros -- and the fixture now ships all ten publish-tree metadata tables rather than six. The generator's file list is a single constant now, so the copy step and the manifest step cannot drift. Three test repairs, each a real consequence of sparsification rather than a number to bump: test-categories.R "assistance" joined the spending subtype vocabulary with the J-prefix codes. test-spending.R The harmonization block counts rows that exist. Broward's E21/F21/G21 were zero-pads and are gone, so the anchor moves to FL state, whose three NA-mapped rows carry $2.83B -- the amount accounting was previously asserted only against 0 and could not have caught a bug. Broward keeps a test of its own, now asserting the zero-pads are absent. test-expenditure-concept.R Coverage-gap suggestions are presence-based. AL state's only FY2011 B47 cell was an explicit zero, so ig_federal_b47_wide stopped being a candidate there; FL state carries a real amount, so the counterpart guard is exercised against a suggestion that fires. test-fixture-vintage.R pins the structural facts that separate this vintage from its predecessor -- the ten metadata tables, the dense/sparse representation contract, zero explicit zeros in FY2011, code_set coverage, and J19's category. Checked against the old fixture: FY2011 carried 2,368,208 explicit zeros, so the assertion discriminates rather than merely passing. Suites: uscogdata 594 pass / 0 fail / 6 skip (was 576/0/6). cog-api 357 pass / 0 fail / 8 skip against the regenerated fixture, unchanged from its baseline. |
||
|
|
e635a1fc9e
|
feat!: require corpus schema_version 4 (Phase P canonical ids)
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across every vintage; corpora built against schema_version 3 are rejected. Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master schema (adds legacy_govs_id, census_geoid, id_source; confidence is renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match. |
||
|
|
874347242b
|
fix(manifest): actionable errors when USCOGDATA_URL is unset or returns non-JSON
R-CMD-check / check (push) Successful in 2m5s
`cog_gov_search()` (and every other verb) used to fail with a cryptic `jsonlite` lexical error when the package's placeholder default URL was hit and the server returned an HTML welcome page that got cached as `manifest.json`. Three guards added: 1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured` when the resolved URL is empty or still contains the `REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both `Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)` remediations and points at the bundled fixture. 2. `.fetch_or_cache_manifest()` parses the response body before persisting it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL, Content-Type, parse error) and never touch the on-disk cache. 3. Cache writes are atomic via a sibling tempfile + `file.rename`, and existing caches with non-JSON content are silently refetched instead of returning a parse error to the caller. Local-path manifests that aren't valid JSON now surface the same `uscogdata_invalid_manifest` class with file context. |
||
|
|
33c0274727 |
docs(news): per-year population denominators (unreleased)
Summarizes the per-capita and peer-cohort behavior changes for users upgrading from earlier 0.1 snapshots. |
||
|
|
24e67791a0
|
docs: roxygen, NEWS, and pkgdown for basket mode
Adds full @description, @details (algorithm), @examples on cog_gov_search(); examples on cog_basket_resolution() / cog_basket_unresolved(); NEWS.md entry covering the new mode and the pattern->name rename; pkgdown reference entries for the two new exports. |