6f98d061a917b78530a4e4eb80259ebb42d91fea
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
af85a23ea7
|
feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean. |
||
|
|
1d553a788f
|
fix: surface ALL-scoped series breaks in provenance (#19)
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`. No row's item_code is ever the literal "ALL", so the four corpus-wide entries could never match and reached no user: SB085 1977 dollar precision across the 1976/1977 boundary SB087 2002 imputation exclusion FY2002-2006 SB194 2012 dense -> sparse representation change SB086 2017 government id scheme change SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv carries an ALL @ 2012 entry describing the representation change, SO cog_explain() surfaces it". The entry shipped; the reader dropped it. A query spanning FY2011 -> FY2012 crosses the boundary where an absent cell stops meaning "Census published $0" and starts meaning "not reported", and nothing said so. Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs() on the break_year window alone -- which codes a result happens to contain is irrelevant to a caveat about the corpus. A separate field rather than more entries in series_break_refs, because an ALL caveat qualifies the whole result and folding the two together invites reading it as a caveat about one series; .build_series_break_refs() now excludes 'ALL' explicitly so the two stay disjoint by construction. cog_explain() prints them under their own "Corpus-wide caveats" heading, and cog-api passes provenance through verbatim, so the field reaches the API with no change there. On the year rule: all four entries are BOUNDARY caveats -- their own join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not being comparable across FY2012, of pre- vs post-2017 ids -- so the same `break_year BETWEEN min(years) AND max(years)` rule the code-specific path uses is the right one, and matches the issue's DoD 1. The issue's DoD 3 also asks that a FY2011 query surface SB085; that cannot hold under DoD 1 and does not hold under any reading of SB085's text, whose boundary is 1976/1977. Tested with a range that actually spans it, and flagged on the issue. Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's bundled fixture, which predates the break being catalogued. Suite: 606 pass / 0 fail / 6 skip (was 594/0/6). cog-api 357 / 0 / 8, unchanged. |
||
|
|
c7260cb20c |
fix: scope total's coverage-gap detection to the Direct leg; require IG category rows (C1, C2)
C1: spending_long/spending_long_harmonized filter NOT is_aggregate but
ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a
legacy aggregate-only family (e.g. Corrections pre-2012) can survive on
the IG leg while Direct is suppressed. expenditure_concept = "total"
then UNIONs an IG-only figure that reads as a plausible Total, and the
coverage-gap suggestion machinery -- fed the UNION'd result -- saw the
surviving IG row as coverage and stayed silent.
(a) .build_suggestions() is now fed a Direct-leg-only view of the
result (IG rows filtered out before the gap-years computation),
so the recipe hints fire for "total" exactly as they do for
"direct".
(b) Any row where IG has dollars but Direct has none for the same
(year, canonical_govid, category) is now flagged: the row's
`notes` name the recovering recipe (drawn from the Direct-leg
suggestions), and provenance gains an explicit
`expenditure_concept_direct_suppressed` boolean plus an appended
warning on `expenditure_concept_note` -- both cheap for a
downstream consumer (cog-api passes provenance through verbatim)
to test, rather than silently asserting Direct + IG when that
arithmetic didn't happen.
Measured before/after on AL state government, Corrections, 2011:
"total" already correctly returns the corpus's actual IG-only figure
($31,358,000, vs. true Direct of $521,651,000 via recipe =
"corrections_combined"), but before this fix it did so with 0
suggestions and an unqualified "Total = Direct + IG" note; after, it
fires 3 recipe hints and both the row notes and provenance say plainly
that Direct is unavailable through this basis.
C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59
with no schema_version bump, so schema_version can't gate "total" --
a pre-#59 corpus can report any supported schema_version and still
have zero M/L category rows, in which case ig_annotated's LEFT JOIN
silently produces NA category/spend_subtype (0 rows for a specific
category, or one invisible NA-subtype group for category = NULL). New
.require_ig_categories() checks summary_categories directly and aborts
with class uscogdata_ig_categories_unsupported, naming PR #59 and
directing the user to a newer corpus.
Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose
synthetic summary_categories carries only one E36 row) by asserting
the new guard fires against that same connection, rather than leaving
the two silently contradictory.
|
||
|
|
7913b0f664 |
feat: record expenditure_concept in provenance and its JSON schema
Always populated, never implicit, so a downstream artifact says which concept produced it. cog-api passes provenance through verbatim. |
||
|
|
7818cd2b1a
|
feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
|
||
|
|
f7035049e7
|
feat: package skeleton — DESCRIPTION, NAMESPACE, session/manifest/cache/views
Minimal skeleton for uscogdata v0.1. Internal session layer with lazy cog_open(), manifest fetch+validate+cache, view registration placeholder. Depends on DuckDB >=1.0, httr2, jsonlite. inst/schemas/provenance-v1.json ships the structured provenance JSON Schema. |