feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
This commit is contained in:
@@ -1,5 +1,37 @@
|
||||
# uscogdata 0.1.0 (development)
|
||||
|
||||
## `complete = TRUE`: absent cells, labelled with why they are absent
|
||||
|
||||
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
|
||||
(today's behaviour). With `complete = TRUE` the requested grid is filled
|
||||
from the corpus's `code_set` table and every row carries a new
|
||||
`value_source` column:
|
||||
|
||||
| `value_source` | meaning | `amt_nominal` |
|
||||
|---|---|---|
|
||||
| `reported` | the corpus carries this cell | as published |
|
||||
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
|
||||
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
|
||||
|
||||
The `NA` is deliberate and is the whole point: filling a modern absence
|
||||
with `0` would invent data, which is precisely the error the corpus's
|
||||
representation contract exists to prevent.
|
||||
* This restores information the reader lost when the corpus was sparsified
|
||||
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
|
||||
had begun returning nothing at all — and improves on what came before it,
|
||||
since the pre-sparsification corpus could not distinguish a published zero
|
||||
from an unreported cell either.
|
||||
* The grid is scoped to each government's **own type**, so a county is never
|
||||
filled with cells only a state can report.
|
||||
* Needs a corpus published from 2026-07-29 onward (when `representation` and
|
||||
`code_set` began shipping); aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
|
||||
listing those tables rather than on `schema_version`, which was never
|
||||
bumped for the change. Not available with `recipe` or
|
||||
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
|
||||
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
|
||||
`absence_means` rule; `cog_explain()` prints a "Completion" section.
|
||||
|
||||
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
|
||||
|
||||
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
|
||||
|
||||
Reference in New Issue
Block a user