feat: complete = TRUE fills absent cells with their meaning (#18) #23

Merged
jared merged 1 commits from feat/complete-argument-18 into main 2026-07-30 12:04:28 -04:00
Owner

Closes #18 — DoD 2, 3 and 4 here; DoD 1 (fixture) landed in #20, and DoD 5 is filed as cog-api#22.

The problem

Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous:

year cell absent means
≤ FY2011 (dense_source) Census published $0
≥ FY2012 (sparse_source) not reported — unknown

A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back — strictly less than the reader exposed before, which is why #64 filed this follow-on.

The fix

complete = TRUE fills the requested grid from code_set and stamps every row with value_source: reported, census_zero (amount 0), or not_reported (amount NA).

The NA is the point. Filling a modern absence with 0 would invent data — the error the representation contract exists to prevent — and it makes this strictly more informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either.

Measured on the fixture, Broward County:

query result
FY2011, complete = TRUE 28 reported + 16 census_zero (amount 0)
FY2019, complete = TRUE 30 reported + 14 not_reported (amount NA)

The five categories walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly: census_zero before, not_reported after.

Three scoping decisions, each of which would invent rows if taken loosely

  • Per government type (code_set.type). Filling against the union of all types would hand a county cells like "state IG transfer to school districts" — indistinguishable from real census zeros.
  • NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0.
  • Filled before per_capita/inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0.

Gating: manifest presence, not schema_version

The two new views (36-representation, 37-code_set) are gated on the manifest listing those tables. Sparsification did not bump the schema version — the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table — so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort (uscogdata_representation_unavailable).

Refused where the fill would be guesswork

Both uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql), so its cells are not the ones code_set describes.

Verification

Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb — verifying what a filter does through that same filter proves nothing.

before after
suite 629 / 0 / 3 658 / 0 / 3
rcmdcheck clean 0 errors / 0 warnings / 0 notes

Note man/cog_spending.Rd carries ~26 lines of reindentation churn: my roxygen is 8.0.0, the repo's RoxygenNote is 7.3.3, and unlike last time I could not revert the file because it genuinely changed. DESCRIPTION is reverted as before. Bumping the repo to roxygen 8 is still worth deciding separately.

Closes #18 — DoD 2, 3 and 4 here; DoD 1 (fixture) landed in #20, and DoD 5 is filed as **cog-api#22**. ## The problem Sparsification (cog_pipeline#64, `SB194`) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: | year | cell absent means | |---|---| | ≤ FY2011 (`dense_source`) | Census published **$0** | | ≥ FY2012 (`sparse_source`) | **not reported** — unknown | A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back — strictly less than the reader exposed before, which is why #64 filed this follow-on. ## The fix `complete = TRUE` fills the requested grid from `code_set` and stamps every row with `value_source`: `reported`, `census_zero` (amount **0**), or `not_reported` (amount **NA**). **The NA is the point.** Filling a modern absence with 0 would invent data — the error the representation contract exists to prevent — and it makes this strictly *more* informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: | query | result | |---|---| | FY2011, `complete = TRUE` | 28 `reported` + **16 `census_zero`** (amount 0) | | FY2019, `complete = TRUE` | 30 `reported` + **14 `not_reported`** (amount NA) | The five categories walkthrough finding **F-006** read as "retired at FY2012" now report themselves correctly: `census_zero` before, `not_reported` after. ## Three scoping decisions, each of which would invent rows if taken loosely - **Per government type** (`code_set.type`). Filling against the union of all types would hand a county cells like *"state IG transfer to school districts"* — indistinguishable from real census zeros. - **`NOT is_aggregate`**, mirroring `spending_long`/`revenue_long`. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - **Filled before `per_capita`/inflation**, so a `census_zero` stays 0 through both and a `not_reported` stays NA rather than becoming 0. ## Gating: manifest presence, not schema_version The two new views (`36-representation`, `37-code_set`) are gated on the manifest **listing** those tables. Sparsification did not bump the schema version — the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table — so a version gate would register a view over a missing file and fail at `CREATE VIEW` time on exactly the corpora the check exists to tolerate. `with_corpus_missing_representation()` models that corpus and asserts the abort (`uscogdata_representation_unavailable`). ## Refused where the fill would be guesswork Both `uscogdata_complete_unsupported`: a **recipe** defines its own component codes and never touches `summary_categories`; the **intergovernmental leg** deliberately keeps aggregate rows (`inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes. ## Verification Expected cell sets in the tests are computed from the corpus parquet **directly, never through the verb** — verifying what a filter does through that same filter proves nothing. | | before | after | |---|---|---| | suite | 629 / 0 / 3 | **658 / 0 / 3** | | `rcmdcheck` | clean | **0 errors / 0 warnings / 0 notes** | Note `man/cog_spending.Rd` carries ~26 lines of reindentation churn: my roxygen is 8.0.0, the repo's `RoxygenNote` is 7.3.3, and unlike last time I could not revert the file because it genuinely changed. `DESCRIPTION` is reverted as before. Bumping the repo to roxygen 8 is still worth deciding separately.
jared added 1 commit 2026-07-30 11:49:02 -04:00
feat: complete = TRUE fills absent cells with their meaning (#18)
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
af85a23ea7
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
jared merged commit 6f98d061a9 into main 2026-07-30 12:04:28 -04:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#23