Closes#18 — DoD 2, 3 and 4 here; DoD 1 (fixture) landed in #20, and DoD 5 is filed as cog-api#22.
The problem
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous:
year
cell absent means
≤ FY2011 (dense_source)
Census published $0
≥ FY2012 (sparse_source)
not reported — unknown
A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back — strictly less than the reader exposed before, which is why #64 filed this follow-on.
The fix
complete = TRUE fills the requested grid from code_set and stamps every row with value_source: reported, census_zero (amount 0), or not_reported (amount NA).
The NA is the point. Filling a modern absence with 0 would invent data — the error the representation contract exists to prevent — and it makes this strictly more informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either.
Measured on the fixture, Broward County:
query
result
FY2011, complete = TRUE
28 reported + 16 census_zero (amount 0)
FY2019, complete = TRUE
30 reported + 14 not_reported (amount NA)
The five categories walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly: census_zero before, not_reported after.
Three scoping decisions, each of which would invent rows if taken loosely
Per government type (code_set.type). Filling against the union of all types would hand a county cells like "state IG transfer to school districts" — indistinguishable from real census zeros.
NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0.
Filled before per_capita/inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0.
Gating: manifest presence, not schema_version
The two new views (36-representation, 37-code_set) are gated on the manifest listing those tables. Sparsification did not bump the schema version — the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table — so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort (uscogdata_representation_unavailable).
Refused where the fill would be guesswork
Both uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql), so its cells are not the ones code_set describes.
Verification
Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb — verifying what a filter does through that same filter proves nothing.
before
after
suite
629 / 0 / 3
658 / 0 / 3
rcmdcheck
clean
0 errors / 0 warnings / 0 notes
Note man/cog_spending.Rd carries ~26 lines of reindentation churn: my roxygen is 8.0.0, the repo's RoxygenNote is 7.3.3, and unlike last time I could not revert the file because it genuinely changed. DESCRIPTION is reverted as before. Bumping the repo to roxygen 8 is still worth deciding separately.
Closes #18 — DoD 2, 3 and 4 here; DoD 1 (fixture) landed in #20, and DoD 5 is filed as **cog-api#22**.
## The problem
Sparsification (cog_pipeline#64, `SB194`) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous:
| year | cell absent means |
|---|---|
| ≤ FY2011 (`dense_source`) | Census published **$0** |
| ≥ FY2012 (`sparse_source`) | **not reported** — unknown |
A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back — strictly less than the reader exposed before, which is why #64 filed this follow-on.
## The fix
`complete = TRUE` fills the requested grid from `code_set` and stamps every row with `value_source`: `reported`, `census_zero` (amount **0**), or `not_reported` (amount **NA**).
**The NA is the point.** Filling a modern absence with 0 would invent data — the error the representation contract exists to prevent — and it makes this strictly *more* informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either.
Measured on the fixture, Broward County:
| query | result |
|---|---|
| FY2011, `complete = TRUE` | 28 `reported` + **16 `census_zero`** (amount 0) |
| FY2019, `complete = TRUE` | 30 `reported` + **14 `not_reported`** (amount NA) |
The five categories walkthrough finding **F-006** read as "retired at FY2012" now report themselves correctly: `census_zero` before, `not_reported` after.
## Three scoping decisions, each of which would invent rows if taken loosely
- **Per government type** (`code_set.type`). Filling against the union of all types would hand a county cells like *"state IG transfer to school districts"* — indistinguishable from real census zeros.
- **`NOT is_aggregate`**, mirroring `spending_long`/`revenue_long`. Without it the grid offers cells those views never return, so each would fill as a phantom $0.
- **Filled before `per_capita`/inflation**, so a `census_zero` stays 0 through both and a `not_reported` stays NA rather than becoming 0.
## Gating: manifest presence, not schema_version
The two new views (`36-representation`, `37-code_set`) are gated on the manifest **listing** those tables. Sparsification did not bump the schema version — the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table — so a version gate would register a view over a missing file and fail at `CREATE VIEW` time on exactly the corpora the check exists to tolerate. `with_corpus_missing_representation()` models that corpus and asserts the abort (`uscogdata_representation_unavailable`).
## Refused where the fill would be guesswork
Both `uscogdata_complete_unsupported`: a **recipe** defines its own component codes and never touches `summary_categories`; the **intergovernmental leg** deliberately keeps aggregate rows (`inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.
## Verification
Expected cell sets in the tests are computed from the corpus parquet **directly, never through the verb** — verifying what a filter does through that same filter proves nothing.
| | before | after |
|---|---|---|
| suite | 629 / 0 / 3 | **658 / 0 / 3** |
| `rcmdcheck` | clean | **0 errors / 0 warnings / 0 notes** |
Note `man/cog_spending.Rd` carries ~26 lines of reindentation churn: my roxygen is 8.0.0, the repo's `RoxygenNote` is 7.3.3, and unlike last time I could not revert the file because it genuinely changed. `DESCRIPTION` is reverted as before. Bumping the repo to roxygen 8 is still worth deciding separately.
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:
<= FY2011 dense_source absent => Census published $0
>= FY2012 sparse_source absent => not reported, unknown
A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.
complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.
Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.
Scoping decisions, each of which would invent rows if taken loosely:
- The grid is per government TYPE (code_set.type). Filling against the
union of all types would give a county cells like "state IG transfer to
school districts", indistinguishable from real census zeros.
- NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
grid offers cells those views never return, so each would fill as a
phantom $0.
- Filling happens BEFORE per_capita and inflation, so a census_zero stays
0 through both and a not_reported stays NA rather than becoming 0.
Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.
Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.
Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.
Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.
Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes #18 — DoD 2, 3 and 4 here; DoD 1 (fixture) landed in #20, and DoD 5 is filed as cog-api#22.
The problem
Sparsification (cog_pipeline#64,
SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous:dense_source)sparse_source)A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back — strictly less than the reader exposed before, which is why #64 filed this follow-on.
The fix
complete = TRUEfills the requested grid fromcode_setand stamps every row withvalue_source:reported,census_zero(amount 0), ornot_reported(amount NA).The NA is the point. Filling a modern absence with 0 would invent data — the error the representation contract exists to prevent — and it makes this strictly more informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either.
Measured on the fixture, Broward County:
complete = TRUEreported+ 16census_zero(amount 0)complete = TRUEreported+ 14not_reported(amount NA)The five categories walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly:
census_zerobefore,not_reportedafter.Three scoping decisions, each of which would invent rows if taken loosely
code_set.type). Filling against the union of all types would hand a county cells like "state IG transfer to school districts" — indistinguishable from real census zeros.NOT is_aggregate, mirroringspending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0.per_capita/inflation, so acensus_zerostays 0 through both and anot_reportedstays NA rather than becoming 0.Gating: manifest presence, not schema_version
The two new views (
36-representation,37-code_set) are gated on the manifest listing those tables. Sparsification did not bump the schema version — the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table — so a version gate would register a view over a missing file and fail atCREATE VIEWtime on exactly the corpora the check exists to tolerate.with_corpus_missing_representation()models that corpus and asserts the abort (uscogdata_representation_unavailable).Refused where the fill would be guesswork
Both
uscogdata_complete_unsupported: a recipe defines its own component codes and never touchessummary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql), so its cells are not the onescode_setdescribes.Verification
Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb — verifying what a filter does through that same filter proves nothing.
rcmdcheckNote
man/cog_spending.Rdcarries ~26 lines of reindentation churn: my roxygen is 8.0.0, the repo'sRoxygenNoteis 7.3.3, and unlike last time I could not revert the file because it genuinely changed.DESCRIPTIONis reverted as before. Bumping the repo to roxygen 8 is still worth deciding separately.