Add complete= so wide-era census zeros are recoverable after sparsification #18

Closed
opened 2026-07-29 19:57:52 -04:00 by jared · 1 comment
Owner

Follow-on from census_of_governments_finance_pipeline#64 DoD 7. Filed after the corpus was sparsified and published on 2026-07-29 (pipeline_commit 83f9715).

What changed under us

The published corpus no longer stores the wide era's explicit zeros. 331.4M rows -> 45.9M (-86%). Absence now means two different things:

year cell absent means
<= FY2011 (dense_source) Census published $0
>= FY2012 (sparse_source) not reported (unknown)

Two new corpus artifacts carry the rule: data/representation.parquet (per year: representation, absence_means, code_set_id) and data/code_set.parquet (code_set_id, year, type, item_code, is_aggregate, n_units -- 84,693 rows across 55 years). The government universe is already in canonical_fips_xwalk. Catalogued as SB194.

The user-visible consequence

A wide-era query whose rows were all $0 now returns nothing instead of $0 rows, and the reader offers no way to get them back. Verified against the live corpus: cog_spending(govid="552025209777", years=2005L) returns 27 rows, 0 of them zero-amount. Before sparsification the all-zero categories came back as explicit $0.

This is strictly less information than the reader used to expose, which is why #64 filed this follow-on.

What to build

complete = TRUE on the money verbs: fill the requested grid from code_set (scoped to each government's type -- filling against the union of all types invents rows like "$0 state IG transfer to school districts" for counties) and stamp each row's value_source:

value meaning
reported a row exists in long
census_zero dense-source year, cell absent
not_reported sparse-source year, cell absent

Strictly more information than the reader had before sparsification, where a wide-era zero and a modern absent cell were different shapes with nothing explaining the difference.

Prerequisite

The bundled fixture corpus is stale and must be regenerated first. inst/extdata/fixture_corpus/ has no J rows in summary_categories (it predates the #65 crosswalk work), no representation.parquet, no code_set.parquet, and a still-dense wide era. Every test in this package and in cog-api runs against it, so it currently does not represent what is published.

Definition of done

  1. Fixture regenerated from the current publish tree.
  2. complete= implemented on cog_spending() / cog_revenue(), defaulting to today's behaviour.
  3. value_source on every returned row when complete = TRUE.
  4. A round-trip test: for a dense-source government-year, complete = TRUE reproduces the pre-sparsification row set exactly, zeros included.
  5. cog-api follow-on filed (threading complete= behind its own pagination fixes).

The pipeline asserts this same round-trip on every build -- see tests/testthat/test-end-to-end.R in that repo for the reference implementation of the densification.

Follow-on from `census_of_governments_finance_pipeline#64` DoD 7. Filed after the corpus was sparsified and published on 2026-07-29 (`pipeline_commit 83f9715`). ## What changed under us The published corpus no longer stores the wide era's explicit zeros. 331.4M rows -> 45.9M (-86%). Absence now means two different things: | year | cell absent means | |---|---| | <= FY2011 (`dense_source`) | Census published **$0** | | >= FY2012 (`sparse_source`) | **not reported** (unknown) | Two new corpus artifacts carry the rule: `data/representation.parquet` (per year: `representation`, `absence_means`, `code_set_id`) and `data/code_set.parquet` (`code_set_id`, `year`, `type`, `item_code`, `is_aggregate`, `n_units` -- 84,693 rows across 55 years). The government universe is already in `canonical_fips_xwalk`. Catalogued as `SB194`. ## The user-visible consequence A wide-era query whose rows were all `$0` now returns **nothing** instead of `$0` rows, and the reader offers no way to get them back. Verified against the live corpus: `cog_spending(govid="552025209777", years=2005L)` returns 27 rows, 0 of them zero-amount. Before sparsification the all-zero categories came back as explicit `$0`. This is strictly less information than the reader used to expose, which is why #64 filed this follow-on. ## What to build `complete = TRUE` on the money verbs: fill the requested grid from `code_set` (scoped to each government's `type` -- filling against the union of all types invents rows like "$0 state IG transfer to school districts" for counties) and stamp each row's `value_source`: | value | meaning | |---|---| | `reported` | a row exists in `long` | | `census_zero` | dense-source year, cell absent | | `not_reported` | sparse-source year, cell absent | Strictly more information than the reader had before sparsification, where a wide-era zero and a modern absent cell were different shapes with nothing explaining the difference. ## Prerequisite **The bundled fixture corpus is stale and must be regenerated first.** `inst/extdata/fixture_corpus/` has no `J` rows in `summary_categories` (it predates the #65 crosswalk work), no `representation.parquet`, no `code_set.parquet`, and a still-dense wide era. Every test in this package and in cog-api runs against it, so it currently does not represent what is published. ## Definition of done 1. Fixture regenerated from the current publish tree. 2. `complete=` implemented on `cog_spending()` / `cog_revenue()`, defaulting to today's behaviour. 3. `value_source` on every returned row when `complete = TRUE`. 4. A round-trip test: for a dense-source government-year, `complete = TRUE` reproduces the pre-sparsification row set exactly, zeros included. 5. `cog-api` follow-on filed (threading `complete=` behind its own pagination fixes). The pipeline asserts this same round-trip on every build -- see `tests/testthat/test-end-to-end.R` in that repo for the reference implementation of the densification.
jared closed this issue 2026-07-30 10:32:40 -04:00
jared reopened this issue 2026-07-30 10:33:57 -04:00
Author
Owner

Reopened. PR #20 landed DoD 1 only — the fixture regeneration this issue names as its prerequisite. My PR body said "closes #18's stated prerequisite" and Gitea's keyword parser took the closes #18 out of it; my mistake.

Still open:

  • 1. Fixture regenerated from the current publish tree (PR #20 — pipeline_commit 83f9715, wide era sparse, representation.parquet + code_set.parquet now shipped).
  • 2. complete= on cog_spending() / cog_revenue(), defaulting to today's behaviour.
  • 3. value_source (reported / census_zero / not_reported) on every row when complete = TRUE.
  • 4. Round-trip test: for a dense-source government-year, complete = TRUE reproduces the pre-sparsification row set exactly, zeros included.
  • 5. cog-api follow-on filed (threads complete= behind its own pagination fixes — cog-api#6).

The fixture now carries everything DoD 2-4 need, so the remaining work is self-contained. Note the reader has no views over representation/code_set yet: the parquets ship in the fixture and the publish tree, but inst/sql/ registers nothing for them, so step 2 starts with two new view definitions.

Reopened. PR #20 landed **DoD 1 only** — the fixture regeneration this issue names as its prerequisite. My PR body said "closes #18's stated prerequisite" and Gitea's keyword parser took the `closes #18` out of it; my mistake. Still open: - [x] **1.** Fixture regenerated from the current publish tree (PR #20 — `pipeline_commit 83f9715`, wide era sparse, `representation.parquet` + `code_set.parquet` now shipped). - [ ] **2.** `complete=` on `cog_spending()` / `cog_revenue()`, defaulting to today's behaviour. - [ ] **3.** `value_source` (`reported` / `census_zero` / `not_reported`) on every row when `complete = TRUE`. - [ ] **4.** Round-trip test: for a dense-source government-year, `complete = TRUE` reproduces the pre-sparsification row set exactly, zeros included. - [ ] **5.** cog-api follow-on filed (threads `complete=` behind its own pagination fixes — cog-api#6). The fixture now carries everything DoD 2-4 need, so the remaining work is self-contained. Note the reader has no views over `representation`/`code_set` yet: the parquets ship in the fixture and the publish tree, but `inst/sql/` registers nothing for them, so step 2 starts with two new view definitions.
jared closed this issue 2026-07-30 12:04:28 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#18