feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
This commit is contained in:
+27
-2
@@ -11,7 +11,8 @@ cog_revenue(
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL
|
||||
recipe = NULL,
|
||||
complete = FALSE
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -55,12 +56,36 @@ argument is ignored and the result's provenance reports
|
||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
}
|
||||
\description{
|
||||
Mirror of [cog_spending()] for revenue categories. One row per
|
||||
|
||||
+56
-29
@@ -12,7 +12,8 @@ cog_spending(
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")
|
||||
expenditure_concept = c("direct", "total"),
|
||||
complete = FALSE
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -58,40 +59,66 @@ FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{expenditure_concept}{`"direct"` (default) returns only the
|
||||
government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
from prior releases. `"total"` additionally UNIONs in the
|
||||
intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
to the state government (`L` codes, excluding the `L--` family-total
|
||||
rollup) -- so results gain rows with `spend_subtype ==
|
||||
"intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
`uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
already defines its own component codes). **Do not sum `"total"`
|
||||
results across levels of government** (e.g. state + county + city):
|
||||
a state's `M12` payment to a school district is the same dollar the
|
||||
district reports as its own direct `E12`, so summing both double-counts
|
||||
it. This matters in particular with [cog_geographic_rollup()], which
|
||||
sums across exactly that kind of multi-layer government set.
|
||||
government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
from prior releases. `"total"` additionally UNIONs in the
|
||||
intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
to the state government (`L` codes, excluding the `L--` family-total
|
||||
rollup) -- so results gain rows with `spend_subtype ==
|
||||
"intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
`uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
already defines its own component codes). **Do not sum `"total"`
|
||||
results across levels of government** (e.g. state + county + city):
|
||||
a state's `M12` payment to a school district is the same dollar the
|
||||
district reports as its own direct `E12`, so summing both double-counts
|
||||
it. This matters in particular with [cog_geographic_rollup()], which
|
||||
sums across exactly that kind of multi-layer government set.
|
||||
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG.}
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG.}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||
whose `completion` block reports `applied`, `rows_filled`, and the
|
||||
per-year `absence_means` rule that was applied.
|
||||
}
|
||||
\description{
|
||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||
|
||||
Reference in New Issue
Block a user