Files
uscogdata/R/provenance.R
jared af85a23ea7
R-CMD-check / check (push) Successful in 3m7s
R-CMD-check / check (pull_request) Successful in 3m9s
feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:

  <= FY2011  dense_source   absent => Census published $0
  >= FY2012  sparse_source  absent => not reported, unknown

A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.

complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.

Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.

Scoping decisions, each of which would invent rows if taken loosely:

  - The grid is per government TYPE (code_set.type). Filling against the
    union of all types would give a county cells like "state IG transfer to
    school districts", indistinguishable from real census zeros.
  - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
    grid offers cells those views never return, so each would fill as a
    phantom $0.
  - Filling happens BEFORE per_capita and inflation, so a census_zero stays
    0 through both and a not_reported stays NA rather than becoming 0.

Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.

Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.

Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.

Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.

Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
2026-07-30 11:47:51 -04:00

145 lines
4.9 KiB
R

# R/provenance.R
# Shared provenance construction. Matches inst/schemas/provenance-v1.json.
#' @noRd
.build_provenance <- function(verb, call, govid, years, category,
per_capita, adjust_to_year, result, sql,
subtype_col, basis = NA_character_,
basis_note = NA_character_,
expenditure_concept = "direct",
expenditure_concept_note = NA_character_,
expenditure_concept_direct_suppressed = FALSE,
harmonization = NULL, recipe = NULL,
suggestions = list(),
completion = NULL) {
manifest <- .uscogdata_env$manifest
codes <- result[["codes_included"]]
codes_observed <- if (length(codes) == 0L) {
character(0)
} else {
sorted <- sort(unique(unlist(strsplit(codes, ",", fixed = TRUE))))
sorted[nzchar(sorted)]
}
agg_flag <- result[["aggregate_fallback"]]
agg_applied <- isTRUE(any(agg_flag, na.rm = TRUE))
agg_years <- if (agg_applied) {
unique(as.integer(result$year[which(agg_flag)]))
} else {
integer(0)
}
gov_names <- if (nrow(result) == 0L) {
character(0)
} else {
unique(result$gov_name)
}
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
con <- .uscogdata_env$con
have_con <- !is.null(con) && DBI::dbIsValid(con)
break_refs <- if (have_con) {
.build_series_break_refs(con, codes_observed, years, schema_version)
} else {
character(0)
}
# Corpus-wide caveats travel separately: they qualify the whole result
# rather than one series, and they do not depend on codes_observed (see
# .build_corpus_break_refs()).
corpus_refs <- if (have_con) {
.build_corpus_break_refs(con, years, schema_version)
} else {
character(0)
}
list(
verb = verb,
call = paste(deparse(call), collapse = " "),
target = list(
canonical_govid = as.character(govid),
gov_name = gov_names
),
years = as.integer(years),
category = category,
basis = basis,
basis_note = basis_note,
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note,
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
harmonization = harmonization %||% list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
note = NA_character_
),
recipe = recipe,
suggestions = suggestions,
scope = list(
gov_types_included = as.integer(unlist(manifest$scope$gov_types_included)),
gov_types_excluded = as.integer(unlist(manifest$scope$gov_types_excluded)),
scope_note = manifest$scope$scope_note %||% ""
),
codes_summed = list(
observed = codes_observed,
subtype_column = subtype_col
),
aggregate_fallback = list(
applied = agg_applied,
years = agg_years
),
transformations = list(
units_conversion = list(
applied = TRUE,
source_unit = "$1,000s (raw Census)",
target_unit = "$USD",
multiplier = 1000L
),
per_capita = list(
applied = isTRUE(per_capita),
denominator_source = if (isTRUE(per_capita)) {
"Census F-33 population (per-year, from long.population)"
} else {
NA_character_
},
popyear_range = if (isTRUE(per_capita)) {
attr(result, ".popyear_range") %||% integer(0)
} else {
integer(0)
},
pop_source_counts = if (isTRUE(per_capita)) {
ps <- result[["pop_source"]]
if (is.null(ps) || length(ps) == 0L) {
list(census_f33 = 0L, unavailable = 0L)
} else {
list(
census_f33 = sum(ps == "census_f33", na.rm = TRUE),
unavailable = sum(ps == "unavailable", na.rm = TRUE)
)
}
} else {
NULL
}
),
inflation = list(
applied = !is.null(adjust_to_year),
base_year = if (is.null(adjust_to_year)) NA_integer_ else as.integer(adjust_to_year),
index = if (is.null(adjust_to_year)) NA_character_ else "CPI-U (BLS CPIAUCSL annual average, bundled)"
)
),
series_break_refs = break_refs,
corpus_break_refs = corpus_refs,
# What `complete = TRUE` filled, and the rule it filled by. Always
# present so a consumer can read `completion$applied` without testing
# for the key -- an absent block and applied = FALSE would otherwise be
# indistinguishable from an older reader version.
completion = completion %||% list(
applied = FALSE, rows_filled = 0L, absence_means = list()
),
manifest = list(
schema_version = as.integer(manifest$schema_version),
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
built_at = manifest$built_at %||% NA_character_
),
sql_query = sql
)
}