Compare commits
11
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
af85a23ea7
|
||
|
|
8db944e4a0 | ||
|
|
2e8383b098
|
||
|
|
d006dea6e4
|
||
|
|
ebac39e6de | ||
|
|
47dc08c4b0 | ||
|
|
1d553a788f
|
||
|
|
c375c55da7
|
||
|
|
82acda6f93 | ||
|
|
9233c3d18e
|
||
|
|
1f257812b6 |
@@ -1,5 +1,83 @@
|
||||
# uscogdata 0.1.0 (development)
|
||||
|
||||
## `complete = TRUE`: absent cells, labelled with why they are absent
|
||||
|
||||
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
|
||||
(today's behaviour). With `complete = TRUE` the requested grid is filled
|
||||
from the corpus's `code_set` table and every row carries a new
|
||||
`value_source` column:
|
||||
|
||||
| `value_source` | meaning | `amt_nominal` |
|
||||
|---|---|---|
|
||||
| `reported` | the corpus carries this cell | as published |
|
||||
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
|
||||
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
|
||||
|
||||
The `NA` is deliberate and is the whole point: filling a modern absence
|
||||
with `0` would invent data, which is precisely the error the corpus's
|
||||
representation contract exists to prevent.
|
||||
* This restores information the reader lost when the corpus was sparsified
|
||||
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
|
||||
had begun returning nothing at all — and improves on what came before it,
|
||||
since the pre-sparsification corpus could not distinguish a published zero
|
||||
from an unreported cell either.
|
||||
* The grid is scoped to each government's **own type**, so a county is never
|
||||
filled with cells only a state can report.
|
||||
* Needs a corpus published from 2026-07-29 onward (when `representation` and
|
||||
`code_set` began shipping); aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
|
||||
listing those tables rather than on `schema_version`, which was never
|
||||
bumped for the change. Not available with `recipe` or
|
||||
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
|
||||
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
|
||||
`absence_means` rule; `cog_explain()` prints a "Completion" section.
|
||||
|
||||
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
|
||||
|
||||
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
|
||||
corpus as a whole rather than about one item code. `series_break_refs` is
|
||||
built by matching `fin_code` against the item codes in the result, and no
|
||||
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
|
||||
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
|
||||
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
|
||||
representation change at FY2012) and `SB086` (the government id scheme
|
||||
change at FY2017).
|
||||
* Provenance gains `corpus_break_refs`, selected on the break-year window
|
||||
alone and disjoint from `series_break_refs` by construction, so a consumer
|
||||
can tell a whole-result caveat from a break in one series. `cog_explain()`
|
||||
prints them under their own "Corpus-wide caveats" heading. cog-api passes
|
||||
provenance through verbatim, so the field appears there without an API
|
||||
change.
|
||||
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
|
||||
crosses the boundary where an absent cell stops meaning "Census published
|
||||
`$0`" and starts meaning "not reported", and until now nothing said so.
|
||||
|
||||
## Bundled fixture regenerated against the sparsified corpus
|
||||
|
||||
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
|
||||
2026-07-29 (`pipeline_commit 83f9715`, schema v6). The wide era no longer
|
||||
stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of
|
||||
which none are `$0`. **Absence now means two different things** — in a
|
||||
`dense_source` year (≤ FY2011) an absent cell means Census published `$0`;
|
||||
in a `sparse_source` year (≥ FY2012) it means not reported. The corpus
|
||||
carries that rule in two new tables the fixture now ships,
|
||||
`representation.parquet` and `code_set.parquet`, alongside
|
||||
`census_collection_coverage.parquet` and `lineage_events.parquet`
|
||||
(all ten publish-tree metadata tables, up from six). Catalogued upstream
|
||||
as series break `SB194`.
|
||||
* `cog_categories()` gains an `assistance` spending subtype: the J-prefix
|
||||
aid/benefit codes (`J19`, `J67`, `J68`, `J85`) are categorised now that
|
||||
the upstream crosswalk covers every flow code carrying dollars.
|
||||
* Two consequences worth knowing about, both visible in provenance rather
|
||||
than in returned dollars. The harmonization block's `na_rows_excluded`
|
||||
counts only rows that exist, so wide-era codes that were zero-padded no
|
||||
longer appear there. Coverage-gap `suggestions` are presence-based for the
|
||||
same reason, so a recipe whose component codes were all `$0` for a given
|
||||
government-year is no longer suggested for it.
|
||||
* `tests/testthat/test-fixture-vintage.R` pins these structural facts, so a
|
||||
fixture left behind by a future publish fails loudly instead of letting the
|
||||
suite pass against a corpus that no longer exists.
|
||||
|
||||
## Breaking: corpus schema_version 4 (Phase P canonical ids)
|
||||
|
||||
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
|
||||
|
||||
+148
@@ -0,0 +1,148 @@
|
||||
# R/complete.R
|
||||
#
|
||||
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
|
||||
# cell the corpus does not carry still appears, labelled with WHY it is
|
||||
# missing.
|
||||
#
|
||||
# The corpus stopped storing the wide era's explicit zeros
|
||||
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
|
||||
#
|
||||
# <= FY2011 dense_source absent => Census published $0 (census_zero)
|
||||
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
|
||||
#
|
||||
# Before sparsification a wide-era query whose cells were all $0 came back as
|
||||
# explicit $0 rows; afterwards it came back empty, with nothing to say which
|
||||
# of the two meanings applied. This restores that -- and improves on it,
|
||||
# because the pre-sparsification corpus could not distinguish the two either.
|
||||
#
|
||||
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
|
||||
# That difference is the entire point: writing 0 into a modern absence would
|
||||
# invent data, which is the error the representation contract exists to stop.
|
||||
|
||||
#' @noRd
|
||||
.abort_complete_unsupported <- function(reason, alternative) {
|
||||
cli::cli_abort(c(
|
||||
"{.code complete = TRUE} is not supported for this query.",
|
||||
x = reason,
|
||||
i = alternative
|
||||
), class = "uscogdata_complete_unsupported")
|
||||
}
|
||||
|
||||
#' @noRd
|
||||
.require_representation <- function(con, manifest) {
|
||||
needed <- c("representation.parquet", "code_set.parquet")
|
||||
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
|
||||
logical(1))]
|
||||
if (length(missing) == 0L) return(invisible(TRUE))
|
||||
cli::cli_abort(c(
|
||||
"This corpus does not publish the representation contract.",
|
||||
x = "Missing: {.file {missing}}.",
|
||||
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
|
||||
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
|
||||
), class = "uscogdata_representation_unavailable")
|
||||
}
|
||||
|
||||
#' The cells a government-year COULD carry: every code in force for that
|
||||
#' government's own type, mapped through `summary_categories`, restricted to
|
||||
#' the calling verb's flow prefixes and (when given) its category filter.
|
||||
#'
|
||||
#' Scoped by `govs_type` deliberately. Filling against the union of all types
|
||||
#' would invent cells that the government can never report -- a county row for
|
||||
#' "state IG transfer to school districts" -- and those inventions would then
|
||||
#' be indistinguishable from real census zeros.
|
||||
#'
|
||||
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
|
||||
#' aggregate rows. Without it the grid would offer cells the verb structurally
|
||||
#' never returns, so every one of them would fill as a phantom $0.
|
||||
#' @noRd
|
||||
.completion_grid_sql <- function(subtype_col, govid, years, category,
|
||||
flow_prefixes) {
|
||||
category_pred <- if (is.null(category)) {
|
||||
""
|
||||
} else {
|
||||
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
|
||||
}
|
||||
sprintf(
|
||||
"SELECT DISTINCT
|
||||
cs.year,
|
||||
x.canonical_govid,
|
||||
x.gov_name,
|
||||
c.%1$s AS subtype_value,
|
||||
c.category,
|
||||
r.absence_means
|
||||
FROM code_set cs
|
||||
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
|
||||
JOIN summary_categories c ON c.item_code = cs.item_code
|
||||
JOIN representation r ON r.year = cs.year
|
||||
WHERE x.canonical_govid IN (%2$s)
|
||||
AND cs.year IN (%3$s)
|
||||
AND NOT cs.is_aggregate
|
||||
AND LEFT(cs.item_code, 1) IN (%4$s)
|
||||
AND c.category IS NOT NULL
|
||||
AND c.%1$s IS NOT NULL
|
||||
%5$s",
|
||||
subtype_col, .sql_lit_chr(govid),
|
||||
paste(as.integer(years), collapse = ","),
|
||||
.sql_lit_chr(flow_prefixes), category_pred
|
||||
)
|
||||
}
|
||||
|
||||
#' Fill `result` out to the full grid, stamping `value_source` on every row.
|
||||
#'
|
||||
#' Returns the completed tibble with a `.completion` attribute carrying the
|
||||
#' provenance block. Reported rows are passed through untouched -- filling
|
||||
#' must never alter or drop what the corpus actually published.
|
||||
#' @noRd
|
||||
.complete_result <- function(result, con, subtype_col, govid, years, category,
|
||||
flow_prefixes) {
|
||||
grid <- tibble::as_tibble(DBI::dbGetQuery(
|
||||
con, .completion_grid_sql(subtype_col, govid, years, category, flow_prefixes)
|
||||
))
|
||||
|
||||
result$value_source <- rep("reported", nrow(result))
|
||||
if (nrow(grid) == 0L) {
|
||||
attr(result, ".completion") <- list(
|
||||
applied = TRUE, rows_filled = 0L, absence_means = list()
|
||||
)
|
||||
return(result)
|
||||
}
|
||||
|
||||
names(grid)[names(grid) == "subtype_value"] <- subtype_col
|
||||
key <- function(d) {
|
||||
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
|
||||
}
|
||||
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
|
||||
|
||||
if (nrow(missing) > 0L) {
|
||||
filled <- tibble::tibble(
|
||||
year = as.integer(missing$year),
|
||||
canonical_govid = as.character(missing$canonical_govid),
|
||||
gov_name = as.character(missing$gov_name),
|
||||
category = as.character(missing$category),
|
||||
# census_zero is a value Census published; not_reported is unknown and
|
||||
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
|
||||
amt_nominal = ifelse(missing$absence_means == "census_zero",
|
||||
0, NA_real_),
|
||||
codes_included = NA_character_,
|
||||
aggregate_fallback = NA,
|
||||
value_source = as.character(missing$absence_means)
|
||||
)
|
||||
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
|
||||
if ("notes" %in% names(result)) filled$notes <- NA_character_
|
||||
|
||||
result <- dplyr::bind_rows(result, filled)
|
||||
result <- result[order(result$year, result$canonical_govid,
|
||||
result[[subtype_col]], result$category), ,
|
||||
drop = FALSE]
|
||||
}
|
||||
|
||||
rules <- unique(grid[, c("year", "absence_means")])
|
||||
attr(result, ".completion") <- list(
|
||||
applied = TRUE,
|
||||
rows_filled = nrow(missing),
|
||||
absence_means = stats::setNames(
|
||||
as.list(as.character(rules$absence_means)), as.character(rules$year)
|
||||
)
|
||||
)
|
||||
result
|
||||
}
|
||||
+26
@@ -118,11 +118,37 @@ cog_explain <- function(result, format = c("print", "list")) {
|
||||
cli::cli_ul(sugg_lines)
|
||||
}
|
||||
|
||||
if (isTRUE(prov$completion$applied)) {
|
||||
cli::cli_h2("Completion")
|
||||
cli::cli_text(
|
||||
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
|
||||
)
|
||||
rules <- prov$completion$absence_means
|
||||
if (length(rules) > 0L) {
|
||||
cli::cli_ul(vapply(names(rules), function(y) {
|
||||
sprintf("%s: an absent cell means %s", y,
|
||||
if (identical(rules[[y]], "census_zero")) {
|
||||
"Census published $0 (filled as 0)"
|
||||
} else {
|
||||
"the government did not report (filled as NA, not 0)"
|
||||
})
|
||||
}, character(1)))
|
||||
}
|
||||
}
|
||||
|
||||
if (length(prov$series_break_refs) > 0L) {
|
||||
cli::cli_h2("Series breaks")
|
||||
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
||||
}
|
||||
|
||||
# Kept in a section of its own: these qualify the whole result, so folding
|
||||
# them in with the per-code breaks above would invite reading them as a
|
||||
# caveat about one series.
|
||||
if (length(prov$corpus_break_refs) > 0L) {
|
||||
cli::cli_h2("Corpus-wide caveats")
|
||||
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
|
||||
}
|
||||
|
||||
cli::cli_h2("Transformations")
|
||||
uc <- prov$transformations$units_conversion
|
||||
if (isTRUE(uc$applied)) {
|
||||
|
||||
@@ -133,7 +133,9 @@ cog_find_peers <- function(target_govid,
|
||||
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||
#' call.
|
||||
#' call. Those summary rows are quantiles **within each category**, not
|
||||
#' quantiles of each peer's total — see the `@return` section before summing
|
||||
#' them.
|
||||
#'
|
||||
#' @param target_govid Character scalar.
|
||||
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
||||
@@ -155,6 +157,32 @@ cog_find_peers <- function(target_govid,
|
||||
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||
#' `cohort_year`, and `cohort_govids`.
|
||||
#'
|
||||
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||
#' Each one is computed **within each `(year, spend_subtype,
|
||||
#' category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||
#' median peer's value in that one category*, not *the value of the median
|
||||
#' peer's total*. The median peer for Police and the median peer for Fire
|
||||
#' are usually different governments, so summing `summary_*` rows across
|
||||
#' categories does not give any peer's total and misstates the band it
|
||||
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||
#' one cohort, with a sign flip at FY2012.
|
||||
#'
|
||||
#' Facet by `role` **and** `category` (the documented use, and what the
|
||||
#' rows are built for). For a genuine "median peer's total spending" line,
|
||||
#' sum each peer's own categories first and take the quantile of those
|
||||
#' per-government totals:
|
||||
#'
|
||||
#' ```r
|
||||
#' library(dplyr)
|
||||
#' cmp |>
|
||||
#' filter(role %in% c("target", "peer")) |>
|
||||
#' group_by(year, role, canonical_govid) |>
|
||||
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||
#' filter(role == "peer") |>
|
||||
#' group_by(year) |>
|
||||
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||
#' ```
|
||||
#' @export
|
||||
cog_peer_compare <- function(target_govid, peers, category, years,
|
||||
per_capita = TRUE, adjust_to_year = NULL,
|
||||
|
||||
+20
-2
@@ -10,7 +10,8 @@
|
||||
expenditure_concept_note = NA_character_,
|
||||
expenditure_concept_direct_suppressed = FALSE,
|
||||
harmonization = NULL, recipe = NULL,
|
||||
suggestions = list()) {
|
||||
suggestions = list(),
|
||||
completion = NULL) {
|
||||
manifest <- .uscogdata_env$manifest
|
||||
|
||||
codes <- result[["codes_included"]]
|
||||
@@ -37,11 +38,20 @@
|
||||
|
||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||
con <- .uscogdata_env$con
|
||||
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) {
|
||||
have_con <- !is.null(con) && DBI::dbIsValid(con)
|
||||
break_refs <- if (have_con) {
|
||||
.build_series_break_refs(con, codes_observed, years, schema_version)
|
||||
} else {
|
||||
character(0)
|
||||
}
|
||||
# Corpus-wide caveats travel separately: they qualify the whole result
|
||||
# rather than one series, and they do not depend on codes_observed (see
|
||||
# .build_corpus_break_refs()).
|
||||
corpus_refs <- if (have_con) {
|
||||
.build_corpus_break_refs(con, years, schema_version)
|
||||
} else {
|
||||
character(0)
|
||||
}
|
||||
|
||||
list(
|
||||
verb = verb,
|
||||
@@ -116,6 +126,14 @@
|
||||
)
|
||||
),
|
||||
series_break_refs = break_refs,
|
||||
corpus_break_refs = corpus_refs,
|
||||
# What `complete = TRUE` filled, and the rule it filled by. Always
|
||||
# present so a consumer can read `completion$applied` without testing
|
||||
# for the key -- an absent block and applied = FALSE would otherwise be
|
||||
# indistinguishable from an older reader version.
|
||||
completion = completion %||% list(
|
||||
applied = FALSE, rows_filled = 0L, absence_means = list()
|
||||
),
|
||||
manifest = list(
|
||||
schema_version = as.integer(manifest$schema_version),
|
||||
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
||||
|
||||
+6
-3
@@ -11,11 +11,13 @@
|
||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
#' and `value_source` when `complete = TRUE`.
|
||||
#' @export
|
||||
cog_revenue <- function(govid, years, category = NULL,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
complete = FALSE) {
|
||||
.verb_spendrev(
|
||||
verb = "cog_revenue",
|
||||
view_base = "revenue_annotated",
|
||||
@@ -28,6 +30,7 @@ cog_revenue <- function(govid, years, category = NULL,
|
||||
per_capita = per_capita,
|
||||
adjust_to_year = adjust_to_year,
|
||||
basis = basis,
|
||||
recipe = recipe
|
||||
recipe = recipe,
|
||||
complete = complete
|
||||
)
|
||||
}
|
||||
|
||||
+19
-7
@@ -6,8 +6,11 @@
|
||||
#' the cross-vintage canonical-government registry. Operates in two modes:
|
||||
#'
|
||||
#' * **Utility mode** (single `name`, the original behavior): returns all
|
||||
#' rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||
#' `population_acs` descending. Useful for exploratory lookups.
|
||||
#' rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||
#' substring**, sorted by `population_acs` descending. Useful for
|
||||
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||
#' government is findable by its own complete name even when that name
|
||||
#' contains parentheses or a period.
|
||||
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||
#' single canonical govid and returns a tibble in input order, suitable
|
||||
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||
@@ -19,7 +22,8 @@
|
||||
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||
#' Single hit -> resolved. Multiple -> step 4.
|
||||
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||
#' 3. **Substring fallback:** case-insensitive literal substring against
|
||||
#' `gov_name` (metacharacters escaped).
|
||||
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||
#' `status = "no_match"`. Multiple hits -> step 4.
|
||||
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||
@@ -48,7 +52,7 @@
|
||||
#' [cog_spending()], [cog_revenue()].
|
||||
#' @examples
|
||||
#' \dontrun{
|
||||
#' # Utility mode — exploratory regex lookup
|
||||
#' # Utility mode — exploratory substring lookup
|
||||
#' cog_gov_search("broward", state = "FL")
|
||||
#'
|
||||
#' # Basket mode — resolve a known cohort
|
||||
@@ -98,9 +102,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
||||
if (!is.character(name) || length(name) != 1L) {
|
||||
cli::cli_abort("`name` must be a length-1 character string.")
|
||||
}
|
||||
# Escaped, so `name` is a literal case-insensitive substring -- the same
|
||||
# treatment basket mode has always given it. Interpolating it raw made a
|
||||
# government unfindable by its own name whenever that name contains a
|
||||
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
|
||||
# match-everything wildcard, and let malformed pattern text reach the
|
||||
# engine as an error -- which cog-api surfaced as a 500, reachable by
|
||||
# typing a real name one character at a time (uscogdata#16, F-025).
|
||||
preds <- c(preds,
|
||||
sprintf("regexp_matches(gov_name, %s, 'i')",
|
||||
.sql_lit_chr(name)))
|
||||
.sql_lit_chr(.escape_regex(name))))
|
||||
}
|
||||
if (!is.null(state)) {
|
||||
st_fips <- .coerce_state_to_fips(state)
|
||||
@@ -136,8 +147,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
||||
#' @noRd
|
||||
.escape_regex <- function(x) {
|
||||
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
||||
# literal substring in the DuckDB regexp_matches call (substring fallback
|
||||
# only; utility-mode intentionally preserves regex behavior).
|
||||
# literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
|
||||
# utility mode used to interpolate raw, which was a defect rather than a
|
||||
# feature -- see the call site and uscogdata#16.
|
||||
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
||||
}
|
||||
|
||||
|
||||
+33
-1
@@ -14,9 +14,41 @@
|
||||
sql <- sprintf(
|
||||
"SELECT DISTINCT break_id
|
||||
FROM series_breaks_pq
|
||||
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d
|
||||
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
|
||||
AND break_year BETWEEN %d AND %d
|
||||
ORDER BY break_id",
|
||||
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
||||
)
|
||||
DBI::dbGetQuery(con, sql)$break_id
|
||||
}
|
||||
|
||||
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
|
||||
#' `"ALL"` rather than an item code. They qualify the whole result, so they
|
||||
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
|
||||
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
|
||||
#' before uscogdata#19. Selection is on the break_year window alone: which
|
||||
#' codes a result happens to contain is irrelevant to a caveat about the
|
||||
#' corpus.
|
||||
#'
|
||||
#' All four catalogued entries are *boundary* caveats (dollar precision
|
||||
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
|
||||
#' representation change at 2012, the id scheme change at 2017), so the same
|
||||
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
|
||||
#' path uses is the right one -- a request that never crosses the boundary
|
||||
#' is not affected by it.
|
||||
#'
|
||||
#' Returned separately from `series_break_refs` so a consumer can tell a
|
||||
#' whole-result caveat from a break in one series; the two are disjoint by
|
||||
#' construction.
|
||||
#' @noRd
|
||||
.build_corpus_break_refs <- function(con, years, schema_version) {
|
||||
if (schema_version < 5L || length(years) == 0L) return(character(0))
|
||||
sql <- sprintf(
|
||||
"SELECT DISTINCT break_id
|
||||
FROM series_breaks_pq
|
||||
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
|
||||
ORDER BY break_id",
|
||||
min(as.integer(years)), max(as.integer(years))
|
||||
)
|
||||
DBI::dbGetQuery(con, sql)$break_id
|
||||
}
|
||||
|
||||
+62
-6
@@ -70,16 +70,42 @@
|
||||
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
#' IG.
|
||||
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
||||
#' corpus does not carry still appears, labelled with **why** it is
|
||||
#' missing, and add a `value_source` column to every row:
|
||||
#'
|
||||
#' * `"reported"` — the corpus carries this cell.
|
||||
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
#' Census published `$0`. `amt_nominal` is `0`.
|
||||
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
#' government did not report, and the value is unknown. `amt_nominal` is
|
||||
#' `NA`, **not** `0` — writing a zero there would invent data.
|
||||
#'
|
||||
#' The grid comes from the corpus's `code_set` table, scoped to each
|
||||
#' government's own type, so a county is never filled with cells only a
|
||||
#' state can report. Reported rows are passed through untouched.
|
||||
#'
|
||||
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
#' when `representation`/`code_set` began shipping; aborts with class
|
||||
#' `uscogdata_representation_unavailable` otherwise. Not available with
|
||||
#' `recipe` or with `expenditure_concept = "total"` (class
|
||||
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
#' `code_set`.
|
||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
#' and `value_source` when `complete = TRUE`.
|
||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
||||
#' per-year `absence_means` rule that was applied.
|
||||
#' @export
|
||||
cog_spending <- function(govid, years, category = NULL,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("direct", "total"),
|
||||
complete = FALSE) {
|
||||
.verb_spendrev(
|
||||
verb = "cog_spending",
|
||||
view_base = "spending_annotated",
|
||||
@@ -93,7 +119,8 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
adjust_to_year = adjust_to_year,
|
||||
basis = basis,
|
||||
recipe = recipe,
|
||||
expenditure_concept = expenditure_concept
|
||||
expenditure_concept = expenditure_concept,
|
||||
complete = complete
|
||||
)
|
||||
}
|
||||
|
||||
@@ -118,7 +145,8 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
govid, years, category,
|
||||
per_capita, adjust_to_year,
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("direct", "total"),
|
||||
complete = FALSE) {
|
||||
basis_explicit <- length(basis) == 1L
|
||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
||||
@@ -165,12 +193,27 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
)
|
||||
}
|
||||
|
||||
complete <- isTRUE(complete)
|
||||
if (complete && !is.null(recipe)) {
|
||||
.abort_complete_unsupported(
|
||||
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
|
||||
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
|
||||
)
|
||||
}
|
||||
if (complete && identical(expenditure_concept, "total")) {
|
||||
.abort_complete_unsupported(
|
||||
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
|
||||
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
||||
)
|
||||
}
|
||||
|
||||
years <- as.integer(years)
|
||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||
|
||||
con <- .ensure_session()
|
||||
manifest <- .uscogdata_env$manifest
|
||||
scope <- .check_govids_in_scope(govid)
|
||||
if (complete) .require_representation(con, manifest)
|
||||
|
||||
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
||||
|
||||
@@ -201,6 +244,18 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
}
|
||||
|
||||
# Fill BEFORE per_capita / inflation so the added cells get the same
|
||||
# treatment as reported ones: a census_zero stays $0 per capita and in real
|
||||
# dollars, and a not_reported stays NA through both rather than becoming a
|
||||
# spurious 0.
|
||||
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
|
||||
if (complete) {
|
||||
result <- .complete_result(result, con, subtype_col, govid, years,
|
||||
category, flow_prefixes)
|
||||
completion <- attr(result, ".completion")
|
||||
attr(result, ".completion") <- NULL
|
||||
}
|
||||
|
||||
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
||||
if (!is.null(adjust_to_year)) {
|
||||
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
||||
@@ -309,7 +364,8 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
expenditure_concept_direct_suppressed = direct_suppressed_flag,
|
||||
harmonization = harmonization,
|
||||
recipe = recipe_block,
|
||||
suggestions = suggestions
|
||||
suggestions = suggestions,
|
||||
completion = completion
|
||||
)
|
||||
prov$scope$govids_found <- scope$found
|
||||
prov$scope$govids_missing <- scope$missing
|
||||
|
||||
@@ -32,6 +32,29 @@
|
||||
"45-ig_annotated_harmonized.sql"
|
||||
)
|
||||
|
||||
# The representation contract (cog_pipeline#64): two parquet tables that say
|
||||
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
|
||||
# on schema_version, because the sparsification that introduced them did not
|
||||
# bump the version -- the pre-sparsification corpus this package shipped
|
||||
# against until 2026-07-30 was already schema v6 and carried neither table.
|
||||
# Keying off the version number would therefore register a view over a file
|
||||
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
|
||||
# check exists to tolerate.
|
||||
.representation_view_files <- c(
|
||||
"36-representation.sql" = "representation.parquet",
|
||||
"37-code_set.sql" = "code_set.parquet"
|
||||
)
|
||||
|
||||
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
|
||||
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
|
||||
#' works identically for a local fixture and a remote share.
|
||||
#' @noRd
|
||||
.corpus_has_table <- function(manifest, file) {
|
||||
paths <- vapply(manifest$files$metadata %||% list(),
|
||||
function(f) as.character(f$path %||% ""), character(1))
|
||||
file %in% basename(paths)
|
||||
}
|
||||
|
||||
#' Register DuckDB views from inst/sql/ SQL files
|
||||
#' @noRd
|
||||
.register_views <- function(con, url, manifest) {
|
||||
@@ -39,7 +62,10 @@
|
||||
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||
for (f in files) {
|
||||
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next
|
||||
base <- basename(f)
|
||||
if (base %in% .harmonization_view_files && schema_version < 5L) next
|
||||
if (base %in% names(.representation_view_files) &&
|
||||
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
|
||||
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
||||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||||
DBI::dbExecute(con, sql)
|
||||
|
||||
@@ -19,6 +19,27 @@ package implements.
|
||||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
||||
```
|
||||
|
||||
## Amounts are in full US dollars
|
||||
|
||||
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
||||
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
|
||||
|
||||
The raw Census source files report **thousands of dollars**, and the corpus's
|
||||
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
||||
you never have to. The conversion is recorded in every result:
|
||||
|
||||
```r
|
||||
r <- cog_spending("552025209777", 2020L)
|
||||
attr(r, "provenance")$transformations$units_conversion
|
||||
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
|
||||
```
|
||||
|
||||
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
||||
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
|
||||
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
|
||||
twice overstates every figure by 1000x, and the result looks plausible rather
|
||||
than obviously wrong.
|
||||
|
||||
## Configuration
|
||||
|
||||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
||||
|
||||
@@ -11,12 +11,14 @@
|
||||
# Each partition is a full year (all states/govs) as published, so
|
||||
# Broward County FL and every other previously-pinned government stay
|
||||
# covered without any per-gov slicing logic.
|
||||
# 2. Copies the full canonical_fips_xwalk.parquet, canonical_alias.parquet,
|
||||
# summary_categories.parquet, harmonization_map.parquet,
|
||||
# harmonization_recipes.parquet, and series_breaks.parquet metadata
|
||||
# tables as-is (these are small cross-vintage registries, not
|
||||
# partitioned by year, so the fixture ships the complete tables rather
|
||||
# than a year-scoped subset).
|
||||
# 2. Copies every metadata parquet the publish tree ships (see
|
||||
# .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
|
||||
# registries, not partitioned by year, so the fixture ships the complete
|
||||
# tables rather than a year-scoped subset. representation.parquet and
|
||||
# code_set.parquet are what make the sparse wide era interpretable --
|
||||
# absence means "$0" in a dense_source year and "not reported" in a
|
||||
# sparse_source one -- so a fixture without them cannot represent the
|
||||
# published corpus.
|
||||
# 3. Resyncs the four reference docs (data_dictionary.md,
|
||||
# reader-specification.md, README.md, series_breaks.md) from the
|
||||
# publish tree's docs/.
|
||||
@@ -38,6 +40,22 @@
|
||||
# source("data-raw/regenerate_fixture_corpus.R")
|
||||
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
|
||||
|
||||
# Every metadata parquet the publish tree ships, in the order they appear in
|
||||
# the corpus manifest. Single source of truth for both the copy step and the
|
||||
# fixture manifest, so the two can never drift apart.
|
||||
.FIXTURE_METADATA_FILES <- c(
|
||||
"canonical_alias.parquet",
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"census_collection_coverage.parquet",
|
||||
"code_set.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"lineage_events.parquet",
|
||||
"representation.parquet",
|
||||
"series_breaks.parquet",
|
||||
"summary_categories.parquet"
|
||||
)
|
||||
|
||||
regenerate_fixture_corpus <- function(
|
||||
publish_cache_dir = file.path(
|
||||
"..", "cog_pipeline", "_targets", "publish_cache"
|
||||
@@ -100,20 +118,11 @@ regenerate_fixture_corpus <- function(
|
||||
invisible(NULL)
|
||||
}
|
||||
|
||||
# Copy the full (not year-scoped) canonical_fips_xwalk, canonical_alias,
|
||||
# summary_categories, and (schema v5+) harmonization_map/
|
||||
# harmonization_recipes/series_breaks parquet tables.
|
||||
# Copy the full (not year-scoped) metadata tables listed in
|
||||
# .FIXTURE_METADATA_FILES.
|
||||
#' @noRd
|
||||
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
|
||||
files <- c(
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"canonical_alias.parquet",
|
||||
"summary_categories.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"series_breaks.parquet"
|
||||
)
|
||||
for (f in files) {
|
||||
for (f in .FIXTURE_METADATA_FILES) {
|
||||
src <- file.path(publish_cache_dir, "data", f)
|
||||
dst <- file.path(fixture_dir, "data", f)
|
||||
if (!file.exists(src)) {
|
||||
@@ -179,15 +188,7 @@ regenerate_fixture_corpus <- function(
|
||||
)
|
||||
})
|
||||
|
||||
metadata_files <- c(
|
||||
"canonical_alias.parquet",
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"summary_categories.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"series_breaks.parquet"
|
||||
)
|
||||
metadata <- lapply(metadata_files, function(f) {
|
||||
metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
|
||||
rel <- file.path("data", f)
|
||||
path <- file.path(fixture_dir, rel)
|
||||
list(
|
||||
@@ -203,13 +204,16 @@ regenerate_fixture_corpus <- function(
|
||||
pipeline_commit = source_manifest$pipeline_commit,
|
||||
fixture_note = paste(
|
||||
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
|
||||
"corpus available via USCOGDATA_URL. Regenerated for Phase R2",
|
||||
"(schema_version 5, harmonization_map/harmonization_recipes/",
|
||||
"series_breaks parquet tables added). 2011/2012 straddle the",
|
||||
"wide-aggregate -> modern-leaf format boundary exercised by basis=",
|
||||
"\"harmonized\" and recipe= queries; 2019/2020 retain the prior",
|
||||
"per-capita/CPI regression anchors. Full canonical_fips_xwalk master",
|
||||
"and canonical_alias lookup table included via",
|
||||
"corpus available via USCOGDATA_URL. Regenerated from the sparsified",
|
||||
"schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
|
||||
"zeros, so FY2011 absence means Census published $0 while FY2012+",
|
||||
"absence means not reported. representation.parquet and",
|
||||
"code_set.parquet carry that rule and ship in full, as do every other",
|
||||
"metadata table in the publish tree. 2011/2012 straddle both the",
|
||||
"wide-aggregate -> modern-leaf format boundary (exercised by",
|
||||
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
|
||||
"representation boundary (SB194); 2019/2020 retain the prior",
|
||||
"per-capita/CPI regression anchors. Regenerated via",
|
||||
"data-raw/regenerate_fixture_corpus.R."
|
||||
),
|
||||
data_vintage = source_manifest$data_vintage,
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+30
-10
@@ -1,8 +1,8 @@
|
||||
{
|
||||
"schema_version": 6,
|
||||
"built_at": "2026-07-27T13:04:05Z",
|
||||
"pipeline_commit": "6098baf",
|
||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.",
|
||||
"built_at": "2026-07-30T14:07:36Z",
|
||||
"pipeline_commit": "83f9715",
|
||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
|
||||
"data_vintage": {
|
||||
"source_vintages": {
|
||||
"2012": "10162019",
|
||||
@@ -36,9 +36,9 @@
|
||||
{
|
||||
"year": 2011,
|
||||
"path": "data/long/year=2011/part-0.parquet",
|
||||
"sha256": "84302ab364dc9fc3b3fbbc3c3f8b826e3508b4d73ff7c42d094d3863cd1e37b5",
|
||||
"row_count": 2864212,
|
||||
"size_bytes": 3845911
|
||||
"sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
|
||||
"row_count": 496004,
|
||||
"size_bytes": 2202455
|
||||
},
|
||||
{
|
||||
"year": 2012,
|
||||
@@ -74,9 +74,14 @@
|
||||
"description": "canonical_fips_xwalk.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/summary_categories.parquet",
|
||||
"sha256": "0985b607f3f35a8dff62c0561261ab6922423b81d11c07b03bcb3e3461f85e33",
|
||||
"description": "summary_categories.parquet"
|
||||
"path": "data/census_collection_coverage.parquet",
|
||||
"sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
|
||||
"description": "census_collection_coverage.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/code_set.parquet",
|
||||
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
|
||||
"description": "code_set.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/harmonization_map.parquet",
|
||||
@@ -88,10 +93,25 @@
|
||||
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
|
||||
"description": "harmonization_recipes.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/lineage_events.parquet",
|
||||
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
|
||||
"description": "lineage_events.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/representation.parquet",
|
||||
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
|
||||
"description": "representation.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/series_breaks.parquet",
|
||||
"sha256": "b0b6794b6887a4f300079adfa10029c2a77109faa4952fbff1c5a270793cc02b",
|
||||
"sha256": "5ae050dd7a76c4d25e5f99e7c2e81c1896482e3504e0443b47ab5d78ba148953",
|
||||
"description": "series_breaks.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/summary_categories.parquet",
|
||||
"sha256": "e71d6d70d767c26c983fe56213baf204355f879582aa94841e62d9aea1877f83",
|
||||
"description": "summary_categories.parquet"
|
||||
}
|
||||
]
|
||||
},
|
||||
|
||||
@@ -33,6 +33,20 @@
|
||||
"aggregate_fallback": { "type": ["object", "null"] },
|
||||
"transformations":{ "type": "object" },
|
||||
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
||||
"completion": {
|
||||
"type": "object",
|
||||
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
|
||||
"properties": {
|
||||
"applied": { "type": "boolean" },
|
||||
"rows_filled": { "type": "integer" },
|
||||
"absence_means": { "type": "object" }
|
||||
}
|
||||
},
|
||||
"corpus_break_refs": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
||||
},
|
||||
"manifest": { "type": "object" },
|
||||
"sql_query": { "type": "string" }
|
||||
}
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
CREATE OR REPLACE VIEW representation AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/representation.parquet');
|
||||
@@ -0,0 +1,3 @@
|
||||
CREATE OR REPLACE VIEW code_set AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/code_set.parquet');
|
||||
@@ -32,8 +32,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
||||
}
|
||||
\details{
|
||||
* **Utility mode** (single `name`, the original behavior): returns all
|
||||
rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||
`population_acs` descending. Useful for exploratory lookups.
|
||||
rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||
substring**, sorted by `population_acs` descending. Useful for
|
||||
exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||
government is findable by its own complete name even when that name
|
||||
contains parentheses or a period.
|
||||
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||
single canonical govid and returns a tibble in input order, suitable
|
||||
for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||
@@ -45,7 +48,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
||||
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||
Single hit -> resolved. Multiple -> step 4.
|
||||
3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||
3. **Substring fallback:** case-insensitive literal substring against
|
||||
`gov_name` (metacharacters escaped).
|
||||
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||
`status = "no_match"`. Multiple hits -> step 4.
|
||||
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||
@@ -58,7 +62,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
|
||||
}
|
||||
\examples{
|
||||
\dontrun{
|
||||
# Utility mode — exploratory regex lookup
|
||||
# Utility mode — exploratory substring lookup
|
||||
cog_gov_search("broward", state = "FL")
|
||||
|
||||
# Basket mode — resolve a known cohort
|
||||
|
||||
+29
-1
@@ -43,11 +43,39 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||
`cohort_year`, and `cohort_govids`.
|
||||
|
||||
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||
Each one is computed **within each `(year, spend_subtype,
|
||||
category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||
median peer's value in that one category*, not *the value of the median
|
||||
peer's total*. The median peer for Police and the median peer for Fire
|
||||
are usually different governments, so summing `summary_*` rows across
|
||||
categories does not give any peer's total and misstates the band it
|
||||
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||
one cohort, with a sign flip at FY2012.
|
||||
|
||||
Facet by `role` **and** `category` (the documented use, and what the
|
||||
rows are built for). For a genuine "median peer's total spending" line,
|
||||
sum each peer's own categories first and take the quantile of those
|
||||
per-government totals:
|
||||
|
||||
```r
|
||||
library(dplyr)
|
||||
cmp |>
|
||||
filter(role %in% c("target", "peer")) |>
|
||||
group_by(year, role, canonical_govid) |>
|
||||
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||
filter(role == "peer") |>
|
||||
group_by(year) |>
|
||||
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||
```
|
||||
}
|
||||
\description{
|
||||
Pulls spending for the target plus a peer set (either a
|
||||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||
call.
|
||||
call. Those summary rows are quantiles **within each category**, not
|
||||
quantiles of each peer's total — see the `@return` section before summing
|
||||
them.
|
||||
}
|
||||
|
||||
+27
-2
@@ -11,7 +11,8 @@ cog_revenue(
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL
|
||||
recipe = NULL,
|
||||
complete = FALSE
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -55,12 +56,36 @@ argument is ignored and the result's provenance reports
|
||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
}
|
||||
\description{
|
||||
Mirror of [cog_spending()] for revenue categories. One row per
|
||||
|
||||
+56
-29
@@ -12,7 +12,8 @@ cog_spending(
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")
|
||||
expenditure_concept = c("direct", "total"),
|
||||
complete = FALSE
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -58,40 +59,66 @@ FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{expenditure_concept}{`"direct"` (default) returns only the
|
||||
government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
from prior releases. `"total"` additionally UNIONs in the
|
||||
intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
to the state government (`L` codes, excluding the `L--` family-total
|
||||
rollup) -- so results gain rows with `spend_subtype ==
|
||||
"intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
`uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
already defines its own component codes). **Do not sum `"total"`
|
||||
results across levels of government** (e.g. state + county + city):
|
||||
a state's `M12` payment to a school district is the same dollar the
|
||||
district reports as its own direct `E12`, so summing both double-counts
|
||||
it. This matters in particular with [cog_geographic_rollup()], which
|
||||
sums across exactly that kind of multi-layer government set.
|
||||
government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
from prior releases. `"total"` additionally UNIONs in the
|
||||
intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
to the state government (`L` codes, excluding the `L--` family-total
|
||||
rollup) -- so results gain rows with `spend_subtype ==
|
||||
"intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
`uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
already defines its own component codes). **Do not sum `"total"`
|
||||
results across levels of government** (e.g. state + county + city):
|
||||
a state's `M12` payment to a school district is the same dollar the
|
||||
district reports as its own direct `E12`, so summing both double-counts
|
||||
it. This matters in particular with [cog_geographic_rollup()], which
|
||||
sums across exactly that kind of multi-layer government set.
|
||||
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG.}
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG.}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||
whose `completion` block reports `applied`, `rows_filled`, and the
|
||||
per-year `absence_means` rule that was applied.
|
||||
}
|
||||
\description{
|
||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||
|
||||
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
|
||||
if (nzchar(p)) paste0(p, "/") else ""
|
||||
}
|
||||
|
||||
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
|
||||
# or "" when it isn't there.
|
||||
#
|
||||
# Tests that assert on documentation content have to read the sources, and the
|
||||
# sources only exist when the suite runs from a checkout. Under R CMD check the
|
||||
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
|
||||
# shipped and `../../README.md` does not resolve -- so those tests must skip
|
||||
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
|
||||
# rcmdcheck, so the assertions are still enforced on every push; this only
|
||||
# stops them from failing a context that structurally cannot satisfy them.
|
||||
source_tree_path <- function(...) {
|
||||
p <- testthat::test_path("..", "..", ...)
|
||||
if (file.exists(p)) p else ""
|
||||
}
|
||||
|
||||
# Skip unless every named source file is present (see source_tree_path()).
|
||||
skip_if_no_source_tree <- function(...) {
|
||||
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
|
||||
character(1))
|
||||
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
|
||||
testthat::skip_if(
|
||||
any(missing),
|
||||
"package source tree not available (running against the installed package)"
|
||||
)
|
||||
invisible(paths)
|
||||
}
|
||||
|
||||
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
||||
skip_if_no_corpus <- function() {
|
||||
p <- fixture_corpus_path()
|
||||
@@ -58,6 +85,42 @@ with_doctored_schema_version <- function(version, code) {
|
||||
force(code)
|
||||
}
|
||||
|
||||
# Copy the bundled fixture to a temp dir with representation.parquet and
|
||||
# code_set.parquet removed (and dropped from the manifest's metadata list),
|
||||
# then run `code` against it. Models a corpus published BEFORE sparsification:
|
||||
# schema_version is left alone deliberately, because it was never bumped for
|
||||
# that change -- the pre-sparsification fixture this package shipped until
|
||||
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
|
||||
# is therefore the only honest signal, and this helper is what proves the
|
||||
# package keys off it rather than off the version number.
|
||||
with_corpus_missing_representation <- function(code) {
|
||||
src <- fixture_corpus_path()
|
||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
||||
|
||||
dropped <- c("representation.parquet", "code_set.parquet")
|
||||
file.remove(file.path(tmp, "data", dropped))
|
||||
|
||||
manifest_path <- file.path(tmp, "manifest.json")
|
||||
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
|
||||
m$files$metadata <- Filter(
|
||||
function(f) !basename(f$path) %in% dropped, m$files$metadata
|
||||
)
|
||||
writeLines(
|
||||
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
|
||||
manifest_path
|
||||
)
|
||||
|
||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
||||
uscogdata:::cog_close()
|
||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
||||
on.exit({
|
||||
uscogdata:::cog_close()
|
||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
||||
}, add = TRUE)
|
||||
force(code)
|
||||
}
|
||||
|
||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
||||
# rewritten to drop every M/L (intergovernmental) row, then run `code`
|
||||
# against it with a clean session (mirrors with_fixture_corpus()/
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# Helper for the Madison-walkthrough finding tests (uscogdata #11-#16).
|
||||
#
|
||||
# Those tests all assert something about what a `cog_*` verb includes or
|
||||
# excludes. The expected amounts must therefore come from the RAW corpus, never
|
||||
# from the verb under test: verifying an absence through the filter that creates
|
||||
# it proves nothing. `wt_raw_*()` opens its own DuckDB connection straight onto
|
||||
# the corpus's `long` parquet partitions, bypassing uscogdata's SQL views (and
|
||||
# therefore its `flow_prefixes` filtering) entirely.
|
||||
|
||||
wt_corpus_glob <- function() {
|
||||
url <- Sys.getenv("USCOGDATA_URL")
|
||||
if (!nzchar(url)) testthat::skip("USCOGDATA_URL is not set")
|
||||
paste0(sub("/$", "", url), "/data/long/**/*.parquet")
|
||||
}
|
||||
|
||||
wt_raw_query <- function(sql) {
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbGetQuery(con, sql)
|
||||
}
|
||||
|
||||
# Sum of `amt` (in $1,000s, as the corpus stores it) for one government-year,
|
||||
# restricted either to an explicit set of item codes or to a set of first-letter
|
||||
# prefixes. Aggregate rows are excluded, matching every published verb.
|
||||
wt_raw_amt <- function(govid, year, codes = NULL, prefixes = NULL) {
|
||||
stopifnot(xor(is.null(codes), is.null(prefixes)))
|
||||
filter_sql <- if (!is.null(codes)) {
|
||||
paste0("item_code IN (", paste0("'", codes, "'", collapse = ", "), ")")
|
||||
} else {
|
||||
paste0("LEFT(item_code, 1) IN (", paste0("'", prefixes, "'", collapse = ", "), ")")
|
||||
}
|
||||
out <- wt_raw_query(paste0(
|
||||
"SELECT COALESCE(SUM(amt), 0) AS amt FROM read_parquet('", wt_corpus_glob(), "') ",
|
||||
"WHERE canonical_govid = '", govid, "' AND year = ", year,
|
||||
" AND NOT is_aggregate AND ", filter_sql
|
||||
))
|
||||
out$amt[[1]]
|
||||
}
|
||||
|
||||
# The item codes a verb reports having summed, flattened out of the
|
||||
# comma-separated `codes_included` column.
|
||||
wt_codes_included <- function(df) {
|
||||
sort(unique(trimws(unlist(strsplit(stats::na.omit(df$codes_included), ",")))))
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
# Madison walkthrough audit -- finding F-004. Tracked as uscogdata#15.
|
||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# The raw Census files report thousands of dollars; this package multiplies by
|
||||
# 1000 and returns full US dollars. That is the friendlier choice and is not
|
||||
# wrong -- but cog_explorer's CLAUDE.md states "All raw `amt` values are in
|
||||
# $1,000s", so a reader who applies that rule to amt_nominal overstates every
|
||||
# figure by 1000x, and gets a plausible-looking number rather than an obvious
|
||||
# error. The audit rates this the highest-consequence definitional gap it found.
|
||||
#
|
||||
# Deliberately NOT asserted here: man/cog_spending.Rd and man/cog_revenue.Rd,
|
||||
# which ALREADY carry the statement in their @return sections (verified
|
||||
# 2026-07-29), as does cog-api's data-dictionary.md (since 2b71b41). The gap is
|
||||
# in the surfaces a reader meets first and in cog_explorer's own conventions
|
||||
# doc -- see uscogdata#15 for the full surface-by-surface table and for the two
|
||||
# secondary tasks (cog_explorer/CLAUDE.md, which has no git remote, and
|
||||
# cog-api's llms.txt, which is silent on units).
|
||||
|
||||
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
||||
|
||||
# README and vignettes ship only in the source tree, not in the installed
|
||||
# package, so these assertions cannot run under R CMD check -- CI's earlier
|
||||
# testthat::test_local() step is what enforces them. See
|
||||
# skip_if_no_source_tree() in helper-fixture.R.
|
||||
docs <- skip_if_no_source_tree(
|
||||
"README.md",
|
||||
c("vignettes", "total-spending.Rmd"),
|
||||
c("vignettes", "population-denominators.Rmd")
|
||||
)
|
||||
|
||||
says_units <- function(path) {
|
||||
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
||||
grepl("full US dollars|full U\\.S\\. dollars", txt, ignore.case = TRUE) &&
|
||||
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
||||
}
|
||||
|
||||
for (path in docs) expect_true(says_units(path))
|
||||
|
||||
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
||||
# The expected raw amount is read straight from the corpus's parquet
|
||||
# partitions -- never through cog_spending(), which is the thing being
|
||||
# described. Madison FY2020: E/F/G = 623,347 ($1,000s) -> $623,347,000.
|
||||
raw_thousands <- wt_raw_amt("552025209777", 2020L, prefixes = c("E", "F", "G"))
|
||||
expect_equal(raw_thousands, 623347)
|
||||
|
||||
returned <- cog_spending(govid = "552025209777", years = 2020L)
|
||||
expect_equal(sum(returned$amt_nominal), raw_thousands * 1000)
|
||||
|
||||
units <- attr(returned, "provenance")$transformations$units_conversion
|
||||
expect_true(units$applied)
|
||||
expect_equal(units$multiplier, 1000)
|
||||
})
|
||||
@@ -15,7 +15,11 @@ test_that("cog_categories(type = 'spending') returns only expenditure rows", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_categories(type = "spending")
|
||||
expect_true(all(r$category_type == "expenditure"))
|
||||
expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental")))
|
||||
# "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
|
||||
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
|
||||
# carrying dollars now maps to a category.
|
||||
expect_true(all(r$subtype %in%
|
||||
c("operations", "capital", "intergovernmental", "assistance")))
|
||||
})
|
||||
|
||||
test_that("cog_categories surfaces the intergovernmental spending subtype", {
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
# tests/testthat/test-complete.R
|
||||
#
|
||||
# uscogdata#18. The published corpus no longer stores the wide era's explicit
|
||||
# zeros (cog_pipeline#64, series break SB194), so absence means two different
|
||||
# things:
|
||||
#
|
||||
# <= FY2011 (dense_source) : cell absent => Census published $0
|
||||
# >= FY2012 (sparse_source): cell absent => not reported, unknown
|
||||
#
|
||||
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
|
||||
# row's `value_source` so the two are distinguishable. Expected row sets here
|
||||
# are built from the corpus parquet directly, never from the verb under test --
|
||||
# verifying what a filter does through that same filter proves nothing.
|
||||
|
||||
# The (subtype, category) cells that SHOULD exist for one government-year:
|
||||
# every code in force for that government's type, mapped through
|
||||
# summary_categories, matching the verb's flow prefixes and excluding
|
||||
# aggregate-flagged codes (which spending_long/revenue_long drop).
|
||||
raw_expected_cells <- function(govid, year, prefixes, subtype_col) {
|
||||
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
|
||||
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
|
||||
wt_raw_query(sprintf(
|
||||
"SELECT DISTINCT c.%s AS subtype, c.category
|
||||
FROM %s cs
|
||||
JOIN %s x ON x.govs_type = cs.type
|
||||
JOIN %s c ON c.item_code = cs.item_code
|
||||
WHERE x.canonical_govid = '%s'
|
||||
AND cs.year = %d
|
||||
AND NOT cs.is_aggregate
|
||||
AND LEFT(cs.item_code, 1) IN (%s)
|
||||
AND c.category IS NOT NULL
|
||||
AND c.%s IS NOT NULL",
|
||||
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
|
||||
q("summary_categories.parquet"), govid, year,
|
||||
paste0("'", prefixes, "'", collapse = ","), subtype_col
|
||||
))
|
||||
}
|
||||
|
||||
test_that("complete = FALSE is the default and changes nothing", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
plain <- cog_spending("121011212191", 2011L)
|
||||
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
|
||||
expect_equal(nrow(plain), nrow(explicit))
|
||||
expect_false("value_source" %in% names(plain))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# FY2011 is dense_source: before sparsification this government carried a
|
||||
# row for every code in force, most of them $0. complete = TRUE must
|
||||
# reproduce that cell set exactly.
|
||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
expected <- raw_expected_cells("121011212191", 2011L,
|
||||
c("E", "F", "G"), "spend_subtype")
|
||||
|
||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||
expect_setequal(key(r$spend_subtype, r$category),
|
||||
key(expected$subtype, expected$category))
|
||||
expect_gt(nrow(expected), 0L)
|
||||
|
||||
# Every filled cell in a dense-source year is a Census-published $0 --
|
||||
# never "unknown", which is what the modern era's absences mean.
|
||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
|
||||
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
plain <- cog_spending("121011212191", 2011L)
|
||||
full <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
|
||||
# Filling adds rows; it must never alter or drop one.
|
||||
expect_gt(nrow(full), nrow(plain))
|
||||
reported <- full[full$value_source == "reported", ]
|
||||
expect_equal(nrow(reported), nrow(plain))
|
||||
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
|
||||
# ... and the total is unchanged, because every added cell is $0.
|
||||
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("a sparse-source year's absences are unknown, not zero", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# FY2019 is sparse_source: an absent cell means the government did not
|
||||
# report, which is NOT a zero. Filling those with 0 would invent data --
|
||||
# the exact error the representation contract exists to prevent.
|
||||
r <- cog_spending("121011212191", 2019L, complete = TRUE)
|
||||
filled <- r[r$value_source != "reported", ]
|
||||
expect_gt(nrow(filled), 0L)
|
||||
expect_true(all(filled$value_source == "not_reported"))
|
||||
expect_true(all(is.na(filled$amt_nominal)))
|
||||
expect_false(any(r$value_source == "census_zero"))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the fill is scoped to each government's own type", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Filling against the union of all types would invent cells for codes a
|
||||
# county can never report. Every filled category must be one that
|
||||
# code_set puts in force for type 1 (county) specifically.
|
||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
county_cells <- raw_expected_cells("121011212191", 2011L,
|
||||
c("E", "F", "G"), "spend_subtype")
|
||||
expect_true(all(r$category %in% county_cells$category))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE respects the category filter", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011L, category = "Police",
|
||||
complete = TRUE)
|
||||
expect_true(all(r$category == "Police"))
|
||||
expect_true("value_source" %in% names(r))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("cog_revenue() completes on its own flow", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
|
||||
expected <- raw_expected_cells("121011212191", 2011L,
|
||||
c("T", "A", "U", "B", "C", "D"),
|
||||
"revenue_subtype")
|
||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||
expect_setequal(key(r$revenue_subtype, r$category),
|
||||
key(expected$subtype, expected$category))
|
||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("provenance records the completion and its absence rule", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
|
||||
"provenance")
|
||||
expect_true(prov$completion$applied)
|
||||
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
|
||||
expect_gt(prov$completion$rows_filled, 0L)
|
||||
|
||||
off <- attr(cog_spending("121011212191", 2011L), "provenance")
|
||||
expect_false(off$completion$applied)
|
||||
expect_equal(off$completion$rows_filled, 0L)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE is refused where the fill would be guesswork", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# A recipe defines its own component codes and does not go through
|
||||
# summary_categories at all, so there is no grid to fill from.
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
# The intergovernmental leg keeps aggregate rows by design
|
||||
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, expenditure_concept = "total",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE aborts on a corpus with no representation contract", {
|
||||
skip_if_no_corpus()
|
||||
# A corpus published before sparsification carries neither table, so there
|
||||
# is nothing to fill from and no rule saying what an absence means. That
|
||||
# must abort rather than guess.
|
||||
with_corpus_missing_representation({
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, complete = TRUE),
|
||||
class = "uscogdata_representation_unavailable"
|
||||
)
|
||||
# ... while an ordinary query on the same corpus still works.
|
||||
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,94 @@
|
||||
# tests/testthat/test-corpus-breaks.R
|
||||
#
|
||||
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
|
||||
# are caveats about the corpus itself rather than about one item code:
|
||||
#
|
||||
# SB085 1977 dollar precision across the 1976/1977 boundary
|
||||
# SB087 2002 imputation exclusion FY2002-2006
|
||||
# SB194 2012 dense -> sparse representation change
|
||||
# SB086 2017 government ID scheme change
|
||||
#
|
||||
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
|
||||
# and no row's item_code is ever the literal "ALL", so none of them could
|
||||
# ever reach a user. They now travel in their own provenance field,
|
||||
# `corpus_break_refs`, which keeps them distinguishable from the
|
||||
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
|
||||
# result, not one series).
|
||||
|
||||
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
|
||||
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
|
||||
# DoD 4 intended to reach users.
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_true("SB194" %in% prov$corpus_break_refs)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# 2019-2020 spans no catalogued corpus-wide break.
|
||||
r <- cog_spending("121011212191", 2019:2020, "Police")
|
||||
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("corpus_break_refs and series_break_refs stay disjoint", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_type(prov$series_break_refs, "character")
|
||||
expect_type(prov$corpus_break_refs, "character")
|
||||
# An ALL caveat must never masquerade as a break in a specific series.
|
||||
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
|
||||
expect_false("SB194" %in% prov$series_break_refs)
|
||||
})
|
||||
})
|
||||
|
||||
test_that(".build_corpus_break_refs matches on the break_year window alone", {
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
|
||||
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
|
||||
# series_breaks table is a full cross-vintage registry, so the matching
|
||||
# logic is testable there even though no long partition covers it.
|
||||
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||
con, years = 1975:1980, schema_version = 6L
|
||||
))
|
||||
# ... and does not fire for a range that misses it, unlike a filter keyed
|
||||
# on the era rather than the boundary.
|
||||
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||
con, years = 1978:1980, schema_version = 6L
|
||||
))
|
||||
|
||||
# Unlike code-specific refs, these do not depend on which codes a result
|
||||
# happens to contain -- that dependency is the whole defect.
|
||||
expect_setequal(
|
||||
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
|
||||
"SB087"
|
||||
)
|
||||
|
||||
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
|
||||
expect_equal(
|
||||
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
|
||||
character(0)
|
||||
)
|
||||
})
|
||||
|
||||
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
out <- paste(c(
|
||||
capture.output(cog_explain(r)),
|
||||
capture.output(cog_explain(r), type = "message")
|
||||
), collapse = "\n")
|
||||
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
|
||||
expect_match(out, "SB194", fixed = TRUE)
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,92 @@
|
||||
# Madison walkthrough audit -- findings F-020 and F-023. Tracked as uscogdata#13.
|
||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# The owner's settled design (2026-07-28): a `coverage` argument on
|
||||
# cog_geographic_rollup(), cog_find_peers()/cog_peer_compare() and their
|
||||
# cog-api equivalents --
|
||||
# "all" every unit that reported that year (today's behaviour, DEFAULT)
|
||||
# "census" census years only (years ending 2 or 7)
|
||||
# "consistent" only units reporting in every requested year (balanced panel)
|
||||
# -- PLUS always-on coverage metadata on every result regardless of mode:
|
||||
# n_units_reporting, n_units_expected, is_census_year.
|
||||
#
|
||||
# Motivating principle: using these verbs correctly must not require the user to
|
||||
# know that the Census of Governments is a complete census only in years ending
|
||||
# in 2 and 7.
|
||||
#
|
||||
# The helper below accepts that metadata either as columns on the returned
|
||||
# tibble or as a per-year table in provenance$coverage -- the design fixes the
|
||||
# three field names and that they reach the caller, not the container.
|
||||
|
||||
wt_coverage <- function(x) {
|
||||
prov <- attr(x, "provenance")
|
||||
cov <- prov$coverage
|
||||
if (is.null(cov)) {
|
||||
needed <- c("year", "n_units_reporting", "n_units_expected", "is_census_year")
|
||||
expect_true(all(needed %in% names(x)))
|
||||
cov <- unique(x[, needed])
|
||||
}
|
||||
cov[order(cov$year), ]
|
||||
}
|
||||
|
||||
test_that("multi-government aggregates disclose reporting coverage on every result", {
|
||||
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
|
||||
|
||||
# -- F-020: geographic rollups -------------------------------------------
|
||||
# Wisconsin's city/village universe is 608 governments. On the bundled
|
||||
# fixture, FY2012 (a census year) has 597 of them reporting while FY2019 and
|
||||
# FY2020 (sample years) have 112 and 114 -- an 18%-98% swing that today's
|
||||
# return value says nothing about. Counts cross-checked against the raw
|
||||
# corpus, not through cog_geographic_rollup(), which is under test.
|
||||
wi <- cog_gov_search(name = NULL, state = "WI", type = "city")
|
||||
expect_equal(nrow(wi), 608L)
|
||||
|
||||
roll <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
|
||||
category = NULL, years = c(2011L, 2012L, 2019L, 2020L))
|
||||
cov <- wt_coverage(roll)
|
||||
|
||||
expect_equal(cov$n_units_expected, rep(608L, 4L))
|
||||
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
|
||||
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
|
||||
|
||||
raw_2012 <- wt_raw_query(paste0(
|
||||
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
|
||||
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ",
|
||||
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate"))
|
||||
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
|
||||
|
||||
# -- F-023: peer cohorts --------------------------------------------------
|
||||
# CHILTON CITY, WI (ACS population 4,017): a 15-peer cohort fixed at FY2012
|
||||
# reports 15 of 15 in FY2012 and only 3 of 15 in FY2019 and FY2020. Nothing
|
||||
# in cog_peer_compare()'s return distinguishes those years today.
|
||||
chilton <- "552015177095"
|
||||
peers <- cog_find_peers(chilton, year = 2012L, max_peers = 15L)
|
||||
expect_equal(nrow(peers), 15L)
|
||||
|
||||
cmp <- cog_peer_compare(target_govid = chilton, peers = peers, category = NULL,
|
||||
years = c(2012L, 2019L, 2020L), per_capita = TRUE)
|
||||
cov_peers <- wt_coverage(cmp)
|
||||
expect_equal(cov_peers$n_units_expected, rep(15L, 3L))
|
||||
expect_equal(cov_peers$n_units_reporting, c(15L, 3L, 3L))
|
||||
expect_equal(cov_peers$is_census_year, c(TRUE, FALSE, FALSE))
|
||||
|
||||
# -- the three coverage modes --------------------------------------------
|
||||
expect_equal(attr(cog_peer_compare(target_govid = chilton, peers = peers,
|
||||
category = NULL, years = c(2012L, 2019L, 2020L),
|
||||
per_capita = TRUE),
|
||||
"provenance")$coverage_mode, "all") # unchanged default
|
||||
|
||||
consistent <- cog_peer_compare(target_govid = chilton, peers = peers,
|
||||
category = NULL, years = c(2012L, 2019L, 2020L),
|
||||
per_capita = TRUE, coverage = "consistent")
|
||||
n_by_year <- tapply(consistent$canonical_govid[consistent$role == "peer"],
|
||||
consistent$year[consistent$role == "peer"],
|
||||
function(g) length(unique(g)))
|
||||
expect_equal(unname(as.integer(n_by_year)), c(3L, 3L, 3L)) # balanced panel
|
||||
|
||||
census_only <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
|
||||
category = NULL,
|
||||
years = c(2011L, 2012L, 2019L, 2020L),
|
||||
coverage = "census")
|
||||
expect_equal(sort(unique(census_only$year)), 2012)
|
||||
})
|
||||
@@ -298,8 +298,16 @@ test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to
|
||||
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
|
||||
# Direct/Total pairing. The flow-family gate in
|
||||
# .attach_ig_counterparts() must keep ig_recipe_id NULL here.
|
||||
#
|
||||
# Anchored on FL state government, not AL. Coverage is presence-based: a
|
||||
# recipe is only suggested when its component codes have rows for the
|
||||
# requested government-year. AL state's only FY2011 B47 cell was an
|
||||
# explicit zero, which the corpus no longer stores after sparsification
|
||||
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
|
||||
# FL state carries a real FY2011 B47 amount, so this exercises the guard
|
||||
# against a suggestion that genuinely fires.
|
||||
r <- suppressMessages(
|
||||
cog_spending("010000226085", years = c(2005, 2011), category = "IG Federal")
|
||||
cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
|
||||
)
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
expect_gt(length(sugg), 0L)
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
# Madison walkthrough audit -- findings F-012, F-017, F-018.
|
||||
# Tracked as uscogdata#11. See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# The owner's settled three-concept model (2026-07-28):
|
||||
# total = primary + interest + intergovernmental transfers
|
||||
# direct = primary + interest (Census's published Direct Expenditure)
|
||||
# primary = direct minus debt service (the NEW DEFAULT)
|
||||
# implemented by reclassifying on the crosswalk's `spend_type` column, NOT on
|
||||
# item-code first letters -- F-018 shows prefix `Y` carries both revenue
|
||||
# (Y01/Y02) and expenditure (Y05/Y06) codes, so no first-letter allowlist can
|
||||
# route them correctly.
|
||||
#
|
||||
# Fixture reproducibility: the finding's headline reconciliation is Madison
|
||||
# FY2022, where the corpus carries I89 = 46,609 (thousands) and Census's
|
||||
# published Direct Expenditure is $654,893,000 against cog_spending()'s
|
||||
# $608,284,000 (-7.1%). FY2022 is outside the bundled fixture's year window
|
||||
# (2011/2012/2019/2020), so the same invariant is asserted on FY2020, where the
|
||||
# fixture carries I89 = 27,704. Anyone running against the full corpus should
|
||||
# also check the FY2022 numbers above.
|
||||
|
||||
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
|
||||
testthat::skip("Blocked on uscogdata#11 (findings F-012, F-017, F-018)")
|
||||
|
||||
mad <- "552025209777" # MADISON CITY, WI
|
||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
||||
|
||||
# -- F-012: `primary` is the new default, and equals today's E/F/G figure ---
|
||||
primary <- cog_spending(govid = mad, years = 2020L)
|
||||
expect_equal(attr(primary, "provenance")$expenditure_concept, "primary")
|
||||
expect_equal(sum(primary$amt_nominal), 623347000)
|
||||
|
||||
# -- F-012: `direct` adds interest on long-term debt ------------------------
|
||||
# Expected interest read from the RAW corpus, never through cog_spending(),
|
||||
# which is the filter under test.
|
||||
interest <- wt_raw_amt(mad, 2020L, prefixes = "I")
|
||||
expect_equal(interest, 27704) # I89, in $1,000s
|
||||
|
||||
direct <- cog_spending(govid = mad, years = 2020L, expenditure_concept = "direct")
|
||||
expect_equal(sum(direct$amt_nominal), 651051000) # 623,347 + 27,704 thousands
|
||||
expect_equal(sum(direct$amt_nominal) - sum(primary$amt_nominal), interest * 1000)
|
||||
expect_true("I89" %in% wt_codes_included(direct))
|
||||
|
||||
# -- F-017: `total` carries Q12/Q18, state IG transfers to school districts --
|
||||
# Wisconsin FY2019: Q12 = 6,431,530 and Q18 = 533,391 (thousands). Today
|
||||
# neither verb's flow_prefixes contains "Q", so both are dropped from the one
|
||||
# concept that is supposed to include intergovernmental transfers.
|
||||
ig_expected <- wt_raw_amt(wi_state, 2019L, prefixes = c("M", "L", "Q"))
|
||||
expect_equal(ig_expected, 11609814) # M 4,644,893 + Q 6,964,921
|
||||
|
||||
wi_direct <- cog_spending(govid = wi_state, years = 2019L,
|
||||
expenditure_concept = "direct")
|
||||
wi_total <- cog_spending(govid = wi_state, years = 2019L,
|
||||
expenditure_concept = "total")
|
||||
|
||||
# total - direct is exactly the intergovernmental component. Asserted as a
|
||||
# delta rather than a grand total so this stays correct however the J and Y
|
||||
# families land inside `primary`.
|
||||
expect_equal(sum(wi_total$amt_nominal) - sum(wi_direct$amt_nominal),
|
||||
ig_expected * 1000)
|
||||
expect_true(all(c("Q12", "Q18") %in% wt_codes_included(wi_total)))
|
||||
|
||||
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
|
||||
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
|
||||
# payments. All four share the first letter `Y` and the spend_type
|
||||
# "Insurance Trust", so this pair of assertions is the concrete proof that
|
||||
# classification is no longer keyed on the first letter.
|
||||
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
|
||||
spend_codes <- wt_codes_included(wi_total)
|
||||
rev_codes <- wt_codes_included(wi_revenue)
|
||||
|
||||
expect_true("Y05" %in% spend_codes)
|
||||
expect_false("Y05" %in% rev_codes)
|
||||
expect_true("Y01" %in% rev_codes)
|
||||
expect_false("Y01" %in% spend_codes)
|
||||
})
|
||||
@@ -0,0 +1,112 @@
|
||||
# tests/testthat/test-fixture-vintage.R
|
||||
#
|
||||
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
|
||||
# every test in this package -- plus the whole cog-api suite -- runs against
|
||||
# it. When the published corpus changes shape and the fixture does not, both
|
||||
# suites stay green against a corpus that no longer exists (uscogdata#18).
|
||||
#
|
||||
# These tests pin the structural facts that distinguish the current published
|
||||
# vintage from its predecessor, so a stale fixture fails loudly instead of
|
||||
# passing quietly. They assert shape, never dollar values: re-running
|
||||
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
|
||||
# keep them green.
|
||||
|
||||
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
|
||||
# not the package session: these assertions are about what the fixture
|
||||
# CONTAINS, and routing them through the reader's own views would let a
|
||||
# filter hide the very absence being checked.
|
||||
fixture_query <- function(sql, ...) {
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
path <- function(rel) {
|
||||
sprintf("read_parquet(%s)",
|
||||
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
|
||||
}
|
||||
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
|
||||
}
|
||||
|
||||
test_that("fixture ships every metadata table the publish tree does", {
|
||||
skip_if_no_corpus()
|
||||
# representation/code_set are what make a sparse corpus interpretable; a
|
||||
# fixture without them predates sparsification (cog_pipeline#64).
|
||||
expected <- c(
|
||||
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
|
||||
"census_collection_coverage.parquet", "code_set.parquet",
|
||||
"harmonization_map.parquet", "harmonization_recipes.parquet",
|
||||
"lineage_events.parquet", "representation.parquet",
|
||||
"series_breaks.parquet", "summary_categories.parquet"
|
||||
)
|
||||
on_disk <- basename(list.files(
|
||||
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
|
||||
))
|
||||
expect_true(all(expected %in% on_disk))
|
||||
|
||||
# The manifest must list them too -- consumers read the manifest, not ls().
|
||||
in_manifest <- with_fixture_corpus(
|
||||
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
|
||||
)
|
||||
expect_true(all(expected %in% in_manifest))
|
||||
})
|
||||
|
||||
test_that("fixture carries the dense/sparse representation contract", {
|
||||
skip_if_no_corpus()
|
||||
rep <- fixture_query(
|
||||
"SELECT year, representation, absence_means FROM %s
|
||||
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
|
||||
"data/representation.parquet"
|
||||
)
|
||||
expect_equal(nrow(rep), 4L)
|
||||
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
|
||||
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
|
||||
})
|
||||
|
||||
test_that("the fixture's wide era is sparse, not zero-padded", {
|
||||
skip_if_no_corpus()
|
||||
# FY2011 is a dense_source year: the corpus publishes only the cells Census
|
||||
# reported non-zero, and an absent cell means Census published $0. Before
|
||||
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
|
||||
# zeros. A single explicit zero here means the fixture predates the change.
|
||||
zeros_2011 <- fixture_query(
|
||||
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
|
||||
"data/long/year=2011/part-0.parquet"
|
||||
)$n
|
||||
expect_equal(zeros_2011, 0L)
|
||||
|
||||
# The modern era is a different regime: a reported zero there is real data
|
||||
# (the government filed $0), so zeros legitimately survive and must not be
|
||||
# asserted away.
|
||||
expect_gt(
|
||||
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
|
||||
0L
|
||||
)
|
||||
})
|
||||
|
||||
test_that("code_set covers every fixture year with the reader-spec columns", {
|
||||
skip_if_no_corpus()
|
||||
cs <- fixture_query(
|
||||
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
|
||||
"data/code_set.parquet"
|
||||
)
|
||||
expect_true(all(
|
||||
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
|
||||
%in% names(cs)
|
||||
))
|
||||
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
|
||||
})
|
||||
|
||||
test_that("every flow code carrying dollars has a category, J-prefix included", {
|
||||
skip_if_no_corpus()
|
||||
# The J (assistance/benefit) codes were uncategorised until the crosswalk
|
||||
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
|
||||
# duplication fix landed). Their absence is how a pre-crosswalk fixture
|
||||
# gives itself away.
|
||||
j <- fixture_query(
|
||||
"SELECT item_code, category, category_type, spend_subtype FROM %s
|
||||
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
|
||||
"data/summary_categories.parquet"
|
||||
)
|
||||
expect_true("J19" %in% j$item_code)
|
||||
expect_true(all(j$category_type == "expenditure"))
|
||||
expect_true(all(j$spend_subtype == "assistance"))
|
||||
expect_false(any(is.na(j$category)))
|
||||
})
|
||||
@@ -0,0 +1,57 @@
|
||||
# Madison walkthrough audit -- finding F-025. Tracked as uscogdata#16.
|
||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# cog_gov_search()'s UTILITY mode interpolates `name` into
|
||||
# regexp_matches(gov_name, <name>, 'i')
|
||||
# unescaped (R/search.R:102), while BASKET mode in the same file already routes
|
||||
# it through .escape_regex() (R/search.R:307) with the comment "so `name` is
|
||||
# treated as a literal substring". Two failure modes result:
|
||||
# correctness -- a real government cannot be found by its own exact name, and
|
||||
# a single "." matches everything (HTTP 200 both ways via the API);
|
||||
# robustness -- malformed regex reaches the engine and errors, which cog-api
|
||||
# surfaces as a 500, reachable by typing a real name one
|
||||
# character at a time.
|
||||
#
|
||||
# NOT asserted here: the finding's `q=St. Louis` example. Under correct literal
|
||||
# matching that search still returns 0 rows, because the stored name is
|
||||
# "ST LOUIS CITY" with no period -- it demonstrates today's over-matching
|
||||
# semantics, not a row the fix makes findable.
|
||||
|
||||
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
||||
|
||||
# -- correctness (1): a government must be findable by its own exact name ---
|
||||
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
||||
# grouping, so its own complete name matches nothing.
|
||||
fredonia <- cog_gov_search(name = "FREDONIA (BRISCOE) CITY")
|
||||
expect_equal(nrow(fredonia), 1L)
|
||||
expect_equal(fredonia$canonical_govid, "052117184386")
|
||||
expect_equal(cog_gov_search(name = "FREDONIA (BRISCOE)")$canonical_govid,
|
||||
"052117184386")
|
||||
|
||||
# -- correctness (2): a metacharacter must not become a wildcard ------------
|
||||
# No Wisconsin city or village name contains a literal period -- established
|
||||
# against the raw registry below, NOT through the verb under test. A literal
|
||||
# search for "." must therefore return nothing; today it returns all 608.
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
xwalk <- paste0(sub("/$", "", Sys.getenv("USCOGDATA_URL")),
|
||||
"/data/canonical_fips_xwalk.parquet")
|
||||
with_dot <- DBI::dbGetQuery(con, paste0(
|
||||
"SELECT COUNT(*) n FROM read_parquet('", xwalk, "') ",
|
||||
"WHERE fips_state = '55' AND govs_type = 2 AND gov_name LIKE '%.%'"))
|
||||
expect_equal(as.integer(with_dot$n[[1]]), 0L)
|
||||
|
||||
expect_equal(nrow(cog_gov_search(name = ".", state = "WI", type = "city")), 0L)
|
||||
expect_equal(nrow(cog_gov_search(name = "M.dison", state = "WI", type = "city")), 0L)
|
||||
expect_equal(nrow(cog_gov_search(name = "Mad(i|o)son", state = "WI", type = "city")), 0L)
|
||||
|
||||
# A metacharacter-free name still resolves exactly as before.
|
||||
expect_equal(nrow(cog_gov_search(name = "Madison", state = "WI", type = "city")), 1L)
|
||||
|
||||
# -- robustness: malformed pattern text returns no rows, and does not error --
|
||||
# "[" alone, and "Athens-Clarke County (bal" -- an in-progress substring of
|
||||
# ATHENS-CLARKE COUNTY (BALANCE), a real government -- both currently raise
|
||||
# (DuckDB: "Invalid Input Error: missing ]").
|
||||
expect_equal(nrow(cog_gov_search(name = "[")), 0L)
|
||||
expect_equal(nrow(cog_gov_search(name = "Athens-Clarke County (bal")), 0L)
|
||||
})
|
||||
@@ -0,0 +1,51 @@
|
||||
# Madison walkthrough audit -- finding F-021. Tracked as uscogdata#14.
|
||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# .peer_summary_rows() computes stats::quantile() separately INSIDE each
|
||||
# (year, spend_subtype, category) cell. A summary_p50 row is therefore "the
|
||||
# median peer's value in that one category", not "the value of the median
|
||||
# peer's total". Summing those rows across categories -- the obvious move for a
|
||||
# caller who wants one peer-median total line and reads only the column names --
|
||||
# misstated a total-spending band by -32.7% to +251.0% across the 24 years the
|
||||
# audit tested, with a sign flip at FY2012.
|
||||
#
|
||||
# The verb is not wrong and its documented use (faceting by role AND category)
|
||||
# is unaffected, so the fix is documentation: one sentence in @return.
|
||||
|
||||
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
|
||||
|
||||
# man/ ships only in the source tree (the installed package carries a
|
||||
# compiled help database instead), so the prose assertions below cannot run
|
||||
# under R CMD check -- CI's earlier testthat::test_local() step enforces
|
||||
# them. The numeric pin further down needs only the corpus, but it lives in
|
||||
# the same test_that() as the sentence it protects, deliberately: they are
|
||||
# one claim, and splitting them would let the prose drift while a separate
|
||||
# test kept passing.
|
||||
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
|
||||
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
|
||||
|
||||
# The @return section must say the quantile is computed within each cell...
|
||||
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
|
||||
# ...and must warn that the rows are not additive across category.
|
||||
expect_match(rd, "not additive|do(es)? not sum|cannot be summed", ignore.case = TRUE)
|
||||
# ...naming the grouping explicitly.
|
||||
expect_match(rd, "spend_subtype", fixed = TRUE)
|
||||
|
||||
# Pin the mechanism numerically so a future refactor that quietly changes the
|
||||
# quantile grouping fails here rather than silently invalidating the sentence
|
||||
# above. Fixture: Madison, 10 peers found at FY2020, category = NULL.
|
||||
peers <- cog_find_peers("552025209777", year = 2020L, max_peers = 10L)
|
||||
cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers,
|
||||
category = NULL, years = 2020L, per_capita = TRUE)
|
||||
|
||||
naive <- sum(cmp$amt_per_capita_nominal[cmp$role == "summary_p50"], na.rm = TRUE)
|
||||
|
||||
peer_rows <- cmp[cmp$role == "peer", ]
|
||||
per_gov <- tapply(peer_rows$amt_per_capita_nominal, peer_rows$canonical_govid,
|
||||
sum, na.rm = TRUE)
|
||||
correct <- unname(stats::quantile(per_gov, 0.5, na.rm = TRUE))
|
||||
|
||||
expect_equal(round(naive), 6180) # summing the built-in summary rows
|
||||
expect_equal(round(correct), 2043) # quantile of each peer's OWN total
|
||||
expect_gt(naive / correct, 2) # a +200% misstatement on this cohort
|
||||
})
|
||||
@@ -0,0 +1,55 @@
|
||||
# Madison walkthrough audit -- finding F-014. Tracked as uscogdata#12.
|
||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
||||
#
|
||||
# cog_revenue()'s flow_prefixes = c("T","A","U","B","C","D") never returns
|
||||
# item-code prefix X (Employee Retirement) or Y (other Insurance Trust). Per
|
||||
# Census's standard identity, Total Revenue = General + Utility + Liquor Store +
|
||||
# Insurance Trust Revenue, and Employee Retirement System contributions and
|
||||
# earnings ARE the Insurance Trust Revenue component -- so prefix X sits inside
|
||||
# a published Census revenue concept exactly the way I89 sits inside Census's
|
||||
# Direct Expenditure concept (finding F-012).
|
||||
#
|
||||
# CAVEAT FOR WHOEVER PICKS THIS UP: the argument name below (`revenue_concept =
|
||||
# "total"`) is this test's *proposal*, not a settled decision. The owner's
|
||||
# 2026-07-28 resolution covers expenditure concepts only; no revenue-side
|
||||
# naming has been ruled on. If the eventual argument is named differently,
|
||||
# change the two calls here -- the asserted dollar invariants are what matter
|
||||
# and are independent of the naming.
|
||||
#
|
||||
# Fixture reproducibility: Madison's own X-prefix revenue (FY1970-FY1986,
|
||||
# $15,098,000 nominal, $0 thereafter) is outside the bundled fixture's year
|
||||
# window (2011/2012/2019/2020), so the same invariant is asserted on Wisconsin
|
||||
# state government FY2012, where the fixture carries nonzero X01/X05/X08.
|
||||
|
||||
test_that("cog_revenue() can return Census Total Revenue including Insurance Trust (prefix X)", {
|
||||
testthat::skip("Blocked on uscogdata#12 (finding F-014)")
|
||||
|
||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
||||
|
||||
# Revenue-shaped Employee Retirement codes, read from the RAW corpus rather
|
||||
# than through cog_revenue(), which is the filter under test:
|
||||
# X01 local employee contribution, X04/X05 contributions and transfers from
|
||||
# other governments, X08 earnings on investments.
|
||||
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X04", "X05", "X08"))
|
||||
expect_equal(x_revenue, 2038800) # 615,835 + 0 + 560,382 + 862,583 ($1,000s)
|
||||
|
||||
general <- cog_revenue(govid = wi_state, years = 2012L)
|
||||
expect_equal(sum(general$amt_nominal), 31338293000)
|
||||
|
||||
total <- cog_revenue(govid = wi_state, years = 2012L, revenue_concept = "total")
|
||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal), x_revenue * 1000)
|
||||
expect_equal(sum(total$amt_nominal), 33377093000)
|
||||
expect_true(all(c("X01", "X05", "X08") %in% wt_codes_included(total)))
|
||||
|
||||
# Sibling codes under the SAME first letter must stay out: X11/X12 are
|
||||
# benefit payments (an expenditure) and X21/X30/X47 are cash and securities
|
||||
# holdings (a balance-sheet stock). This is the F-018 point restated on the
|
||||
# revenue side -- the split has to come from the crosswalk's spend_type, not
|
||||
# from the letter X.
|
||||
expect_false(any(c("X11", "X12", "X21", "X30", "X47") %in% wt_codes_included(total)))
|
||||
|
||||
# Every returned row still resolves to a category. summary_categories has
|
||||
# zero rows for prefix X today, so relaxing the prefix filter alone would
|
||||
# produce category = NA rows -- see census_of_governments_finance_pipeline#60.
|
||||
expect_false(any(is.na(total$category)))
|
||||
})
|
||||
@@ -80,8 +80,11 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
|
||||
|
||||
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
||||
skip_if_no_corpus()
|
||||
fl_state <- cog_gov_search("^FLORIDA$", type = "state")
|
||||
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||
# Unanchored: utility mode matches literally now, so "^...$" would be
|
||||
# searched for as characters rather than read as anchors (uscogdata#16).
|
||||
# Both still resolve to exactly one row once scoped by type/state.
|
||||
fl_state <- cog_gov_search("FLORIDA", type = "state")
|
||||
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||
r <- cog_geographic_rollup(
|
||||
govids = list(state = fl_state, county = broward),
|
||||
category = "Police", years = 2020L
|
||||
|
||||
@@ -98,7 +98,11 @@ test_that("cog_spending rejects invalid inputs", {
|
||||
|
||||
test_that("cog_spending accepts a cog_gov_search result directly", {
|
||||
skip_if_no_corpus()
|
||||
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||
# Unanchored: utility mode matches `name` as a literal substring now, so
|
||||
# "^...$" would be searched for as those characters rather than read as
|
||||
# anchors (uscogdata#16). Scoped by state and type, the bare name still
|
||||
# resolves to exactly one row.
|
||||
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||
expect_gt(nrow(picks), 0L)
|
||||
r <- cog_spending(picks, 2020L, "Corrections")
|
||||
expect_equal(unique(r$canonical_govid), "121011212191")
|
||||
@@ -244,24 +248,42 @@ test_that("basis defaults to 'harmonized' when not passed", {
|
||||
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
||||
# FL state government. The harmonization block is scoped by government,
|
||||
# year and flow prefix -- NOT by category -- so a Corrections query still
|
||||
# counts every E/F/G-prefixed row the harmonized basis drops for having
|
||||
# no harmonized_code. The three that apply here are E21/F21/G21
|
||||
# (Education NEC, SB184-186, "discontinued_na", wide-era window ending
|
||||
# FY2011); the other discontinued_na rulings live outside E/F/G.
|
||||
# See docs/phase_r_harmonization_review.md § 1.3/1.4 and cog_pipeline
|
||||
# data/harmonization_map.csv.
|
||||
r <- cog_spending("120000226351", 2011:2012, "Corrections")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_equal(prov$basis, "harmonized")
|
||||
expect_true(prov$harmonization$applied)
|
||||
expect_true(prov$harmonization$na_rows_excluded >= 0L)
|
||||
expect_true(prov$harmonization$na_amount_excluded >= 0)
|
||||
# Data-verified for the v6 fixture (corpus 2026-07-22). The Task 18 map
|
||||
# extension added E/F/G-prefix discontinued_na rulings the earlier pin's
|
||||
# comment predated: E21/F21/G21 (Education NEC local, SB184-186,
|
||||
# "trivial; explicit-NA, full wide-era window"). Broward's 2011 legacy
|
||||
# partition zero-pads exactly those three codes, so this query now
|
||||
# excludes 3 NA-harmonized rows -- all with amt = 0, hence the excluded
|
||||
# AMOUNT stays exactly zero. (The other discontinued_na rulings -- S74,
|
||||
# Z61, X04, X06, the debt-detail family, L24 -- remain outside the
|
||||
# E/F/G/K prefixes.) See docs/phase_r_harmonization_review.md § 1.3/1.4
|
||||
# and cog_pipeline data/harmonization_map.csv E21/F21/G21 rows.
|
||||
expect_equal(prov$harmonization$na_rows_excluded, 3L)
|
||||
expect_equal(prov$harmonization$na_amount_excluded, 0)
|
||||
# $2,825,439 thousands of FY2011 E21 + F21 + G21, reported in full USD.
|
||||
# Pinning a non-zero amount is the point: the earlier Broward anchor's
|
||||
# three rows were all explicit zeros, so the AMOUNT accounting was
|
||||
# asserted only against 0 and could not have caught a bug.
|
||||
expect_equal(prov$harmonization$na_amount_excluded, 2825439 * 1000)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("sparsification removed the wide era's zero-pads from the exclusion count", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Broward County FY2011 used to carry E21/F21/G21 rows of exactly $0 --
|
||||
# the wide era stored every government x every code, zeros included. The
|
||||
# published corpus no longer does (SB194, cog_pipeline#64), so there is
|
||||
# now nothing for the harmonized basis to exclude. Absence in a
|
||||
# dense_source year means Census published $0; it does not mean the
|
||||
# exclusion machinery stopped working, which the FL state anchor above
|
||||
# proves independently.
|
||||
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
||||
h <- attr(r, "provenance")$harmonization
|
||||
expect_true(h$applied)
|
||||
expect_equal(h$na_rows_excluded, 0L)
|
||||
expect_equal(h$na_amount_excluded, 0)
|
||||
})
|
||||
})
|
||||
|
||||
@@ -303,11 +325,13 @@ test_that("provenance$series_break_refs is a populated-when-applicable character
|
||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
||||
refs <- attr(r, "provenance")$series_break_refs
|
||||
expect_type(refs, "character")
|
||||
# No catalogued series_breaks_pq row falls inside this fixture's
|
||||
# 2011/2012/2019/2020 window for the codes this query touches (E04/G04)
|
||||
# -- data-verified; the mechanism itself is what's under test here, via
|
||||
# a query-shaped unit test in test-views.R since the fixture has no
|
||||
# positive case to pin against.
|
||||
# No catalogued code-specific series_breaks_pq row falls inside this
|
||||
# fixture's 2011/2012/2019/2020 window for the codes this query touches
|
||||
# (E04/G04) -- data-verified; the mechanism itself is what's under test
|
||||
# here, via a query-shaped unit test in test-views.R since the fixture
|
||||
# has no positive case to pin against. Corpus-wide ("ALL") entries never
|
||||
# appear in this field by construction -- they travel in
|
||||
# corpus_break_refs; see test-corpus-breaks.R.
|
||||
expect_equal(refs, character(0))
|
||||
})
|
||||
})
|
||||
|
||||
@@ -177,11 +177,13 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
|
||||
})
|
||||
|
||||
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
||||
# No series_breaks_pq row falls inside the bundled fixture's 2011-2020
|
||||
# window (data-verified; see the "series_break_refs" test in
|
||||
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
|
||||
# 2011-2020 window (data-verified; see the "series_break_refs" test in
|
||||
# test-spending.R), so this proves the matching logic itself against the
|
||||
# live view + a synthetic year window that DOES hit a cataloged break
|
||||
# (SB075, fin_code E62, break_year 2005).
|
||||
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
|
||||
# separate path with its own coverage -- SB194 does sit at 2012, inside
|
||||
# the fixture window; see test-corpus-breaks.R.
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
|
||||
@@ -13,6 +13,8 @@ knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
|
||||
|
||||
# Why per-year population matters
|
||||
|
||||
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
|
||||
|
||||
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
||||
|
||||
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
||||
|
||||
@@ -29,6 +29,13 @@ controls which of these a query answers. This vignette walks through both
|
||||
questions with code that actually runs against the package's bundled fixture
|
||||
corpus, then explains why the second question refuses `"total"` outright.
|
||||
|
||||
Before any of the numbers below: every amount column here — `amt_nominal`,
|
||||
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
|
||||
dollars**. The raw Census files report **thousands of dollars** and the
|
||||
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
|
||||
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
|
||||
billion. Do not scale it again.
|
||||
|
||||
```{r}
|
||||
library(uscogdata)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user