Compare commits
8
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d95c9032c5
|
||
|
|
af85a23ea7
|
||
|
|
8db944e4a0 | ||
|
|
2e8383b098
|
||
|
|
d006dea6e4
|
||
|
|
ebac39e6de | ||
|
|
47dc08c4b0 | ||
|
|
1d553a788f
|
@@ -1,5 +1,91 @@
|
|||||||
# uscogdata 0.1.0 (development)
|
# uscogdata 0.1.0 (development)
|
||||||
|
|
||||||
|
## Multi-government aggregates now disclose their reporting coverage
|
||||||
|
|
||||||
|
* The Census of Governments is a **complete census only in years ending in 2
|
||||||
|
and 7**; every other year is a sample, and the sample varies enormously. On
|
||||||
|
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
|
||||||
|
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
|
||||||
|
return value said nothing about, so a statewide total resting on a fifth of
|
||||||
|
the universe looked exactly like one resting on all of it.
|
||||||
|
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
|
||||||
|
`coverage`:
|
||||||
|
|
||||||
|
| value | effect |
|
||||||
|
|---|---|
|
||||||
|
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
|
||||||
|
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
|
||||||
|
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
|
||||||
|
|
||||||
|
* **Regardless of mode**, every result now carries `provenance$coverage` with
|
||||||
|
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
|
||||||
|
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
|
||||||
|
section. So the default mode can no longer mislead silently.
|
||||||
|
* `is_census_year` is a statement about the **survey calendar**, never a claim
|
||||||
|
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
|
||||||
|
cities report. `n_units_reporting` is the number that tells the truth.
|
||||||
|
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
|
||||||
|
it is the subject of the comparison, not a member of the cohort — and the
|
||||||
|
`summary_*` quantiles are computed after the filter, so they describe the
|
||||||
|
cohort actually returned. `n_units_reporting` counts peers only, against the
|
||||||
|
cohort size.
|
||||||
|
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
|
||||||
|
is `NULL`: `"census"` snaps to the most recent census year with an observed
|
||||||
|
population, so a cohort is not built from a sample year in which most of the
|
||||||
|
candidate universe is absent.
|
||||||
|
|
||||||
|
## `complete = TRUE`: absent cells, labelled with why they are absent
|
||||||
|
|
||||||
|
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
|
||||||
|
(today's behaviour). With `complete = TRUE` the requested grid is filled
|
||||||
|
from the corpus's `code_set` table and every row carries a new
|
||||||
|
`value_source` column:
|
||||||
|
|
||||||
|
| `value_source` | meaning | `amt_nominal` |
|
||||||
|
|---|---|---|
|
||||||
|
| `reported` | the corpus carries this cell | as published |
|
||||||
|
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
|
||||||
|
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
|
||||||
|
|
||||||
|
The `NA` is deliberate and is the whole point: filling a modern absence
|
||||||
|
with `0` would invent data, which is precisely the error the corpus's
|
||||||
|
representation contract exists to prevent.
|
||||||
|
* This restores information the reader lost when the corpus was sparsified
|
||||||
|
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
|
||||||
|
had begun returning nothing at all — and improves on what came before it,
|
||||||
|
since the pre-sparsification corpus could not distinguish a published zero
|
||||||
|
from an unreported cell either.
|
||||||
|
* The grid is scoped to each government's **own type**, so a county is never
|
||||||
|
filled with cells only a state can report.
|
||||||
|
* Needs a corpus published from 2026-07-29 onward (when `representation` and
|
||||||
|
`code_set` began shipping); aborts with class
|
||||||
|
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
|
||||||
|
listing those tables rather than on `schema_version`, which was never
|
||||||
|
bumped for the change. Not available with `recipe` or
|
||||||
|
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
|
||||||
|
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
|
||||||
|
`absence_means` rule; `cog_explain()` prints a "Completion" section.
|
||||||
|
|
||||||
|
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
|
||||||
|
|
||||||
|
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
|
||||||
|
corpus as a whole rather than about one item code. `series_break_refs` is
|
||||||
|
built by matching `fin_code` against the item codes in the result, and no
|
||||||
|
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
|
||||||
|
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
|
||||||
|
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
|
||||||
|
representation change at FY2012) and `SB086` (the government id scheme
|
||||||
|
change at FY2017).
|
||||||
|
* Provenance gains `corpus_break_refs`, selected on the break-year window
|
||||||
|
alone and disjoint from `series_break_refs` by construction, so a consumer
|
||||||
|
can tell a whole-result caveat from a break in one series. `cog_explain()`
|
||||||
|
prints them under their own "Corpus-wide caveats" heading. cog-api passes
|
||||||
|
provenance through verbatim, so the field appears there without an API
|
||||||
|
change.
|
||||||
|
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
|
||||||
|
crosses the boundary where an absent cell stops meaning "Census published
|
||||||
|
`$0`" and starts meaning "not reported", and until now nothing said so.
|
||||||
|
|
||||||
## Bundled fixture regenerated against the sparsified corpus
|
## Bundled fixture regenerated against the sparsified corpus
|
||||||
|
|
||||||
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
|
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
|
||||||
|
|||||||
+148
@@ -0,0 +1,148 @@
|
|||||||
|
# R/complete.R
|
||||||
|
#
|
||||||
|
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
|
||||||
|
# cell the corpus does not carry still appears, labelled with WHY it is
|
||||||
|
# missing.
|
||||||
|
#
|
||||||
|
# The corpus stopped storing the wide era's explicit zeros
|
||||||
|
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
|
||||||
|
#
|
||||||
|
# <= FY2011 dense_source absent => Census published $0 (census_zero)
|
||||||
|
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
|
||||||
|
#
|
||||||
|
# Before sparsification a wide-era query whose cells were all $0 came back as
|
||||||
|
# explicit $0 rows; afterwards it came back empty, with nothing to say which
|
||||||
|
# of the two meanings applied. This restores that -- and improves on it,
|
||||||
|
# because the pre-sparsification corpus could not distinguish the two either.
|
||||||
|
#
|
||||||
|
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
|
||||||
|
# That difference is the entire point: writing 0 into a modern absence would
|
||||||
|
# invent data, which is the error the representation contract exists to stop.
|
||||||
|
|
||||||
|
#' @noRd
|
||||||
|
.abort_complete_unsupported <- function(reason, alternative) {
|
||||||
|
cli::cli_abort(c(
|
||||||
|
"{.code complete = TRUE} is not supported for this query.",
|
||||||
|
x = reason,
|
||||||
|
i = alternative
|
||||||
|
), class = "uscogdata_complete_unsupported")
|
||||||
|
}
|
||||||
|
|
||||||
|
#' @noRd
|
||||||
|
.require_representation <- function(con, manifest) {
|
||||||
|
needed <- c("representation.parquet", "code_set.parquet")
|
||||||
|
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
|
||||||
|
logical(1))]
|
||||||
|
if (length(missing) == 0L) return(invisible(TRUE))
|
||||||
|
cli::cli_abort(c(
|
||||||
|
"This corpus does not publish the representation contract.",
|
||||||
|
x = "Missing: {.file {missing}}.",
|
||||||
|
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
|
||||||
|
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
|
||||||
|
), class = "uscogdata_representation_unavailable")
|
||||||
|
}
|
||||||
|
|
||||||
|
#' The cells a government-year COULD carry: every code in force for that
|
||||||
|
#' government's own type, mapped through `summary_categories`, restricted to
|
||||||
|
#' the calling verb's flow prefixes and (when given) its category filter.
|
||||||
|
#'
|
||||||
|
#' Scoped by `govs_type` deliberately. Filling against the union of all types
|
||||||
|
#' would invent cells that the government can never report -- a county row for
|
||||||
|
#' "state IG transfer to school districts" -- and those inventions would then
|
||||||
|
#' be indistinguishable from real census zeros.
|
||||||
|
#'
|
||||||
|
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
|
||||||
|
#' aggregate rows. Without it the grid would offer cells the verb structurally
|
||||||
|
#' never returns, so every one of them would fill as a phantom $0.
|
||||||
|
#' @noRd
|
||||||
|
.completion_grid_sql <- function(subtype_col, govid, years, category,
|
||||||
|
flow_prefixes) {
|
||||||
|
category_pred <- if (is.null(category)) {
|
||||||
|
""
|
||||||
|
} else {
|
||||||
|
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
|
||||||
|
}
|
||||||
|
sprintf(
|
||||||
|
"SELECT DISTINCT
|
||||||
|
cs.year,
|
||||||
|
x.canonical_govid,
|
||||||
|
x.gov_name,
|
||||||
|
c.%1$s AS subtype_value,
|
||||||
|
c.category,
|
||||||
|
r.absence_means
|
||||||
|
FROM code_set cs
|
||||||
|
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
|
||||||
|
JOIN summary_categories c ON c.item_code = cs.item_code
|
||||||
|
JOIN representation r ON r.year = cs.year
|
||||||
|
WHERE x.canonical_govid IN (%2$s)
|
||||||
|
AND cs.year IN (%3$s)
|
||||||
|
AND NOT cs.is_aggregate
|
||||||
|
AND LEFT(cs.item_code, 1) IN (%4$s)
|
||||||
|
AND c.category IS NOT NULL
|
||||||
|
AND c.%1$s IS NOT NULL
|
||||||
|
%5$s",
|
||||||
|
subtype_col, .sql_lit_chr(govid),
|
||||||
|
paste(as.integer(years), collapse = ","),
|
||||||
|
.sql_lit_chr(flow_prefixes), category_pred
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Fill `result` out to the full grid, stamping `value_source` on every row.
|
||||||
|
#'
|
||||||
|
#' Returns the completed tibble with a `.completion` attribute carrying the
|
||||||
|
#' provenance block. Reported rows are passed through untouched -- filling
|
||||||
|
#' must never alter or drop what the corpus actually published.
|
||||||
|
#' @noRd
|
||||||
|
.complete_result <- function(result, con, subtype_col, govid, years, category,
|
||||||
|
flow_prefixes) {
|
||||||
|
grid <- tibble::as_tibble(DBI::dbGetQuery(
|
||||||
|
con, .completion_grid_sql(subtype_col, govid, years, category, flow_prefixes)
|
||||||
|
))
|
||||||
|
|
||||||
|
result$value_source <- rep("reported", nrow(result))
|
||||||
|
if (nrow(grid) == 0L) {
|
||||||
|
attr(result, ".completion") <- list(
|
||||||
|
applied = TRUE, rows_filled = 0L, absence_means = list()
|
||||||
|
)
|
||||||
|
return(result)
|
||||||
|
}
|
||||||
|
|
||||||
|
names(grid)[names(grid) == "subtype_value"] <- subtype_col
|
||||||
|
key <- function(d) {
|
||||||
|
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
|
||||||
|
}
|
||||||
|
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
|
||||||
|
|
||||||
|
if (nrow(missing) > 0L) {
|
||||||
|
filled <- tibble::tibble(
|
||||||
|
year = as.integer(missing$year),
|
||||||
|
canonical_govid = as.character(missing$canonical_govid),
|
||||||
|
gov_name = as.character(missing$gov_name),
|
||||||
|
category = as.character(missing$category),
|
||||||
|
# census_zero is a value Census published; not_reported is unknown and
|
||||||
|
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
|
||||||
|
amt_nominal = ifelse(missing$absence_means == "census_zero",
|
||||||
|
0, NA_real_),
|
||||||
|
codes_included = NA_character_,
|
||||||
|
aggregate_fallback = NA,
|
||||||
|
value_source = as.character(missing$absence_means)
|
||||||
|
)
|
||||||
|
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
|
||||||
|
if ("notes" %in% names(result)) filled$notes <- NA_character_
|
||||||
|
|
||||||
|
result <- dplyr::bind_rows(result, filled)
|
||||||
|
result <- result[order(result$year, result$canonical_govid,
|
||||||
|
result[[subtype_col]], result$category), ,
|
||||||
|
drop = FALSE]
|
||||||
|
}
|
||||||
|
|
||||||
|
rules <- unique(grid[, c("year", "absence_means")])
|
||||||
|
attr(result, ".completion") <- list(
|
||||||
|
applied = TRUE,
|
||||||
|
rows_filled = nrow(missing),
|
||||||
|
absence_means = stats::setNames(
|
||||||
|
as.list(as.character(rules$absence_means)), as.character(rules$year)
|
||||||
|
)
|
||||||
|
)
|
||||||
|
result
|
||||||
|
}
|
||||||
+107
@@ -0,0 +1,107 @@
|
|||||||
|
# R/coverage.R
|
||||||
|
#
|
||||||
|
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
|
||||||
|
# findings F-020 and F-023).
|
||||||
|
#
|
||||||
|
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
|
||||||
|
# 7. Every other year is a sample, and the sample varies enormously: on the
|
||||||
|
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
|
||||||
|
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
|
||||||
|
# what the verbs have always done -- correctly -- but the return value said
|
||||||
|
# nothing about it, so a statewide total resting on 18% of the universe looked
|
||||||
|
# exactly like one resting on 98%.
|
||||||
|
#
|
||||||
|
# Owner's settled design: a `coverage` argument selecting WHICH units to
|
||||||
|
# include, plus always-on metadata saying how many there were either way. The
|
||||||
|
# principle behind it: using these verbs correctly must not require the caller
|
||||||
|
# to know the survey calendar.
|
||||||
|
|
||||||
|
# Years ending in 2 or 7 are full censuses of every government; all others are
|
||||||
|
# samples.
|
||||||
|
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
|
||||||
|
|
||||||
|
#' @noRd
|
||||||
|
.is_census_year <- function(years) {
|
||||||
|
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
|
||||||
|
}
|
||||||
|
|
||||||
|
#' @noRd
|
||||||
|
.validate_coverage <- function(coverage) {
|
||||||
|
tryCatch(
|
||||||
|
match.arg(coverage, c("all", "census", "consistent")),
|
||||||
|
error = function(e) {
|
||||||
|
cli::cli_abort(
|
||||||
|
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
|
||||||
|
class = "uscogdata_invalid_coverage", parent = e
|
||||||
|
)
|
||||||
|
}
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Restrict `years` to census years for `coverage = "census"`.
|
||||||
|
#'
|
||||||
|
#' Aborts rather than returning an empty result when the requested range holds
|
||||||
|
#' no census year: silently handing back zero rows for a query the caller
|
||||||
|
#' believes they made is the failure mode this whole issue is about.
|
||||||
|
#' @noRd
|
||||||
|
.apply_census_years <- function(years, coverage, verb) {
|
||||||
|
if (!identical(coverage, "census")) return(as.integer(years))
|
||||||
|
keep <- as.integer(years)[.is_census_year(years)]
|
||||||
|
if (length(keep) == 0L) {
|
||||||
|
cli::cli_abort(c(
|
||||||
|
"{.code coverage = \"census\"} leaves no years to query.",
|
||||||
|
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
|
||||||
|
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
|
||||||
|
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
|
||||||
|
), class = "uscogdata_no_census_years")
|
||||||
|
}
|
||||||
|
sort(keep)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Keep only units that report in EVERY requested year (a balanced panel).
|
||||||
|
#'
|
||||||
|
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
|
||||||
|
#' filter (the peer-comparison target, which is the subject of the comparison
|
||||||
|
#' rather than a member of the cohort being balanced).
|
||||||
|
#' @noRd
|
||||||
|
.filter_consistent <- function(result, years, id_col = "canonical_govid",
|
||||||
|
keep_ids = character(0)) {
|
||||||
|
years <- unique(as.integer(years))
|
||||||
|
if (nrow(result) == 0L || length(years) <= 1L) return(result)
|
||||||
|
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
|
||||||
|
present <- vapply(ids, function(g) {
|
||||||
|
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
|
||||||
|
}, logical(1))
|
||||||
|
consistent <- c(ids[present], keep_ids)
|
||||||
|
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
|
||||||
|
drop = FALSE]
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Per-year coverage metadata, always attached regardless of mode.
|
||||||
|
#'
|
||||||
|
#' Built from the REQUESTED years rather than the years present in the result,
|
||||||
|
#' so a year in which nothing reported still appears -- with
|
||||||
|
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
|
||||||
|
#' missing year fails to make.
|
||||||
|
#'
|
||||||
|
#' `n_units_reporting` describes the result the caller actually received, so
|
||||||
|
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
|
||||||
|
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
|
||||||
|
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
|
||||||
|
#' `n_units_reporting` is the number that tells the truth.
|
||||||
|
#' @noRd
|
||||||
|
.coverage_table <- function(result, years, n_expected,
|
||||||
|
id_col = "canonical_govid", rows = NULL) {
|
||||||
|
years <- sort(unique(as.integer(years)))
|
||||||
|
src <- if (is.null(rows)) result else rows
|
||||||
|
reporting <- vapply(years, function(y) {
|
||||||
|
ids <- src[[id_col]][as.integer(src$year) == y]
|
||||||
|
length(unique(ids[!is.na(ids)]))
|
||||||
|
}, integer(1))
|
||||||
|
tibble::tibble(
|
||||||
|
year = years,
|
||||||
|
n_units_reporting = as.integer(reporting),
|
||||||
|
n_units_expected = rep(as.integer(n_expected), length(years)),
|
||||||
|
is_census_year = .is_census_year(years)
|
||||||
|
)
|
||||||
|
}
|
||||||
+43
@@ -118,11 +118,54 @@ cog_explain <- function(result, format = c("print", "list")) {
|
|||||||
cli::cli_ul(sugg_lines)
|
cli::cli_ul(sugg_lines)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
|
||||||
|
cli::cli_h2("Reporting coverage")
|
||||||
|
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
|
||||||
|
cov <- prov$coverage
|
||||||
|
cli::cli_ul(sprintf(
|
||||||
|
"%d: %d of %d units reporting (%.0f%%) -- %s year",
|
||||||
|
cov$year, cov$n_units_reporting, cov$n_units_expected,
|
||||||
|
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
|
||||||
|
ifelse(cov$is_census_year, "census", "sample")
|
||||||
|
))
|
||||||
|
if (any(!cov$is_census_year)) {
|
||||||
|
cli::cli_text(
|
||||||
|
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (isTRUE(prov$completion$applied)) {
|
||||||
|
cli::cli_h2("Completion")
|
||||||
|
cli::cli_text(
|
||||||
|
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
|
||||||
|
)
|
||||||
|
rules <- prov$completion$absence_means
|
||||||
|
if (length(rules) > 0L) {
|
||||||
|
cli::cli_ul(vapply(names(rules), function(y) {
|
||||||
|
sprintf("%s: an absent cell means %s", y,
|
||||||
|
if (identical(rules[[y]], "census_zero")) {
|
||||||
|
"Census published $0 (filled as 0)"
|
||||||
|
} else {
|
||||||
|
"the government did not report (filled as NA, not 0)"
|
||||||
|
})
|
||||||
|
}, character(1)))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
if (length(prov$series_break_refs) > 0L) {
|
if (length(prov$series_break_refs) > 0L) {
|
||||||
cli::cli_h2("Series breaks")
|
cli::cli_h2("Series breaks")
|
||||||
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Kept in a section of its own: these qualify the whole result, so folding
|
||||||
|
# them in with the per-code breaks above would invite reading them as a
|
||||||
|
# caveat about one series.
|
||||||
|
if (length(prov$corpus_break_refs) > 0L) {
|
||||||
|
cli::cli_h2("Corpus-wide caveats")
|
||||||
|
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
|
||||||
|
}
|
||||||
|
|
||||||
cli::cli_h2("Transformations")
|
cli::cli_h2("Transformations")
|
||||||
uc <- prov$transformations$units_conversion
|
uc <- prov$transformations$units_conversion
|
||||||
if (isTRUE(uc$applied)) {
|
if (isTRUE(uc$applied)) {
|
||||||
|
|||||||
@@ -19,6 +19,13 @@
|
|||||||
#' target's population at `year` to produce absolute bounds. If `FALSE`,
|
#' target's population at `year` to produce absolute bounds. If `FALSE`,
|
||||||
#' `pop_range` is interpreted as absolute population counts.
|
#' `pop_range` is interpreted as absolute population counts.
|
||||||
#' @param max_peers Integer cap on the number of peers returned.
|
#' @param max_peers Integer cap on the number of peers returned.
|
||||||
|
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
|
||||||
|
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
||||||
|
#' most recent census year with an observed population, so a cohort is not
|
||||||
|
#' built from a sample year in which most of the candidate universe is
|
||||||
|
#' absent. `"consistent"` needs a year range, which cohort selection does not
|
||||||
|
#' have, so it selects like `"all"` and is carried on the result as
|
||||||
|
#' `attr(x, "coverage")` for [cog_peer_compare()].
|
||||||
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||||
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
|
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
|
||||||
#' `attr(x, "cohort_year")`.
|
#' `attr(x, "cohort_year")`.
|
||||||
@@ -29,7 +36,9 @@ cog_find_peers <- function(target_govid,
|
|||||||
same_state = FALSE,
|
same_state = FALSE,
|
||||||
pop_range = c(0.7, 1.3),
|
pop_range = c(0.7, 1.3),
|
||||||
is_ratio = TRUE,
|
is_ratio = TRUE,
|
||||||
max_peers = 10L) {
|
max_peers = 10L,
|
||||||
|
coverage = c("all", "census", "consistent")) {
|
||||||
|
coverage <- .validate_coverage(coverage)
|
||||||
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
||||||
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
||||||
}
|
}
|
||||||
@@ -59,7 +68,7 @@ cog_find_peers <- function(target_govid,
|
|||||||
))
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
cohort_year <- .resolve_cohort_year(con, target_govid, year)
|
cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
|
||||||
|
|
||||||
pop_sql <- sprintf(
|
pop_sql <- sprintf(
|
||||||
"SELECT population FROM gov_population_yearly
|
"SELECT population FROM gov_population_yearly
|
||||||
@@ -107,12 +116,34 @@ cog_find_peers <- function(target_govid,
|
|||||||
attr(peers, "cohort_year") <- as.integer(cohort_year)
|
attr(peers, "cohort_year") <- as.integer(cohort_year)
|
||||||
attr(peers, "pop_range") <- as.numeric(pop_range)
|
attr(peers, "pop_range") <- as.numeric(pop_range)
|
||||||
attr(peers, "is_ratio") <- isTRUE(is_ratio)
|
attr(peers, "is_ratio") <- isTRUE(is_ratio)
|
||||||
|
attr(peers, "coverage") <- coverage
|
||||||
|
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
|
||||||
peers
|
peers
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# `coverage` picks the cohort vintage when the caller did not name one.
|
||||||
|
# "census" snaps to the most recent CENSUS year with an observed population,
|
||||||
|
# so a cohort is not silently built from a sample year in which most of the
|
||||||
|
# candidate universe is absent. "consistent" is a comparison-time concept --
|
||||||
|
# it needs a year RANGE, which cohort selection does not have -- so it selects
|
||||||
|
# like "all" here and is carried on the result for cog_peer_compare().
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.resolve_cohort_year <- function(con, target_govid, year) {
|
.resolve_cohort_year <- function(con, target_govid, year,
|
||||||
|
coverage = "all") {
|
||||||
if (!is.null(year)) return(as.integer(year))
|
if (!is.null(year)) return(as.integer(year))
|
||||||
|
if (identical(coverage, "census")) {
|
||||||
|
sql <- sprintf(
|
||||||
|
"SELECT MAX(year) AS y FROM gov_population_yearly
|
||||||
|
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
|
||||||
|
.sql_lit_chr(target_govid)
|
||||||
|
)
|
||||||
|
y <- DBI::dbGetQuery(con, sql)$y
|
||||||
|
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
|
||||||
|
cli::cli_abort(c(
|
||||||
|
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
|
||||||
|
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
|
||||||
|
), class = "uscogdata_no_census_years")
|
||||||
|
}
|
||||||
sql <- sprintf(
|
sql <- sprintf(
|
||||||
"SELECT MAX(year) AS y FROM gov_population_yearly
|
"SELECT MAX(year) AS y FROM gov_population_yearly
|
||||||
WHERE canonical_govid = %s",
|
WHERE canonical_govid = %s",
|
||||||
@@ -133,7 +164,9 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||||
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||||
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||||
#' call.
|
#' call. Those summary rows are quantiles **within each category**, not
|
||||||
|
#' quantiles of each peer's total — see the `@return` section before summing
|
||||||
|
#' them.
|
||||||
#'
|
#'
|
||||||
#' @param target_govid Character scalar.
|
#' @param target_govid Character scalar.
|
||||||
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
||||||
@@ -147,6 +180,31 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||||
#' single-government queries but cannot be used here because combining Total
|
#' single-government queries but cannot be used here because combining Total
|
||||||
#' across peer sets counts intergovernmental transfers twice.
|
#' across peer sets counts intergovernmental transfers twice.
|
||||||
|
#' @param coverage How to handle the Census of Governments survey cycle,
|
||||||
|
#' which is a **complete census only in years ending in 2 and 7** -- every
|
||||||
|
#' other year is a sample, and the sample varies enormously (on the bundled
|
||||||
|
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||||
|
#' and 112 in FY2019).
|
||||||
|
#'
|
||||||
|
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
||||||
|
#' behaviour, so existing code keeps working.
|
||||||
|
#' * `"census"` -- census years only. Aborts if the requested range holds
|
||||||
|
#' none, rather than silently returning nothing.
|
||||||
|
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
||||||
|
#' a balanced panel.
|
||||||
|
#'
|
||||||
|
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||||
|
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||||
|
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||||
|
#' statement about the **survey calendar**, never a claim of completeness:
|
||||||
|
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||||
|
#' report. `n_units_reporting` is the number that tells the truth.
|
||||||
|
#'
|
||||||
|
#' The comparison target is exempt from `"consistent"` balancing -- it is the
|
||||||
|
#' subject of the comparison, not a member of the cohort -- and the
|
||||||
|
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
|
||||||
|
#' cohort actually returned. `n_units_reporting` counts peers only, against
|
||||||
|
#' the cohort size: "3 of your 15 peers reported in FY2019".
|
||||||
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
|
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||||
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
||||||
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
||||||
@@ -155,12 +213,40 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||||
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||||
#' `cohort_year`, and `cohort_govids`.
|
#' `cohort_year`, and `cohort_govids`.
|
||||||
|
#'
|
||||||
|
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||||
|
#' Each one is computed **within each `(year, spend_subtype,
|
||||||
|
#' category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||||
|
#' median peer's value in that one category*, not *the value of the median
|
||||||
|
#' peer's total*. The median peer for Police and the median peer for Fire
|
||||||
|
#' are usually different governments, so summing `summary_*` rows across
|
||||||
|
#' categories does not give any peer's total and misstates the band it
|
||||||
|
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||||
|
#' one cohort, with a sign flip at FY2012.
|
||||||
|
#'
|
||||||
|
#' Facet by `role` **and** `category` (the documented use, and what the
|
||||||
|
#' rows are built for). For a genuine "median peer's total spending" line,
|
||||||
|
#' sum each peer's own categories first and take the quantile of those
|
||||||
|
#' per-government totals:
|
||||||
|
#'
|
||||||
|
#' ```r
|
||||||
|
#' library(dplyr)
|
||||||
|
#' cmp |>
|
||||||
|
#' filter(role %in% c("target", "peer")) |>
|
||||||
|
#' group_by(year, role, canonical_govid) |>
|
||||||
|
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||||
|
#' filter(role == "peer") |>
|
||||||
|
#' group_by(year) |>
|
||||||
|
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||||
|
#' ```
|
||||||
#' @export
|
#' @export
|
||||||
cog_peer_compare <- function(target_govid, peers, category, years,
|
cog_peer_compare <- function(target_govid, peers, category, years,
|
||||||
per_capita = TRUE, adjust_to_year = NULL,
|
per_capita = TRUE, adjust_to_year = NULL,
|
||||||
expenditure_concept = c("direct", "total")) {
|
expenditure_concept = c("direct", "total"),
|
||||||
|
coverage = c("all", "census", "consistent")) {
|
||||||
call <- match.call()
|
call <- match.call()
|
||||||
expenditure_concept <- match.arg(expenditure_concept)
|
expenditure_concept <- match.arg(expenditure_concept)
|
||||||
|
coverage <- .validate_coverage(coverage)
|
||||||
if (identical(expenditure_concept, "total")) {
|
if (identical(expenditure_concept, "total")) {
|
||||||
.abort_concept_not_aggregatable("cog_peer_compare")
|
.abort_concept_not_aggregatable("cog_peer_compare")
|
||||||
}
|
}
|
||||||
@@ -183,9 +269,20 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
|||||||
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
|
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
|
||||||
all_govids <- unique(c(target_govid, peer_govids))
|
all_govids <- unique(c(target_govid, peer_govids))
|
||||||
|
|
||||||
|
years <- .apply_census_years(years, coverage, "cog_peer_compare")
|
||||||
|
|
||||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||||
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
|
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
|
||||||
|
|
||||||
|
# The target is exempt from balancing: it is the subject of the comparison,
|
||||||
|
# not a member of the cohort being balanced, and dropping it would leave a
|
||||||
|
# peer comparison with nothing to compare. Filtering happens BEFORE the
|
||||||
|
# quantiles below, so a "consistent" cohort's summary rows describe that
|
||||||
|
# cohort rather than the unbalanced one.
|
||||||
|
if (identical(coverage, "consistent")) {
|
||||||
|
r <- .filter_consistent(r, years, keep_ids = target_govid)
|
||||||
|
}
|
||||||
|
|
||||||
value_col <- .peer_value_col(per_capita, adjust_to_year)
|
value_col <- .peer_value_col(per_capita, adjust_to_year)
|
||||||
|
|
||||||
summary_rows <- .peer_summary_rows(r, value_col)
|
summary_rows <- .peer_summary_rows(r, value_col)
|
||||||
@@ -206,6 +303,14 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
|||||||
canonical_govid = target_govid,
|
canonical_govid = target_govid,
|
||||||
gov_name = unique(r$gov_name[r$role == "target"])
|
gov_name = unique(r$gov_name[r$role == "target"])
|
||||||
)
|
)
|
||||||
|
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
|
||||||
|
# reported in FY2019". Including the target would inflate every count by one
|
||||||
|
# and make a cohort that has entirely stopped reporting look non-empty.
|
||||||
|
prov$coverage_mode <- coverage
|
||||||
|
prov$coverage <- .coverage_table(
|
||||||
|
out, years, length(peer_govids),
|
||||||
|
rows = r[r$role == "peer", , drop = FALSE]
|
||||||
|
)
|
||||||
attr(out, "provenance") <- prov
|
attr(out, "provenance") <- prov
|
||||||
out
|
out
|
||||||
}
|
}
|
||||||
|
|||||||
+20
-2
@@ -10,7 +10,8 @@
|
|||||||
expenditure_concept_note = NA_character_,
|
expenditure_concept_note = NA_character_,
|
||||||
expenditure_concept_direct_suppressed = FALSE,
|
expenditure_concept_direct_suppressed = FALSE,
|
||||||
harmonization = NULL, recipe = NULL,
|
harmonization = NULL, recipe = NULL,
|
||||||
suggestions = list()) {
|
suggestions = list(),
|
||||||
|
completion = NULL) {
|
||||||
manifest <- .uscogdata_env$manifest
|
manifest <- .uscogdata_env$manifest
|
||||||
|
|
||||||
codes <- result[["codes_included"]]
|
codes <- result[["codes_included"]]
|
||||||
@@ -37,11 +38,20 @@
|
|||||||
|
|
||||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||||
con <- .uscogdata_env$con
|
con <- .uscogdata_env$con
|
||||||
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) {
|
have_con <- !is.null(con) && DBI::dbIsValid(con)
|
||||||
|
break_refs <- if (have_con) {
|
||||||
.build_series_break_refs(con, codes_observed, years, schema_version)
|
.build_series_break_refs(con, codes_observed, years, schema_version)
|
||||||
} else {
|
} else {
|
||||||
character(0)
|
character(0)
|
||||||
}
|
}
|
||||||
|
# Corpus-wide caveats travel separately: they qualify the whole result
|
||||||
|
# rather than one series, and they do not depend on codes_observed (see
|
||||||
|
# .build_corpus_break_refs()).
|
||||||
|
corpus_refs <- if (have_con) {
|
||||||
|
.build_corpus_break_refs(con, years, schema_version)
|
||||||
|
} else {
|
||||||
|
character(0)
|
||||||
|
}
|
||||||
|
|
||||||
list(
|
list(
|
||||||
verb = verb,
|
verb = verb,
|
||||||
@@ -116,6 +126,14 @@
|
|||||||
)
|
)
|
||||||
),
|
),
|
||||||
series_break_refs = break_refs,
|
series_break_refs = break_refs,
|
||||||
|
corpus_break_refs = corpus_refs,
|
||||||
|
# What `complete = TRUE` filled, and the rule it filled by. Always
|
||||||
|
# present so a consumer can read `completion$applied` without testing
|
||||||
|
# for the key -- an absent block and applied = FALSE would otherwise be
|
||||||
|
# indistinguishable from an older reader version.
|
||||||
|
completion = completion %||% list(
|
||||||
|
applied = FALSE, rows_filled = 0L, absence_means = list()
|
||||||
|
),
|
||||||
manifest = list(
|
manifest = list(
|
||||||
schema_version = as.integer(manifest$schema_version),
|
schema_version = as.integer(manifest$schema_version),
|
||||||
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
||||||
|
|||||||
+6
-3
@@ -11,11 +11,13 @@
|
|||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||||
|
#' and `value_source` when `complete = TRUE`.
|
||||||
#' @export
|
#' @export
|
||||||
cog_revenue <- function(govid, years, category = NULL,
|
cog_revenue <- function(govid, years, category = NULL,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL) {
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
|
complete = FALSE) {
|
||||||
.verb_spendrev(
|
.verb_spendrev(
|
||||||
verb = "cog_revenue",
|
verb = "cog_revenue",
|
||||||
view_base = "revenue_annotated",
|
view_base = "revenue_annotated",
|
||||||
@@ -28,6 +30,7 @@ cog_revenue <- function(govid, years, category = NULL,
|
|||||||
per_capita = per_capita,
|
per_capita = per_capita,
|
||||||
adjust_to_year = adjust_to_year,
|
adjust_to_year = adjust_to_year,
|
||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe
|
recipe = recipe,
|
||||||
|
complete = complete
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|||||||
+36
-1
@@ -31,6 +31,25 @@
|
|||||||
#' across multiple layers of government double-counts intergovernmental
|
#' across multiple layers of government double-counts intergovernmental
|
||||||
#' transfers (a state's payment to a school district is the same dollar the
|
#' transfers (a state's payment to a school district is the same dollar the
|
||||||
#' district reports as its own Direct spending).
|
#' district reports as its own Direct spending).
|
||||||
|
#' @param coverage How to handle the Census of Governments survey cycle,
|
||||||
|
#' which is a **complete census only in years ending in 2 and 7** -- every
|
||||||
|
#' other year is a sample, and the sample varies enormously (on the bundled
|
||||||
|
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||||
|
#' and 112 in FY2019).
|
||||||
|
#'
|
||||||
|
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
||||||
|
#' behaviour, so existing code keeps working.
|
||||||
|
#' * `"census"` -- census years only. Aborts if the requested range holds
|
||||||
|
#' none, rather than silently returning nothing.
|
||||||
|
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
||||||
|
#' a balanced panel.
|
||||||
|
#'
|
||||||
|
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||||
|
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||||
|
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||||
|
#' statement about the **survey calendar**, never a claim of completeness:
|
||||||
|
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||||
|
#' report. `n_units_reporting` is the number that tells the truth.
|
||||||
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
||||||
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
||||||
@@ -40,9 +59,11 @@
|
|||||||
#' @export
|
#' @export
|
||||||
cog_geographic_rollup <- function(govids, category, years,
|
cog_geographic_rollup <- function(govids, category, years,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
expenditure_concept = c("direct", "total")) {
|
expenditure_concept = c("direct", "total"),
|
||||||
|
coverage = c("all", "census", "consistent")) {
|
||||||
call <- match.call()
|
call <- match.call()
|
||||||
expenditure_concept <- match.arg(expenditure_concept)
|
expenditure_concept <- match.arg(expenditure_concept)
|
||||||
|
coverage <- .validate_coverage(coverage)
|
||||||
if (identical(expenditure_concept, "total")) {
|
if (identical(expenditure_concept, "total")) {
|
||||||
.abort_concept_not_aggregatable("cog_geographic_rollup")
|
.abort_concept_not_aggregatable("cog_geographic_rollup")
|
||||||
}
|
}
|
||||||
@@ -59,11 +80,20 @@ cog_geographic_rollup <- function(govids, category, years,
|
|||||||
layer = rep(layer_names, lengths(govids))
|
layer = rep(layer_names, lengths(govids))
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# coverage = "census" drops non-census years BEFORE the query rather than
|
||||||
|
# after: a sample year's rows are not wanted at all, and fetching them only
|
||||||
|
# to discard them would also let them into the coverage table.
|
||||||
|
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
|
||||||
|
|
||||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||||
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
|
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
|
||||||
relationship = "many-to-many")
|
relationship = "many-to-many")
|
||||||
r$scope_note <- .rollup_scope_note(r$layer)
|
r$scope_note <- .rollup_scope_note(r$layer)
|
||||||
|
|
||||||
|
if (identical(coverage, "consistent")) {
|
||||||
|
r <- .filter_consistent(r, years)
|
||||||
|
}
|
||||||
|
|
||||||
excluded <- character(0)
|
excluded <- character(0)
|
||||||
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
|
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
|
||||||
drop <- r$pop_source == "unavailable"
|
drop <- r$pop_source == "unavailable"
|
||||||
@@ -82,6 +112,11 @@ cog_geographic_rollup <- function(govids, category, years,
|
|||||||
included_govids = included,
|
included_govids = included,
|
||||||
excluded_govids = excluded
|
excluded_govids = excluded
|
||||||
)
|
)
|
||||||
|
# n_units_expected is the universe the CALLER named -- the govids passed in
|
||||||
|
# -- not the national universe. That is what makes the ratio meaningful:
|
||||||
|
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
|
||||||
|
prov$coverage_mode <- coverage
|
||||||
|
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
|
||||||
attr(r, "provenance") <- prov
|
attr(r, "provenance") <- prov
|
||||||
|
|
||||||
r
|
r
|
||||||
|
|||||||
+19
-7
@@ -6,8 +6,11 @@
|
|||||||
#' the cross-vintage canonical-government registry. Operates in two modes:
|
#' the cross-vintage canonical-government registry. Operates in two modes:
|
||||||
#'
|
#'
|
||||||
#' * **Utility mode** (single `name`, the original behavior): returns all
|
#' * **Utility mode** (single `name`, the original behavior): returns all
|
||||||
#' rows whose `gov_name` matches the regex case-insensitively, sorted by
|
#' rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||||
#' `population_acs` descending. Useful for exploratory lookups.
|
#' substring**, sorted by `population_acs` descending. Useful for
|
||||||
|
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||||
|
#' government is findable by its own complete name even when that name
|
||||||
|
#' contains parentheses or a period.
|
||||||
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||||
#' single canonical govid and returns a tibble in input order, suitable
|
#' single canonical govid and returns a tibble in input order, suitable
|
||||||
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||||
@@ -19,7 +22,8 @@
|
|||||||
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||||
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||||
#' Single hit -> resolved. Multiple -> step 4.
|
#' Single hit -> resolved. Multiple -> step 4.
|
||||||
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
#' 3. **Substring fallback:** case-insensitive literal substring against
|
||||||
|
#' `gov_name` (metacharacters escaped).
|
||||||
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||||
#' `status = "no_match"`. Multiple hits -> step 4.
|
#' `status = "no_match"`. Multiple hits -> step 4.
|
||||||
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||||
@@ -48,7 +52,7 @@
|
|||||||
#' [cog_spending()], [cog_revenue()].
|
#' [cog_spending()], [cog_revenue()].
|
||||||
#' @examples
|
#' @examples
|
||||||
#' \dontrun{
|
#' \dontrun{
|
||||||
#' # Utility mode — exploratory regex lookup
|
#' # Utility mode — exploratory substring lookup
|
||||||
#' cog_gov_search("broward", state = "FL")
|
#' cog_gov_search("broward", state = "FL")
|
||||||
#'
|
#'
|
||||||
#' # Basket mode — resolve a known cohort
|
#' # Basket mode — resolve a known cohort
|
||||||
@@ -98,9 +102,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
|||||||
if (!is.character(name) || length(name) != 1L) {
|
if (!is.character(name) || length(name) != 1L) {
|
||||||
cli::cli_abort("`name` must be a length-1 character string.")
|
cli::cli_abort("`name` must be a length-1 character string.")
|
||||||
}
|
}
|
||||||
|
# Escaped, so `name` is a literal case-insensitive substring -- the same
|
||||||
|
# treatment basket mode has always given it. Interpolating it raw made a
|
||||||
|
# government unfindable by its own name whenever that name contains a
|
||||||
|
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
|
||||||
|
# match-everything wildcard, and let malformed pattern text reach the
|
||||||
|
# engine as an error -- which cog-api surfaced as a 500, reachable by
|
||||||
|
# typing a real name one character at a time (uscogdata#16, F-025).
|
||||||
preds <- c(preds,
|
preds <- c(preds,
|
||||||
sprintf("regexp_matches(gov_name, %s, 'i')",
|
sprintf("regexp_matches(gov_name, %s, 'i')",
|
||||||
.sql_lit_chr(name)))
|
.sql_lit_chr(.escape_regex(name))))
|
||||||
}
|
}
|
||||||
if (!is.null(state)) {
|
if (!is.null(state)) {
|
||||||
st_fips <- .coerce_state_to_fips(state)
|
st_fips <- .coerce_state_to_fips(state)
|
||||||
@@ -136,8 +147,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
|||||||
#' @noRd
|
#' @noRd
|
||||||
.escape_regex <- function(x) {
|
.escape_regex <- function(x) {
|
||||||
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
||||||
# literal substring in the DuckDB regexp_matches call (substring fallback
|
# literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
|
||||||
# only; utility-mode intentionally preserves regex behavior).
|
# utility mode used to interpolate raw, which was a defect rather than a
|
||||||
|
# feature -- see the call site and uscogdata#16.
|
||||||
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+33
-1
@@ -14,9 +14,41 @@
|
|||||||
sql <- sprintf(
|
sql <- sprintf(
|
||||||
"SELECT DISTINCT break_id
|
"SELECT DISTINCT break_id
|
||||||
FROM series_breaks_pq
|
FROM series_breaks_pq
|
||||||
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d
|
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
|
||||||
|
AND break_year BETWEEN %d AND %d
|
||||||
ORDER BY break_id",
|
ORDER BY break_id",
|
||||||
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
||||||
)
|
)
|
||||||
DBI::dbGetQuery(con, sql)$break_id
|
DBI::dbGetQuery(con, sql)$break_id
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
|
||||||
|
#' `"ALL"` rather than an item code. They qualify the whole result, so they
|
||||||
|
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
|
||||||
|
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
|
||||||
|
#' before uscogdata#19. Selection is on the break_year window alone: which
|
||||||
|
#' codes a result happens to contain is irrelevant to a caveat about the
|
||||||
|
#' corpus.
|
||||||
|
#'
|
||||||
|
#' All four catalogued entries are *boundary* caveats (dollar precision
|
||||||
|
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
|
||||||
|
#' representation change at 2012, the id scheme change at 2017), so the same
|
||||||
|
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
|
||||||
|
#' path uses is the right one -- a request that never crosses the boundary
|
||||||
|
#' is not affected by it.
|
||||||
|
#'
|
||||||
|
#' Returned separately from `series_break_refs` so a consumer can tell a
|
||||||
|
#' whole-result caveat from a break in one series; the two are disjoint by
|
||||||
|
#' construction.
|
||||||
|
#' @noRd
|
||||||
|
.build_corpus_break_refs <- function(con, years, schema_version) {
|
||||||
|
if (schema_version < 5L || length(years) == 0L) return(character(0))
|
||||||
|
sql <- sprintf(
|
||||||
|
"SELECT DISTINCT break_id
|
||||||
|
FROM series_breaks_pq
|
||||||
|
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
|
||||||
|
ORDER BY break_id",
|
||||||
|
min(as.integer(years)), max(as.integer(years))
|
||||||
|
)
|
||||||
|
DBI::dbGetQuery(con, sql)$break_id
|
||||||
|
}
|
||||||
|
|||||||
+62
-6
@@ -70,16 +70,42 @@
|
|||||||
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||||
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
||||||
#' IG.
|
#' IG.
|
||||||
|
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
||||||
|
#' corpus does not carry still appears, labelled with **why** it is
|
||||||
|
#' missing, and add a `value_source` column to every row:
|
||||||
|
#'
|
||||||
|
#' * `"reported"` — the corpus carries this cell.
|
||||||
|
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||||
|
#' Census published `$0`. `amt_nominal` is `0`.
|
||||||
|
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||||
|
#' government did not report, and the value is unknown. `amt_nominal` is
|
||||||
|
#' `NA`, **not** `0` — writing a zero there would invent data.
|
||||||
|
#'
|
||||||
|
#' The grid comes from the corpus's `code_set` table, scoped to each
|
||||||
|
#' government's own type, so a county is never filled with cells only a
|
||||||
|
#' state can report. Reported rows are passed through untouched.
|
||||||
|
#'
|
||||||
|
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||||
|
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||||
|
#' when `representation`/`code_set` began shipping; aborts with class
|
||||||
|
#' `uscogdata_representation_unavailable` otherwise. Not available with
|
||||||
|
#' `recipe` or with `expenditure_concept = "total"` (class
|
||||||
|
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
|
#' `code_set`.
|
||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
#' and `value_source` when `complete = TRUE`.
|
||||||
|
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||||
|
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
||||||
|
#' per-year `absence_means` rule that was applied.
|
||||||
#' @export
|
#' @export
|
||||||
cog_spending <- function(govid, years, category = NULL,
|
cog_spending <- function(govid, years, category = NULL,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
expenditure_concept = c("direct", "total")) {
|
expenditure_concept = c("direct", "total"),
|
||||||
|
complete = FALSE) {
|
||||||
.verb_spendrev(
|
.verb_spendrev(
|
||||||
verb = "cog_spending",
|
verb = "cog_spending",
|
||||||
view_base = "spending_annotated",
|
view_base = "spending_annotated",
|
||||||
@@ -93,7 +119,8 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
adjust_to_year = adjust_to_year,
|
adjust_to_year = adjust_to_year,
|
||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe,
|
recipe = recipe,
|
||||||
expenditure_concept = expenditure_concept
|
expenditure_concept = expenditure_concept,
|
||||||
|
complete = complete
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -118,7 +145,8 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
govid, years, category,
|
govid, years, category,
|
||||||
per_capita, adjust_to_year,
|
per_capita, adjust_to_year,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
expenditure_concept = c("direct", "total")) {
|
expenditure_concept = c("direct", "total"),
|
||||||
|
complete = FALSE) {
|
||||||
basis_explicit <- length(basis) == 1L
|
basis_explicit <- length(basis) == 1L
|
||||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||||
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
||||||
@@ -165,12 +193,27 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
complete <- isTRUE(complete)
|
||||||
|
if (complete && !is.null(recipe)) {
|
||||||
|
.abort_complete_unsupported(
|
||||||
|
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
|
||||||
|
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
|
||||||
|
)
|
||||||
|
}
|
||||||
|
if (complete && identical(expenditure_concept, "total")) {
|
||||||
|
.abort_complete_unsupported(
|
||||||
|
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
|
||||||
|
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
years <- as.integer(years)
|
years <- as.integer(years)
|
||||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||||
|
|
||||||
con <- .ensure_session()
|
con <- .ensure_session()
|
||||||
manifest <- .uscogdata_env$manifest
|
manifest <- .uscogdata_env$manifest
|
||||||
scope <- .check_govids_in_scope(govid)
|
scope <- .check_govids_in_scope(govid)
|
||||||
|
if (complete) .require_representation(con, manifest)
|
||||||
|
|
||||||
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
||||||
|
|
||||||
@@ -201,6 +244,18 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Fill BEFORE per_capita / inflation so the added cells get the same
|
||||||
|
# treatment as reported ones: a census_zero stays $0 per capita and in real
|
||||||
|
# dollars, and a not_reported stays NA through both rather than becoming a
|
||||||
|
# spurious 0.
|
||||||
|
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
|
||||||
|
if (complete) {
|
||||||
|
result <- .complete_result(result, con, subtype_col, govid, years,
|
||||||
|
category, flow_prefixes)
|
||||||
|
completion <- attr(result, ".completion")
|
||||||
|
attr(result, ".completion") <- NULL
|
||||||
|
}
|
||||||
|
|
||||||
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
||||||
if (!is.null(adjust_to_year)) {
|
if (!is.null(adjust_to_year)) {
|
||||||
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
||||||
@@ -309,7 +364,8 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
expenditure_concept_direct_suppressed = direct_suppressed_flag,
|
expenditure_concept_direct_suppressed = direct_suppressed_flag,
|
||||||
harmonization = harmonization,
|
harmonization = harmonization,
|
||||||
recipe = recipe_block,
|
recipe = recipe_block,
|
||||||
suggestions = suggestions
|
suggestions = suggestions,
|
||||||
|
completion = completion
|
||||||
)
|
)
|
||||||
prov$scope$govids_found <- scope$found
|
prov$scope$govids_found <- scope$found
|
||||||
prov$scope$govids_missing <- scope$missing
|
prov$scope$govids_missing <- scope$missing
|
||||||
|
|||||||
@@ -32,6 +32,29 @@
|
|||||||
"45-ig_annotated_harmonized.sql"
|
"45-ig_annotated_harmonized.sql"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# The representation contract (cog_pipeline#64): two parquet tables that say
|
||||||
|
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
|
||||||
|
# on schema_version, because the sparsification that introduced them did not
|
||||||
|
# bump the version -- the pre-sparsification corpus this package shipped
|
||||||
|
# against until 2026-07-30 was already schema v6 and carried neither table.
|
||||||
|
# Keying off the version number would therefore register a view over a file
|
||||||
|
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
|
||||||
|
# check exists to tolerate.
|
||||||
|
.representation_view_files <- c(
|
||||||
|
"36-representation.sql" = "representation.parquet",
|
||||||
|
"37-code_set.sql" = "code_set.parquet"
|
||||||
|
)
|
||||||
|
|
||||||
|
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
|
||||||
|
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
|
||||||
|
#' works identically for a local fixture and a remote share.
|
||||||
|
#' @noRd
|
||||||
|
.corpus_has_table <- function(manifest, file) {
|
||||||
|
paths <- vapply(manifest$files$metadata %||% list(),
|
||||||
|
function(f) as.character(f$path %||% ""), character(1))
|
||||||
|
file %in% basename(paths)
|
||||||
|
}
|
||||||
|
|
||||||
#' Register DuckDB views from inst/sql/ SQL files
|
#' Register DuckDB views from inst/sql/ SQL files
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.register_views <- function(con, url, manifest) {
|
.register_views <- function(con, url, manifest) {
|
||||||
@@ -39,7 +62,10 @@
|
|||||||
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
||||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||||
for (f in files) {
|
for (f in files) {
|
||||||
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next
|
base <- basename(f)
|
||||||
|
if (base %in% .harmonization_view_files && schema_version < 5L) next
|
||||||
|
if (base %in% names(.representation_view_files) &&
|
||||||
|
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
|
||||||
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
||||||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||||||
DBI::dbExecute(con, sql)
|
DBI::dbExecute(con, sql)
|
||||||
|
|||||||
@@ -19,6 +19,27 @@ package implements.
|
|||||||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Amounts are in full US dollars
|
||||||
|
|
||||||
|
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
||||||
|
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
|
||||||
|
|
||||||
|
The raw Census source files report **thousands of dollars**, and the corpus's
|
||||||
|
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
||||||
|
you never have to. The conversion is recorded in every result:
|
||||||
|
|
||||||
|
```r
|
||||||
|
r <- cog_spending("552025209777", 2020L)
|
||||||
|
attr(r, "provenance")$transformations$units_conversion
|
||||||
|
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
|
||||||
|
```
|
||||||
|
|
||||||
|
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
||||||
|
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
|
||||||
|
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
|
||||||
|
twice overstates every figure by 1000x, and the result looks plausible rather
|
||||||
|
than obviously wrong.
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
||||||
|
|||||||
@@ -33,6 +33,20 @@
|
|||||||
"aggregate_fallback": { "type": ["object", "null"] },
|
"aggregate_fallback": { "type": ["object", "null"] },
|
||||||
"transformations":{ "type": "object" },
|
"transformations":{ "type": "object" },
|
||||||
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
||||||
|
"completion": {
|
||||||
|
"type": "object",
|
||||||
|
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
|
||||||
|
"properties": {
|
||||||
|
"applied": { "type": "boolean" },
|
||||||
|
"rows_filled": { "type": "integer" },
|
||||||
|
"absence_means": { "type": "object" }
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"corpus_break_refs": {
|
||||||
|
"type": "array",
|
||||||
|
"items": { "type": "string" },
|
||||||
|
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
||||||
|
},
|
||||||
"manifest": { "type": "object" },
|
"manifest": { "type": "object" },
|
||||||
"sql_query": { "type": "string" }
|
"sql_query": { "type": "string" }
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
CREATE OR REPLACE VIEW representation AS
|
||||||
|
SELECT *
|
||||||
|
FROM read_parquet('{url}data/representation.parquet');
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
CREATE OR REPLACE VIEW code_set AS
|
||||||
|
SELECT *
|
||||||
|
FROM read_parquet('{url}data/code_set.parquet');
|
||||||
+10
-1
@@ -11,7 +11,8 @@ cog_find_peers(
|
|||||||
same_state = FALSE,
|
same_state = FALSE,
|
||||||
pop_range = c(0.7, 1.3),
|
pop_range = c(0.7, 1.3),
|
||||||
is_ratio = TRUE,
|
is_ratio = TRUE,
|
||||||
max_peers = 10L
|
max_peers = 10L,
|
||||||
|
coverage = c("all", "census", "consistent")
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -34,6 +35,14 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
|
|||||||
`pop_range` is interpreted as absolute population counts.}
|
`pop_range` is interpreted as absolute population counts.}
|
||||||
|
|
||||||
\item{max_peers}{Integer cap on the number of peers returned.}
|
\item{max_peers}{Integer cap on the number of peers returned.}
|
||||||
|
|
||||||
|
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
|
||||||
|
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
||||||
|
most recent census year with an observed population, so a cohort is not
|
||||||
|
built from a sample year in which most of the candidate universe is
|
||||||
|
absent. `"consistent"` needs a year range, which cohort selection does not
|
||||||
|
have, so it selects like `"all"` and is carried on the result as
|
||||||
|
`attr(x, "coverage")` for [cog_peer_compare()].}
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||||
|
|||||||
@@ -10,7 +10,8 @@ cog_geographic_rollup(
|
|||||||
years,
|
years,
|
||||||
per_capita = FALSE,
|
per_capita = FALSE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
expenditure_concept = c("direct", "total")
|
expenditure_concept = c("direct", "total"),
|
||||||
|
coverage = c("all", "census", "consistent")
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -35,6 +36,26 @@ single-government queries but cannot be used here because combining Total
|
|||||||
across multiple layers of government double-counts intergovernmental
|
across multiple layers of government double-counts intergovernmental
|
||||||
transfers (a state's payment to a school district is the same dollar the
|
transfers (a state's payment to a school district is the same dollar the
|
||||||
district reports as its own Direct spending).}
|
district reports as its own Direct spending).}
|
||||||
|
|
||||||
|
\item{coverage}{How to handle the Census of Governments survey cycle,
|
||||||
|
which is a **complete census only in years ending in 2 and 7** -- every
|
||||||
|
other year is a sample, and the sample varies enormously (on the bundled
|
||||||
|
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||||
|
and 112 in FY2019).
|
||||||
|
|
||||||
|
* `"all"` (default) -- every unit that reported that year. Unchanged
|
||||||
|
behaviour, so existing code keeps working.
|
||||||
|
* `"census"` -- census years only. Aborts if the requested range holds
|
||||||
|
none, rather than silently returning nothing.
|
||||||
|
* `"consistent"` -- only units reporting in *every* requested year, giving
|
||||||
|
a balanced panel.
|
||||||
|
|
||||||
|
Regardless of mode, `provenance$coverage` always carries per-year
|
||||||
|
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||||
|
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||||
|
statement about the **survey calendar**, never a claim of completeness:
|
||||||
|
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||||
|
report. `n_units_reporting` is the number that tells the truth.}
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||||
|
|||||||
@@ -32,8 +32,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
|||||||
}
|
}
|
||||||
\details{
|
\details{
|
||||||
* **Utility mode** (single `name`, the original behavior): returns all
|
* **Utility mode** (single `name`, the original behavior): returns all
|
||||||
rows whose `gov_name` matches the regex case-insensitively, sorted by
|
rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||||
`population_acs` descending. Useful for exploratory lookups.
|
substring**, sorted by `population_acs` descending. Useful for
|
||||||
|
exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||||
|
government is findable by its own complete name even when that name
|
||||||
|
contains parentheses or a period.
|
||||||
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||||
single canonical govid and returns a tibble in input order, suitable
|
single canonical govid and returns a tibble in input order, suitable
|
||||||
for piping straight into [cog_spending()] / [cog_revenue()] /
|
for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||||
@@ -45,7 +48,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
|||||||
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||||
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||||
Single hit -> resolved. Multiple -> step 4.
|
Single hit -> resolved. Multiple -> step 4.
|
||||||
3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
3. **Substring fallback:** case-insensitive literal substring against
|
||||||
|
`gov_name` (metacharacters escaped).
|
||||||
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||||
`status = "no_match"`. Multiple hits -> step 4.
|
`status = "no_match"`. Multiple hits -> step 4.
|
||||||
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||||
@@ -58,7 +62,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
|
|||||||
}
|
}
|
||||||
\examples{
|
\examples{
|
||||||
\dontrun{
|
\dontrun{
|
||||||
# Utility mode — exploratory regex lookup
|
# Utility mode — exploratory substring lookup
|
||||||
cog_gov_search("broward", state = "FL")
|
cog_gov_search("broward", state = "FL")
|
||||||
|
|
||||||
# Basket mode — resolve a known cohort
|
# Basket mode — resolve a known cohort
|
||||||
|
|||||||
+57
-2
@@ -11,7 +11,8 @@ cog_peer_compare(
|
|||||||
years,
|
years,
|
||||||
per_capita = TRUE,
|
per_capita = TRUE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
expenditure_concept = c("direct", "total")
|
expenditure_concept = c("direct", "total"),
|
||||||
|
coverage = c("all", "census", "consistent")
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -33,6 +34,32 @@ population.}
|
|||||||
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||||
single-government queries but cannot be used here because combining Total
|
single-government queries but cannot be used here because combining Total
|
||||||
across peer sets counts intergovernmental transfers twice.}
|
across peer sets counts intergovernmental transfers twice.}
|
||||||
|
|
||||||
|
\item{coverage}{How to handle the Census of Governments survey cycle,
|
||||||
|
which is a **complete census only in years ending in 2 and 7** -- every
|
||||||
|
other year is a sample, and the sample varies enormously (on the bundled
|
||||||
|
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||||
|
and 112 in FY2019).
|
||||||
|
|
||||||
|
* `"all"` (default) -- every unit that reported that year. Unchanged
|
||||||
|
behaviour, so existing code keeps working.
|
||||||
|
* `"census"` -- census years only. Aborts if the requested range holds
|
||||||
|
none, rather than silently returning nothing.
|
||||||
|
* `"consistent"` -- only units reporting in *every* requested year, giving
|
||||||
|
a balanced panel.
|
||||||
|
|
||||||
|
Regardless of mode, `provenance$coverage` always carries per-year
|
||||||
|
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||||
|
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||||
|
statement about the **survey calendar**, never a claim of completeness:
|
||||||
|
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||||
|
report. `n_units_reporting` is the number that tells the truth.
|
||||||
|
|
||||||
|
The comparison target is exempt from `"consistent"` balancing -- it is the
|
||||||
|
subject of the comparison, not a member of the cohort -- and the
|
||||||
|
`summary_*` quantiles are computed AFTER the filter, so they describe the
|
||||||
|
cohort actually returned. `n_units_reporting` counts peers only, against
|
||||||
|
the cohort size: "3 of your 15 peers reported in FY2019".}
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble matching [cog_spending()]'s columns, plus a `role`
|
Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||||
@@ -43,11 +70,39 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
|
|||||||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||||
`cohort_year`, and `cohort_govids`.
|
`cohort_year`, and `cohort_govids`.
|
||||||
|
|
||||||
|
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||||
|
Each one is computed **within each `(year, spend_subtype,
|
||||||
|
category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||||
|
median peer's value in that one category*, not *the value of the median
|
||||||
|
peer's total*. The median peer for Police and the median peer for Fire
|
||||||
|
are usually different governments, so summing `summary_*` rows across
|
||||||
|
categories does not give any peer's total and misstates the band it
|
||||||
|
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||||
|
one cohort, with a sign flip at FY2012.
|
||||||
|
|
||||||
|
Facet by `role` **and** `category` (the documented use, and what the
|
||||||
|
rows are built for). For a genuine "median peer's total spending" line,
|
||||||
|
sum each peer's own categories first and take the quantile of those
|
||||||
|
per-government totals:
|
||||||
|
|
||||||
|
```r
|
||||||
|
library(dplyr)
|
||||||
|
cmp |>
|
||||||
|
filter(role %in% c("target", "peer")) |>
|
||||||
|
group_by(year, role, canonical_govid) |>
|
||||||
|
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||||
|
filter(role == "peer") |>
|
||||||
|
group_by(year) |>
|
||||||
|
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||||
|
```
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
Pulls spending for the target plus a peer set (either a
|
Pulls spending for the target plus a peer set (either a
|
||||||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||||
call.
|
call. Those summary rows are quantiles **within each category**, not
|
||||||
|
quantiles of each peer's total — see the `@return` section before summing
|
||||||
|
them.
|
||||||
}
|
}
|
||||||
|
|||||||
+27
-2
@@ -11,7 +11,8 @@ cog_revenue(
|
|||||||
per_capita = FALSE,
|
per_capita = FALSE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL
|
recipe = NULL,
|
||||||
|
complete = FALSE
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -55,12 +56,36 @@ argument is ignored and the result's provenance reports
|
|||||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||||
|
|
||||||
|
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||||
|
corpus does not carry still appears, labelled with **why** it is
|
||||||
|
missing, and add a `value_source` column to every row:
|
||||||
|
|
||||||
|
* `"reported"` — the corpus carries this cell.
|
||||||
|
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||||
|
Census published `$0`. `amt_nominal` is `0`.
|
||||||
|
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||||
|
government did not report, and the value is unknown. `amt_nominal` is
|
||||||
|
`NA`, **not** `0` — writing a zero there would invent data.
|
||||||
|
|
||||||
|
The grid comes from the corpus's `code_set` table, scoped to each
|
||||||
|
government's own type, so a county is never filled with cells only a
|
||||||
|
state can report. Reported rows are passed through untouched.
|
||||||
|
|
||||||
|
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||||
|
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||||
|
when `representation`/`code_set` began shipping; aborts with class
|
||||||
|
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||||
|
`recipe` or with `expenditure_concept = "total"` (class
|
||||||
|
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
|
`code_set`.}
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||||
|
and `value_source` when `complete = TRUE`.
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
Mirror of [cog_spending()] for revenue categories. One row per
|
Mirror of [cog_spending()] for revenue categories. One row per
|
||||||
|
|||||||
+30
-3
@@ -12,7 +12,8 @@ cog_spending(
|
|||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL,
|
recipe = NULL,
|
||||||
expenditure_concept = c("direct", "total")
|
expenditure_concept = c("direct", "total"),
|
||||||
|
complete = FALSE
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -85,13 +86,39 @@ component (when one exists), and
|
|||||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||||
IG.}
|
IG.}
|
||||||
|
|
||||||
|
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||||
|
corpus does not carry still appears, labelled with **why** it is
|
||||||
|
missing, and add a `value_source` column to every row:
|
||||||
|
|
||||||
|
* `"reported"` — the corpus carries this cell.
|
||||||
|
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||||
|
Census published `$0`. `amt_nominal` is `0`.
|
||||||
|
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||||
|
government did not report, and the value is unknown. `amt_nominal` is
|
||||||
|
`NA`, **not** `0` — writing a zero there would invent data.
|
||||||
|
|
||||||
|
The grid comes from the corpus's `code_set` table, scoped to each
|
||||||
|
government's own type, so a county is never filled with cells only a
|
||||||
|
state can report. Reported rows are passed through untouched.
|
||||||
|
|
||||||
|
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||||
|
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||||
|
when `representation`/`code_set` began shipping; aborts with class
|
||||||
|
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||||
|
`recipe` or with `expenditure_concept = "total"` (class
|
||||||
|
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
|
`code_set`.}
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
and `value_source` when `complete = TRUE`.
|
||||||
|
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||||
|
whose `completion` block reports `applied`, `rows_filled`, and the
|
||||||
|
per-year `absence_means` rule that was applied.
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||||
|
|||||||
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
|
|||||||
if (nzchar(p)) paste0(p, "/") else ""
|
if (nzchar(p)) paste0(p, "/") else ""
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
|
||||||
|
# or "" when it isn't there.
|
||||||
|
#
|
||||||
|
# Tests that assert on documentation content have to read the sources, and the
|
||||||
|
# sources only exist when the suite runs from a checkout. Under R CMD check the
|
||||||
|
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
|
||||||
|
# shipped and `../../README.md` does not resolve -- so those tests must skip
|
||||||
|
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
|
||||||
|
# rcmdcheck, so the assertions are still enforced on every push; this only
|
||||||
|
# stops them from failing a context that structurally cannot satisfy them.
|
||||||
|
source_tree_path <- function(...) {
|
||||||
|
p <- testthat::test_path("..", "..", ...)
|
||||||
|
if (file.exists(p)) p else ""
|
||||||
|
}
|
||||||
|
|
||||||
|
# Skip unless every named source file is present (see source_tree_path()).
|
||||||
|
skip_if_no_source_tree <- function(...) {
|
||||||
|
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
|
||||||
|
character(1))
|
||||||
|
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
|
||||||
|
testthat::skip_if(
|
||||||
|
any(missing),
|
||||||
|
"package source tree not available (running against the installed package)"
|
||||||
|
)
|
||||||
|
invisible(paths)
|
||||||
|
}
|
||||||
|
|
||||||
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
||||||
skip_if_no_corpus <- function() {
|
skip_if_no_corpus <- function() {
|
||||||
p <- fixture_corpus_path()
|
p <- fixture_corpus_path()
|
||||||
@@ -58,6 +85,42 @@ with_doctored_schema_version <- function(version, code) {
|
|||||||
force(code)
|
force(code)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Copy the bundled fixture to a temp dir with representation.parquet and
|
||||||
|
# code_set.parquet removed (and dropped from the manifest's metadata list),
|
||||||
|
# then run `code` against it. Models a corpus published BEFORE sparsification:
|
||||||
|
# schema_version is left alone deliberately, because it was never bumped for
|
||||||
|
# that change -- the pre-sparsification fixture this package shipped until
|
||||||
|
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
|
||||||
|
# is therefore the only honest signal, and this helper is what proves the
|
||||||
|
# package keys off it rather than off the version number.
|
||||||
|
with_corpus_missing_representation <- function(code) {
|
||||||
|
src <- fixture_corpus_path()
|
||||||
|
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
||||||
|
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
||||||
|
|
||||||
|
dropped <- c("representation.parquet", "code_set.parquet")
|
||||||
|
file.remove(file.path(tmp, "data", dropped))
|
||||||
|
|
||||||
|
manifest_path <- file.path(tmp, "manifest.json")
|
||||||
|
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
|
||||||
|
m$files$metadata <- Filter(
|
||||||
|
function(f) !basename(f$path) %in% dropped, m$files$metadata
|
||||||
|
)
|
||||||
|
writeLines(
|
||||||
|
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
|
||||||
|
manifest_path
|
||||||
|
)
|
||||||
|
|
||||||
|
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
||||||
|
uscogdata:::cog_close()
|
||||||
|
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
||||||
|
on.exit({
|
||||||
|
uscogdata:::cog_close()
|
||||||
|
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
||||||
|
}, add = TRUE)
|
||||||
|
force(code)
|
||||||
|
}
|
||||||
|
|
||||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
||||||
# rewritten to drop every M/L (intergovernmental) row, then run `code`
|
# rewritten to drop every M/L (intergovernmental) row, then run `code`
|
||||||
# against it with a clean session (mirrors with_fixture_corpus()/
|
# against it with a clean session (mirrors with_fixture_corpus()/
|
||||||
|
|||||||
@@ -17,7 +17,16 @@
|
|||||||
# cog-api's llms.txt, which is silent on units).
|
# cog-api's llms.txt, which is silent on units).
|
||||||
|
|
||||||
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
||||||
testthat::skip("Blocked on uscogdata#15 (finding F-004)")
|
|
||||||
|
# README and vignettes ship only in the source tree, not in the installed
|
||||||
|
# package, so these assertions cannot run under R CMD check -- CI's earlier
|
||||||
|
# testthat::test_local() step is what enforces them. See
|
||||||
|
# skip_if_no_source_tree() in helper-fixture.R.
|
||||||
|
docs <- skip_if_no_source_tree(
|
||||||
|
"README.md",
|
||||||
|
c("vignettes", "total-spending.Rmd"),
|
||||||
|
c("vignettes", "population-denominators.Rmd")
|
||||||
|
)
|
||||||
|
|
||||||
says_units <- function(path) {
|
says_units <- function(path) {
|
||||||
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
||||||
@@ -25,10 +34,7 @@ test_that("returned amounts are documented as full US dollars where readers meet
|
|||||||
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
||||||
}
|
}
|
||||||
|
|
||||||
expect_true(says_units(testthat::test_path("..", "..", "README.md")))
|
for (path in docs) expect_true(says_units(path))
|
||||||
expect_true(says_units(testthat::test_path("..", "..", "vignettes", "total-spending.Rmd")))
|
|
||||||
expect_true(says_units(testthat::test_path("..", "..", "vignettes",
|
|
||||||
"population-denominators.Rmd")))
|
|
||||||
|
|
||||||
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
||||||
# The expected raw amount is read straight from the corpus's parquet
|
# The expected raw amount is read straight from the corpus's parquet
|
||||||
|
|||||||
@@ -0,0 +1,188 @@
|
|||||||
|
# tests/testthat/test-complete.R
|
||||||
|
#
|
||||||
|
# uscogdata#18. The published corpus no longer stores the wide era's explicit
|
||||||
|
# zeros (cog_pipeline#64, series break SB194), so absence means two different
|
||||||
|
# things:
|
||||||
|
#
|
||||||
|
# <= FY2011 (dense_source) : cell absent => Census published $0
|
||||||
|
# >= FY2012 (sparse_source): cell absent => not reported, unknown
|
||||||
|
#
|
||||||
|
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
|
||||||
|
# row's `value_source` so the two are distinguishable. Expected row sets here
|
||||||
|
# are built from the corpus parquet directly, never from the verb under test --
|
||||||
|
# verifying what a filter does through that same filter proves nothing.
|
||||||
|
|
||||||
|
# The (subtype, category) cells that SHOULD exist for one government-year:
|
||||||
|
# every code in force for that government's type, mapped through
|
||||||
|
# summary_categories, matching the verb's flow prefixes and excluding
|
||||||
|
# aggregate-flagged codes (which spending_long/revenue_long drop).
|
||||||
|
raw_expected_cells <- function(govid, year, prefixes, subtype_col) {
|
||||||
|
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
|
||||||
|
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
|
||||||
|
wt_raw_query(sprintf(
|
||||||
|
"SELECT DISTINCT c.%s AS subtype, c.category
|
||||||
|
FROM %s cs
|
||||||
|
JOIN %s x ON x.govs_type = cs.type
|
||||||
|
JOIN %s c ON c.item_code = cs.item_code
|
||||||
|
WHERE x.canonical_govid = '%s'
|
||||||
|
AND cs.year = %d
|
||||||
|
AND NOT cs.is_aggregate
|
||||||
|
AND LEFT(cs.item_code, 1) IN (%s)
|
||||||
|
AND c.category IS NOT NULL
|
||||||
|
AND c.%s IS NOT NULL",
|
||||||
|
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
|
||||||
|
q("summary_categories.parquet"), govid, year,
|
||||||
|
paste0("'", prefixes, "'", collapse = ","), subtype_col
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
test_that("complete = FALSE is the default and changes nothing", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
plain <- cog_spending("121011212191", 2011L)
|
||||||
|
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
|
||||||
|
expect_equal(nrow(plain), nrow(explicit))
|
||||||
|
expect_false("value_source" %in% names(plain))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# FY2011 is dense_source: before sparsification this government carried a
|
||||||
|
# row for every code in force, most of them $0. complete = TRUE must
|
||||||
|
# reproduce that cell set exactly.
|
||||||
|
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||||
|
expected <- raw_expected_cells("121011212191", 2011L,
|
||||||
|
c("E", "F", "G"), "spend_subtype")
|
||||||
|
|
||||||
|
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||||
|
expect_setequal(key(r$spend_subtype, r$category),
|
||||||
|
key(expected$subtype, expected$category))
|
||||||
|
expect_gt(nrow(expected), 0L)
|
||||||
|
|
||||||
|
# Every filled cell in a dense-source year is a Census-published $0 --
|
||||||
|
# never "unknown", which is what the modern era's absences mean.
|
||||||
|
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||||
|
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
|
||||||
|
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
plain <- cog_spending("121011212191", 2011L)
|
||||||
|
full <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||||
|
|
||||||
|
# Filling adds rows; it must never alter or drop one.
|
||||||
|
expect_gt(nrow(full), nrow(plain))
|
||||||
|
reported <- full[full$value_source == "reported", ]
|
||||||
|
expect_equal(nrow(reported), nrow(plain))
|
||||||
|
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
|
||||||
|
# ... and the total is unchanged, because every added cell is $0.
|
||||||
|
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("a sparse-source year's absences are unknown, not zero", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# FY2019 is sparse_source: an absent cell means the government did not
|
||||||
|
# report, which is NOT a zero. Filling those with 0 would invent data --
|
||||||
|
# the exact error the representation contract exists to prevent.
|
||||||
|
r <- cog_spending("121011212191", 2019L, complete = TRUE)
|
||||||
|
filled <- r[r$value_source != "reported", ]
|
||||||
|
expect_gt(nrow(filled), 0L)
|
||||||
|
expect_true(all(filled$value_source == "not_reported"))
|
||||||
|
expect_true(all(is.na(filled$amt_nominal)))
|
||||||
|
expect_false(any(r$value_source == "census_zero"))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("the fill is scoped to each government's own type", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# Filling against the union of all types would invent cells for codes a
|
||||||
|
# county can never report. Every filled category must be one that
|
||||||
|
# code_set puts in force for type 1 (county) specifically.
|
||||||
|
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||||
|
county_cells <- raw_expected_cells("121011212191", 2011L,
|
||||||
|
c("E", "F", "G"), "spend_subtype")
|
||||||
|
expect_true(all(r$category %in% county_cells$category))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("complete = TRUE respects the category filter", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
r <- cog_spending("121011212191", 2011L, category = "Police",
|
||||||
|
complete = TRUE)
|
||||||
|
expect_true(all(r$category == "Police"))
|
||||||
|
expect_true("value_source" %in% names(r))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("cog_revenue() completes on its own flow", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
|
||||||
|
expected <- raw_expected_cells("121011212191", 2011L,
|
||||||
|
c("T", "A", "U", "B", "C", "D"),
|
||||||
|
"revenue_subtype")
|
||||||
|
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||||
|
expect_setequal(key(r$revenue_subtype, r$category),
|
||||||
|
key(expected$subtype, expected$category))
|
||||||
|
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("provenance records the completion and its absence rule", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
|
||||||
|
"provenance")
|
||||||
|
expect_true(prov$completion$applied)
|
||||||
|
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
|
||||||
|
expect_gt(prov$completion$rows_filled, 0L)
|
||||||
|
|
||||||
|
off <- attr(cog_spending("121011212191", 2011L), "provenance")
|
||||||
|
expect_false(off$completion$applied)
|
||||||
|
expect_equal(off$completion$rows_filled, 0L)
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("complete = TRUE is refused where the fill would be guesswork", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# A recipe defines its own component codes and does not go through
|
||||||
|
# summary_categories at all, so there is no grid to fill from.
|
||||||
|
expect_error(
|
||||||
|
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
|
||||||
|
complete = TRUE),
|
||||||
|
class = "uscogdata_complete_unsupported"
|
||||||
|
)
|
||||||
|
# The intergovernmental leg keeps aggregate rows by design
|
||||||
|
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
|
||||||
|
expect_error(
|
||||||
|
cog_spending("121011212191", 2011L, expenditure_concept = "total",
|
||||||
|
complete = TRUE),
|
||||||
|
class = "uscogdata_complete_unsupported"
|
||||||
|
)
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("complete = TRUE aborts on a corpus with no representation contract", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# A corpus published before sparsification carries neither table, so there
|
||||||
|
# is nothing to fill from and no rule saying what an absence means. That
|
||||||
|
# must abort rather than guess.
|
||||||
|
with_corpus_missing_representation({
|
||||||
|
expect_error(
|
||||||
|
cog_spending("121011212191", 2011L, complete = TRUE),
|
||||||
|
class = "uscogdata_representation_unavailable"
|
||||||
|
)
|
||||||
|
# ... while an ordinary query on the same corpus still works.
|
||||||
|
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
|
||||||
|
})
|
||||||
|
})
|
||||||
@@ -0,0 +1,94 @@
|
|||||||
|
# tests/testthat/test-corpus-breaks.R
|
||||||
|
#
|
||||||
|
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
|
||||||
|
# are caveats about the corpus itself rather than about one item code:
|
||||||
|
#
|
||||||
|
# SB085 1977 dollar precision across the 1976/1977 boundary
|
||||||
|
# SB087 2002 imputation exclusion FY2002-2006
|
||||||
|
# SB194 2012 dense -> sparse representation change
|
||||||
|
# SB086 2017 government ID scheme change
|
||||||
|
#
|
||||||
|
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
|
||||||
|
# and no row's item_code is ever the literal "ALL", so none of them could
|
||||||
|
# ever reach a user. They now travel in their own provenance field,
|
||||||
|
# `corpus_break_refs`, which keeps them distinguishable from the
|
||||||
|
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
|
||||||
|
# result, not one series).
|
||||||
|
|
||||||
|
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
|
||||||
|
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
|
||||||
|
# DoD 4 intended to reach users.
|
||||||
|
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
|
expect_true("SB194" %in% prov$corpus_break_refs)
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
# 2019-2020 spans no catalogued corpus-wide break.
|
||||||
|
r <- cog_spending("121011212191", 2019:2020, "Police")
|
||||||
|
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("corpus_break_refs and series_break_refs stay disjoint", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
|
expect_type(prov$series_break_refs, "character")
|
||||||
|
expect_type(prov$corpus_break_refs, "character")
|
||||||
|
# An ALL caveat must never masquerade as a break in a specific series.
|
||||||
|
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
|
||||||
|
expect_false("SB194" %in% prov$series_break_refs)
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that(".build_corpus_break_refs matches on the break_year window alone", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
con <- cog_open()
|
||||||
|
on.exit(cog_close())
|
||||||
|
|
||||||
|
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
|
||||||
|
# series_breaks table is a full cross-vintage registry, so the matching
|
||||||
|
# logic is testable there even though no long partition covers it.
|
||||||
|
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||||
|
con, years = 1975:1980, schema_version = 6L
|
||||||
|
))
|
||||||
|
# ... and does not fire for a range that misses it, unlike a filter keyed
|
||||||
|
# on the era rather than the boundary.
|
||||||
|
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||||
|
con, years = 1978:1980, schema_version = 6L
|
||||||
|
))
|
||||||
|
|
||||||
|
# Unlike code-specific refs, these do not depend on which codes a result
|
||||||
|
# happens to contain -- that dependency is the whole defect.
|
||||||
|
expect_setequal(
|
||||||
|
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
|
||||||
|
"SB087"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
|
||||||
|
expect_equal(
|
||||||
|
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
|
||||||
|
character(0)
|
||||||
|
)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
with_fixture_corpus({
|
||||||
|
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||||
|
out <- paste(c(
|
||||||
|
capture.output(cog_explain(r)),
|
||||||
|
capture.output(cog_explain(r), type = "message")
|
||||||
|
), collapse = "\n")
|
||||||
|
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
|
||||||
|
expect_match(out, "SB194", fixed = TRUE)
|
||||||
|
})
|
||||||
|
})
|
||||||
@@ -30,7 +30,6 @@ wt_coverage <- function(x) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
test_that("multi-government aggregates disclose reporting coverage on every result", {
|
test_that("multi-government aggregates disclose reporting coverage on every result", {
|
||||||
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
|
|
||||||
|
|
||||||
# -- F-020: geographic rollups -------------------------------------------
|
# -- F-020: geographic rollups -------------------------------------------
|
||||||
# Wisconsin's city/village universe is 608 governments. On the bundled
|
# Wisconsin's city/village universe is 608 governments. On the bundled
|
||||||
@@ -49,10 +48,22 @@ test_that("multi-government aggregates disclose reporting coverage on every resu
|
|||||||
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
|
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
|
||||||
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
|
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
|
||||||
|
|
||||||
|
# Cross-check against the raw partitions, scoped to the SAME universe the
|
||||||
|
# rollup was given -- the 608 govids above. Scoping instead on the long
|
||||||
|
# table's own `type`/`fips_state` asks a different question and answers 595:
|
||||||
|
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
|
||||||
|
# identity, when they were townships) while the xwalk lists them as
|
||||||
|
# govs_type = 2 (their present identity, as villages). Schema v6 made the
|
||||||
|
# long table's geography present-harmonized and moved as-of-year to the
|
||||||
|
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
|
||||||
|
# in R/manifest.R. n_units_reporting counts against the requested universe,
|
||||||
|
# so 597 is the number that answers "how many of the governments I asked
|
||||||
|
# about reported".
|
||||||
raw_2012 <- wt_raw_query(paste0(
|
raw_2012 <- wt_raw_query(paste0(
|
||||||
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
|
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
|
||||||
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ",
|
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
|
||||||
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate"))
|
"AND canonical_govid IN (",
|
||||||
|
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
|
||||||
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
|
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
|
||||||
|
|
||||||
# -- F-023: peer cohorts --------------------------------------------------
|
# -- F-023: peer cohorts --------------------------------------------------
|
||||||
|
|||||||
@@ -18,7 +18,6 @@
|
|||||||
# semantics, not a row the fix makes findable.
|
# semantics, not a row the fix makes findable.
|
||||||
|
|
||||||
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
||||||
testthat::skip("Blocked on uscogdata#16 (finding F-025)")
|
|
||||||
|
|
||||||
# -- correctness (1): a government must be findable by its own exact name ---
|
# -- correctness (1): a government must be findable by its own exact name ---
|
||||||
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
||||||
|
|||||||
@@ -13,10 +13,16 @@
|
|||||||
# is unaffected, so the fix is documentation: one sentence in @return.
|
# is unaffected, so the fix is documentation: one sentence in @return.
|
||||||
|
|
||||||
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
|
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
|
||||||
testthat::skip("Blocked on uscogdata#14 (finding F-021)")
|
|
||||||
|
|
||||||
rd <- paste(readLines(testthat::test_path("..", "..", "man", "cog_peer_compare.Rd"),
|
# man/ ships only in the source tree (the installed package carries a
|
||||||
warn = FALSE), collapse = " ")
|
# compiled help database instead), so the prose assertions below cannot run
|
||||||
|
# under R CMD check -- CI's earlier testthat::test_local() step enforces
|
||||||
|
# them. The numeric pin further down needs only the corpus, but it lives in
|
||||||
|
# the same test_that() as the sentence it protects, deliberately: they are
|
||||||
|
# one claim, and splitting them would let the prose drift while a separate
|
||||||
|
# test kept passing.
|
||||||
|
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
|
||||||
|
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
|
||||||
|
|
||||||
# The @return section must say the quantile is computed within each cell...
|
# The @return section must say the quantile is computed within each cell...
|
||||||
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
|
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
|
||||||
|
|||||||
@@ -80,8 +80,11 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
|
|||||||
|
|
||||||
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
fl_state <- cog_gov_search("^FLORIDA$", type = "state")
|
# Unanchored: utility mode matches literally now, so "^...$" would be
|
||||||
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
# searched for as characters rather than read as anchors (uscogdata#16).
|
||||||
|
# Both still resolve to exactly one row once scoped by type/state.
|
||||||
|
fl_state <- cog_gov_search("FLORIDA", type = "state")
|
||||||
|
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||||
r <- cog_geographic_rollup(
|
r <- cog_geographic_rollup(
|
||||||
govids = list(state = fl_state, county = broward),
|
govids = list(state = fl_state, county = broward),
|
||||||
category = "Police", years = 2020L
|
category = "Police", years = 2020L
|
||||||
|
|||||||
@@ -98,7 +98,11 @@ test_that("cog_spending rejects invalid inputs", {
|
|||||||
|
|
||||||
test_that("cog_spending accepts a cog_gov_search result directly", {
|
test_that("cog_spending accepts a cog_gov_search result directly", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
# Unanchored: utility mode matches `name` as a literal substring now, so
|
||||||
|
# "^...$" would be searched for as those characters rather than read as
|
||||||
|
# anchors (uscogdata#16). Scoped by state and type, the bare name still
|
||||||
|
# resolves to exactly one row.
|
||||||
|
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||||
expect_gt(nrow(picks), 0L)
|
expect_gt(nrow(picks), 0L)
|
||||||
r <- cog_spending(picks, 2020L, "Corrections")
|
r <- cog_spending(picks, 2020L, "Corrections")
|
||||||
expect_equal(unique(r$canonical_govid), "121011212191")
|
expect_equal(unique(r$canonical_govid), "121011212191")
|
||||||
@@ -321,11 +325,13 @@ test_that("provenance$series_break_refs is a populated-when-applicable character
|
|||||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
r <- cog_spending("121011212191", 2020L, "Corrections")
|
||||||
refs <- attr(r, "provenance")$series_break_refs
|
refs <- attr(r, "provenance")$series_break_refs
|
||||||
expect_type(refs, "character")
|
expect_type(refs, "character")
|
||||||
# No catalogued series_breaks_pq row falls inside this fixture's
|
# No catalogued code-specific series_breaks_pq row falls inside this
|
||||||
# 2011/2012/2019/2020 window for the codes this query touches (E04/G04)
|
# fixture's 2011/2012/2019/2020 window for the codes this query touches
|
||||||
# -- data-verified; the mechanism itself is what's under test here, via
|
# (E04/G04) -- data-verified; the mechanism itself is what's under test
|
||||||
# a query-shaped unit test in test-views.R since the fixture has no
|
# here, via a query-shaped unit test in test-views.R since the fixture
|
||||||
# positive case to pin against.
|
# has no positive case to pin against. Corpus-wide ("ALL") entries never
|
||||||
|
# appear in this field by construction -- they travel in
|
||||||
|
# corpus_break_refs; see test-corpus-breaks.R.
|
||||||
expect_equal(refs, character(0))
|
expect_equal(refs, character(0))
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|||||||
@@ -177,11 +177,13 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
|
|||||||
})
|
})
|
||||||
|
|
||||||
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
||||||
# No series_breaks_pq row falls inside the bundled fixture's 2011-2020
|
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
|
||||||
# window (data-verified; see the "series_break_refs" test in
|
# 2011-2020 window (data-verified; see the "series_break_refs" test in
|
||||||
# test-spending.R), so this proves the matching logic itself against the
|
# test-spending.R), so this proves the matching logic itself against the
|
||||||
# live view + a synthetic year window that DOES hit a cataloged break
|
# live view + a synthetic year window that DOES hit a cataloged break
|
||||||
# (SB075, fin_code E62, break_year 2005).
|
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
|
||||||
|
# separate path with its own coverage -- SB194 does sit at 2012, inside
|
||||||
|
# the fixture window; see test-corpus-breaks.R.
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
con <- cog_open()
|
con <- cog_open()
|
||||||
on.exit(cog_close())
|
on.exit(cog_close())
|
||||||
|
|||||||
@@ -13,6 +13,8 @@ knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
|
|||||||
|
|
||||||
# Why per-year population matters
|
# Why per-year population matters
|
||||||
|
|
||||||
|
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
|
||||||
|
|
||||||
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
||||||
|
|
||||||
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
||||||
|
|||||||
@@ -29,6 +29,13 @@ controls which of these a query answers. This vignette walks through both
|
|||||||
questions with code that actually runs against the package's bundled fixture
|
questions with code that actually runs against the package's bundled fixture
|
||||||
corpus, then explains why the second question refuses `"total"` outright.
|
corpus, then explains why the second question refuses `"total"` outright.
|
||||||
|
|
||||||
|
Before any of the numbers below: every amount column here — `amt_nominal`,
|
||||||
|
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
|
||||||
|
dollars**. The raw Census files report **thousands of dollars** and the
|
||||||
|
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
|
||||||
|
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
|
||||||
|
billion. Do not scale it again.
|
||||||
|
|
||||||
```{r}
|
```{r}
|
||||||
library(uscogdata)
|
library(uscogdata)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user