Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4bedf857e9
|
||
|
|
a9de5ba5ab
|
||
|
|
d0b4bae3cc
|
+1
-1
@@ -1,7 +1,7 @@
|
|||||||
Package: uscogdata
|
Package: uscogdata
|
||||||
Type: Package
|
Type: Package
|
||||||
Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus
|
Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus
|
||||||
Version: 0.2.0
|
Version: 0.1.0
|
||||||
Authors@R:
|
Authors@R:
|
||||||
person("Civilytics", , , "jknowles@gmail.com", role = c("aut", "cre"))
|
person("Civilytics", , , "jknowles@gmail.com", role = c("aut", "cre"))
|
||||||
Description: Curated R verbs over the Civilytics US Census of Governments
|
Description: Curated R verbs over the Civilytics US Census of Governments
|
||||||
|
|||||||
@@ -1,80 +1,5 @@
|
|||||||
# uscogdata 0.2.0
|
|
||||||
|
|
||||||
## New features
|
|
||||||
|
|
||||||
* `cog_spending()` and `cog_revenue()` accept the reserved category
|
|
||||||
`"All Categories"`, returning one summed row per
|
|
||||||
`(year, canonical_govid, subtype)` across every category inside the
|
|
||||||
requested concept's subtype scope. Filtering the result to
|
|
||||||
`spend_subtype == "operations"` gives an operating-expenditure total.
|
|
||||||
`cog_geographic_rollup()` inherits it,
|
|
||||||
which is the efficient way to build a geographic total — previously a
|
|
||||||
caller had to issue one rollup per category and sum the results
|
|
||||||
(cog-api#37).
|
|
||||||
|
|
||||||
`"All Categories"` is not the same thing as `expenditure_concept = "total"`.
|
|
||||||
The concept chooses which subtypes are in scope; `"All Categories"` chooses
|
|
||||||
whether the rows inside that scope are broken out or summed.
|
|
||||||
|
|
||||||
* `cog_categories()` advertises `"All Categories"` for the expenditure and
|
|
||||||
revenue vocabularies, so the reserved value is discoverable.
|
|
||||||
|
|
||||||
* Coverage signposting (see "Signposting now catches partially-suppressed
|
|
||||||
categories" below) now also works in `category = "All Categories"` mode.
|
|
||||||
The recipe-suggestion candidate query used to be scoped by `category`,
|
|
||||||
which is never a match for the reserved `"All Categories"` value, so
|
|
||||||
`provenance$suggestions` always came back empty there — the one mode whose
|
|
||||||
whole point is "you cannot sum the wrong scope" was silently unable to
|
|
||||||
signal a wrong scope. The candidate query is now scoped by the concept's
|
|
||||||
subtype allowlist instead, symmetric with how `.build_verb_sql()` itself
|
|
||||||
scopes the summed total: Los Angeles County FY2011, `category = "All
|
|
||||||
Categories"` still excludes $271,589,000 of aggregate-published Public
|
|
||||||
Welfare (`E68`), but now names `recipe = "welfare_cash_e68_wide"` to
|
|
||||||
recover it instead of reporting zero suggestions.
|
|
||||||
|
|
||||||
## Documentation
|
|
||||||
|
|
||||||
* `cog_geographic_rollup()` and `cog_peer_compare()` now document that
|
|
||||||
`provenance$coverage`'s `n_units_reporting` is **category-conditional** and
|
|
||||||
is not a response rate: a government that was surveyed and genuinely spends
|
|
||||||
nothing in the requested category is indistinguishable from one never
|
|
||||||
surveyed (uscogdata#36).
|
|
||||||
|
|
||||||
# uscogdata 0.1.0 (development)
|
# uscogdata 0.1.0 (development)
|
||||||
|
|
||||||
## Signposting now catches partially-suppressed categories
|
|
||||||
|
|
||||||
* A coverage suggestion used to fire only when a category returned **no rows
|
|
||||||
at all** in a requested year. That missed the more dangerous case: a
|
|
||||||
category that still returns rows while silently dropping component codes
|
|
||||||
the wide era publishes only as aggregates (#9). `cog_spending(category =
|
|
||||||
"Public Welfare")` for FY2011 returned a plausible figure that omitted
|
|
||||||
`E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true
|
|
||||||
$5,261,404,000, a 39% understatement, with `provenance$suggestions` empty.
|
|
||||||
* Suggestions now also fire on **partial** coverage, and every suggestion
|
|
||||||
carries `trigger` (`"empty_year"` or `"suppressed_component"`),
|
|
||||||
`suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller
|
|
||||||
can see how much is missing and decide whether to re-run with the recipe.
|
|
||||||
* `cog_revenue()` gets the same fix through the shared verb path. Alaska's
|
|
||||||
FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping
|
|
||||||
$1,899,995,000 of aggregate-published `U4-` rents and royalties.
|
|
||||||
* The trigger stays recipe-driven, so it only fires where a harmonization
|
|
||||||
recipe actually exists to name the fix. `higher_ed_e18_wide` and
|
|
||||||
`general_gov_e89_wide` stay silent in every year measured on the bundled
|
|
||||||
fixture, because their components are ordinary classified leaves even
|
|
||||||
pre-2012.
|
|
||||||
* The `suppressed_component` trigger (and any `suppressed_amount`/
|
|
||||||
`suppressed_codes` an `empty_year` fire also carries) is scoped to the
|
|
||||||
calling verb's own flow family: `cog_spending()` only ever measures E/F/G
|
|
||||||
component dollars, `cog_revenue()` only T/A/U/B/C/D. A component from the
|
|
||||||
OTHER flow family reports `suppressed_amount = 0` rather than a fabricated
|
|
||||||
claim. The `empty_year` trigger itself is not flow-scoped -- a category
|
|
||||||
belonging to the other flow (e.g. `cog_spending(category = "IG Local")`)
|
|
||||||
still returns zero rows and can still fire, in any year including modern
|
|
||||||
ones, naming the recipe whose own generic join finds real data for this
|
|
||||||
government. That is a mis-scoped query, not a corpus-format gap, so its
|
|
||||||
`suppressed_amount` is correctly 0.
|
|
||||||
|
|
||||||
## New: `cog_balances()` for cash-and-security holdings
|
## New: `cog_balances()` for cash-and-security holdings
|
||||||
|
|
||||||
* New `cog_balances()` exposes the 14 cash-and-security holding codes
|
* New `cog_balances()` exposes the 14 cash-and-security holding codes
|
||||||
|
|||||||
+1
-13
@@ -28,12 +28,7 @@
|
|||||||
#' every combination would be either redundant or empty.
|
#' every combination would be either redundant or empty.
|
||||||
#' `category = "Fund Balances"` is exactly the `general` family
|
#' `category = "Fund Balances"` is exactly the `general` family
|
||||||
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
||||||
#' one `dplyr::filter()` away. The reserved pseudo-category
|
#' one `dplyr::filter()` away.
|
||||||
#' `"All Categories"` (see [cog_spending()]) is **not** supported here and
|
|
||||||
#' errors with class `uscogdata_all_categories_unsupported`: it sums a
|
|
||||||
#' concept's subtype scope, and holdings are a stock with no concept
|
|
||||||
#' vocabulary to sum across. Omit `category` to get every category broken
|
|
||||||
#' out instead.
|
|
||||||
#' @param per_capita Divide holdings by population. Note this is a **stock per
|
#' @param per_capita Divide holdings by population. Note this is a **stock per
|
||||||
#' resident** (reserves per person), which is *not* comparable to
|
#' resident** (reserves per person), which is *not* comparable to
|
||||||
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
|
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
|
||||||
@@ -79,13 +74,6 @@ cog_balances <- function(govid, years, category = NULL,
|
|||||||
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
|
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
|
||||||
# route the verb through .verb_spendrev(), which stays deliberately unused
|
# route the verb through .verb_spendrev(), which stays deliberately unused
|
||||||
# here because its flow vocabulary is meaningless for a stock.
|
# here because its flow vocabulary is meaningless for a stock.
|
||||||
#
|
|
||||||
# allow_all_categories is left at its FALSE default (contrast
|
|
||||||
# .verb_spendrev(), which passes TRUE): the all-categories mode's "sum"
|
|
||||||
# only means something in terms of a concept's subtype scope, and holdings
|
|
||||||
# have no concept vocabulary. The reuse above is exactly why this can be a
|
|
||||||
# one-line default rather than a second bespoke check -- see the
|
|
||||||
# validator's own doc comment for the incident that made that matter.
|
|
||||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
||||||
recipe)
|
recipe)
|
||||||
years <- as.integer(years)
|
years <- as.integer(years)
|
||||||
|
|||||||
+2
-28
@@ -22,10 +22,7 @@
|
|||||||
#' `category` column (e.g. `"Police"` or `"Tax"`).
|
#' `category` column (e.g. `"Police"` or `"Tax"`).
|
||||||
#' @return Tibble with columns `category`, `category_type`, `subtype`,
|
#' @return Tibble with columns `category`, `category_type`, `subtype`,
|
||||||
#' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
#' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
||||||
#' `category_type`, `category`, `subtype`. Includes one row per flow for the
|
#' `category_type`, `category`, `subtype`.
|
||||||
#' reserved pseudo-category `"All Categories"`, which carries `NA` for
|
|
||||||
#' `subtype`, `n_codes` and `item_codes` because it is a query mode rather
|
|
||||||
#' than a crosswalk entry — see [cog_spending()]'s `category` argument.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_categories <- function(type = NULL, pattern = NULL) {
|
cog_categories <- function(type = NULL, pattern = NULL) {
|
||||||
if (!is.null(type)) {
|
if (!is.null(type)) {
|
||||||
@@ -66,28 +63,5 @@ cog_categories <- function(type = NULL, pattern = NULL) {
|
|||||||
"GROUP BY category, category_type, subtype
|
"GROUP BY category, category_type, subtype
|
||||||
ORDER BY category_type, category, subtype"
|
ORDER BY category_type, category, subtype"
|
||||||
)
|
)
|
||||||
out <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||||
|
|
||||||
# The reserved pseudo-category is a query mode, not a crosswalk row, so it
|
|
||||||
# has no item codes to report -- hence NA rather than 0 for n_codes. It is
|
|
||||||
# emitted for the two FLOW vocabularies only: cog_balances() returns a stock
|
|
||||||
# and has no concept argument to sum within.
|
|
||||||
pseudo <- tibble::tibble(
|
|
||||||
category = .ALL_CATEGORIES,
|
|
||||||
category_type = c("expenditure", "revenue"),
|
|
||||||
subtype = NA_character_,
|
|
||||||
n_codes = NA_integer_,
|
|
||||||
item_codes = NA_character_
|
|
||||||
)
|
|
||||||
if (!is.null(type)) {
|
|
||||||
db_type <- if (type == "spending") "expenditure" else type
|
|
||||||
pseudo <- pseudo[pseudo$category_type == db_type, , drop = FALSE]
|
|
||||||
}
|
|
||||||
if (!is.null(pattern) && nrow(pseudo) > 0L) {
|
|
||||||
keep <- grepl(pattern, pseudo$category, ignore.case = TRUE)
|
|
||||||
pseudo <- pseudo[keep, , drop = FALSE]
|
|
||||||
}
|
|
||||||
if (nrow(pseudo) == 0L) return(out)
|
|
||||||
out <- rbind(out, pseudo)
|
|
||||||
out[order(out$category_type, out$category, out$subtype), , drop = FALSE]
|
|
||||||
}
|
}
|
||||||
|
|||||||
+1
-8
@@ -125,15 +125,8 @@ cog_explain <- function(result, format = c("print", "list")) {
|
|||||||
if (length(prov$suggestions) > 0L) {
|
if (length(prov$suggestions) > 0L) {
|
||||||
cli::cli_h2("Suggestions")
|
cli::cli_h2("Suggestions")
|
||||||
sugg_lines <- vapply(prov$suggestions, function(s) {
|
sugg_lines <- vapply(prov$suggestions, function(s) {
|
||||||
line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
|
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
|
||||||
s$available_years[1], s$available_years[2], s$hint)
|
s$available_years[1], s$available_years[2], s$hint)
|
||||||
if (isTRUE(s$suppressed_amount > 0)) {
|
|
||||||
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
|
|
||||||
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
|
|
||||||
paste0("FY", s$suppressed_years, collapse = ", "),
|
|
||||||
paste(s$suppressed_codes, collapse = ", ")))
|
|
||||||
}
|
|
||||||
line
|
|
||||||
}, character(1))
|
}, character(1))
|
||||||
cli::cli_ul(sugg_lines)
|
cli::cli_ul(sugg_lines)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -240,23 +240,6 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' group_by(year) |>
|
#' group_by(year) |>
|
||||||
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||||
#' ```
|
#' ```
|
||||||
#' @section Reading `coverage`:
|
|
||||||
#' `provenance$coverage` reports `n_units_reporting` against
|
|
||||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
|
||||||
#' it counts cohort members with rows for the category you asked for, not
|
|
||||||
#' cohort members collected that year.** A government that was surveyed and
|
|
||||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
|
||||||
#' that was never surveyed.
|
|
||||||
#'
|
|
||||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
|
||||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
|
||||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
|
||||||
#' contract policing to the county sheriff, not non-response.
|
|
||||||
#'
|
|
||||||
#' The comparison that *is* valid is the same category across a census year
|
|
||||||
#' (ending in 2 or 7) and a sample year, where the real-zero component is
|
|
||||||
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
|
|
||||||
#' marks which is which.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_peer_compare <- function(target_govid, peers, category, years,
|
cog_peer_compare <- function(target_govid, peers, category, years,
|
||||||
per_capita = TRUE, adjust_to_year = NULL,
|
per_capita = TRUE, adjust_to_year = NULL,
|
||||||
|
|||||||
+1
-9
@@ -67,15 +67,7 @@
|
|||||||
basis_note = basis_note,
|
basis_note = basis_note,
|
||||||
expenditure_concept = expenditure_concept,
|
expenditure_concept = expenditure_concept,
|
||||||
expenditure_concept_note = expenditure_concept_note,
|
expenditure_concept_note = expenditure_concept_note,
|
||||||
# isTRUE() alone would collapse a deliberate NA (all-categories mode,
|
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
|
||||||
# where suppression detection cannot run -- see .verb_spendrev()) down to
|
|
||||||
# FALSE, turning "we don't know" back into the false claim this field
|
|
||||||
# exists to avoid. Preserve NA; otherwise normalize to a strict logical.
|
|
||||||
expenditure_concept_direct_suppressed = if (isTRUE(is.na(expenditure_concept_direct_suppressed))) {
|
|
||||||
NA
|
|
||||||
} else {
|
|
||||||
isTRUE(expenditure_concept_direct_suppressed)
|
|
||||||
},
|
|
||||||
revenue_concept = revenue_concept,
|
revenue_concept = revenue_concept,
|
||||||
harmonization = harmonization %||% list(
|
harmonization = harmonization %||% list(
|
||||||
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
|
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
|
||||||
|
|||||||
+2
-15
@@ -8,17 +8,6 @@
|
|||||||
#' multiplies by 1000 and records the conversion in `provenance`).
|
#' multiplies by 1000 and records the conversion in `provenance`).
|
||||||
#'
|
#'
|
||||||
#' @inheritParams cog_spending
|
#' @inheritParams cog_spending
|
||||||
#' @param category Character vector of category names (from
|
|
||||||
#' `summary_categories.category`), or `NULL` for all categories broken out
|
|
||||||
#' one row each. The reserved value `"All Categories"` instead returns a
|
|
||||||
#' single summed row per `(year, canonical_govid, subtype)`, covering every
|
|
||||||
#' category inside the requested concept's subtype scope. It cannot be
|
|
||||||
#' combined with other category names, and it is not the same thing as
|
|
||||||
#' `revenue_concept = "total"`: the concept chooses which subtypes are in
|
|
||||||
#' scope, `"All Categories"` chooses whether rows inside that scope are
|
|
||||||
#' broken out or summed. Because the result keeps one row per
|
|
||||||
#' `revenue_subtype`, filtering the returned frame to
|
|
||||||
#' `revenue_subtype == "own_source"` gives an own-source revenue total.
|
|
||||||
#' @param revenue_concept Which of Census's two published revenue concepts to
|
#' @param revenue_concept Which of Census's two published revenue concepts to
|
||||||
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
||||||
#' values -- never as item-code first letters, which cannot classify
|
#' values -- never as item-code first letters, which cannot classify
|
||||||
@@ -54,7 +43,7 @@ cog_revenue <- function(govid, years, category = NULL,
|
|||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
revenue_concept = c("general", "total"),
|
revenue_concept = c("general", "total"),
|
||||||
complete = FALSE, limit = NULL, offset = NULL) {
|
complete = FALSE) {
|
||||||
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
|
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
|
||||||
# membership does -- General Revenue, i.e. everything except
|
# membership does -- General Revenue, i.e. everything except
|
||||||
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
|
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
|
||||||
@@ -73,8 +62,6 @@ cog_revenue <- function(govid, years, category = NULL,
|
|||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe,
|
recipe = recipe,
|
||||||
revenue_concept = revenue_concept,
|
revenue_concept = revenue_concept,
|
||||||
complete = complete,
|
complete = complete
|
||||||
limit = limit,
|
|
||||||
offset = offset
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|||||||
+1
-22
@@ -19,11 +19,7 @@
|
|||||||
#' `state`, `county`, `city`. Each element is a character vector of
|
#' `state`, `county`, `city`. Each element is a character vector of
|
||||||
#' `canonical_govid` values. At least one layer required.
|
#' `canonical_govid` values. At least one layer required.
|
||||||
#' @param category Single category name or character vector (passed through
|
#' @param category Single category name or character vector (passed through
|
||||||
#' to [cog_spending()]), or the reserved `"All Categories"` for one summed
|
#' to [cog_spending()]).
|
||||||
#' row per `(year, canonical_govid, subtype)` covering every category in the
|
|
||||||
#' concept's scope. `"All Categories"` is the efficient way to build a
|
|
||||||
#' geographic total: without it a caller must issue one rollup per category
|
|
||||||
#' and sum the results themselves.
|
|
||||||
#' @param years Integer vector of years.
|
#' @param years Integer vector of years.
|
||||||
#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
|
#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
|
||||||
#' population from `gov_population_yearly`. Govs with missing population
|
#' population from `gov_population_yearly`. Govs with missing population
|
||||||
@@ -60,23 +56,6 @@
|
|||||||
#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
|
#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
|
||||||
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
|
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
|
||||||
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
||||||
#' @section Reading `coverage`:
|
|
||||||
#' `provenance$coverage` reports `n_units_reporting` against
|
|
||||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
|
||||||
#' it counts governments with rows for the category you asked for, not
|
|
||||||
#' governments collected that year.** A government that was surveyed and
|
|
||||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
|
||||||
#' that was never surveyed.
|
|
||||||
#'
|
|
||||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
|
||||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
|
||||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
|
||||||
#' contract policing to the county sheriff, not non-response.
|
|
||||||
#'
|
|
||||||
#' The comparison that *is* valid is the same category across a census year
|
|
||||||
#' (ending in 2 or 7) and a sample year, where the real-zero component is
|
|
||||||
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
|
|
||||||
#' marks which is which.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_geographic_rollup <- function(govids, category, years,
|
cog_geographic_rollup <- function(govids, category, years,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
|
|||||||
+21
-246
@@ -19,14 +19,6 @@
|
|||||||
.spend_subtypes_primary <- c("operations", "capital", "assistance")
|
.spend_subtypes_primary <- c("operations", "capital", "assistance")
|
||||||
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
|
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
|
||||||
|
|
||||||
# The reserved pseudo-category. Deliberately NOT "Total": `category = "Total"`
|
|
||||||
# would sit one argument away from `expenditure_concept = "total"` and mean
|
|
||||||
# something different -- the concept selects WHICH SUBTYPES are in scope, this
|
|
||||||
# selects whether the rows inside that scope are broken out by category or
|
|
||||||
# summed. "All Categories" states the operation and cannot be misread as the
|
|
||||||
# concept.
|
|
||||||
.ALL_CATEGORIES <- "All Categories"
|
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.expenditure_concept_subtypes <- function(concept) {
|
.expenditure_concept_subtypes <- function(concept) {
|
||||||
switch(concept,
|
switch(concept,
|
||||||
@@ -71,16 +63,7 @@
|
|||||||
#' @param govid Character vector of `canonical_govid` values.
|
#' @param govid Character vector of `canonical_govid` values.
|
||||||
#' @param years Integer vector of years.
|
#' @param years Integer vector of years.
|
||||||
#' @param category Character vector of category names (from
|
#' @param category Character vector of category names (from
|
||||||
#' `summary_categories.category`), or `NULL` for all categories broken out
|
#' `summary_categories.category`), or `NULL` for all categories.
|
||||||
#' one row each. The reserved value `"All Categories"` instead returns a
|
|
||||||
#' single summed row per `(year, canonical_govid, subtype)`, covering every
|
|
||||||
#' category inside the requested concept's subtype scope. It cannot be
|
|
||||||
#' combined with other category names, and it is not the same thing as
|
|
||||||
#' `expenditure_concept = "total"`: the concept chooses which subtypes are in
|
|
||||||
#' scope, `"All Categories"` chooses whether rows inside that scope are
|
|
||||||
#' broken out or summed. Because the result keeps one row per
|
|
||||||
#' `spend_subtype`, filtering the returned frame to
|
|
||||||
#' `spend_subtype == "operations"` gives an operating-expenditure total.
|
|
||||||
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
|
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||||
#' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
#' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||||
#' Census F-33 population from `gov_population_yearly`. Result also gains
|
#' Census F-33 population from `gov_population_yearly`. Result also gains
|
||||||
@@ -150,12 +133,7 @@
|
|||||||
#' component (when one exists), and
|
#' component (when one exists), and
|
||||||
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||||
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
||||||
#' IG. When `category = "All Categories"` is combined with
|
#' IG.
|
||||||
#' `expenditure_concept = "total"`, this detection cannot run (it keys on
|
|
||||||
#' per-category rows, which all-categories mode collapses to one literal
|
|
||||||
#' value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
|
|
||||||
#' possibly-false `FALSE`; query an explicit `category` to get a real
|
|
||||||
#' answer.
|
|
||||||
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
||||||
#' corpus does not carry still appears, labelled with **why** it is
|
#' corpus does not carry still appears, labelled with **why** it is
|
||||||
#' missing, and add a `value_source` column to every row:
|
#' missing, and add a `value_source` column to every row:
|
||||||
@@ -178,13 +156,6 @@
|
|||||||
#' `recipe` or with `expenditure_concept = "total"` (class
|
#' `recipe` or with `expenditure_concept = "total"` (class
|
||||||
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
#' `code_set`.
|
#' `code_set`.
|
||||||
#' @param limit Maximum number of result rows to return, pushed into the SQL
|
|
||||||
#' query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
|
||||||
#' result is materialized. `NULL` (the default) returns every matching row,
|
|
||||||
#' exactly as before this parameter existed. Mutually exclusive with
|
|
||||||
#' `recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.
|
|
||||||
#' @param offset Rows to skip before `limit` starts counting (0-based).
|
|
||||||
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
|
|
||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
@@ -192,17 +163,13 @@
|
|||||||
#' and `value_source` when `complete = TRUE`.
|
#' and `value_source` when `complete = TRUE`.
|
||||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||||
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
||||||
#' per-year `absence_means` rule that was applied. When `limit` is set,
|
#' per-year `absence_means` rule that was applied.
|
||||||
#' also carries a `total_rows` attribute: the full unpaginated row count,
|
|
||||||
#' computed by the same query (`COUNT(*) OVER()`) rather than a second
|
|
||||||
#' round trip -- so a caller walking pages never has to ask "how many are
|
|
||||||
#' there" separately.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_spending <- function(govid, years, category = NULL,
|
cog_spending <- function(govid, years, category = NULL,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
expenditure_concept = c("primary", "direct", "total"),
|
||||||
complete = FALSE, limit = NULL, offset = NULL) {
|
complete = FALSE) {
|
||||||
# flow_prefixes no longer classifies rows (crosswalk subtype membership
|
# flow_prefixes no longer classifies rows (crosswalk subtype membership
|
||||||
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
|
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
|
||||||
# machinery to this verb's recipe families (see R/suggestions.R; the
|
# machinery to this verb's recipe families (see R/suggestions.R; the
|
||||||
@@ -221,9 +188,7 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe,
|
recipe = recipe,
|
||||||
expenditure_concept = expenditure_concept,
|
expenditure_concept = expenditure_concept,
|
||||||
complete = complete,
|
complete = complete
|
||||||
limit = limit,
|
|
||||||
offset = offset
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -251,7 +216,7 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL,
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
expenditure_concept = c("primary", "direct", "total"),
|
||||||
revenue_concept = c("general", "total"),
|
revenue_concept = c("general", "total"),
|
||||||
complete = FALSE, limit = NULL, offset = NULL) {
|
complete = FALSE) {
|
||||||
basis_explicit <- length(basis) == 1L
|
basis_explicit <- length(basis) == 1L
|
||||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||||
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
||||||
@@ -292,24 +257,8 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
}
|
}
|
||||||
|
|
||||||
govid <- .coerce_govid_input(govid, arg = "govid")
|
govid <- .coerce_govid_input(govid, arg = "govid")
|
||||||
# allow_all_categories = TRUE: cog_spending()/cog_revenue() are the two
|
|
||||||
# verbs the reserved pseudo-category is defined for. cog_balances() shares
|
|
||||||
# this validator but leaves the argument at its FALSE default, so it
|
|
||||||
# rejects "All Categories" instead of silently returning zero rows
|
|
||||||
# (finding 3, all-categories review).
|
|
||||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
||||||
recipe, allow_all_categories = TRUE)
|
recipe)
|
||||||
|
|
||||||
# Recognize the reserved pseudo-category. Detected after type validation so a
|
|
||||||
# non-character `category` still fails with the ordinary type error.
|
|
||||||
all_categories <- !is.null(category) && .ALL_CATEGORIES %in% category
|
|
||||||
if (all_categories && length(category) > 1L) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.val {(.ALL_CATEGORIES)}} cannot be combined with other categories.",
|
|
||||||
"i" = "It already sums every category in the requested concept's scope.",
|
|
||||||
"*" = "Ask for it alone, or list the specific categories you want."
|
|
||||||
), class = "uscogdata_all_categories_not_combinable")
|
|
||||||
}
|
|
||||||
|
|
||||||
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
|
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
|
||||||
cli::cli_abort(c(
|
cli::cli_abort(c(
|
||||||
@@ -350,44 +299,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
if (complete && all_categories) {
|
|
||||||
.abort_complete_unsupported(
|
|
||||||
"`category = \"All Categories\"` collapses the category dimension that `code_set` grids over (see `.completion_grid_sql()`), so there is no per-category grid left to fill -- filling a summed row has no defined semantics.",
|
|
||||||
"Drop `complete`, or use `complete = TRUE` with an explicit `category` (or `category = NULL` for every category)."
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
# limit/offset push the page into the SQL itself (see .build_verb_sql()),
|
|
||||||
# so the two things that would make "a page of what" ambiguous are refused
|
|
||||||
# up front rather than silently ignored: complete = TRUE fills a grid over
|
|
||||||
# the FULL requested (year, category) space, and a recipe's result comes
|
|
||||||
# from .run_recipe()'s own query, which this function does not touch.
|
|
||||||
if (!is.null(limit)) {
|
|
||||||
limit <- as.integer(limit)
|
|
||||||
if (length(limit) != 1L || is.na(limit) || limit < 0L) {
|
|
||||||
cli::cli_abort("`limit` must be a single non-negative integer.",
|
|
||||||
class = "uscogdata_invalid_pagination")
|
|
||||||
}
|
|
||||||
offset <- if (is.null(offset)) 0L else as.integer(offset)
|
|
||||||
if (length(offset) != 1L || is.na(offset) || offset < 0L) {
|
|
||||||
cli::cli_abort("`offset` must be a single non-negative integer.",
|
|
||||||
class = "uscogdata_invalid_pagination")
|
|
||||||
}
|
|
||||||
if (complete) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"`limit`/`offset` cannot be combined with `complete = TRUE`.",
|
|
||||||
"i" = "`complete` fills a grid over the FULL requested (year, category) space; paginating a slice of already-grouped rows has no defined meaning for the cells it would fill.",
|
|
||||||
"*" = "Drop `limit`/`offset`, or drop `complete`."
|
|
||||||
), class = "uscogdata_complete_pagination_conflict")
|
|
||||||
}
|
|
||||||
if (!is.null(recipe)) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"`limit`/`offset` cannot be combined with `recipe`.",
|
|
||||||
"i" = "A recipe's result comes from a separate query (`.run_recipe()`) that pagination is not wired into yet.",
|
|
||||||
"*" = "Drop `limit`/`offset`, or drop `recipe`."
|
|
||||||
), class = "uscogdata_recipe_pagination_conflict")
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
years <- as.integer(years)
|
years <- as.integer(years)
|
||||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||||
@@ -401,7 +312,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
|
|
||||||
recipe_block <- NULL
|
recipe_block <- NULL
|
||||||
category_for_prov <- category
|
category_for_prov <- category
|
||||||
total_rows <- NULL # set below only when limit is non-NULL (non-recipe path)
|
|
||||||
if (!is.null(recipe)) {
|
if (!is.null(recipe)) {
|
||||||
.require_schema_v5(con, manifest, "recipe =")
|
.require_schema_v5(con, manifest, "recipe =")
|
||||||
.validate_recipe_id(con, recipe)
|
.validate_recipe_id(con, recipe)
|
||||||
@@ -423,32 +333,9 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
} else {
|
} else {
|
||||||
NULL
|
NULL
|
||||||
}
|
}
|
||||||
sql <- .build_verb_sql(view, subtype_col, govid, years,
|
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
|
||||||
if (all_categories) NULL else category,
|
subtype_scope)
|
||||||
ig_view, subtype_scope,
|
|
||||||
all_categories = all_categories,
|
|
||||||
limit = limit, offset = offset)
|
|
||||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||||
if (!is.null(limit)) {
|
|
||||||
# COUNT(*) OVER() rides along as an ordinary column so the total comes
|
|
||||||
# from the same scan when this page has any rows -- see
|
|
||||||
# .build_verb_sql(). An empty page (offset past the end) carries no
|
|
||||||
# such row to read it from, so that one case falls back to a second,
|
|
||||||
# unpaginated COUNT(*) query rather than reporting a wrong zero.
|
|
||||||
if (nrow(result) > 0L) {
|
|
||||||
total_rows <- result$pagination_total_rows[[1]]
|
|
||||||
result$pagination_total_rows <- NULL
|
|
||||||
} else {
|
|
||||||
count_sql <- sprintf(
|
|
||||||
"SELECT COUNT(*) AS n FROM (%s) AS _uncounted",
|
|
||||||
.build_verb_sql(view, subtype_col, govid, years,
|
|
||||||
if (all_categories) NULL else category,
|
|
||||||
ig_view, subtype_scope,
|
|
||||||
all_categories = all_categories)
|
|
||||||
)
|
|
||||||
total_rows <- as.integer(DBI::dbGetQuery(con, count_sql)$n[[1]])
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
# Fill BEFORE per_capita / inflation so the added cells get the same
|
# Fill BEFORE per_capita / inflation so the added cells get the same
|
||||||
@@ -508,11 +395,7 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
}
|
}
|
||||||
suggestions <- .build_suggestions(con, govid, years, category,
|
suggestions <- .build_suggestions(con, govid, years, category,
|
||||||
direct_leg_result,
|
direct_leg_result,
|
||||||
resolved$basis, flow_prefixes,
|
resolved$basis, flow_prefixes)
|
||||||
.select_long_view(view_base, resolved$basis),
|
|
||||||
all_categories = all_categories,
|
|
||||||
subtype_col = subtype_col,
|
|
||||||
subtype_scope = subtype_scope)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
# C1(b): when expenditure_concept = "total", flag any row where the IG
|
# C1(b): when expenditure_concept = "total", flag any row where the IG
|
||||||
@@ -524,32 +407,13 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
# direct spending in that category, which is correct, ordinary data). When
|
# direct spending in that category, which is correct, ordinary data). When
|
||||||
# a covering recipe is found, both the row-level notes and the provenance
|
# a covering recipe is found, both the row-level notes and the provenance
|
||||||
# say so rather than pass silently as a plausible Total.
|
# say so rather than pass silently as a plausible Total.
|
||||||
#
|
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
|
||||||
# In all-categories mode this cannot run at all: .detect_direct_suppressed()
|
|
||||||
# keys on (year, canonical_govid, category), and every row shares the same
|
|
||||||
# literal "All Categories" value, so the key collides across every real
|
|
||||||
# category for that (year, govid) -- an IG-only row for a suppressed
|
|
||||||
# category becomes indistinguishable from one sharing a key with an
|
|
||||||
# unrelated category's ordinary Direct row. `has_direct` would then read
|
|
||||||
# TRUE whenever the government has ANY direct spending at all, and the
|
|
||||||
# detector could never fire. Rather than run it and report a false FALSE,
|
|
||||||
# skip it and record NA -- the provenance must stop making a claim it
|
|
||||||
# cannot support (finding 1, all-categories review).
|
|
||||||
suppression_unavailable <- all_categories &&
|
|
||||||
identical(expenditure_concept, "total")
|
|
||||||
direct_suppressed_info <- if (suppression_unavailable) {
|
|
||||||
list(flag = rep(NA, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
|
||||||
} else if (identical(expenditure_concept, "total")) {
|
|
||||||
.detect_direct_suppressed(con, result, subtype_col)
|
.detect_direct_suppressed(con, result, subtype_col)
|
||||||
} else {
|
} else {
|
||||||
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
||||||
}
|
}
|
||||||
direct_suppressed <- direct_suppressed_info$flag
|
direct_suppressed <- direct_suppressed_info$flag
|
||||||
direct_suppressed_flag <- if (suppression_unavailable) {
|
direct_suppressed_flag <- isTRUE(any(direct_suppressed))
|
||||||
NA
|
|
||||||
} else {
|
|
||||||
isTRUE(any(direct_suppressed))
|
|
||||||
}
|
|
||||||
|
|
||||||
result$notes <- .notes_column(result, direct_suppressed_info$notes)
|
result$notes <- .notes_column(result, direct_suppressed_info$notes)
|
||||||
|
|
||||||
@@ -558,20 +422,9 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
# leg is suppressed for at least one requested (year, category), append an
|
# leg is suppressed for at least one requested (year, category), append an
|
||||||
# explicit warning rather than let the base note's "Total = Direct + IG"
|
# explicit warning rather than let the base note's "Total = Direct + IG"
|
||||||
# framing stand unqualified for rows where that arithmetic didn't happen.
|
# framing stand unqualified for rows where that arithmetic didn't happen.
|
||||||
# When suppression detection itself is unavailable (all-categories mode),
|
|
||||||
# say so instead of silently reusing the unqualified base note.
|
|
||||||
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
|
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
|
||||||
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
|
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
|
||||||
if (suppression_unavailable) {
|
if (direct_suppressed_flag) {
|
||||||
paste0(
|
|
||||||
base_note,
|
|
||||||
" NOTE: direct-leg-suppression detection is unavailable when ",
|
|
||||||
"`category = \"All Categories\"` -- it keys on per-category rows, ",
|
|
||||||
"which this mode collapses. `expenditure_concept_direct_suppressed` ",
|
|
||||||
"is NA here rather than a possibly-false FALSE; query an explicit ",
|
|
||||||
"`category` (or `category = NULL`) to get a real answer."
|
|
||||||
)
|
|
||||||
} else if (isTRUE(direct_suppressed_flag)) {
|
|
||||||
paste0(
|
paste0(
|
||||||
base_note,
|
base_note,
|
||||||
" NOTE: for at least one requested (year, category) the Direct leg ",
|
" NOTE: for at least one requested (year, category) the Direct leg ",
|
||||||
@@ -613,32 +466,15 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
prov$scope$govids_missing <- scope$missing
|
prov$scope$govids_missing <- scope$missing
|
||||||
attr(result, "provenance") <- prov
|
attr(result, "provenance") <- prov
|
||||||
attr(result, ".popyear_range") <- NULL
|
attr(result, ".popyear_range") <- NULL
|
||||||
# Attached here, after every downstream transform (per_capita/real-dollar
|
|
||||||
# joins, notes, subtype filtering), the same way provenance is -- an
|
|
||||||
# attribute set before those runs is not guaranteed to survive them.
|
|
||||||
if (!is.null(limit)) attr(result, "total_rows") <- total_rows
|
|
||||||
|
|
||||||
if (length(suggestions) > 0L) .inform_suggestions(suggestions)
|
if (length(suggestions) > 0L) .inform_suggestions(suggestions)
|
||||||
|
|
||||||
result
|
result
|
||||||
}
|
}
|
||||||
|
|
||||||
#' Shared input validation for the money/holdings verbs.
|
|
||||||
#'
|
|
||||||
#' `allow_all_categories` gates the reserved pseudo-category
|
|
||||||
#' `.ALL_CATEGORIES` ("All Categories"). It is meaningful only where a
|
|
||||||
#' concept's subtype scope defines what "all" sums over --
|
|
||||||
#' `cog_spending()`/`cog_revenue()`, via `.verb_spendrev()`, pass `TRUE`.
|
|
||||||
#' `cog_balances()` leaves it at the `FALSE` default: holdings are a stock
|
|
||||||
#' with no concept vocabulary to sum across (see R/balances.R), and before
|
|
||||||
#' this guard existed `cog_balances(category = "All Categories")` silently
|
|
||||||
#' matched zero crosswalk rows and returned an empty result with no error
|
|
||||||
#' (finding 3, all-categories review). This validator is shared specifically
|
|
||||||
#' so the three verbs cannot drift apart on this again.
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.validate_verb_inputs <- function(govid, years, category,
|
.validate_verb_inputs <- function(govid, years, category,
|
||||||
per_capita, adjust_to_year, recipe = NULL,
|
per_capita, adjust_to_year, recipe = NULL) {
|
||||||
allow_all_categories = FALSE) {
|
|
||||||
if (!is.character(govid) || length(govid) == 0L) {
|
if (!is.character(govid) || length(govid) == 0L) {
|
||||||
cli::cli_abort("`govid` must be a non-empty character vector.")
|
cli::cli_abort("`govid` must be a non-empty character vector.")
|
||||||
}
|
}
|
||||||
@@ -648,14 +484,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
if (!is.null(category) && !is.character(category)) {
|
if (!is.null(category) && !is.character(category)) {
|
||||||
cli::cli_abort("`category` must be character or NULL.")
|
cli::cli_abort("`category` must be character or NULL.")
|
||||||
}
|
}
|
||||||
if (!allow_all_categories && !is.null(category) &&
|
|
||||||
.ALL_CATEGORIES %in% category) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.val {(.ALL_CATEGORIES)}} is not supported here.",
|
|
||||||
i = "It sums a spending or revenue concept's subtype scope; this verb has no concept vocabulary to sum across.",
|
|
||||||
i = "Use {.fn cog_spending} or {.fn cog_revenue} for an all-categories total."
|
|
||||||
), class = "uscogdata_all_categories_unsupported")
|
|
||||||
}
|
|
||||||
if (!is.logical(per_capita) || length(per_capita) != 1L) {
|
if (!is.logical(per_capita) || length(per_capita) != 1L) {
|
||||||
cli::cli_abort("`per_capita` must be a length-1 logical.")
|
cli::cli_abort("`per_capita` must be a length-1 logical.")
|
||||||
}
|
}
|
||||||
@@ -684,18 +512,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
|
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
|
||||||
}
|
}
|
||||||
|
|
||||||
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
|
|
||||||
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
|
|
||||||
#' anti-joins the LONG view rather than the annotated one: they have identical
|
|
||||||
#' row membership (the annotated views are the long views plus LEFT JOINs, see
|
|
||||||
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
|
|
||||||
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
|
|
||||||
#' the suppression test is asking about.
|
|
||||||
#' @noRd
|
|
||||||
.select_long_view <- function(view_base, basis) {
|
|
||||||
.select_view(sub("_annotated$", "_long", view_base), basis)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.select_ig_view <- function(basis) {
|
.select_ig_view <- function(basis) {
|
||||||
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
|
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
|
||||||
@@ -741,16 +557,10 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.build_verb_sql <- function(view, subtype_col, govid, years, category,
|
.build_verb_sql <- function(view, subtype_col, govid, years, category,
|
||||||
ig_view = NULL, subtype_scope = NULL,
|
ig_view = NULL, subtype_scope = NULL) {
|
||||||
all_categories = FALSE, limit = NULL, offset = NULL) {
|
|
||||||
govid_lit <- .sql_lit_chr(govid)
|
govid_lit <- .sql_lit_chr(govid)
|
||||||
years_lit <- paste(as.integer(years), collapse = ",")
|
years_lit <- paste(as.integer(years), collapse = ",")
|
||||||
# In all-categories mode there is no category filter: the sum is defined by
|
category_pred <- if (is.null(category)) {
|
||||||
# the concept's SUBTYPE allowlist (subtype_pred below), which is the real
|
|
||||||
# concept boundary. Filtering by category as well would be a no-op at best
|
|
||||||
# and, if the crosswalk ever gained an uncategorized code, a silent
|
|
||||||
# under-count of the very total this mode exists to guarantee.
|
|
||||||
category_pred <- if (all_categories || is.null(category)) {
|
|
||||||
""
|
""
|
||||||
} else {
|
} else {
|
||||||
sprintf("AND category IN (%s)", .sql_lit_chr(category))
|
sprintf("AND category IN (%s)", .sql_lit_chr(category))
|
||||||
@@ -789,26 +599,13 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
# though its dollars came entirely from an aggregate row, silently
|
# though its dollars came entirely from an aggregate row, silently
|
||||||
# suppressing the "Aggregate fallback applied" note on exactly the rows
|
# suppressing the "Aggregate fallback applied" note on exactly the rows
|
||||||
# this feature exists to surface.
|
# this feature exists to surface.
|
||||||
|
sprintf(
|
||||||
# Collapse the category dimension. subtype is deliberately KEPT: it is what
|
|
||||||
# lets a caller filter the result to `spend_subtype == "operations"` and
|
|
||||||
# get an operating-expenditure total, the measure a fiscal comparison
|
|
||||||
# actually wants. (There is no `subtype` argument -- this is a post-hoc
|
|
||||||
# filter on the returned column, not a query parameter.)
|
|
||||||
category_select <- if (all_categories) {
|
|
||||||
sprintf("%s AS category", .sql_lit_chr(.ALL_CATEGORIES))
|
|
||||||
} else {
|
|
||||||
"category"
|
|
||||||
}
|
|
||||||
category_group <- if (all_categories) "" else ", category"
|
|
||||||
|
|
||||||
base_sql <- sprintf(
|
|
||||||
"SELECT
|
"SELECT
|
||||||
year,
|
year,
|
||||||
canonical_govid,
|
canonical_govid,
|
||||||
COALESCE(xwalk_gov_name, gov_name) AS gov_name,
|
COALESCE(xwalk_gov_name, gov_name) AS gov_name,
|
||||||
%1$s,
|
%1$s,
|
||||||
%7$s,
|
category,
|
||||||
SUM(amt) * 1000.0 AS amt_nominal,
|
SUM(amt) * 1000.0 AS amt_nominal,
|
||||||
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
|
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
|
||||||
bool_or(is_aggregate) AS aggregate_fallback
|
bool_or(is_aggregate) AS aggregate_fallback
|
||||||
@@ -817,32 +614,10 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
AND year IN (%4$s)
|
AND year IN (%4$s)
|
||||||
%5$s
|
%5$s
|
||||||
%6$s
|
%6$s
|
||||||
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s%8$s
|
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
|
||||||
ORDER BY year, canonical_govid, %1$s%8$s",
|
ORDER BY year, canonical_govid, %1$s, category",
|
||||||
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred,
|
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
|
||||||
category_select, category_group
|
|
||||||
)
|
)
|
||||||
|
|
||||||
# limit/offset push the page into the query itself instead of pulling every
|
|
||||||
# matching row across the network only to slice and discard most of it
|
|
||||||
# afterward (the pattern behind the 2026-08-06 production incident: a
|
|
||||||
# 193,105-row/194-page sweep re-ran the full query and re-listified every
|
|
||||||
# row on EVERY page). COUNT(*) OVER() rides along as an ordinary column so
|
|
||||||
# the caller gets the true total from this same scan -- see the call site
|
|
||||||
# in .verb_spendrev(), which reads it off row 1 and strips it back out.
|
|
||||||
# The outer SELECT * wrapping (rather than appending LIMIT/OFFSET directly
|
|
||||||
# to base_sql) is what makes COUNT(*) OVER() see the post-GROUP-BY row
|
|
||||||
# count, not the pre-aggregation one.
|
|
||||||
if (is.null(limit)) {
|
|
||||||
base_sql
|
|
||||||
} else {
|
|
||||||
sprintf(
|
|
||||||
"SELECT *, COUNT(*) OVER() AS pagination_total_rows
|
|
||||||
FROM (%s) AS _paged
|
|
||||||
LIMIT %d OFFSET %d",
|
|
||||||
base_sql, limit, offset
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
|
|||||||
+18
-127
@@ -1,16 +1,8 @@
|
|||||||
# R/suggestions.R
|
# R/suggestions.R
|
||||||
# Recipe-component-driven signposting. When a basis = "harmonized" query for
|
# Recipe-component-driven signposting: when a basis = "harmonized" query for
|
||||||
# a category comes back incomplete in some requested year -- and a
|
# a category comes back with a coverage gap in some requested years (the
|
||||||
# harmonization recipe would actually fill it for this government -- surface
|
# result has no rows at all in that year) that a harmonization recipe would
|
||||||
# that recipe as a suggestion. "Incomplete" has two forms, and a recipe
|
# actually fill for this government, surface that recipe as a suggestion.
|
||||||
# qualifies on either:
|
|
||||||
# 1. empty_year -- the result has no rows at all in that year.
|
|
||||||
# 2. suppressed_component -- the result HAS rows, but a component code
|
|
||||||
# carries dollars the verb's own long view structurally excludes
|
|
||||||
# (aggregate-published, or absent from summary_categories). This is
|
|
||||||
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
|
|
||||||
# aggregate-only E67/E68, so form 1 never fired and the caller got a
|
|
||||||
# number a third too low with no signpost at all.
|
|
||||||
#
|
#
|
||||||
# This is deliberately keyed off the recipe catalog's component codes, not
|
# This is deliberately keyed off the recipe catalog's component codes, not
|
||||||
# off harmonization_map rows: no live map row carries a non-blank
|
# off harmonization_map rows: no live map row carries a non-blank
|
||||||
@@ -56,39 +48,11 @@
|
|||||||
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
|
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
|
||||||
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
|
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
|
||||||
#' lookup scoped to the calling verb's own flow family.
|
#' lookup scoped to the calling verb's own flow family.
|
||||||
#' @param long_view Name of the verb's own long view (from
|
|
||||||
#' `.select_long_view()`), passed through to `.suppressed_components()` to
|
|
||||||
#' measure the second qualifying path (uscogdata#9).
|
|
||||||
#' @param all_categories `TRUE` when the caller's `category` is the reserved
|
|
||||||
#' pseudo-category (`.ALL_CATEGORIES`). Defaults to `FALSE` so no other
|
|
||||||
#' caller's behaviour changes. When `TRUE`, the candidate-recipe sub-select
|
|
||||||
#' is scoped by `subtype_col`/`subtype_scope` instead of by `category` --
|
|
||||||
#' symmetric with `.build_verb_sql()`'s own all-categories branch (see
|
|
||||||
#' R/spending.R): the concept's subtype allowlist is the real scope
|
|
||||||
#' boundary, not any literal category value, and
|
|
||||||
#' `.ALL_CATEGORIES` ("All Categories") is never itself a row in
|
|
||||||
#' `summary_categories.category`, so leaving the category-keyed sub-select
|
|
||||||
#' in place here always returned zero candidates and silently disabled
|
|
||||||
#' signposting in all-categories mode (final whole-branch review, finding
|
|
||||||
#' 6).
|
|
||||||
#' @param subtype_col Name of the `summary_categories` subtype column to
|
|
||||||
#' scope by when `all_categories = TRUE` (`"spend_subtype"` or
|
|
||||||
#' `"revenue_subtype"` -- the same value `.build_verb_sql()` already
|
|
||||||
#' receives as its own `subtype_col`). Ignored when `all_categories =
|
|
||||||
#' FALSE`. `NULL` by default.
|
|
||||||
#' @param subtype_scope Character vector of subtype values to scope by when
|
|
||||||
#' `all_categories = TRUE` (the same value `.build_verb_sql()` already
|
|
||||||
#' receives as its own `subtype_scope` -- the concept's subtype allowlist,
|
|
||||||
#' e.g. `.expenditure_concept_subtypes(expenditure_concept)`). Ignored when
|
|
||||||
#' `all_categories = FALSE`. `NULL` by default.
|
|
||||||
#' @return List of `list(recipe_id, label, available_years, hint,
|
#' @return List of `list(recipe_id, label, available_years, hint,
|
||||||
#' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
|
#' ig_recipe_id)`, possibly empty.
|
||||||
#' suppressed_codes)`, possibly empty.
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.build_suggestions <- function(con, govid, years, category, result, basis,
|
.build_suggestions <- function(con, govid, years, category, result, basis,
|
||||||
flow_prefixes, long_view,
|
flow_prefixes) {
|
||||||
all_categories = FALSE,
|
|
||||||
subtype_col = NULL, subtype_scope = NULL) {
|
|
||||||
if (!identical(basis, "harmonized") || is.null(category)) return(list())
|
if (!identical(basis, "harmonized") || is.null(category)) return(list())
|
||||||
|
|
||||||
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
|
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
|
||||||
@@ -103,35 +67,16 @@
|
|||||||
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
|
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
|
||||||
# should never be suggested as a coverage-gap filler for EITHER verb, not
|
# should never be suggested as a coverage-gap filler for EITHER verb, not
|
||||||
# just kept from being named as the *counterpart* of another suggestion.
|
# just kept from being named as the *counterpart* of another suggestion.
|
||||||
#
|
|
||||||
# The inner sub-select is the concept boundary (finding 6, final
|
|
||||||
# whole-branch review): in all-categories mode it is scoped by
|
|
||||||
# `subtype_col`/`subtype_scope` -- the same allowlist `.build_verb_sql()`
|
|
||||||
# applies as a WHERE predicate to make the summed result a *concept*, not
|
|
||||||
# by `category` (`.ALL_CATEGORIES` is never a row in
|
|
||||||
# `summary_categories.category`, so a category-keyed sub-select always
|
|
||||||
# came back empty here). The M/L exclusion below is unchanged either way.
|
|
||||||
candidate_scope_sql <- if (isTRUE(all_categories)) {
|
|
||||||
sprintf(
|
|
||||||
"SELECT DISTINCT item_code FROM summary_categories WHERE %s IN (%s)",
|
|
||||||
subtype_col, .sql_lit_chr(subtype_scope)
|
|
||||||
)
|
|
||||||
} else {
|
|
||||||
sprintf(
|
|
||||||
"SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)",
|
|
||||||
.sql_lit_chr(category)
|
|
||||||
)
|
|
||||||
}
|
|
||||||
candidates <- DBI::dbGetQuery(con, sprintf(
|
candidates <- DBI::dbGetQuery(con, sprintf(
|
||||||
"SELECT DISTINCT recipe_id FROM harmonization_recipes
|
"SELECT DISTINCT recipe_id FROM harmonization_recipes
|
||||||
WHERE component_code IN (
|
WHERE component_code IN (
|
||||||
%s
|
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)
|
||||||
)
|
)
|
||||||
AND recipe_id NOT IN (
|
AND recipe_id NOT IN (
|
||||||
SELECT DISTINCT recipe_id FROM harmonization_recipes
|
SELECT DISTINCT recipe_id FROM harmonization_recipes
|
||||||
WHERE LEFT(component_code, 1) IN ('M', 'L')
|
WHERE LEFT(component_code, 1) IN ('M', 'L')
|
||||||
)",
|
)",
|
||||||
candidate_scope_sql
|
.sql_lit_chr(category)
|
||||||
))$recipe_id
|
))$recipe_id
|
||||||
if (length(candidates) == 0L) return(list())
|
if (length(candidates) == 0L) return(list())
|
||||||
|
|
||||||
@@ -141,28 +86,7 @@
|
|||||||
unique(as.integer(result$year))
|
unique(as.integer(result$year))
|
||||||
}
|
}
|
||||||
gap_years <- setdiff(as.integer(years), result_years)
|
gap_years <- setdiff(as.integer(years), result_years)
|
||||||
|
if (length(gap_years) == 0L) return(list())
|
||||||
# Path 2 (uscogdata#9): component dollars this government holds that the
|
|
||||||
# verb's own view structurally excludes. Measured across ALL requested
|
|
||||||
# years, not just gap years -- the whole point is that a year with rows can
|
|
||||||
# still be missing dollars. Scoped to the calling verb's own flow_prefixes
|
|
||||||
# (I1) -- see `.suppressed_components()`'s own roxygen for why.
|
|
||||||
#
|
|
||||||
# This runs unconditionally whenever there are candidates -- an earlier
|
|
||||||
# revision of this fix wave tried a free, in-memory pre-check
|
|
||||||
# (`.needs_suppression_query()`) to skip the round trip on an already-
|
|
||||||
# covered path, but a scoped re-review measured it against the fixture and
|
|
||||||
# found it didn't pay for itself (it skipped ~3% of healthy calls, ~0% of
|
|
||||||
# the multi-govid batch shape it was meant to help, at a net cost increase
|
|
||||||
# once its own always-run metadata query was counted) while adding an
|
|
||||||
# untested exactness invariant -- that `result$codes_included` and this
|
|
||||||
# anti-join share the harmonized `item_code` space -- whose silent
|
|
||||||
# violation would kill signposting, the exact failure class uscogdata#9
|
|
||||||
# exists to prevent. Owner's call: keep this simple; a batch-aware
|
|
||||||
# optimization, if one is worth building, is a separate issue.
|
|
||||||
supp <- .suppressed_components(con, candidates, govid, years, long_view, flow_prefixes)
|
|
||||||
|
|
||||||
if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
|
|
||||||
|
|
||||||
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
|
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
|
||||||
"SELECT recipe_id, any_value(label) AS label,
|
"SELECT recipe_id, any_value(label) AS label,
|
||||||
@@ -173,12 +97,11 @@
|
|||||||
.sql_lit_chr(candidates)
|
.sql_lit_chr(candidates)
|
||||||
)))
|
)))
|
||||||
|
|
||||||
# Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
|
# Which (recipe_id, year) pairs the recipe's own generic join actually
|
||||||
# covers for this government, restricted to the gap years.
|
# covers for this government, restricted to the gap years -- the same
|
||||||
covered <- if (length(gap_years) == 0L) {
|
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
|
||||||
data.frame(recipe_id = character(0), year = integer(0))
|
# no is_aggregate filter), just checking existence instead of summing.
|
||||||
} else {
|
covered <- DBI::dbGetQuery(con, sprintf(
|
||||||
DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT DISTINCT r.recipe_id, l.year
|
"SELECT DISTINCT r.recipe_id, l.year
|
||||||
FROM long l
|
FROM long l
|
||||||
JOIN harmonization_recipes r
|
JOIN harmonization_recipes r
|
||||||
@@ -193,35 +116,16 @@
|
|||||||
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
||||||
paste(gap_years, collapse = ",")
|
paste(gap_years, collapse = ",")
|
||||||
))
|
))
|
||||||
}
|
|
||||||
|
|
||||||
suggestions <- list()
|
suggestions <- list()
|
||||||
for (rid in candidates) {
|
for (rid in candidates) {
|
||||||
empty_hit <- rid %in% covered$recipe_id
|
if (!rid %in% covered$recipe_id) next
|
||||||
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
|
|
||||||
supp_hit <- nrow(s_rows) > 0L
|
|
||||||
if (!empty_hit && !supp_hit) next
|
|
||||||
m <- meta[meta$recipe_id == rid, ]
|
m <- meta[meta$recipe_id == rid, ]
|
||||||
suggestions[[length(suggestions) + 1L]] <- list(
|
suggestions[[length(suggestions) + 1L]] <- list(
|
||||||
recipe_id = rid,
|
recipe_id = rid,
|
||||||
label = m$label[[1]],
|
label = m$label[[1]],
|
||||||
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
|
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
|
||||||
hint = sprintf("re-run with recipe = '%s'", rid),
|
hint = sprintf("re-run with recipe = '%s'", rid)
|
||||||
# An empty year is the stronger claim -- the category returned nothing
|
|
||||||
# at all -- so it wins when both paths qualify. The suppressed_* fields
|
|
||||||
# are still populated, so an empty_year fire also reports its dollars.
|
|
||||||
trigger = if (empty_hit) "empty_year" else "suppressed_component",
|
|
||||||
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
|
|
||||||
suppressed_years = if (supp_hit) {
|
|
||||||
sort(unique(as.integer(s_rows$year)))
|
|
||||||
} else {
|
|
||||||
integer(0)
|
|
||||||
},
|
|
||||||
suppressed_codes = if (supp_hit) {
|
|
||||||
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
|
|
||||||
} else {
|
|
||||||
character(0)
|
|
||||||
}
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
.attach_ig_counterparts(con, suggestions, flow_prefixes)
|
.attach_ig_counterparts(con, suggestions, flow_prefixes)
|
||||||
@@ -328,25 +232,12 @@
|
|||||||
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
|
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
|
||||||
#' continuation line is appended naming the intergovernmental counterpart
|
#' continuation line is appended naming the intergovernmental counterpart
|
||||||
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
|
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
|
||||||
#' same bullet under cli, not a new bullet). Same treatment for
|
#' same bullet under cli, not a new bullet).
|
||||||
#' `suppressed_amount` (uscogdata#9): only present when dollars were
|
|
||||||
#' actually measured as excluded (an `empty_year` fire can carry them too --
|
|
||||||
#' see `.build_suggestions()` -- so this keys off the amount, not `trigger`).
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.inform_suggestions <- function(suggestions) {
|
.inform_suggestions <- function(suggestions) {
|
||||||
bullets <- vapply(suggestions, function(s) {
|
bullets <- vapply(suggestions, function(s) {
|
||||||
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
|
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
|
||||||
s$available_years[1], s$available_years[2], s$hint)
|
s$available_years[1], s$available_years[2], s$hint)
|
||||||
# Only present when dollars were actually measured as excluded. An
|
|
||||||
# empty_year fire can carry them too -- the year had no rows AND the
|
|
||||||
# component was suppressed -- which is strictly more informative.
|
|
||||||
if (isTRUE(s$suppressed_amount > 0)) {
|
|
||||||
bullet <- paste0(bullet, sprintf(
|
|
||||||
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
|
|
||||||
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
|
|
||||||
paste0("FY", s$suppressed_years, collapse = ", "),
|
|
||||||
paste(s$suppressed_codes, collapse = ", ")))
|
|
||||||
}
|
|
||||||
if (!is.null(s$ig_recipe_id)) {
|
if (!is.null(s$ig_recipe_id)) {
|
||||||
bullet <- paste0(bullet, sprintf(
|
bullet <- paste0(bullet, sprintf(
|
||||||
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
|
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
|
||||||
@@ -354,7 +245,7 @@
|
|||||||
bullet
|
bullet
|
||||||
}, character(1))
|
}, character(1))
|
||||||
cli::cli_inform(c(
|
cli::cli_inform(c(
|
||||||
i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
|
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
|
||||||
stats::setNames(bullets, rep("*", length(bullets)))
|
stats::setNames(bullets, rep("*", length(bullets)))
|
||||||
))
|
))
|
||||||
}
|
}
|
||||||
|
|||||||
-115
@@ -1,115 +0,0 @@
|
|||||||
# R/suppression.R
|
|
||||||
# Split out of R/suggestions.R (2026-08-05) to keep files under the project's
|
|
||||||
# 400-line limit. Owns the second qualifying path for coverage signposting
|
|
||||||
# (uscogdata#9): measuring, per government, the component dollars the
|
|
||||||
# calling verb's own long view structurally excludes (aggregate-published,
|
|
||||||
# or absent from summary_categories). See R/suggestions.R for the
|
|
||||||
# orchestrator (`.build_suggestions()`) that calls this and the full
|
|
||||||
# uscogdata#9 background.
|
|
||||||
|
|
||||||
#' Measure, per (recipe, year), the component dollars this government holds
|
|
||||||
#' that the calling verb's own long view structurally excludes.
|
|
||||||
#'
|
|
||||||
#' This is the second qualifying path for a suggestion (uscogdata#9). The
|
|
||||||
#' first -- row absence -- only fires when a category returns NOTHING in a
|
|
||||||
#' requested year, which is how Corrections behaves in the wide era. Public
|
|
||||||
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
|
|
||||||
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
|
|
||||||
#' 2011, and absent from `summary_categories` entirely) are dropped. The
|
|
||||||
#' caller gets a plausible number a third too low, silently.
|
|
||||||
#'
|
|
||||||
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
|
|
||||||
#' view rather than restating its WHERE clause, so this stays correct if
|
|
||||||
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
|
|
||||||
#' anti-join is keyed on `item_code`, which is sound only because
|
|
||||||
#' harmonization never renames a recipe component -- asserted by the "no
|
|
||||||
#' recipe component is ever renamed by harmonization" test in
|
|
||||||
#' tests/testthat/test-recipes.R.
|
|
||||||
#'
|
|
||||||
#' Note what this deliberately does NOT count as suppressed: a component
|
|
||||||
#' excluded from the RESULT for scoping reasons -- because it belongs to a
|
|
||||||
#' different `category`, or because `expenditure_concept` narrowed the
|
|
||||||
#' subtypes -- is still present in the view, so it never fires. Suggesting a
|
|
||||||
#' recipe is a coverage fix, not a category redefinition.
|
|
||||||
#'
|
|
||||||
#' `flow_prefixes` (uscogdata#9 review, finding I1) restricts the measured
|
|
||||||
#' components to the CALLING VERB's own flow family (`c("E","F","G")` for
|
|
||||||
#' spending, `c("T","A","U","B","C","D")` for revenue). Without this, a
|
|
||||||
#' candidate recipe belonging to the OTHER flow family is always absent from
|
|
||||||
#' this verb's view (by construction -- `cog_revenue()`'s view never carries
|
|
||||||
#' an E-coded row) and so was always reported as "suppressed", fabricating a
|
|
||||||
#' dollar claim across flow families (`cog_revenue(category = "Corrections")`
|
|
||||||
#' claimed $3.63B excluded that `cog_spending()` reports and fully accounts
|
|
||||||
#' for). Filtering on `LEFT(r.component_code, 1)` also drops M/L-prefixed
|
|
||||||
#' components from measurement under `cog_spending()` (`flow_prefixes` never
|
|
||||||
#' includes "M"/"L") -- harmless today, because a recipe's own M/L components
|
|
||||||
#' (e.g. `corrections_ig_local_combined`'s M04/M05) are present in the view
|
|
||||||
#' in every year they exist and so never fired as suppressed anyway, but
|
|
||||||
#' worth recording since this filter is now the thing relied on to prevent
|
|
||||||
#' it.
|
|
||||||
#'
|
|
||||||
#' @param con Active DuckDB connection.
|
|
||||||
#' @param candidates Character vector of recipe ids to measure.
|
|
||||||
#' @param govid Character vector of canonical_govid values.
|
|
||||||
#' @param years Integer vector of requested years.
|
|
||||||
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
|
|
||||||
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
|
|
||||||
#' `.build_suggestions()`). Only recipe components whose first character is
|
|
||||||
#' in this set are measured.
|
|
||||||
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
|
|
||||||
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
|
|
||||||
#' nothing is suppressed.
|
|
||||||
#' @noRd
|
|
||||||
.suppressed_components <- function(con, candidates, govid, years, long_view,
|
|
||||||
flow_prefixes) {
|
|
||||||
empty <- tibble::tibble(
|
|
||||||
recipe_id = character(0), year = numeric(0),
|
|
||||||
suppressed_amount = numeric(0), suppressed_codes = character(0)
|
|
||||||
)
|
|
||||||
if (length(candidates) == 0L) return(empty)
|
|
||||||
|
|
||||||
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
|
|
||||||
# never be quoted safely. It is always internally derived from a fixed
|
|
||||||
# view_base, so an off-allowlist value is a programming error, not input.
|
|
||||||
if (!long_view %in% c("spending_long", "spending_long_harmonized",
|
|
||||||
"revenue_long", "revenue_long_harmonized")) {
|
|
||||||
cli::cli_abort(
|
|
||||||
"Internal error: unexpected `long_view` {.val {long_view}}.",
|
|
||||||
class = "uscogdata_internal_error"
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
sql <- sprintf(
|
|
||||||
"SELECT r.recipe_id,
|
|
||||||
l.year,
|
|
||||||
SUM(l.amt) * 1000.0 AS suppressed_amount,
|
|
||||||
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
|
|
||||||
AS suppressed_codes
|
|
||||||
FROM long l
|
|
||||||
JOIN harmonization_recipes r
|
|
||||||
ON l.item_code = r.component_code
|
|
||||||
AND l.year BETWEEN r.year_min AND r.year_max
|
|
||||||
AND (r.gov_type_scope = 'all'
|
|
||||||
OR (r.gov_type_scope = 'state' AND l.type = 0)
|
|
||||||
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
|
|
||||||
WHERE r.recipe_id IN (%1$s)
|
|
||||||
AND l.canonical_govid IN (%2$s)
|
|
||||||
AND l.year IN (%3$s)
|
|
||||||
AND l.amt <> 0
|
|
||||||
AND LEFT(r.component_code, 1) IN (%5$s)
|
|
||||||
AND NOT EXISTS (
|
|
||||||
SELECT 1 FROM %4$s v
|
|
||||||
WHERE v.canonical_govid = l.canonical_govid
|
|
||||||
AND v.year = l.year
|
|
||||||
AND v.item_code = l.item_code
|
|
||||||
AND v.year IN (%3$s) -- restated: enables partition pruning (I3a)
|
|
||||||
AND v.canonical_govid IN (%2$s) -- restated: pushes the govid filter (I3a)
|
|
||||||
)
|
|
||||||
GROUP BY 1, 2
|
|
||||||
ORDER BY 1, 2",
|
|
||||||
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
|
||||||
paste(as.integer(years), collapse = ","), long_view,
|
|
||||||
.sql_lit_chr(flow_prefixes)
|
|
||||||
)
|
|
||||||
tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
|
||||||
}
|
|
||||||
Binary file not shown.
+3
-3
@@ -1,7 +1,7 @@
|
|||||||
{
|
{
|
||||||
"schema_version": 6,
|
"schema_version": 6,
|
||||||
"built_at": "2026-08-03T16:51:32Z",
|
"built_at": "2026-07-31T00:47:27Z",
|
||||||
"pipeline_commit": "e7394a4",
|
"pipeline_commit": "aadb46b",
|
||||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
|
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
|
||||||
"data_vintage": {
|
"data_vintage": {
|
||||||
"source_vintages": {
|
"source_vintages": {
|
||||||
@@ -105,7 +105,7 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"path": "data/series_breaks.parquet",
|
"path": "data/series_breaks.parquet",
|
||||||
"sha256": "731998516cd802f63fcf7fb66053c7a62b7be955ab0794cad4a4979cb7628b87",
|
"sha256": "06dcc995ff533e57cc65fa25086cc9bf83ba592c58bf7cc99269dc2576f69944",
|
||||||
"description": "series_breaks.parquet"
|
"description": "series_breaks.parquet"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -22,8 +22,8 @@
|
|||||||
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
|
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
|
||||||
},
|
},
|
||||||
"expenditure_concept_direct_suppressed": {
|
"expenditure_concept_direct_suppressed": {
|
||||||
"type": ["boolean", "null"],
|
"type": "boolean",
|
||||||
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. null (NA) when expenditure_concept = 'total' AND category = 'All Categories': the detector keys on per-category rows, which that mode collapses, so suppression cannot be computed -- see `expenditure_concept_note`. See the affected rows' `notes` for the recovering recipe, if any."
|
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
|
||||||
},
|
},
|
||||||
"revenue_concept": {
|
"revenue_concept": {
|
||||||
"type": "string",
|
"type": "string",
|
||||||
@@ -33,48 +33,7 @@
|
|||||||
},
|
},
|
||||||
"harmonization": { "type": "object" },
|
"harmonization": { "type": "object" },
|
||||||
"recipe": { "type": ["object", "null"] },
|
"recipe": { "type": ["object", "null"] },
|
||||||
"suggestions": {
|
"suggestions": { "type": "array" },
|
||||||
"type": "array",
|
|
||||||
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
|
|
||||||
"items": {
|
|
||||||
"type": "object",
|
|
||||||
"required": ["recipe_id", "label", "available_years", "hint", "ig_recipe_id",
|
|
||||||
"trigger", "suppressed_amount", "suppressed_years", "suppressed_codes"],
|
|
||||||
"properties": {
|
|
||||||
"recipe_id": { "type": "string" },
|
|
||||||
"label": { "type": "string" },
|
|
||||||
"available_years": {
|
|
||||||
"type": "array",
|
|
||||||
"items": { "type": "integer" },
|
|
||||||
"description": "[year_min, year_max] of the recipe's component coverage."
|
|
||||||
},
|
|
||||||
"hint": { "type": "string" },
|
|
||||||
"ig_recipe_id": {
|
|
||||||
"type": ["string", "null"],
|
|
||||||
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
|
|
||||||
},
|
|
||||||
"trigger": {
|
|
||||||
"type": "string",
|
|
||||||
"enum": ["empty_year", "suppressed_component"],
|
|
||||||
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars this government reports in the requested years that the verb's underlying long view structurally excludes -- aggregate-published, carrying no harmonized code, or absent from summary_categories. This is NOT the same thing as 'excluded from the result': a component present in the view under a different category (a scoping choice, e.g. a different `category` or a narrower `expenditure_concept`) contributes 0 and never fires. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way, using the same underlying-view measurement, and can be 0 even on an 'empty_year' fire."
|
|
||||||
},
|
|
||||||
"suppressed_amount": {
|
|
||||||
"type": "number",
|
|
||||||
"description": "Full US dollars this government reports, in the recipe's component codes, in the requested years, that the verb's underlying long view structurally excludes (aggregate-published, carrying no harmonized code, or absent from summary_categories) -- summed across those years. This is NOT the same quantity as 'what the result excludes': a component present in the view under a different category or a narrower `expenditure_concept` is scoped out on purpose, counts as 0 here, and is not suppression. 0 does not always mean full coverage -- see 'trigger' and 'empty_year'. May be negative where Census publishes a negative `amt` for the excluded rows."
|
|
||||||
},
|
|
||||||
"suppressed_years": {
|
|
||||||
"type": "array",
|
|
||||||
"items": { "type": "integer" },
|
|
||||||
"description": "The requested years contributing to suppressed_amount."
|
|
||||||
},
|
|
||||||
"suppressed_codes": {
|
|
||||||
"type": "array",
|
|
||||||
"items": { "type": "string" },
|
|
||||||
"description": "The excluded component item codes, sorted."
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
},
|
|
||||||
"scope": { "type": "object" },
|
"scope": { "type": "object" },
|
||||||
"codes_summed": { "type": "object" },
|
"codes_summed": { "type": "object" },
|
||||||
"aggregate_fallback": { "type": ["object", "null"] },
|
"aggregate_fallback": { "type": ["object", "null"] },
|
||||||
|
|||||||
+1
-6
@@ -28,12 +28,7 @@ argument: for holdings, `category` is a strict coarsening of
|
|||||||
every combination would be either redundant or empty.
|
every combination would be either redundant or empty.
|
||||||
`category = "Fund Balances"` is exactly the `general` family
|
`category = "Fund Balances"` is exactly the `general` family
|
||||||
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
||||||
one `dplyr::filter()` away. The reserved pseudo-category
|
one `dplyr::filter()` away.}
|
||||||
`"All Categories"` (see [cog_spending()]) is **not** supported here and
|
|
||||||
errors with class `uscogdata_all_categories_unsupported`: it sums a
|
|
||||||
concept's subtype scope, and holdings are a stock with no concept
|
|
||||||
vocabulary to sum across. Omit `category` to get every category broken
|
|
||||||
out instead.}
|
|
||||||
|
|
||||||
\item{per_capita}{Divide holdings by population. Note this is a **stock per
|
\item{per_capita}{Divide holdings by population. Note this is a **stock per
|
||||||
resident** (reserves per person), which is *not* comparable to
|
resident** (reserves per person), which is *not* comparable to
|
||||||
|
|||||||
@@ -16,10 +16,7 @@ balance), `"spending"`, `"revenue"`, or `"balance"`.}
|
|||||||
\value{
|
\value{
|
||||||
Tibble with columns `category`, `category_type`, `subtype`,
|
Tibble with columns `category`, `category_type`, `subtype`,
|
||||||
`n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
`n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
||||||
`category_type`, `category`, `subtype`. Includes one row per flow for the
|
`category_type`, `category`, `subtype`.
|
||||||
reserved pseudo-category `"All Categories"`, which carries `NA` for
|
|
||||||
`subtype`, `n_codes` and `item_codes` because it is a query mode rather
|
|
||||||
than a crosswalk entry — see [cog_spending()]'s `category` argument.
|
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
Returns the category taxonomy exposed by the corpus's
|
Returns the category taxonomy exposed by the corpus's
|
||||||
|
|||||||
@@ -20,11 +20,7 @@ cog_geographic_rollup(
|
|||||||
`canonical_govid` values. At least one layer required.}
|
`canonical_govid` values. At least one layer required.}
|
||||||
|
|
||||||
\item{category}{Single category name or character vector (passed through
|
\item{category}{Single category name or character vector (passed through
|
||||||
to [cog_spending()]), or the reserved `"All Categories"` for one summed
|
to [cog_spending()]).}
|
||||||
row per `(year, canonical_govid, subtype)` covering every category in the
|
|
||||||
concept's scope. `"All Categories"` is the efficient way to build a
|
|
||||||
geographic total: without it a caller must issue one rollup per category
|
|
||||||
and sum the results themselves.}
|
|
||||||
|
|
||||||
\item{years}{Integer vector of years.}
|
\item{years}{Integer vector of years.}
|
||||||
|
|
||||||
@@ -84,23 +80,3 @@ the result. The dropped govids are recorded in
|
|||||||
(gov type 4) and school districts (gov type 5) from per-capita rollups
|
(gov type 4) and school districts (gov type 5) from per-capita rollups
|
||||||
by design — see `vignette('population-denominators')`.
|
by design — see `vignette('population-denominators')`.
|
||||||
}
|
}
|
||||||
\section{Reading `coverage`}{
|
|
||||||
|
|
||||||
`provenance$coverage` reports `n_units_reporting` against
|
|
||||||
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
|
||||||
it counts governments with rows for the category you asked for, not
|
|
||||||
governments collected that year.** A government that was surveyed and
|
|
||||||
genuinely spends nothing in that category is indistinguishable here from one
|
|
||||||
that was never surveyed.
|
|
||||||
|
|
||||||
The ratio is therefore **not a response rate** and must not be used as one.
|
|
||||||
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
|
||||||
`category = "Police"`; the 174-city gap is overwhelmingly cities that
|
|
||||||
contract policing to the county sheriff, not non-response.
|
|
||||||
|
|
||||||
The comparison that *is* valid is the same category across a census year
|
|
||||||
(ending in 2 or 7) and a sample year, where the real-zero component is
|
|
||||||
roughly constant and the difference reflects the survey cycle. `is_census_year`
|
|
||||||
marks which is which.
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|||||||
@@ -107,23 +107,3 @@ call. Those summary rows are quantiles **within each category**, not
|
|||||||
quantiles of each peer's total — see the `@return` section before summing
|
quantiles of each peer's total — see the `@return` section before summing
|
||||||
them.
|
them.
|
||||||
}
|
}
|
||||||
\section{Reading `coverage`}{
|
|
||||||
|
|
||||||
`provenance$coverage` reports `n_units_reporting` against
|
|
||||||
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
|
||||||
it counts cohort members with rows for the category you asked for, not
|
|
||||||
cohort members collected that year.** A government that was surveyed and
|
|
||||||
genuinely spends nothing in that category is indistinguishable here from one
|
|
||||||
that was never surveyed.
|
|
||||||
|
|
||||||
The ratio is therefore **not a response rate** and must not be used as one.
|
|
||||||
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
|
||||||
`category = "Police"`; the 174-city gap is overwhelmingly cities that
|
|
||||||
contract policing to the county sheriff, not non-response.
|
|
||||||
|
|
||||||
The comparison that *is* valid is the same category across a census year
|
|
||||||
(ending in 2 or 7) and a sample year, where the real-zero component is
|
|
||||||
roughly constant and the difference reflects the survey cycle. `is_census_year`
|
|
||||||
marks which is which.
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|||||||
+2
-22
@@ -13,9 +13,7 @@ cog_revenue(
|
|||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL,
|
recipe = NULL,
|
||||||
revenue_concept = c("general", "total"),
|
revenue_concept = c("general", "total"),
|
||||||
complete = FALSE,
|
complete = FALSE
|
||||||
limit = NULL,
|
|
||||||
offset = NULL
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -24,16 +22,7 @@ cog_revenue(
|
|||||||
\item{years}{Integer vector of years.}
|
\item{years}{Integer vector of years.}
|
||||||
|
|
||||||
\item{category}{Character vector of category names (from
|
\item{category}{Character vector of category names (from
|
||||||
`summary_categories.category`), or `NULL` for all categories broken out
|
`summary_categories.category`), or `NULL` for all categories.}
|
||||||
one row each. The reserved value `"All Categories"` instead returns a
|
|
||||||
single summed row per `(year, canonical_govid, subtype)`, covering every
|
|
||||||
category inside the requested concept's subtype scope. It cannot be
|
|
||||||
combined with other category names, and it is not the same thing as
|
|
||||||
`revenue_concept = "total"`: the concept chooses which subtypes are in
|
|
||||||
scope, `"All Categories"` chooses whether rows inside that scope are
|
|
||||||
broken out or summed. Because the result keeps one row per
|
|
||||||
`revenue_subtype`, filtering the returned frame to
|
|
||||||
`revenue_subtype == "own_source"` gives an own-source revenue total.}
|
|
||||||
|
|
||||||
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||||
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||||
@@ -117,15 +106,6 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
|
|||||||
`recipe` or with `expenditure_concept = "total"` (class
|
`recipe` or with `expenditure_concept = "total"` (class
|
||||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
`code_set`.}
|
`code_set`.}
|
||||||
|
|
||||||
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
|
||||||
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
|
||||||
result is materialized. `NULL` (the default) returns every matching row,
|
|
||||||
exactly as before this parameter existed. Mutually exclusive with
|
|
||||||
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
|
|
||||||
|
|
||||||
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
|
||||||
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
|
|||||||
+4
-33
@@ -13,9 +13,7 @@ cog_spending(
|
|||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL,
|
recipe = NULL,
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
expenditure_concept = c("primary", "direct", "total"),
|
||||||
complete = FALSE,
|
complete = FALSE
|
||||||
limit = NULL,
|
|
||||||
offset = NULL
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -24,16 +22,7 @@ cog_spending(
|
|||||||
\item{years}{Integer vector of years.}
|
\item{years}{Integer vector of years.}
|
||||||
|
|
||||||
\item{category}{Character vector of category names (from
|
\item{category}{Character vector of category names (from
|
||||||
`summary_categories.category`), or `NULL` for all categories broken out
|
`summary_categories.category`), or `NULL` for all categories.}
|
||||||
one row each. The reserved value `"All Categories"` instead returns a
|
|
||||||
single summed row per `(year, canonical_govid, subtype)`, covering every
|
|
||||||
category inside the requested concept's subtype scope. It cannot be
|
|
||||||
combined with other category names, and it is not the same thing as
|
|
||||||
`expenditure_concept = "total"`: the concept chooses which subtypes are in
|
|
||||||
scope, `"All Categories"` chooses whether rows inside that scope are
|
|
||||||
broken out or summed. Because the result keeps one row per
|
|
||||||
`spend_subtype`, filtering the returned frame to
|
|
||||||
`spend_subtype == "operations"` gives an operating-expenditure total.}
|
|
||||||
|
|
||||||
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||||
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||||
@@ -108,12 +97,7 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
|
|||||||
component (when one exists), and
|
component (when one exists), and
|
||||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||||
IG. When `category = "All Categories"` is combined with
|
IG.}
|
||||||
`expenditure_concept = "total"`, this detection cannot run (it keys on
|
|
||||||
per-category rows, which all-categories mode collapses to one literal
|
|
||||||
value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
|
|
||||||
possibly-false `FALSE`; query an explicit `category` to get a real
|
|
||||||
answer.}
|
|
||||||
|
|
||||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||||
corpus does not carry still appears, labelled with **why** it is
|
corpus does not carry still appears, labelled with **why** it is
|
||||||
@@ -137,15 +121,6 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
|
|||||||
`recipe` or with `expenditure_concept = "total"` (class
|
`recipe` or with `expenditure_concept = "total"` (class
|
||||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||||
`code_set`.}
|
`code_set`.}
|
||||||
|
|
||||||
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
|
||||||
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
|
||||||
result is materialized. `NULL` (the default) returns every matching row,
|
|
||||||
exactly as before this parameter existed. Mutually exclusive with
|
|
||||||
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
|
|
||||||
|
|
||||||
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
|
||||||
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
@@ -155,11 +130,7 @@ Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
|||||||
and `value_source` when `complete = TRUE`.
|
and `value_source` when `complete = TRUE`.
|
||||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||||
whose `completion` block reports `applied`, `rows_filled`, and the
|
whose `completion` block reports `applied`, `rows_filled`, and the
|
||||||
per-year `absence_means` rule that was applied. When `limit` is set,
|
per-year `absence_means` rule that was applied.
|
||||||
also carries a `total_rows` attribute: the full unpaginated row count,
|
|
||||||
computed by the same query (`COUNT(*) OVER()`) rather than a second
|
|
||||||
round trip -- so a caller walking pages never has to ask "how many are
|
|
||||||
there" separately.
|
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,286 @@
|
|||||||
|
# `uscogdata` 0.1.0 — public release
|
||||||
|
|
||||||
|
**Date:** 2026-08-08 · **Status:** design, awaiting approval
|
||||||
|
**Scope:** release-readiness, README, NEWS. Distribution mechanics recorded here as
|
||||||
|
decided, sequenced after the package is clean.
|
||||||
|
|
||||||
|
`uscogdata` is feature-complete and the corpus it reads has been public on
|
||||||
|
HuggingFace since 2026-08-07 (294 downloads as of this writing). The API built on
|
||||||
|
it is live. What does not exist is a public *package*: the repo is private, there
|
||||||
|
is no install path, and — measured, not assumed — **a stranger who installed it
|
||||||
|
today could not read the corpus at all.**
|
||||||
|
|
||||||
|
This spec covers making that untrue.
|
||||||
|
|
||||||
|
## Decisions locked
|
||||||
|
|
||||||
|
| Decision | Choice |
|
||||||
|
|---|---|
|
||||||
|
| Canonical source | `gitea.civilytics.org/Civilytics/uscogdata`, flipped public |
|
||||||
|
| Public mirror | `github.com/civilytics/uscogdata` — issues, PRs, multi-OS check, CDN |
|
||||||
|
| Mirror mechanism | Gitea Actions non-force `git push` (not a push mirror) |
|
||||||
|
| Binaries | `civilytics.r-universe.dev`, registry pinned to a release tag |
|
||||||
|
| Author of record | Jared E. Knowles `<jared@civilytics.com>`, ORCID `0000-0003-0005-9478` |
|
||||||
|
| Copyright | Civilytics Consulting LLC (`cph`, `fnd`) |
|
||||||
|
| License | MIT (package) · CC-BY-4.0 (corpus) |
|
||||||
|
| Corrections intake | Deferred — see *Out of scope* |
|
||||||
|
| Other packages | Parked until this one walks the path end to end |
|
||||||
|
|
||||||
|
## P0 — the corpus is unreachable
|
||||||
|
|
||||||
|
Two independent faults, either of which alone is fatal.
|
||||||
|
|
||||||
|
**No corpus URL exists.** `R/config.R` defaults to the literal
|
||||||
|
`REPLACE_WITH_SHARE_TOKEN` sentinel, and no file in the repo supplies a working
|
||||||
|
one. A new user calling any verb gets `uscogdata_url_not_configured` with no path
|
||||||
|
to resolution.
|
||||||
|
|
||||||
|
**Remote reads are broken regardless.** Every partitioned view globs:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true)
|
||||||
|
```
|
||||||
|
|
||||||
|
DuckDB 1.5.5 refuses globs over generic HTTP. Its suggested
|
||||||
|
`allow_asterisks_in_http_paths` escape hatch does not help — it forwards the
|
||||||
|
literal `**/*` as a filename and 404s, because plain HTTP exposes no directory
|
||||||
|
listing to expand against.
|
||||||
|
|
||||||
|
The package therefore works only against a **local path**. That is how the API
|
||||||
|
runs it (`CORPUS_HOST_PATH` is a host mount on maxwell) and how the tests run
|
||||||
|
(bundled fixture), which is why the fault went unnoticed. The README's headline
|
||||||
|
claim — *"Reads the published corpus directly from Nextcloud via DuckDB httpfs —
|
||||||
|
no local bulk downloads required"* — is currently false.
|
||||||
|
|
||||||
|
### Fix: enumerate from the manifest, do not glob
|
||||||
|
|
||||||
|
`manifest.json` already lists every partition under `files.long_partitions[]`
|
||||||
|
with `path`, `year`, `sha256`, `row_count` and `size_bytes` — 56 of them.
|
||||||
|
Substituting an explicit file list for the glob was measured against the
|
||||||
|
published corpus on 2026-08-08:
|
||||||
|
|
||||||
|
| Path | Result |
|
||||||
|
|---|---|
|
||||||
|
| `https://…/data/long/**/*.parquet` (default) | error — globs unsupported over HTTP |
|
||||||
|
| same, `allow_asterisks_in_http_paths = true` | error — literal `**/*` 404s |
|
||||||
|
| `hf://datasets/civilytics/us-cog-finance/…` glob | 46,148,034 rows |
|
||||||
|
| **explicit list over plain https** | **46,148,034 rows** |
|
||||||
|
|
||||||
|
`hive_partitioning = true` still recovers `year` from the paths under
|
||||||
|
enumeration, so no downstream view or verb changes.
|
||||||
|
|
||||||
|
Enumeration is preferred over `hf://` deliberately. It is **host-agnostic** —
|
||||||
|
Nextcloud, HuggingFace, or any static server take the same code path — where
|
||||||
|
`hf://` would tie the default to one vendor's protocol and still need
|
||||||
|
special-casing, since manifest fetching goes through `httr2`, which cannot speak
|
||||||
|
`hf://`. Enumeration also *removes* a dependency (globbing) rather than adding
|
||||||
|
one, and the manifest's per-file `sha256` becomes available for integrity
|
||||||
|
checking later.
|
||||||
|
|
||||||
|
Views are registered from `inst/sql/` with `{url}` substitution in
|
||||||
|
`R/views.R:.register_views()`. The list must be built once per session from the
|
||||||
|
already-fetched manifest and substituted the same way, so the change is confined
|
||||||
|
to view registration and does not touch verb code.
|
||||||
|
|
||||||
|
### Fix: ship a working default
|
||||||
|
|
||||||
|
`R/config.R`'s default becomes the public HuggingFace `resolve/main/` URL:
|
||||||
|
CC-BY-4.0, no token to publish, CDN-backed, and it keeps maxwell's uplink out of
|
||||||
|
the path — the same reasoning behind the GitHub mirror and r-universe.
|
||||||
|
|
||||||
|
This means `library(uscogdata)` followed by a verb works with **zero
|
||||||
|
configuration**, which is what makes the package demonstrable in a README and
|
||||||
|
later in a post. `USCOGDATA_URL` and `options(uscogdata.url=)` continue to
|
||||||
|
override, so the Nextcloud copy and local mirrors are unaffected.
|
||||||
|
|
||||||
|
The `uscogdata_url_not_configured` error class stays — it still fires for an
|
||||||
|
explicitly-set empty or placeholder URL — but ceases to be the default
|
||||||
|
experience.
|
||||||
|
|
||||||
|
### Consequence: `cog_mirror()` is promoted
|
||||||
|
|
||||||
|
Measured cost of the remote default, from efron on a good connection:
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| Whole corpus | **190.6 MB**, 56 partitions, 46,148,034 rows, FY1967–FY2024 |
|
||||||
|
| One government, one year | 1.5 s |
|
||||||
|
| One government, all 56 years | 2.8 s |
|
||||||
|
| Disk written | **0.00 MB** — range requests only; `external_file_cache` is in-memory |
|
||||||
|
|
||||||
|
Nothing persists locally beyond the shared `httpfs` extension in `~/.duckdb` (a
|
||||||
|
few MB, once per machine, across all DuckDB use). Costs are RAM and per-query
|
||||||
|
bandwidth, since nothing caches between sessions.
|
||||||
|
|
||||||
|
Those timings are raw scans. Real verbs additionally join crosswalks, resolve
|
||||||
|
categories and assemble provenance, so end-to-end verb latency will be higher and
|
||||||
|
**must be re-measured once the fix lands** — it cannot be measured today.
|
||||||
|
|
||||||
|
The corpus being only 190.6 MB makes `cog_mirror()` a first-class option rather
|
||||||
|
than a developer footnote. The README presents **both paths**:
|
||||||
|
|
||||||
|
- **Remote (default, zero setup)** — trying it out, teaching, one-off questions.
|
||||||
|
- **Mirrored (`cog_mirror()`, 190 MB once)** — repeated or heavy analysis,
|
||||||
|
offline work, reproducibility, or preferring not to depend on HuggingFace.
|
||||||
|
|
||||||
|
The second is also the honest answer to the vendor-dependency question raised by
|
||||||
|
defaulting to HuggingFace: **the escape hatch is one function call and 190 MB**,
|
||||||
|
after which no analysis touches an external service. The README says so
|
||||||
|
explicitly. That is the difference between a convenience default and lock-in.
|
||||||
|
|
||||||
|
## Release-readiness fixes
|
||||||
|
|
||||||
|
| # | Issue | Fix |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | `MaxCorpusSchema: 5` in DESCRIPTION; `.validate_schema()` accepts `4,5,6,7`; published corpus is **7** | `MaxCorpusSchema: 7` |
|
||||||
|
| 2 | `^vignettes$` in `.Rbuildignore` — both vignettes absent from the installed package, while README tells users to run `vignette("total-spending")` | Remove `^vignettes$`, `^doc$`, `^Meta$`. Both vignettes build offline (`total-spending` reads the bundled fixture; `population-denominators` is `eval = FALSE`) |
|
||||||
|
| 3 | `_pkgdown.yml` reference index covers 6 of 14 exports — pkgdown errors on missing topics | Add `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`; set `url:` |
|
||||||
|
| 4 | No `URL:` / `BugReports:` in DESCRIPTION | Add both, pointing at the GitHub mirror |
|
||||||
|
| 5 | No `LICENSE.md`; `LICENSE` holder reads `Civilytics` | `usethis::use_mit_license("Civilytics Consulting LLC")` |
|
||||||
|
| 6 | README instructs stripping the fixture at release | Delete that section — see below |
|
||||||
|
| 7 | `Authors@R` is an org with no human | Jared E. Knowles `aut`/`cre` + ORCID; Civilytics Consulting LLC `cph`/`fnd` |
|
||||||
|
|
||||||
|
**On #6.** The advice to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`
|
||||||
|
is CRAN-sized thinking (5 MB limit) and this package is not going to CRAN.
|
||||||
|
Stripping the 15 MB fixture would break `total-spending.Rmd`, which reads from
|
||||||
|
it, and would leave r-universe and GitHub Actions unable to run the 28 test files
|
||||||
|
without a corpus credential. **The fixture is what lets `R CMD check` pass
|
||||||
|
anywhere with zero secrets** — precisely what public CI needs. It ships.
|
||||||
|
|
||||||
|
## README
|
||||||
|
|
||||||
|
The current README addresses someone standing inside the repo tree: status reads
|
||||||
|
"Under active development (Phase 2 of the cog_pipeline project)", it points at
|
||||||
|
`../cog_pipeline/docs/reader-specification.md`, the install line is commented
|
||||||
|
out, and developer, testing and release sections sit above anything a user needs.
|
||||||
|
|
||||||
|
Restructured around a stranger, in this order:
|
||||||
|
|
||||||
|
1. **What this is** — one paragraph, and what the corpus covers (types 0–3,
|
||||||
|
FY1967–FY2024, 46M rows, 190.6 MB).
|
||||||
|
2. **Install** — r-universe first (binaries), git second.
|
||||||
|
3. **Quickstart that actually runs** — resolve a government, get its history,
|
||||||
|
print provenance. No configuration step.
|
||||||
|
4. **Two ways to read the corpus** — remote default vs `cog_mirror()`, with the
|
||||||
|
measured numbers and the independence note.
|
||||||
|
5. **Amounts are in full US dollars** — kept near the top. This is the errata
|
||||||
|
most likely to produce a wrong answer that looks plausible.
|
||||||
|
6. **Concepts** — primary/direct/total spending, general/total revenue,
|
||||||
|
coverage. Condensed, linking to the vignettes for the full treatment.
|
||||||
|
7. **How to cite** — `citation("uscogdata")`, corpus CC-BY-4.0 attribution.
|
||||||
|
8. **Contributing** — canonical-on-Gitea PR flow.
|
||||||
|
|
||||||
|
Developer notes, testing instructions and release procedure move to
|
||||||
|
`CONTRIBUTING.md`. Every path reference to a sibling repo is removed or replaced
|
||||||
|
with a URL that resolves for someone who has only this repo.
|
||||||
|
|
||||||
|
## NEWS.md
|
||||||
|
|
||||||
|
The current NEWS is a pre-release churn log: changes described relative to states
|
||||||
|
no user has seen ("Breaking: corpus schema_version 4", "the package now
|
||||||
|
requires…"), newest-first across the package's entire pre-release development
|
||||||
|
(2026-04-23 to 2026-08-04, 140 commits). To a newcomer evaluating whether to
|
||||||
|
depend on the package, it reads as instability.
|
||||||
|
|
||||||
|
**0.1.0 is rewritten as an initial release**: what the package does, what the
|
||||||
|
corpus covers, and the caveats that are genuinely load-bearing. The pre-release
|
||||||
|
history is not preserved in NEWS — it is in git, where it belongs.
|
||||||
|
|
||||||
|
The substantive content is migrated, not deleted. These are hard-won and belong
|
||||||
|
in documentation rather than buried in a changelog:
|
||||||
|
|
||||||
|
| Content | Destination |
|
||||||
|
|---|---|
|
||||||
|
| Coverage disclosure on multi-government aggregates (census vs sample years) | README concepts + `cog_geographic_rollup()` docs |
|
||||||
|
| `complete = TRUE` three-way absence semantics (`reported` / `census_zero` / `not_reported`) | `cog_spending()` / `cog_revenue()` docs |
|
||||||
|
| Series-break and corpus-break surfacing | README + `cog_explain()` docs |
|
||||||
|
| $1,000s → full dollars conversion | README, already prominent |
|
||||||
|
| Per-year F-33 population denominators | `population-denominators` vignette, already there |
|
||||||
|
|
||||||
|
This also makes NEWS reusable as raw material for the release announcement,
|
||||||
|
which is the stated downstream purpose.
|
||||||
|
|
||||||
|
## Distribution mechanics
|
||||||
|
|
||||||
|
Recorded as decided; executed after the package is clean and checks are green.
|
||||||
|
|
||||||
|
**Sequence matters.** r-universe publishes check results the moment a package is
|
||||||
|
registered. Registering before the fixes above land means a red badge on day one,
|
||||||
|
which is a worse first impression than a week's delay.
|
||||||
|
|
||||||
|
1. `gitleaks` over full history. A coarse grep found nothing across 140 commits
|
||||||
|
and the default corpus URL is still the placeholder sentinel, but a proper
|
||||||
|
scan is the gate on an irreversible action.
|
||||||
|
2. Flip the Gitea repo public. Disable Gitea issues on it, so there is exactly
|
||||||
|
one inbox.
|
||||||
|
3. Create `github.com/civilytics/uscogdata`. Add `.github/workflows/` for the
|
||||||
|
Windows/macOS/Linux `R CMD check` matrix — the platforms the Gitea runner
|
||||||
|
cannot provide, and which this package has never been tested on despite
|
||||||
|
depending on duckdb and httr2. Gitea reads `.gitea/workflows`, GitHub reads
|
||||||
|
`.github/workflows`; both live in one tree without colliding.
|
||||||
|
4. Gitea Actions workflow pushing to GitHub **without `--force`**, so divergence
|
||||||
|
fails loudly in CI rather than silently overwriting.
|
||||||
|
5. Add `jared@civilytics.com` as a verified secondary email on the GitHub
|
||||||
|
account — r-universe links maintainer identity by matching DESCRIPTION's email
|
||||||
|
against registered GitHub emails, and the association only takes effect on the
|
||||||
|
next build.
|
||||||
|
6. Tag `v0.1.0`. Create `github.com/civilytics/civilytics.r-universe.dev` with a
|
||||||
|
`packages.json` pinned to the tag, pointing at the GitHub mirror rather than
|
||||||
|
Gitea so clone traffic stays off maxwell. Install the r-universe app.
|
||||||
|
|
||||||
|
### PR flow
|
||||||
|
|
||||||
|
Never press Merge on GitHub. A merge there is overwritten by the next sync, the
|
||||||
|
PR still displays "Merged", and nothing says otherwise.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git remote add github https://github.com/civilytics/uscogdata.git
|
||||||
|
git config --add remote.github.fetch '+refs/pull/*/head:refs/remotes/github/pr/*'
|
||||||
|
git fetch github
|
||||||
|
git switch -c pr-42 github/pr/42 # test
|
||||||
|
git switch main && git merge --no-ff pr-42
|
||||||
|
git push origin main # Gitea -> mirror -> GitHub
|
||||||
|
```
|
||||||
|
|
||||||
|
GitHub auto-closes a PR as merged once its head commit becomes an ancestor of the
|
||||||
|
base branch, so `--no-ff` — which preserves the contributor's SHAs — makes the PR
|
||||||
|
close itself when the mirror pushes. **For external PRs, merge; do not squash or
|
||||||
|
rebase.** Squashing rewrites the SHAs, the auto-close never fires, and closing by
|
||||||
|
hand reads to a first-time contributor as rejection.
|
||||||
|
|
||||||
|
`CONTRIBUTING.md` states this, and a GitHub Action comments it on incoming PRs.
|
||||||
|
No CLA; no DCO.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
The release is not done until all of these pass:
|
||||||
|
|
||||||
|
1. `R CMD check --as-cran` clean on Linux, and on Windows and macOS via the
|
||||||
|
GitHub matrix. This package has never been checked on the latter two.
|
||||||
|
2. Full test suite (28 files) green against the **bundled fixture**, offline,
|
||||||
|
with no credentials — the property public CI depends on.
|
||||||
|
3. Full test suite green against the **live corpus**, which additionally
|
||||||
|
exercises the enumeration fix that the fixture's local path cannot.
|
||||||
|
4. `pkgdown::build_site()` completes.
|
||||||
|
5. Both vignettes present in the built tarball and
|
||||||
|
`vignette("total-spending", package = "uscogdata")` resolves from an
|
||||||
|
installed copy.
|
||||||
|
6. **Cold-start check on a machine that has never seen this package:** install
|
||||||
|
from r-universe, `library(uscogdata)`, run the README quickstart verbatim with
|
||||||
|
no environment variables set. This is the only test that catches the P0 class
|
||||||
|
of fault, and its absence is why the fault survived.
|
||||||
|
7. End-to-end verb latency re-measured against the live corpus and the README's
|
||||||
|
numbers updated if they moved.
|
||||||
|
|
||||||
|
## Out of scope
|
||||||
|
|
||||||
|
- **Corrections intake.** Deferred by decision. Consequence: the release cannot
|
||||||
|
invite data-error reports or make the "traceable and correctable" claim that
|
||||||
|
most distinguishes this corpus from Census's own files. `BugReports:` points at
|
||||||
|
package issues only. A verified correction should eventually terminate as a
|
||||||
|
`lineage_event` or `series_break` row so it propagates through provenance to
|
||||||
|
every consumer — that design is unstarted.
|
||||||
|
- **Announcement posts.** Deferred. The API announcement is gated on corrections
|
||||||
|
landing and merits a Civic Pulse edition.
|
||||||
|
- **The rest of the R package backlog.** Parked until this one completes the path.
|
||||||
|
- **`cog_pipeline` publication.** Stays private.
|
||||||
@@ -1,58 +0,0 @@
|
|||||||
test_that('cog_geographic_rollup() accepts "All Categories" and agrees with per-category sums', {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
|
||||||
expect_gt(nrow(govs), 1L)
|
|
||||||
ids <- list(city = utils::head(govs$canonical_govid, 25L))
|
|
||||||
|
|
||||||
by_cat <- cog_geographic_rollup(ids, category = NULL, years = 2019L)
|
|
||||||
total <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L)
|
|
||||||
|
|
||||||
expect_setequal(unique(total$category), "All Categories")
|
|
||||||
# one row per (govid, subtype) that appears in the per-category result
|
|
||||||
key_by_cat <- unique(paste(by_cat$canonical_govid, by_cat$spend_subtype))
|
|
||||||
key_total <- paste(total$canonical_govid, total$spend_subtype)
|
|
||||||
expect_setequal(key_total, key_by_cat)
|
|
||||||
|
|
||||||
lhs <- tapply(by_cat$amt_nominal, paste(by_cat$canonical_govid, by_cat$spend_subtype), sum)
|
|
||||||
rhs <- tapply(total$amt_nominal, key_total, sum)
|
|
||||||
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" survives per_capita and inflation adjustment through the rollup', {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
|
||||||
ids <- list(city = utils::head(govs$canonical_govid, 10L))
|
|
||||||
r <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
|
|
||||||
per_capita = TRUE, adjust_to_year = 2020L)
|
|
||||||
expect_true(all(c("amt_per_capita_nominal", "amt_real", "amt_per_capita_real") %in% names(r)))
|
|
||||||
expect_setequal(unique(r$category), "All Categories")
|
|
||||||
expect_true(all(is.finite(r$amt_real)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_geographic_rollup() still refuses expenditure_concept = "total" with "All Categories"', {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
|
||||||
ids <- list(city = utils::head(govs$canonical_govid, 5L))
|
|
||||||
expect_error(
|
|
||||||
cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
|
|
||||||
expenditure_concept = "total")
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("n_units_reporting is category-conditional, not a response rate", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
|
||||||
ids <- list(city = govs$canonical_govid)
|
|
||||||
|
|
||||||
police <- cog_geographic_rollup(ids, category = "Police", years = 2012L)
|
|
||||||
allcat <- cog_geographic_rollup(ids, category = "All Categories", years = 2012L)
|
|
||||||
|
|
||||||
cov_police <- cog_explain(police, format = "list")$coverage
|
|
||||||
cov_all <- cog_explain(allcat, format = "list")$coverage
|
|
||||||
|
|
||||||
# Same year, same requested govids, same collection -- yet a single category
|
|
||||||
# reports fewer units than the all-categories query. That gap is real zeros,
|
|
||||||
# not non-response, which is exactly why the ratio is not a response rate.
|
|
||||||
expect_lte(cov_police$n_units_reporting, cov_all$n_units_reporting)
|
|
||||||
expect_identical(cov_police$n_units_expected, cov_all$n_units_expected)
|
|
||||||
})
|
|
||||||
@@ -1,267 +0,0 @@
|
|||||||
# Baseline at branch point: 843 PASS / 0 FAIL / 0 SKIP / 0 WARN (2026-08-05, origin/main 2fc9e75)
|
|
||||||
|
|
||||||
test_that(".build_verb_sql emits a literal category and no category filter in all-categories mode", {
|
|
||||||
sql <- uscogdata:::.build_verb_sql(
|
|
||||||
view = "spending_annotated",
|
|
||||||
subtype_col = "spend_subtype",
|
|
||||||
govid = "552025209777",
|
|
||||||
years = 2019L,
|
|
||||||
category = NULL,
|
|
||||||
subtype_scope = c("operations", "capital"),
|
|
||||||
all_categories = TRUE
|
|
||||||
)
|
|
||||||
|
|
||||||
expect_match(sql, "'All Categories' AS category", fixed = TRUE)
|
|
||||||
# no category filter of any kind
|
|
||||||
expect_false(grepl("AND category IN", sql, fixed = TRUE))
|
|
||||||
# category is not a grouping key
|
|
||||||
expect_false(grepl("GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
|
|
||||||
sql, fixed = TRUE))
|
|
||||||
# the subtype allowlist still applies -- this is what makes the sum a concept
|
|
||||||
expect_match(sql, "AND spend_subtype IN ('operations','capital')", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".build_verb_sql is unchanged when all_categories is FALSE", {
|
|
||||||
args <- list(
|
|
||||||
view = "spending_annotated", subtype_col = "spend_subtype",
|
|
||||||
govid = "552025209777", years = 2019L, category = NULL,
|
|
||||||
subtype_scope = c("operations", "capital")
|
|
||||||
)
|
|
||||||
old <- do.call(uscogdata:::.build_verb_sql, args)
|
|
||||||
new <- do.call(uscogdata:::.build_verb_sql, c(args, list(all_categories = FALSE)))
|
|
||||||
expect_identical(old, new)
|
|
||||||
expect_match(new, "GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
|
|
||||||
fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".ALL_CATEGORIES is the exact reserved string", {
|
|
||||||
expect_identical(uscogdata:::.ALL_CATEGORIES, "All Categories")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_spending(category = "All Categories") sums to the per-category total', {
|
|
||||||
gov <- "552025209777"
|
|
||||||
by_cat <- cog_spending(gov, 2019L)
|
|
||||||
total <- cog_spending(gov, 2019L, category = "All Categories")
|
|
||||||
|
|
||||||
expect_true(nrow(total) > 0L)
|
|
||||||
expect_setequal(unique(total$category), "All Categories")
|
|
||||||
# one row per subtype present in the by-category result
|
|
||||||
expect_setequal(unique(total$spend_subtype), unique(by_cat$spend_subtype))
|
|
||||||
expect_equal(nrow(total), length(unique(by_cat$spend_subtype)))
|
|
||||||
|
|
||||||
# the dollars agree, per subtype
|
|
||||||
lhs <- tapply(by_cat$amt_nominal, by_cat$spend_subtype, sum)
|
|
||||||
rhs <- tapply(total$amt_nominal, total$spend_subtype, sum)
|
|
||||||
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" respects expenditure_concept', {
|
|
||||||
gov <- "552025209777"
|
|
||||||
prim <- cog_spending(gov, 2019L, category = "All Categories",
|
|
||||||
expenditure_concept = "primary")
|
|
||||||
dir <- cog_spending(gov, 2019L, category = "All Categories",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
# direct = primary plus interest and insurance benefits, so it is never smaller
|
|
||||||
expect_gte(sum(dir$amt_nominal), sum(prim$amt_nominal))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" works on revenue and respects revenue_concept', {
|
|
||||||
gov <- "552025209777"
|
|
||||||
gen <- cog_revenue(gov, 2019L, category = "All Categories",
|
|
||||||
revenue_concept = "general")
|
|
||||||
tot <- cog_revenue(gov, 2019L, category = "All Categories",
|
|
||||||
revenue_concept = "total")
|
|
||||||
expect_setequal(unique(gen$category), "All Categories")
|
|
||||||
expect_gte(sum(tot$amt_nominal), sum(gen$amt_nominal))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" cannot be combined with another category', {
|
|
||||||
expect_error(
|
|
||||||
cog_spending("552025209777", 2019L, category = c("All Categories", "Police")),
|
|
||||||
class = "uscogdata_all_categories_not_combinable"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" is recorded in provenance', {
|
|
||||||
r <- cog_spending("552025209777", 2019L, category = "All Categories")
|
|
||||||
expect_identical(cog_explain(r, format = "list")$category, "All Categories")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" combines with subtype to give operating totals', {
|
|
||||||
gov <- "552025209777"
|
|
||||||
ops_by_cat <- cog_spending(gov, 2019L)
|
|
||||||
ops_by_cat <- ops_by_cat[ops_by_cat$spend_subtype == "operations", ]
|
|
||||||
ops_total <- cog_spending(gov, 2019L, category = "All Categories")
|
|
||||||
ops_total <- ops_total[ops_total$spend_subtype == "operations", ]
|
|
||||||
expect_equal(sum(ops_total$amt_nominal), sum(ops_by_cat$amt_nominal),
|
|
||||||
tolerance = 1e-8)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_categories() advertises "All Categories" for both flows', {
|
|
||||||
all <- cog_categories()
|
|
||||||
rows <- all[all$category == "All Categories", ]
|
|
||||||
expect_setequal(rows$category_type, c("expenditure", "revenue"))
|
|
||||||
expect_true(all(is.na(rows$subtype)))
|
|
||||||
expect_true(all(is.na(rows$n_codes)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_categories(type=) still scopes, including the pseudo-category', {
|
|
||||||
sp <- cog_categories(type = "spending")
|
|
||||||
expect_setequal(unique(sp$category_type), "expenditure")
|
|
||||||
expect_true("All Categories" %in% sp$category)
|
|
||||||
|
|
||||||
rev <- cog_categories(type = "revenue")
|
|
||||||
expect_setequal(unique(rev$category_type), "revenue")
|
|
||||||
expect_true("All Categories" %in% rev$category)
|
|
||||||
|
|
||||||
# balances have no concept vocabulary, so no pseudo-category
|
|
||||||
bal <- cog_categories(type = "balance")
|
|
||||||
expect_false("All Categories" %in% bal$category)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_categories(pattern=) matches the pseudo-category', {
|
|
||||||
hit <- cog_categories(pattern = "^All Categories$")
|
|
||||||
expect_equal(nrow(hit), 2L)
|
|
||||||
})
|
|
||||||
|
|
||||||
# --- final whole-branch review fixes ---------------------------------------
|
|
||||||
|
|
||||||
test_that('complete = TRUE is refused when combined with "All Categories"', {
|
|
||||||
# .completion_grid_sql() would emit `AND c.category IN ('All Categories')`,
|
|
||||||
# match zero crosswalk rows, and the early return in .complete_result()
|
|
||||||
# would stamp completion$applied = TRUE, rows_filled = 0 -- reading as "the
|
|
||||||
# grid was checked and nothing was missing" when nothing was actually
|
|
||||||
# checked. Filling a summed row has no defined semantics, so the verb must
|
|
||||||
# refuse the combination outright (finding 2).
|
|
||||||
expect_error(
|
|
||||||
cog_spending("552025209777", 2019L, category = "All Categories",
|
|
||||||
complete = TRUE),
|
|
||||||
class = "uscogdata_complete_unsupported"
|
|
||||||
)
|
|
||||||
expect_error(
|
|
||||||
cog_revenue("552025209777", 2019L, category = "All Categories",
|
|
||||||
complete = TRUE),
|
|
||||||
class = "uscogdata_complete_unsupported"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('cog_balances() rejects "All Categories" instead of silently returning zero rows', {
|
|
||||||
# cog_balances() reuses .validate_verb_inputs() but did not pass
|
|
||||||
# allow_all_categories = TRUE, so "All Categories" used to become
|
|
||||||
# `AND category IN ('All Categories')` against balance_annotated -- 0
|
|
||||||
# matching crosswalk rows, 0 rows back, no error (finding 3). Holdings are
|
|
||||||
# a stock with no concept vocabulary to sum across, so the honest answer is
|
|
||||||
# to refuse, the same way cog_spending()/cog_revenue() refuse other
|
|
||||||
# nonsensical combinations.
|
|
||||||
expect_error(
|
|
||||||
cog_balances("552025209777", 2019L, category = "All Categories"),
|
|
||||||
class = "uscogdata_all_categories_unsupported"
|
|
||||||
)
|
|
||||||
# An ordinary category still works -- this is not a blanket regression.
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_balances("552025209777", 2019L, category = "Fund Balances")
|
|
||||||
)
|
|
||||||
expect_gt(nrow(r), 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('expenditure_concept_direct_suppressed is NA, not FALSE, when categories are collapsed', {
|
|
||||||
# .detect_direct_suppressed() keys on
|
|
||||||
# paste(year, canonical_govid, category, sep = "\r"). In all-categories
|
|
||||||
# mode every row carries the literal "All Categories" value, so an IG-only
|
|
||||||
# row's key collides with any ordinary Direct row for the same
|
|
||||||
# (year, govid) -- has_direct reads TRUE whenever the government has ANY
|
|
||||||
# direct spending at all, candidate is always empty, and the detector can
|
|
||||||
# never fire. Before the fix this silently reported FALSE, an affirmative
|
|
||||||
# claim the code did not actually compute (finding 1). NA is the honest
|
|
||||||
# answer: cog_explain(x, format = "list") is required here, since without
|
|
||||||
# format = "list" it returns the result tibble, not the provenance list.
|
|
||||||
gov <- "552025209777"
|
|
||||||
t <- cog_spending(gov, 2019L, category = "All Categories",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
prov <- cog_explain(t, format = "list")
|
|
||||||
expect_true(is.na(prov$expenditure_concept_direct_suppressed))
|
|
||||||
expect_false(isTRUE(prov$expenditure_concept_direct_suppressed))
|
|
||||||
expect_match(prov$expenditure_concept_note, "unavailable", fixed = TRUE)
|
|
||||||
|
|
||||||
# A per-category "total" query on the same government/year is unaffected --
|
|
||||||
# the detector can still key correctly and reports a strict logical.
|
|
||||||
t_by_cat <- cog_spending(gov, 2019L, expenditure_concept = "total")
|
|
||||||
prov_by_cat <- cog_explain(t_by_cat, format = "list")
|
|
||||||
expect_false(is.na(prov_by_cat$expenditure_concept_direct_suppressed))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" still signposts coverage gaps (finding 6, final whole-branch review)', {
|
|
||||||
# .build_suggestions()'s candidate sub-select used to be keyed on
|
|
||||||
# `category`, e.g. `WHERE category IN ('All Categories')`. Since
|
|
||||||
# .ALL_CATEGORIES is never itself a row in summary_categories.category,
|
|
||||||
# that sub-select always came back empty in all-categories mode, so
|
|
||||||
# `candidates` was empty and .build_suggestions() short-circuited to
|
|
||||||
# list() -- coverage signposting was structurally impossible for the one
|
|
||||||
# mode whose whole selling point is "you cannot sum the wrong scope"
|
|
||||||
# (uscogdata#9's entire point, silently defeated).
|
|
||||||
#
|
|
||||||
# AL state government, FY2011, category = "Corrections": this category has
|
|
||||||
# no legacy leaf rows in FY2011 (aggregate-flagged E04/E05 family), so the
|
|
||||||
# per-category query returns 0 rows and 3 recipe-hint suggestions fire
|
|
||||||
# (empty_year path). All-categories mode does not have an empty year --
|
|
||||||
# the government has other primary spending in FY2011 -- but the same
|
|
||||||
# suppressed Corrections dollars are still excluded from the summed total,
|
|
||||||
# so the fix (scoping the candidate sub-select by subtype_col/subtype_scope
|
|
||||||
# instead of by category, symmetric with .build_verb_sql()) must still
|
|
||||||
# surface them via the suppressed_component path.
|
|
||||||
gov <- "010000226085"
|
|
||||||
|
|
||||||
by_cat <- suppressMessages(cog_spending(gov, 2011L, category = "Corrections"))
|
|
||||||
sugg_by_cat <- cog_explain(by_cat, format = "list")$suggestions
|
|
||||||
expect_gt(length(sugg_by_cat), 0L)
|
|
||||||
|
|
||||||
all_cat <- suppressMessages(cog_spending(gov, 2011L, category = "All Categories"))
|
|
||||||
sugg_all_cat <- cog_explain(all_cat, format = "list")$suggestions
|
|
||||||
expect_gt(length(sugg_all_cat), 0L)
|
|
||||||
|
|
||||||
# The same Corrections recipe that fired per-category must also fire in
|
|
||||||
# all-categories mode -- not just some unrelated recipe.
|
|
||||||
ids_by_cat <- vapply(sugg_by_cat, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
ids_all_cat <- vapply(sugg_all_cat, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
expect_true("corrections_combined" %in% ids_by_cat)
|
|
||||||
expect_true("corrections_combined" %in% ids_all_cat)
|
|
||||||
|
|
||||||
# In all-categories mode the government DOES have other primary spending
|
|
||||||
# in FY2011 (the year itself is not a gap), so the suggestion can only have
|
|
||||||
# fired via the suppressed_component path, not empty_year.
|
|
||||||
corr_all <- sugg_all_cat[[which(ids_all_cat == "corrections_combined")]]
|
|
||||||
expect_identical(corr_all$trigger, "suppressed_component")
|
|
||||||
expect_gt(corr_all$suppressed_amount, 0)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('"All Categories" candidate scoping is symmetric with .build_verb_sql() -- subtype, not category', {
|
|
||||||
# Direct assertion on the mechanism itself (finding 6): in all-categories
|
|
||||||
# mode .build_suggestions() must scope its candidate recipe sub-select by
|
|
||||||
# subtype_col/subtype_scope, not by the literal "All Categories" value.
|
|
||||||
# Passing all_categories = FALSE with the identical category value proves
|
|
||||||
# the branch -- not merely the subtype_col/subtype_scope arguments' mere
|
|
||||||
# presence -- is what changes the query.
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
|
|
||||||
none <- uscogdata:::.build_suggestions(
|
|
||||||
con, govid = "010000226085", years = 2011L,
|
|
||||||
category = "All Categories", result = NULL, basis = "harmonized",
|
|
||||||
flow_prefixes = c("E", "F", "G"),
|
|
||||||
long_view = "spending_long_harmonized",
|
|
||||||
all_categories = FALSE,
|
|
||||||
subtype_col = "spend_subtype",
|
|
||||||
subtype_scope = c("operations", "capital", "assistance")
|
|
||||||
)
|
|
||||||
expect_length(none, 0L)
|
|
||||||
|
|
||||||
scoped <- uscogdata:::.build_suggestions(
|
|
||||||
con, govid = "010000226085", years = 2011L,
|
|
||||||
category = "All Categories", result = NULL, basis = "harmonized",
|
|
||||||
flow_prefixes = c("E", "F", "G"),
|
|
||||||
long_view = "spending_long_harmonized",
|
|
||||||
all_categories = TRUE,
|
|
||||||
subtype_col = "spend_subtype",
|
|
||||||
subtype_scope = c("operations", "capital", "assistance")
|
|
||||||
)
|
|
||||||
expect_gt(length(scoped), 0L)
|
|
||||||
})
|
|
||||||
@@ -29,9 +29,7 @@ test_that("cog_categories(type = 'spending') returns only expenditure rows", {
|
|||||||
# joined with the I/Q/Y flow batch -- the last two characters of Census's
|
# joined with the I/Q/Y flow batch -- the last two characters of Census's
|
||||||
# expenditure taxonomy. `interest` is what makes the three-concept model
|
# expenditure taxonomy. `interest` is what makes the three-concept model
|
||||||
# computable: primary = direct minus debt service.
|
# computable: primary = direct minus debt service.
|
||||||
# Exclude pseudo-category which has NA for subtype
|
expect_true(all(r$subtype %in%
|
||||||
r_crosswalk <- r[r$category != "All Categories", ]
|
|
||||||
expect_true(all(r_crosswalk$subtype %in%
|
|
||||||
c("operations", "capital", "intergovernmental", "assistance",
|
c("operations", "capital", "intergovernmental", "assistance",
|
||||||
"interest", "insurance_benefits")))
|
"interest", "insurance_benefits")))
|
||||||
})
|
})
|
||||||
@@ -56,9 +54,7 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
|
|||||||
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
|
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
|
||||||
# (A90) revenue by definition, which is what makes both of its published
|
# (A90) revenue by definition, which is what makes both of its published
|
||||||
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
|
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
|
||||||
# Exclude pseudo-category which has NA for subtype
|
expect_true(all(r$subtype %in%
|
||||||
r_crosswalk <- r[r$category != "All Categories", ]
|
|
||||||
expect_true(all(r_crosswalk$subtype %in%
|
|
||||||
c("own_source", "federal", "state", "local_aid",
|
c("own_source", "federal", "state", "local_aid",
|
||||||
"insurance_trust", "utility", "liquor_store")))
|
"insurance_trust", "utility", "liquor_store")))
|
||||||
})
|
})
|
||||||
@@ -73,8 +69,6 @@ test_that("cog_categories(pattern = ...) filters case-insensitively", {
|
|||||||
test_that("cog_categories has one row per (category, subtype)", {
|
test_that("cog_categories has one row per (category, subtype)", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
r <- cog_categories()
|
r <- cog_categories()
|
||||||
# Exclude pseudo-category which is not a crosswalk entry
|
|
||||||
r <- r[r$category != "All Categories", ]
|
|
||||||
key <- paste(r$category, r$subtype, sep = "|")
|
key <- paste(r$category, r$subtype, sep = "|")
|
||||||
expect_equal(length(key), length(unique(key)))
|
expect_equal(length(key), length(unique(key)))
|
||||||
})
|
})
|
||||||
@@ -82,8 +76,6 @@ test_that("cog_categories has one row per (category, subtype)", {
|
|||||||
test_that("cog_categories item_codes is non-empty comma-separated string", {
|
test_that("cog_categories item_codes is non-empty comma-separated string", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
r <- cog_categories()
|
r <- cog_categories()
|
||||||
# Exclude pseudo-category which has NA for n_codes and item_codes
|
|
||||||
r <- r[r$category != "All Categories", ]
|
|
||||||
expect_true(all(nzchar(r$item_codes)))
|
expect_true(all(nzchar(r$item_codes)))
|
||||||
expect_true(all(r$n_codes >= 1L))
|
expect_true(all(r$n_codes >= 1L))
|
||||||
# n_codes should equal count of commas + 1
|
# n_codes should equal count of commas + 1
|
||||||
|
|||||||
@@ -214,255 +214,3 @@ test_that("no signposting under basis = 'raw'", {
|
|||||||
prov <- attr(r, "provenance")
|
prov <- attr(r, "provenance")
|
||||||
expect_length(prov$suggestions, 0L)
|
expect_length(prov$suggestions, 0L)
|
||||||
})
|
})
|
||||||
|
|
||||||
# --- uscogdata#9: partial-coverage signposting ------------------------------
|
|
||||||
|
|
||||||
test_that("no recipe component is ever renamed by harmonization", {
|
|
||||||
# The suppression trigger anti-joins the verb's long view on item_code.
|
|
||||||
# That is only sound because harmonization never rewrites a recipe
|
|
||||||
# component's code -- every component whose harmonized_code differs has
|
|
||||||
# harmonized_code IS NULL (and is aggregate-flagged). If this ever fails,
|
|
||||||
# .suppressed_components() would report reachable dollars as suppressed.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS renamed FROM long
|
|
||||||
WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes)
|
|
||||||
AND harmonized_code IS NOT NULL
|
|
||||||
AND harmonized_code <> item_code")$renamed
|
|
||||||
expect_equal(as.integer(n), 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".select_long_view maps annotated view bases to their long views", {
|
|
||||||
expect_equal(
|
|
||||||
uscogdata:::.select_long_view("spending_annotated", "harmonized"),
|
|
||||||
"spending_long_harmonized")
|
|
||||||
expect_equal(
|
|
||||||
uscogdata:::.select_long_view("revenue_annotated", "harmonized"),
|
|
||||||
"revenue_long_harmonized")
|
|
||||||
expect_equal(
|
|
||||||
uscogdata:::.select_long_view("spending_annotated", "raw"),
|
|
||||||
"spending_long")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
s <- uscogdata:::.suppressed_components(
|
|
||||||
con,
|
|
||||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
|
||||||
govid = "061037123085", years = 2011L,
|
|
||||||
long_view = "spending_long_harmonized",
|
|
||||||
flow_prefixes = c("E", "F", "G"))
|
|
||||||
|
|
||||||
expect_s3_class(s, "tbl_df")
|
|
||||||
expect_equal(nrow(s), 2L)
|
|
||||||
s <- s[order(s$recipe_id), ]
|
|
||||||
expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
|
|
||||||
expect_equal(s$suppressed_amount, c(1803872000, 271589000))
|
|
||||||
expect_equal(s$suppressed_codes, c("E67", "E68"))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".suppressed_components finds nothing in a modern year", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
s <- uscogdata:::.suppressed_components(
|
|
||||||
con,
|
|
||||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
|
||||||
govid = "061037123085", years = 2019L,
|
|
||||||
long_view = "spending_long_harmonized",
|
|
||||||
flow_prefixes = c("E", "F", "G"))
|
|
||||||
expect_equal(nrow(s), 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".suppressed_components rejects a long_view outside the allowlist", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
expect_error(
|
|
||||||
uscogdata:::.suppressed_components(
|
|
||||||
con, candidates = "welfare_cash_e67_wide", govid = "061037123085",
|
|
||||||
years = 2011L, long_view = "long; DROP TABLE x",
|
|
||||||
flow_prefixes = c("E", "F", "G")),
|
|
||||||
class = "uscogdata_internal_error")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".suppressed_components never measures a component from the other flow family (I1)", {
|
|
||||||
# uscogdata#9 review, finding I1: without the flow_prefixes filter, a
|
|
||||||
# candidate recipe entirely outside the calling verb's own flow family is
|
|
||||||
# ALWAYS absent from that verb's view (by construction), so it was always
|
|
||||||
# reported as "suppressed" -- fabricating a dollar claim. E67/E68 are
|
|
||||||
# Public Welfare EXPENDITURE codes; scoping the measurement to revenue's
|
|
||||||
# own flow_prefixes must find nothing for them.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
s <- uscogdata:::.suppressed_components(
|
|
||||||
con,
|
|
||||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
|
||||||
govid = "061037123085", years = 2011L,
|
|
||||||
long_view = "revenue_long_harmonized",
|
|
||||||
flow_prefixes = c("T", "A", "U", "B", "C", "D"))
|
|
||||||
expect_equal(nrow(s), 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", {
|
|
||||||
# The bug: E74/E79 return rows for FY2011, so there is no row-absence gap,
|
|
||||||
# so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were
|
|
||||||
# dropped for being aggregate-published. LA County reports $3,185,943,000
|
|
||||||
# and omits $2,075,461,000, a 39% understatement, silently.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
|
|
||||||
expect_length(sugg, 2L)
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
|
||||||
expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
|
|
||||||
|
|
||||||
e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]]
|
|
||||||
expect_equal(e67$trigger, "suppressed_component")
|
|
||||||
expect_equal(e67$suppressed_amount, 1803872000)
|
|
||||||
expect_equal(e67$suppressed_years, 2011L)
|
|
||||||
expect_equal(e67$suppressed_codes, "E67")
|
|
||||||
expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'")
|
|
||||||
|
|
||||||
e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]]
|
|
||||||
expect_equal(e68$trigger, "suppressed_component")
|
|
||||||
expect_equal(e68$suppressed_amount, 271589000)
|
|
||||||
expect_equal(e68$suppressed_codes, "E68")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", {
|
|
||||||
# Corrections is the case that already worked: zero rows in FY2011, so the
|
|
||||||
# row-absence path fires. It must keep firing, keep trigger = "empty_year",
|
|
||||||
# keep its IG counterpart -- and now also report what was suppressed.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Corrections"))
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
|
|
||||||
expect_length(sugg, 3L)
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
|
||||||
expect_setequal(ids, c("corrections_combined", "corrections_capital_combined",
|
|
||||||
"corrections_other_capital_combined"))
|
|
||||||
expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year"))
|
|
||||||
|
|
||||||
cc <- sugg[[which(ids == "corrections_combined")]]
|
|
||||||
expect_equal(cc$suppressed_amount, 1371460000)
|
|
||||||
expect_equal(cc$suppressed_codes, "E05")
|
|
||||||
expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: the revenue verb inherits the same trigger", {
|
|
||||||
# Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from
|
|
||||||
# U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-`
|
|
||||||
# rents and royalties -- the omission is LARGER than the reported figure.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_revenue("020000227749", years = 2011L,
|
|
||||||
category = "Miscellaneous Revenue"))
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
|
|
||||||
expect_length(sugg, 1L)
|
|
||||||
expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide")
|
|
||||||
expect_equal(sugg[[1]]$trigger, "suppressed_component")
|
|
||||||
expect_equal(sugg[[1]]$suppressed_amount, 1899995000)
|
|
||||||
expect_equal(sugg[[1]]$suppressed_codes, "U4-")
|
|
||||||
# A revenue recipe must never be handed an M/L expenditure counterpart.
|
|
||||||
expect_null(sugg[[1]]$ig_recipe_id)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("I1: cog_revenue never fabricates suppressed dollars for an expenditure-only recipe", {
|
|
||||||
# uscogdata#9 review, finding I1: Corrections is an expenditure-only
|
|
||||||
# category (E04/E05). cog_revenue() naturally returns zero rows for it, so
|
|
||||||
# corrections_combined still fires as an empty_year suggestion (its own
|
|
||||||
# generic join finds real E04/E05 data for this government) -- but before
|
|
||||||
# the flow_prefixes fix, .suppressed_components() measured E04/E05 against
|
|
||||||
# cog_revenue()'s OWN view (which can never contain an E-coded row by
|
|
||||||
# construction) and reported the full $3,631,945,000 as "suppressed",
|
|
||||||
# when cog_spending() for the same gov/years/category actually returns
|
|
||||||
# $3,691,029,000 -- nothing was suppressed at all.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_revenue("061037123085", years = 2019:2020, category = "Corrections"))
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
|
||||||
expect_true("corrections_combined" %in% ids)
|
|
||||||
|
|
||||||
hit <- sugg[[which(ids == "corrections_combined")]]
|
|
||||||
expect_equal(hit$suppressed_amount, 0)
|
|
||||||
expect_equal(hit$suppressed_years, integer(0))
|
|
||||||
expect_equal(hit$suppressed_codes, character(0))
|
|
||||||
|
|
||||||
# And cog_spending() for the identical gov/years/category is unaffected --
|
|
||||||
# it actually finds the E04/E05 dollars the buggy measurement claimed were
|
|
||||||
# excluded.
|
|
||||||
sp <- suppressMessages(
|
|
||||||
cog_spending("061037123085", years = 2019:2020, category = "Corrections"))
|
|
||||||
expect_equal(sum(sp$amt_nominal), 3691029000)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: no partial-coverage fire in a modern year", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare")
|
|
||||||
expect_length(attr(r, "provenance")$suggestions, 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: leaf-and-classified wide-era families never fire", {
|
|
||||||
# higher_ed_e18_wide and general_gov_e89_wide are the control group: their
|
|
||||||
# components (E16/E18, E85/E89) are ordinary classified leaves even in the
|
|
||||||
# wide era, so widening the trigger must leave them silent. This is the
|
|
||||||
# measurement that refutes "it would fire on every category in every legacy
|
|
||||||
# year" -- corpus-wide on the fixture, these two produce zero suppressed rows.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n
|
|
||||||
FROM long l
|
|
||||||
JOIN harmonization_recipes r
|
|
||||||
ON l.item_code = r.component_code
|
|
||||||
AND l.year BETWEEN r.year_min AND r.year_max
|
|
||||||
WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide')
|
|
||||||
AND l.amt <> 0
|
|
||||||
AND NOT EXISTS (
|
|
||||||
SELECT 1 FROM spending_long_harmonized v
|
|
||||||
WHERE v.canonical_govid = l.canonical_govid
|
|
||||||
AND v.year = l.year AND v.item_code = l.item_code)")$n
|
|
||||||
expect_equal(as.integer(n), 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: the cli message reports the suppressed dollars", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
expect_message(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
|
||||||
"1,803,872,000", fixed = TRUE)
|
|
||||||
expect_message(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
|
||||||
"FY2011", fixed = TRUE)
|
|
||||||
expect_message(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
|
||||||
"E67", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("uscogdata#9: cog_explain() reports the suppressed dollars", {
|
|
||||||
# cog_explain()'s whole "print" output -- including the Suggestions
|
|
||||||
# section built from cli::cli_ul() -- is emitted on the message stream
|
|
||||||
# (verified empirically 2026-08-04: capture.output(..., type = "output")
|
|
||||||
# returns character(0) for this call; testthat::capture_messages() is what
|
|
||||||
# actually carries it), so that is the stream this test captures.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
|
|
||||||
out <- paste(testthat::capture_messages(cog_explain(r)), collapse = "")
|
|
||||||
expect_match(out, "271,589,000", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the provenance schema documents the suggestion trigger fields", {
|
|
||||||
sch <- jsonlite::fromJSON(
|
|
||||||
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
|
|
||||||
simplifyVector = FALSE)
|
|
||||||
props <- sch$properties$suggestions$items$properties
|
|
||||||
expect_true(all(c("trigger", "suppressed_amount", "suppressed_years",
|
|
||||||
"suppressed_codes") %in% names(props)))
|
|
||||||
expect_setequal(unlist(props$trigger$enum),
|
|
||||||
c("empty_year", "suppressed_component"))
|
|
||||||
})
|
|
||||||
|
|||||||
@@ -1,29 +0,0 @@
|
|||||||
# Mirror of test-spending-pagination.R for cog_revenue(), which shares the
|
|
||||||
# same .verb_spendrev()/.build_verb_sql() pushdown -- see that file for the
|
|
||||||
# incident this fixes.
|
|
||||||
|
|
||||||
test_that("cog_revenue limit/offset page correctly and report total_rows", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_revenue("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
page <- cog_revenue("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = 5L, offset = 3L)
|
|
||||||
expect_equal(nrow(page), 5L)
|
|
||||||
expect_equal(page[c("year", "canonical_govid", "revenue_subtype", "category")],
|
|
||||||
full[4:8, c("year", "canonical_govid", "revenue_subtype", "category")],
|
|
||||||
ignore_attr = TRUE)
|
|
||||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_revenue limit unset by default leaves total_rows absent", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- cog_revenue("121011212191", 2020L, "Property Tax")
|
|
||||||
expect_null(attr(r, "total_rows"))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_revenue complete + limit conflict aborts the same way as cog_spending", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
expect_error(
|
|
||||||
cog_revenue("121011212191", 2020L, "Property Tax", complete = TRUE, limit = 5L),
|
|
||||||
class = "uscogdata_complete_pagination_conflict"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
@@ -1,95 +0,0 @@
|
|||||||
# cog-api's paginate() used to slice an already-fully-materialized result:
|
|
||||||
# every page of a deep sweep re-ran the whole query and re-listified every
|
|
||||||
# row, just to keep 1000 and discard the rest. For a 193,105-row fleet-wide
|
|
||||||
# query walked 194 pages deep, that repeated the full cost 194 times and
|
|
||||||
# wedged the production server for hours (2026-08-06 incident). limit/offset
|
|
||||||
# here push the slice into the SQL itself, so a page costs O(limit), not
|
|
||||||
# O(full result).
|
|
||||||
|
|
||||||
test_that("limit without offset returns the first page, matching the unpaginated head", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = 10L)
|
|
||||||
expect_equal(nrow(page), 10L)
|
|
||||||
expect_equal(page[c("year", "canonical_govid", "spend_subtype", "category")],
|
|
||||||
full[1:10, c("year", "canonical_govid", "spend_subtype", "category")],
|
|
||||||
ignore_attr = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("offset skips ahead without gaps or overlap", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
page2 <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = 10L, offset = 10L)
|
|
||||||
expect_equal(nrow(page2), 10L)
|
|
||||||
expect_equal(page2[c("year", "canonical_govid", "spend_subtype", "category")],
|
|
||||||
full[11:20, c("year", "canonical_govid", "spend_subtype", "category")],
|
|
||||||
ignore_attr = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("walking every page reconstructs the unpaginated result exactly", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
n <- nrow(full)
|
|
||||||
limit <- 7L
|
|
||||||
pages <- list()
|
|
||||||
offset <- 0L
|
|
||||||
repeat {
|
|
||||||
p <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = limit, offset = offset)
|
|
||||||
if (nrow(p) == 0L) break
|
|
||||||
pages[[length(pages) + 1L]] <- p
|
|
||||||
offset <- offset + limit
|
|
||||||
if (offset > n + limit) stop("test runaway: paging did not terminate")
|
|
||||||
}
|
|
||||||
walked <- dplyr::bind_rows(pages)
|
|
||||||
expect_equal(nrow(walked), n)
|
|
||||||
key_cols <- c("year", "canonical_govid", "spend_subtype", "category", "amt_nominal")
|
|
||||||
expect_equal(walked[key_cols], full[key_cols], ignore_attr = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("total_rows attribute reports the full unpaginated count", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = 5L, offset = 0L)
|
|
||||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("offset past the end returns zero rows, not an error", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
|
||||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
|
||||||
limit = 10L, offset = nrow(full) + 100L)
|
|
||||||
expect_equal(nrow(page), 0L)
|
|
||||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("limit is unset by default -- unpaginated calls are unaffected", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
|
||||||
expect_null(attr(r, "total_rows"))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("per_capita and adjust_to_year still apply correctly within a page", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
full <- cog_spending("121011212191", years = 2020L, category = NULL,
|
|
||||||
per_capita = TRUE, adjust_to_year = 2022L)
|
|
||||||
page <- cog_spending("121011212191", years = 2020L, category = NULL,
|
|
||||||
per_capita = TRUE, adjust_to_year = 2022L,
|
|
||||||
limit = 3L, offset = 2L)
|
|
||||||
expect_equal(page[c("amt_nominal", "amt_real", "amt_per_capita_nominal",
|
|
||||||
"amt_per_capita_real")],
|
|
||||||
full[3:5, c("amt_nominal", "amt_real", "amt_per_capita_nominal",
|
|
||||||
"amt_per_capita_real")],
|
|
||||||
ignore_attr = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE with limit aborts -- pagination over a partial grid is undefined", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
expect_error(
|
|
||||||
cog_spending("121011212191", 2020L, "Corrections", complete = TRUE, limit = 5L),
|
|
||||||
class = "uscogdata_complete_pagination_conflict"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
Reference in New Issue
Block a user