Compare commits

..
Author SHA1 Message Date
jared 785f3af16d chore: regenerate fixture against pipeline e7394a4 (SB203-SB209)
R-CMD-check / check (pull_request) Successful in 3m50s
R-CMD-check / check (push) Successful in 4m9s
Picks up seven new catalogued series breaks, all break_year 2017,
recording that the employee-retirement X-codes (X21, X30, X42, X44, X47,
Z77, Z78) were last collected in the annual finance file at FY2016 before
those systems moved to the Annual Survey of Public Pensions.

series_breaks.parquet 202 -> 209 rows; nothing removed. manifest built_at
and pipeline_commit updated, and the series_breaks sha256 with them.
2026-08-08 17:02:14 -04:00
jared c587c8ba87 Merge pull request 'fix: push limit/offset into the query instead of materializing then slicing' (#39) from fix/pushdown-pagination into main
R-CMD-check / check (push) Successful in 3m35s
Reviewed-on: #39
2026-08-06 14:19:42 -04:00
jared 6a06302036 fix: push limit/offset into the query instead of materializing then slicing
R-CMD-check / check (push) Successful in 4m6s
R-CMD-check / check (pull_request) Successful in 3m33s
cog-api's paginate() sliced an already-fully-materialized result: every
page of a deep sweep re-ran the whole cog_spending()/cog_revenue() query
and re-listified every row, just to keep up to 1000 and discard the
rest. A 193,105-row/194-page fleet-wide sweep (cog_explorer's Southern
guide, corpus summary build) repeated that full cost 194 times and
wedged the production server for hours on 2026-08-06 -- single request,
CPU-bound, single-threaded plumber process, no other request could get
through, not even /health.

cog_spending()/cog_revenue() gain optional limit/offset, pushed into
.build_verb_sql() as SQL LIMIT/OFFSET behind the existing (already
deterministic) ORDER BY. The full unpaginated row count rides along via
COUNT(*) OVER() in the same scan -- exposed as a total_rows attribute --
so a caller walking pages never needs a second round trip to ask how
many there are. A page now costs O(limit), not O(full result).

Mutually exclusive with complete = TRUE (which fills a grid over the
FULL requested (year, category) space -- pagination over a partial slice
of already-grouped rows has no defined meaning for the cells it would
fill) and with recipe (whose result comes from a separate, not-yet-wired
query path). Both abort with a clear classed condition rather than
silently ignoring the parameter.

limit/offset default to NULL; every existing call site is unaffected.
2026-08-06 13:56:27 -04:00
jared e3ab26c3e6 Merge pull request 'feat: 'All Categories' pseudo-category + n_units_reporting semantics' (#37) from feat/all-categories-37 into main
R-CMD-check / check (push) Successful in 3m50s
Reviewed-on: #37
2026-08-05 12:50:46 -04:00
jared 498950afa6 fix: scope all-categories suggestion candidates by subtype, not category (finding 6)
R-CMD-check / check (push) Successful in 3m41s
R-CMD-check / check (pull_request) Successful in 4m52s
.build_suggestions()'s recipe-candidate sub-select was keyed on
`WHERE category IN (<category>)`. The reserved pseudo-category
"All Categories" is never itself a row in summary_categories.category, so
in all-categories mode `candidates` always came back empty and coverage
signposting (uscogdata#9) was structurally impossible for the one mode
whose entire premise is "you cannot sum the wrong scope" -- measured on
Los Angeles County FY2011: category = "Public Welfare" reports 2
suggestions (incl. $271,589,000 excluded E68), category = "All Categories"
reported 0, silently losing that same signal.

Apply the branch's own design principle: the concept boundary is subtype,
not category. .build_suggestions() now accepts all_categories/subtype_col/
subtype_scope (all optional, default off, so no other caller's behaviour
changes) and, when all-categories mode is active, scopes the candidate
sub-select by `<subtype_col> IN (<subtype_scope>)` instead -- symmetric
with .build_verb_sql()'s own WHERE predicate. The M/L recipe exclusion and
the is.null(category) early return are unchanged.

After the fix, LA County FY2011 "All Categories" reports 5 suggestions,
including welfare_cash_e68_wide for the exact $271,589,000 gap.

Adds two covering tests to test-all-categories.R using the bundled fixture
(AL state gov, FY2011, "Corrections"): one end-to-end (per-category and
all-categories both signpost the same recipe) and one direct on
.build_suggestions() proving the subtype-vs-category branch is what
changes the query. Updates the 0.2.0 NEWS entry.
2026-08-05 12:38:12 -04:00
jared a5f86d87b3 fix: close five final-review gaps in all-categories mode
- .detect_direct_suppressed() keys on (year, canonical_govid, category);
  all-categories mode collapses category to one literal value, so the key
  collides and the detector silently reports FALSE instead of "unknown".
  Report NA there instead, and stop isTRUE() in .build_provenance() from
  collapsing that NA back to FALSE. Schema widened to allow null.
- Refuse complete = TRUE + category = "All Categories": the completion grid
  has no per-category cells left to fill once categories are collapsed,
  so the prior silent 0-rows-filled result was never actually checked.
- cog_balances(category = "All Categories") returned zero rows with no
  error. .validate_verb_inputs() gains allow_all_categories (default
  FALSE); .verb_spendrev() passes TRUE, cog_balances() does not, so the
  three verbs share one place to reject it instead of drifting again.
- Fix the false `subtype = "operations"` argument claim (no such argument
  exists) in NEWS.md and an internal spending.R comment.

Adds three covering tests to test-all-categories.R for the three
behaviour changes above.
2026-08-05 12:30:23 -04:00
jared 61b9c95731 chore: release 0.2.0
R-CMD-check / check (push) Successful in 3m34s
R-CMD-check / check (pull_request) Successful in 6m30s
Bumps the minor version because 'All Categories' adds public surface
without breaking any existing call.

The bump is load-bearing, not cosmetic: cog-api installs this package
with install_local(), which no-ops when the version already matches.
Without it, Phase 1 would silently test against the 0.1.0 reader and
pass while proving nothing.
2026-08-05 12:04:24 -04:00
jared 44e9b40b86 docs: n_units_reporting is category-conditional, not a response rate
Closes uscogdata#36. It counts governments with rows for the requested
category, so a surveyed government that genuinely spends nothing there
is indistinguishable from one never surveyed. In FY2022, a complete
census year, Georgia reports 393 of 567 cities for Police -- the gap is
cities that contract to the sheriff.

Documents the comparison that IS valid: same category, census year vs
sample year.
2026-08-05 11:58:25 -04:00
jared f1e9aa383a test: prove 'All Categories' passes through the geographic rollup
Geographic totals are the expensive case cog-api#37 was filed about --
without this a caller issues one rollup per category and sums them.
The pass-through was expected to work by construction; this asserts it
rather than assuming it, including under per_capita and inflation
adjustment.
2026-08-05 11:53:27 -04:00
jaredandClaude Opus 5 12a9be110f feat: advertise 'All Categories' from cog_categories()
A reserved value nobody can discover is a trap, and this is the view
the API's /categories endpoint is built from. Emitted for the two flow
vocabularies only -- cog_balances() returns a stock and has no concept
to sum within.

Also fix test-categories.R to exclude pseudo-category rows from
crosswalk-specific assertions (one row per (category, subtype) pair,
non-empty item_codes, valid subtypes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:43:02 -04:00
jared 503fa6562f docs: fix false subtype= claim in All Categories roxygen (F1)
The @param category text on cog_spending()/cog_revenue() told users to
"Combine with subtype = ..." but neither verb has a subtype argument.
Replace with accurate guidance: filter the returned frame's
spend_subtype/revenue_subtype column.
2026-08-05 11:35:47 -04:00
jared 11ae99c382 feat: accept category = 'All Categories' on cog_spending/cog_revenue
Returns one summed row per (year, govid, subtype) across every
category in the requested concept's subtype scope, so a caller never
sums categories client-side and cannot sum the wrong scope.

Combining it with other category names is an error rather than a
silent partial sum.
2026-08-05 11:32:47 -04:00
jared 5e22e940e7 feat: all-categories mode in .build_verb_sql()
The concept boundary in this package is subtype, not category, so a
total is the existing query with the category dimension collapsed and
no category predicate applied. subtype is deliberately kept in the
grouping: subtype=operations plus all-categories is 'operating
expenditure', which is the measure a fiscal comparison wants.

Named 'All Categories' rather than 'Total' because category='Total'
would sit one argument from expenditure_concept='total' and mean
something different.
2026-08-05 11:21:35 -04:00
jared 2fc9e7585b Merge pull request 'fix: signpost partially-suppressed categories (#9)' (#32) from fix/partial-coverage-signposting-9 into main
R-CMD-check / check (push) Successful in 3m24s
2026-08-05 08:15:46 -04:00
24 changed files with 1000 additions and 48 deletions
+1 -1
View File
@@ -1,7 +1,7 @@
Package: uscogdata
Type: Package
Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus
Version: 0.1.0
Version: 0.2.0
Authors@R:
person("Civilytics", , , "jknowles@gmail.com", role = c("aut", "cre"))
Description: Curated R verbs over the Civilytics US Census of Governments
+42
View File
@@ -1,3 +1,45 @@
# uscogdata 0.2.0
## New features
* `cog_spending()` and `cog_revenue()` accept the reserved category
`"All Categories"`, returning one summed row per
`(year, canonical_govid, subtype)` across every category inside the
requested concept's subtype scope. Filtering the result to
`spend_subtype == "operations"` gives an operating-expenditure total.
`cog_geographic_rollup()` inherits it,
which is the efficient way to build a geographic total — previously a
caller had to issue one rollup per category and sum the results
(cog-api#37).
`"All Categories"` is not the same thing as `expenditure_concept = "total"`.
The concept chooses which subtypes are in scope; `"All Categories"` chooses
whether the rows inside that scope are broken out or summed.
* `cog_categories()` advertises `"All Categories"` for the expenditure and
revenue vocabularies, so the reserved value is discoverable.
* Coverage signposting (see "Signposting now catches partially-suppressed
categories" below) now also works in `category = "All Categories"` mode.
The recipe-suggestion candidate query used to be scoped by `category`,
which is never a match for the reserved `"All Categories"` value, so
`provenance$suggestions` always came back empty there — the one mode whose
whole point is "you cannot sum the wrong scope" was silently unable to
signal a wrong scope. The candidate query is now scoped by the concept's
subtype allowlist instead, symmetric with how `.build_verb_sql()` itself
scopes the summed total: Los Angeles County FY2011, `category = "All
Categories"` still excludes $271,589,000 of aggregate-published Public
Welfare (`E68`), but now names `recipe = "welfare_cash_e68_wide"` to
recover it instead of reporting zero suggestions.
## Documentation
* `cog_geographic_rollup()` and `cog_peer_compare()` now document that
`provenance$coverage`'s `n_units_reporting` is **category-conditional** and
is not a response rate: a government that was surveyed and genuinely spends
nothing in the requested category is indistinguishable from one never
surveyed (uscogdata#36).
# uscogdata 0.1.0 (development)
## Signposting now catches partially-suppressed categories
+13 -1
View File
@@ -28,7 +28,12 @@
#' every combination would be either redundant or empty.
#' `category = "Fund Balances"` is exactly the `general` family
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
#' one `dplyr::filter()` away.
#' one `dplyr::filter()` away. The reserved pseudo-category
#' `"All Categories"` (see [cog_spending()]) is **not** supported here and
#' errors with class `uscogdata_all_categories_unsupported`: it sums a
#' concept's subtype scope, and holdings are a stock with no concept
#' vocabulary to sum across. Omit `category` to get every category broken
#' out instead.
#' @param per_capita Divide holdings by population. Note this is a **stock per
#' resident** (reserves per person), which is *not* comparable to
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
@@ -74,6 +79,13 @@ cog_balances <- function(govid, years, category = NULL,
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
# route the verb through .verb_spendrev(), which stays deliberately unused
# here because its flow vocabulary is meaningless for a stock.
#
# allow_all_categories is left at its FALSE default (contrast
# .verb_spendrev(), which passes TRUE): the all-categories mode's "sum"
# only means something in terms of a concept's subtype scope, and holdings
# have no concept vocabulary. The reuse above is exactly why this can be a
# one-line default rather than a second bespoke check -- see the
# validator's own doc comment for the incident that made that matter.
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
years <- as.integer(years)
+28 -2
View File
@@ -22,7 +22,10 @@
#' `category` column (e.g. `"Police"` or `"Tax"`).
#' @return Tibble with columns `category`, `category_type`, `subtype`,
#' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
#' `category_type`, `category`, `subtype`.
#' `category_type`, `category`, `subtype`. Includes one row per flow for the
#' reserved pseudo-category `"All Categories"`, which carries `NA` for
#' `subtype`, `n_codes` and `item_codes` because it is a query mode rather
#' than a crosswalk entry — see [cog_spending()]'s `category` argument.
#' @export
cog_categories <- function(type = NULL, pattern = NULL) {
if (!is.null(type)) {
@@ -63,5 +66,28 @@ cog_categories <- function(type = NULL, pattern = NULL) {
"GROUP BY category, category_type, subtype
ORDER BY category_type, category, subtype"
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
out <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
# The reserved pseudo-category is a query mode, not a crosswalk row, so it
# has no item codes to report -- hence NA rather than 0 for n_codes. It is
# emitted for the two FLOW vocabularies only: cog_balances() returns a stock
# and has no concept argument to sum within.
pseudo <- tibble::tibble(
category = .ALL_CATEGORIES,
category_type = c("expenditure", "revenue"),
subtype = NA_character_,
n_codes = NA_integer_,
item_codes = NA_character_
)
if (!is.null(type)) {
db_type <- if (type == "spending") "expenditure" else type
pseudo <- pseudo[pseudo$category_type == db_type, , drop = FALSE]
}
if (!is.null(pattern) && nrow(pseudo) > 0L) {
keep <- grepl(pattern, pseudo$category, ignore.case = TRUE)
pseudo <- pseudo[keep, , drop = FALSE]
}
if (nrow(pseudo) == 0L) return(out)
out <- rbind(out, pseudo)
out[order(out$category_type, out$category, out$subtype), , drop = FALSE]
}
+17
View File
@@ -240,6 +240,23 @@ cog_find_peers <- function(target_govid,
#' group_by(year) |>
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
#' ```
#' @section Reading `coverage`:
#' `provenance$coverage` reports `n_units_reporting` against
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
#' it counts cohort members with rows for the category you asked for, not
#' cohort members collected that year.** A government that was surveyed and
#' genuinely spends nothing in that category is indistinguishable here from one
#' that was never surveyed.
#'
#' The ratio is therefore **not a response rate** and must not be used as one.
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
#' contract policing to the county sheriff, not non-response.
#'
#' The comparison that *is* valid is the same category across a census year
#' (ending in 2 or 7) and a sample year, where the real-zero component is
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
#' marks which is which.
#' @export
cog_peer_compare <- function(target_govid, peers, category, years,
per_capita = TRUE, adjust_to_year = NULL,
+9 -1
View File
@@ -67,7 +67,15 @@
basis_note = basis_note,
expenditure_concept = expenditure_concept,
expenditure_concept_note = expenditure_concept_note,
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
# isTRUE() alone would collapse a deliberate NA (all-categories mode,
# where suppression detection cannot run -- see .verb_spendrev()) down to
# FALSE, turning "we don't know" back into the false claim this field
# exists to avoid. Preserve NA; otherwise normalize to a strict logical.
expenditure_concept_direct_suppressed = if (isTRUE(is.na(expenditure_concept_direct_suppressed))) {
NA
} else {
isTRUE(expenditure_concept_direct_suppressed)
},
revenue_concept = revenue_concept,
harmonization = harmonization %||% list(
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
+15 -2
View File
@@ -8,6 +8,17 @@
#' multiplies by 1000 and records the conversion in `provenance`).
#'
#' @inheritParams cog_spending
#' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories broken out
#' one row each. The reserved value `"All Categories"` instead returns a
#' single summed row per `(year, canonical_govid, subtype)`, covering every
#' category inside the requested concept's subtype scope. It cannot be
#' combined with other category names, and it is not the same thing as
#' `revenue_concept = "total"`: the concept chooses which subtypes are in
#' scope, `"All Categories"` chooses whether rows inside that scope are
#' broken out or summed. Because the result keeps one row per
#' `revenue_subtype`, filtering the returned frame to
#' `revenue_subtype == "own_source"` gives an own-source revenue total.
#' @param revenue_concept Which of Census's two published revenue concepts to
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
#' values -- never as item-code first letters, which cannot classify
@@ -43,7 +54,7 @@ cog_revenue <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE) {
complete = FALSE, limit = NULL, offset = NULL) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
@@ -62,6 +73,8 @@ cog_revenue <- function(govid, years, category = NULL,
basis = basis,
recipe = recipe,
revenue_concept = revenue_concept,
complete = complete
complete = complete,
limit = limit,
offset = offset
)
}
+22 -1
View File
@@ -19,7 +19,11 @@
#' `state`, `county`, `city`. Each element is a character vector of
#' `canonical_govid` values. At least one layer required.
#' @param category Single category name or character vector (passed through
#' to [cog_spending()]).
#' to [cog_spending()]), or the reserved `"All Categories"` for one summed
#' row per `(year, canonical_govid, subtype)` covering every category in the
#' concept's scope. `"All Categories"` is the efficient way to build a
#' geographic total: without it a caller must issue one rollup per category
#' and sum the results themselves.
#' @param years Integer vector of years.
#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
#' population from `gov_population_yearly`. Govs with missing population
@@ -56,6 +60,23 @@
#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
#' and `rollup$included_govids` / `rollup$excluded_govids`.
#' @section Reading `coverage`:
#' `provenance$coverage` reports `n_units_reporting` against
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
#' it counts governments with rows for the category you asked for, not
#' governments collected that year.** A government that was surveyed and
#' genuinely spends nothing in that category is indistinguishable here from one
#' that was never surveyed.
#'
#' The ratio is therefore **not a response rate** and must not be used as one.
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
#' contract policing to the county sheriff, not non-response.
#'
#' The comparison that *is* valid is the same category across a census year
#' (ending in 2 or 7) and a sample year, where the real-zero component is
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
#' marks which is which.
#' @export
cog_geographic_rollup <- function(govids, category, years,
per_capita = FALSE, adjust_to_year = NULL,
+233 -21
View File
@@ -19,6 +19,14 @@
.spend_subtypes_primary <- c("operations", "capital", "assistance")
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
# The reserved pseudo-category. Deliberately NOT "Total": `category = "Total"`
# would sit one argument away from `expenditure_concept = "total"` and mean
# something different -- the concept selects WHICH SUBTYPES are in scope, this
# selects whether the rows inside that scope are broken out by category or
# summed. "All Categories" states the operation and cannot be misread as the
# concept.
.ALL_CATEGORIES <- "All Categories"
#' @noRd
.expenditure_concept_subtypes <- function(concept) {
switch(concept,
@@ -63,7 +71,16 @@
#' @param govid Character vector of `canonical_govid` values.
#' @param years Integer vector of years.
#' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories.
#' `summary_categories.category`), or `NULL` for all categories broken out
#' one row each. The reserved value `"All Categories"` instead returns a
#' single summed row per `(year, canonical_govid, subtype)`, covering every
#' category inside the requested concept's subtype scope. It cannot be
#' combined with other category names, and it is not the same thing as
#' `expenditure_concept = "total"`: the concept chooses which subtypes are in
#' scope, `"All Categories"` chooses whether rows inside that scope are
#' broken out or summed. Because the result keeps one row per
#' `spend_subtype`, filtering the returned frame to
#' `spend_subtype == "operations"` gives an operating-expenditure total.
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
#' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
#' Census F-33 population from `gov_population_yearly`. Result also gains
@@ -133,7 +150,12 @@
#' component (when one exists), and
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
#' figure in those rows is the intergovernmental leg alone, not Direct +
#' IG.
#' IG. When `category = "All Categories"` is combined with
#' `expenditure_concept = "total"`, this detection cannot run (it keys on
#' per-category rows, which all-categories mode collapses to one literal
#' value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
#' possibly-false `FALSE`; query an explicit `category` to get a real
#' answer.
#' @param complete If `TRUE`, fill the requested grid so that a cell the
#' corpus does not carry still appears, labelled with **why** it is
#' missing, and add a `value_source` column to every row:
@@ -156,6 +178,13 @@
#' `recipe` or with `expenditure_concept = "total"` (class
#' `uscogdata_complete_unsupported`) — neither draws its cells from
#' `code_set`.
#' @param limit Maximum number of result rows to return, pushed into the SQL
#' query itself (`LIMIT`/`OFFSET`) rather than applied after the full
#' result is materialized. `NULL` (the default) returns every matching row,
#' exactly as before this parameter existed. Mutually exclusive with
#' `recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.
#' @param offset Rows to skip before `limit` starts counting (0-based).
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
@@ -163,13 +192,17 @@
#' and `value_source` when `complete = TRUE`.
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
#' whose `completion` block reports `applied`, `rows_filled`, and the
#' per-year `absence_means` rule that was applied.
#' per-year `absence_means` rule that was applied. When `limit` is set,
#' also carries a `total_rows` attribute: the full unpaginated row count,
#' computed by the same query (`COUNT(*) OVER()`) rather than a second
#' round trip -- so a caller walking pages never has to ask "how many are
#' there" separately.
#' @export
cog_spending <- function(govid, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE) {
complete = FALSE, limit = NULL, offset = NULL) {
# flow_prefixes no longer classifies rows (crosswalk subtype membership
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
# machinery to this verb's recipe families (see R/suggestions.R; the
@@ -188,7 +221,9 @@ cog_spending <- function(govid, years, category = NULL,
basis = basis,
recipe = recipe,
expenditure_concept = expenditure_concept,
complete = complete
complete = complete,
limit = limit,
offset = offset
)
}
@@ -216,7 +251,7 @@ cog_spending <- function(govid, years, category = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
revenue_concept = c("general", "total"),
complete = FALSE) {
complete = FALSE, limit = NULL, offset = NULL) {
basis_explicit <- length(basis) == 1L
basis <- match.arg(basis, c("harmonized", "raw"))
# match.arg() itself throws a base `simpleError`, not an rlang-classed
@@ -257,8 +292,24 @@ cog_spending <- function(govid, years, category = NULL,
}
govid <- .coerce_govid_input(govid, arg = "govid")
# allow_all_categories = TRUE: cog_spending()/cog_revenue() are the two
# verbs the reserved pseudo-category is defined for. cog_balances() shares
# this validator but leaves the argument at its FALSE default, so it
# rejects "All Categories" instead of silently returning zero rows
# (finding 3, all-categories review).
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
recipe)
recipe, allow_all_categories = TRUE)
# Recognize the reserved pseudo-category. Detected after type validation so a
# non-character `category` still fails with the ordinary type error.
all_categories <- !is.null(category) && .ALL_CATEGORIES %in% category
if (all_categories && length(category) > 1L) {
cli::cli_abort(c(
"{.val {(.ALL_CATEGORIES)}} cannot be combined with other categories.",
"i" = "It already sums every category in the requested concept's scope.",
"*" = "Ask for it alone, or list the specific categories you want."
), class = "uscogdata_all_categories_not_combinable")
}
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
cli::cli_abort(c(
@@ -299,6 +350,44 @@ cog_spending <- function(govid, years, category = NULL,
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
)
}
if (complete && all_categories) {
.abort_complete_unsupported(
"`category = \"All Categories\"` collapses the category dimension that `code_set` grids over (see `.completion_grid_sql()`), so there is no per-category grid left to fill -- filling a summed row has no defined semantics.",
"Drop `complete`, or use `complete = TRUE` with an explicit `category` (or `category = NULL` for every category)."
)
}
# limit/offset push the page into the SQL itself (see .build_verb_sql()),
# so the two things that would make "a page of what" ambiguous are refused
# up front rather than silently ignored: complete = TRUE fills a grid over
# the FULL requested (year, category) space, and a recipe's result comes
# from .run_recipe()'s own query, which this function does not touch.
if (!is.null(limit)) {
limit <- as.integer(limit)
if (length(limit) != 1L || is.na(limit) || limit < 0L) {
cli::cli_abort("`limit` must be a single non-negative integer.",
class = "uscogdata_invalid_pagination")
}
offset <- if (is.null(offset)) 0L else as.integer(offset)
if (length(offset) != 1L || is.na(offset) || offset < 0L) {
cli::cli_abort("`offset` must be a single non-negative integer.",
class = "uscogdata_invalid_pagination")
}
if (complete) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with `complete = TRUE`.",
"i" = "`complete` fills a grid over the FULL requested (year, category) space; paginating a slice of already-grouped rows has no defined meaning for the cells it would fill.",
"*" = "Drop `limit`/`offset`, or drop `complete`."
), class = "uscogdata_complete_pagination_conflict")
}
if (!is.null(recipe)) {
cli::cli_abort(c(
"`limit`/`offset` cannot be combined with `recipe`.",
"i" = "A recipe's result comes from a separate query (`.run_recipe()`) that pagination is not wired into yet.",
"*" = "Drop `limit`/`offset`, or drop `recipe`."
), class = "uscogdata_recipe_pagination_conflict")
}
}
years <- as.integer(years)
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
@@ -312,6 +401,7 @@ cog_spending <- function(govid, years, category = NULL,
recipe_block <- NULL
category_for_prov <- category
total_rows <- NULL # set below only when limit is non-NULL (non-recipe path)
if (!is.null(recipe)) {
.require_schema_v5(con, manifest, "recipe =")
.validate_recipe_id(con, recipe)
@@ -333,9 +423,32 @@ cog_spending <- function(govid, years, category = NULL,
} else {
NULL
}
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
subtype_scope)
sql <- .build_verb_sql(view, subtype_col, govid, years,
if (all_categories) NULL else category,
ig_view, subtype_scope,
all_categories = all_categories,
limit = limit, offset = offset)
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
if (!is.null(limit)) {
# COUNT(*) OVER() rides along as an ordinary column so the total comes
# from the same scan when this page has any rows -- see
# .build_verb_sql(). An empty page (offset past the end) carries no
# such row to read it from, so that one case falls back to a second,
# unpaginated COUNT(*) query rather than reporting a wrong zero.
if (nrow(result) > 0L) {
total_rows <- result$pagination_total_rows[[1]]
result$pagination_total_rows <- NULL
} else {
count_sql <- sprintf(
"SELECT COUNT(*) AS n FROM (%s) AS _uncounted",
.build_verb_sql(view, subtype_col, govid, years,
if (all_categories) NULL else category,
ig_view, subtype_scope,
all_categories = all_categories)
)
total_rows <- as.integer(DBI::dbGetQuery(con, count_sql)$n[[1]])
}
}
}
# Fill BEFORE per_capita / inflation so the added cells get the same
@@ -396,7 +509,10 @@ cog_spending <- function(govid, years, category = NULL,
suggestions <- .build_suggestions(con, govid, years, category,
direct_leg_result,
resolved$basis, flow_prefixes,
.select_long_view(view_base, resolved$basis))
.select_long_view(view_base, resolved$basis),
all_categories = all_categories,
subtype_col = subtype_col,
subtype_scope = subtype_scope)
}
# C1(b): when expenditure_concept = "total", flag any row where the IG
@@ -408,13 +524,32 @@ cog_spending <- function(govid, years, category = NULL,
# direct spending in that category, which is correct, ordinary data). When
# a covering recipe is found, both the row-level notes and the provenance
# say so rather than pass silently as a plausible Total.
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
#
# In all-categories mode this cannot run at all: .detect_direct_suppressed()
# keys on (year, canonical_govid, category), and every row shares the same
# literal "All Categories" value, so the key collides across every real
# category for that (year, govid) -- an IG-only row for a suppressed
# category becomes indistinguishable from one sharing a key with an
# unrelated category's ordinary Direct row. `has_direct` would then read
# TRUE whenever the government has ANY direct spending at all, and the
# detector could never fire. Rather than run it and report a false FALSE,
# skip it and record NA -- the provenance must stop making a claim it
# cannot support (finding 1, all-categories review).
suppression_unavailable <- all_categories &&
identical(expenditure_concept, "total")
direct_suppressed_info <- if (suppression_unavailable) {
list(flag = rep(NA, nrow(result)), notes = rep(NA_character_, nrow(result)))
} else if (identical(expenditure_concept, "total")) {
.detect_direct_suppressed(con, result, subtype_col)
} else {
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
}
direct_suppressed <- direct_suppressed_info$flag
direct_suppressed_flag <- isTRUE(any(direct_suppressed))
direct_suppressed_flag <- if (suppression_unavailable) {
NA
} else {
isTRUE(any(direct_suppressed))
}
result$notes <- .notes_column(result, direct_suppressed_info$notes)
@@ -423,9 +558,20 @@ cog_spending <- function(govid, years, category = NULL,
# leg is suppressed for at least one requested (year, category), append an
# explicit warning rather than let the base note's "Total = Direct + IG"
# framing stand unqualified for rows where that arithmetic didn't happen.
# When suppression detection itself is unavailable (all-categories mode),
# say so instead of silently reusing the unqualified base note.
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
if (direct_suppressed_flag) {
if (suppression_unavailable) {
paste0(
base_note,
" NOTE: direct-leg-suppression detection is unavailable when ",
"`category = \"All Categories\"` -- it keys on per-category rows, ",
"which this mode collapses. `expenditure_concept_direct_suppressed` ",
"is NA here rather than a possibly-false FALSE; query an explicit ",
"`category` (or `category = NULL`) to get a real answer."
)
} else if (isTRUE(direct_suppressed_flag)) {
paste0(
base_note,
" NOTE: for at least one requested (year, category) the Direct leg ",
@@ -467,15 +613,32 @@ cog_spending <- function(govid, years, category = NULL,
prov$scope$govids_missing <- scope$missing
attr(result, "provenance") <- prov
attr(result, ".popyear_range") <- NULL
# Attached here, after every downstream transform (per_capita/real-dollar
# joins, notes, subtype filtering), the same way provenance is -- an
# attribute set before those runs is not guaranteed to survive them.
if (!is.null(limit)) attr(result, "total_rows") <- total_rows
if (length(suggestions) > 0L) .inform_suggestions(suggestions)
result
}
#' Shared input validation for the money/holdings verbs.
#'
#' `allow_all_categories` gates the reserved pseudo-category
#' `.ALL_CATEGORIES` ("All Categories"). It is meaningful only where a
#' concept's subtype scope defines what "all" sums over --
#' `cog_spending()`/`cog_revenue()`, via `.verb_spendrev()`, pass `TRUE`.
#' `cog_balances()` leaves it at the `FALSE` default: holdings are a stock
#' with no concept vocabulary to sum across (see R/balances.R), and before
#' this guard existed `cog_balances(category = "All Categories")` silently
#' matched zero crosswalk rows and returned an empty result with no error
#' (finding 3, all-categories review). This validator is shared specifically
#' so the three verbs cannot drift apart on this again.
#' @noRd
.validate_verb_inputs <- function(govid, years, category,
per_capita, adjust_to_year, recipe = NULL) {
per_capita, adjust_to_year, recipe = NULL,
allow_all_categories = FALSE) {
if (!is.character(govid) || length(govid) == 0L) {
cli::cli_abort("`govid` must be a non-empty character vector.")
}
@@ -485,6 +648,14 @@ cog_spending <- function(govid, years, category = NULL,
if (!is.null(category) && !is.character(category)) {
cli::cli_abort("`category` must be character or NULL.")
}
if (!allow_all_categories && !is.null(category) &&
.ALL_CATEGORIES %in% category) {
cli::cli_abort(c(
"{.val {(.ALL_CATEGORIES)}} is not supported here.",
i = "It sums a spending or revenue concept's subtype scope; this verb has no concept vocabulary to sum across.",
i = "Use {.fn cog_spending} or {.fn cog_revenue} for an all-categories total."
), class = "uscogdata_all_categories_unsupported")
}
if (!is.logical(per_capita) || length(per_capita) != 1L) {
cli::cli_abort("`per_capita` must be a length-1 logical.")
}
@@ -570,10 +741,16 @@ cog_spending <- function(govid, years, category = NULL,
#' @noRd
.build_verb_sql <- function(view, subtype_col, govid, years, category,
ig_view = NULL, subtype_scope = NULL) {
ig_view = NULL, subtype_scope = NULL,
all_categories = FALSE, limit = NULL, offset = NULL) {
govid_lit <- .sql_lit_chr(govid)
years_lit <- paste(as.integer(years), collapse = ",")
category_pred <- if (is.null(category)) {
# In all-categories mode there is no category filter: the sum is defined by
# the concept's SUBTYPE allowlist (subtype_pred below), which is the real
# concept boundary. Filtering by category as well would be a no-op at best
# and, if the crosswalk ever gained an uncategorized code, a silent
# under-count of the very total this mode exists to guarantee.
category_pred <- if (all_categories || is.null(category)) {
""
} else {
sprintf("AND category IN (%s)", .sql_lit_chr(category))
@@ -612,13 +789,26 @@ cog_spending <- function(govid, years, category = NULL,
# though its dollars came entirely from an aggregate row, silently
# suppressing the "Aggregate fallback applied" note on exactly the rows
# this feature exists to surface.
sprintf(
# Collapse the category dimension. subtype is deliberately KEPT: it is what
# lets a caller filter the result to `spend_subtype == "operations"` and
# get an operating-expenditure total, the measure a fiscal comparison
# actually wants. (There is no `subtype` argument -- this is a post-hoc
# filter on the returned column, not a query parameter.)
category_select <- if (all_categories) {
sprintf("%s AS category", .sql_lit_chr(.ALL_CATEGORIES))
} else {
"category"
}
category_group <- if (all_categories) "" else ", category"
base_sql <- sprintf(
"SELECT
year,
canonical_govid,
COALESCE(xwalk_gov_name, gov_name) AS gov_name,
%1$s,
category,
%7$s,
SUM(amt) * 1000.0 AS amt_nominal,
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
bool_or(is_aggregate) AS aggregate_fallback
@@ -627,10 +817,32 @@ cog_spending <- function(govid, years, category = NULL,
AND year IN (%4$s)
%5$s
%6$s
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
ORDER BY year, canonical_govid, %1$s, category",
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s%8$s
ORDER BY year, canonical_govid, %1$s%8$s",
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred,
category_select, category_group
)
# limit/offset push the page into the query itself instead of pulling every
# matching row across the network only to slice and discard most of it
# afterward (the pattern behind the 2026-08-06 production incident: a
# 193,105-row/194-page sweep re-ran the full query and re-listified every
# row on EVERY page). COUNT(*) OVER() rides along as an ordinary column so
# the caller gets the true total from this same scan -- see the call site
# in .verb_spendrev(), which reads it off row 1 and strips it back out.
# The outer SELECT * wrapping (rather than appending LIMIT/OFFSET directly
# to base_sql) is what makes COUNT(*) OVER() see the post-GROUP-BY row
# count, not the pre-aggregation one.
if (is.null(limit)) {
base_sql
} else {
sprintf(
"SELECT *, COUNT(*) OVER() AS pagination_total_rows
FROM (%s) AS _paged
LIMIT %d OFFSET %d",
base_sql, limit, offset
)
}
}
#' @noRd
+46 -3
View File
@@ -59,12 +59,36 @@
#' @param long_view Name of the verb's own long view (from
#' `.select_long_view()`), passed through to `.suppressed_components()` to
#' measure the second qualifying path (uscogdata#9).
#' @param all_categories `TRUE` when the caller's `category` is the reserved
#' pseudo-category (`.ALL_CATEGORIES`). Defaults to `FALSE` so no other
#' caller's behaviour changes. When `TRUE`, the candidate-recipe sub-select
#' is scoped by `subtype_col`/`subtype_scope` instead of by `category` --
#' symmetric with `.build_verb_sql()`'s own all-categories branch (see
#' R/spending.R): the concept's subtype allowlist is the real scope
#' boundary, not any literal category value, and
#' `.ALL_CATEGORIES` ("All Categories") is never itself a row in
#' `summary_categories.category`, so leaving the category-keyed sub-select
#' in place here always returned zero candidates and silently disabled
#' signposting in all-categories mode (final whole-branch review, finding
#' 6).
#' @param subtype_col Name of the `summary_categories` subtype column to
#' scope by when `all_categories = TRUE` (`"spend_subtype"` or
#' `"revenue_subtype"` -- the same value `.build_verb_sql()` already
#' receives as its own `subtype_col`). Ignored when `all_categories =
#' FALSE`. `NULL` by default.
#' @param subtype_scope Character vector of subtype values to scope by when
#' `all_categories = TRUE` (the same value `.build_verb_sql()` already
#' receives as its own `subtype_scope` -- the concept's subtype allowlist,
#' e.g. `.expenditure_concept_subtypes(expenditure_concept)`). Ignored when
#' `all_categories = FALSE`. `NULL` by default.
#' @return List of `list(recipe_id, label, available_years, hint,
#' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
#' suppressed_codes)`, possibly empty.
#' @noRd
.build_suggestions <- function(con, govid, years, category, result, basis,
flow_prefixes, long_view) {
flow_prefixes, long_view,
all_categories = FALSE,
subtype_col = NULL, subtype_scope = NULL) {
if (!identical(basis, "harmonized") || is.null(category)) return(list())
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
@@ -79,16 +103,35 @@
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
# should never be suggested as a coverage-gap filler for EITHER verb, not
# just kept from being named as the *counterpart* of another suggestion.
#
# The inner sub-select is the concept boundary (finding 6, final
# whole-branch review): in all-categories mode it is scoped by
# `subtype_col`/`subtype_scope` -- the same allowlist `.build_verb_sql()`
# applies as a WHERE predicate to make the summed result a *concept*, not
# by `category` (`.ALL_CATEGORIES` is never a row in
# `summary_categories.category`, so a category-keyed sub-select always
# came back empty here). The M/L exclusion below is unchanged either way.
candidate_scope_sql <- if (isTRUE(all_categories)) {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories WHERE %s IN (%s)",
subtype_col, .sql_lit_chr(subtype_scope)
)
} else {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)",
.sql_lit_chr(category)
)
}
candidates <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE component_code IN (
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)
%s
)
AND recipe_id NOT IN (
SELECT DISTINCT recipe_id FROM harmonization_recipes
WHERE LEFT(component_code, 1) IN ('M', 'L')
)",
.sql_lit_chr(category)
candidate_scope_sql
))$recipe_id
if (length(candidates) == 0L) return(list())
Binary file not shown.
+3 -3
View File
@@ -1,7 +1,7 @@
{
"schema_version": 6,
"built_at": "2026-07-31T00:47:27Z",
"pipeline_commit": "aadb46b",
"built_at": "2026-08-03T16:51:32Z",
"pipeline_commit": "e7394a4",
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
"data_vintage": {
"source_vintages": {
@@ -105,7 +105,7 @@
},
{
"path": "data/series_breaks.parquet",
"sha256": "06dcc995ff533e57cc65fa25086cc9bf83ba592c58bf7cc99269dc2576f69944",
"sha256": "731998516cd802f63fcf7fb66053c7a62b7be955ab0794cad4a4979cb7628b87",
"description": "series_breaks.parquet"
},
{
+2 -2
View File
@@ -22,8 +22,8 @@
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
},
"expenditure_concept_direct_suppressed": {
"type": "boolean",
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
"type": ["boolean", "null"],
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. null (NA) when expenditure_concept = 'total' AND category = 'All Categories': the detector keys on per-category rows, which that mode collapses, so suppression cannot be computed -- see `expenditure_concept_note`. See the affected rows' `notes` for the recovering recipe, if any."
},
"revenue_concept": {
"type": "string",
+6 -1
View File
@@ -28,7 +28,12 @@ argument: for holdings, `category` is a strict coarsening of
every combination would be either redundant or empty.
`category = "Fund Balances"` is exactly the `general` family
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
one `dplyr::filter()` away.}
one `dplyr::filter()` away. The reserved pseudo-category
`"All Categories"` (see [cog_spending()]) is **not** supported here and
errors with class `uscogdata_all_categories_unsupported`: it sums a
concept's subtype scope, and holdings are a stock with no concept
vocabulary to sum across. Omit `category` to get every category broken
out instead.}
\item{per_capita}{Divide holdings by population. Note this is a **stock per
resident** (reserves per person), which is *not* comparable to
+4 -1
View File
@@ -16,7 +16,10 @@ balance), `"spending"`, `"revenue"`, or `"balance"`.}
\value{
Tibble with columns `category`, `category_type`, `subtype`,
`n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
`category_type`, `category`, `subtype`.
`category_type`, `category`, `subtype`. Includes one row per flow for the
reserved pseudo-category `"All Categories"`, which carries `NA` for
`subtype`, `n_codes` and `item_codes` because it is a query mode rather
than a crosswalk entry — see [cog_spending()]'s `category` argument.
}
\description{
Returns the category taxonomy exposed by the corpus's
+25 -1
View File
@@ -20,7 +20,11 @@ cog_geographic_rollup(
`canonical_govid` values. At least one layer required.}
\item{category}{Single category name or character vector (passed through
to [cog_spending()]).}
to [cog_spending()]), or the reserved `"All Categories"` for one summed
row per `(year, canonical_govid, subtype)` covering every category in the
concept's scope. `"All Categories"` is the efficient way to build a
geographic total: without it a caller must issue one rollup per category
and sum the results themselves.}
\item{years}{Integer vector of years.}
@@ -80,3 +84,23 @@ the result. The dropped govids are recorded in
(gov type 4) and school districts (gov type 5) from per-capita rollups
by design — see `vignette('population-denominators')`.
}
\section{Reading `coverage`}{
`provenance$coverage` reports `n_units_reporting` against
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
it counts governments with rows for the category you asked for, not
governments collected that year.** A government that was surveyed and
genuinely spends nothing in that category is indistinguishable here from one
that was never surveyed.
The ratio is therefore **not a response rate** and must not be used as one.
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
`category = "Police"`; the 174-city gap is overwhelmingly cities that
contract policing to the county sheriff, not non-response.
The comparison that *is* valid is the same category across a census year
(ending in 2 or 7) and a sample year, where the real-zero component is
roughly constant and the difference reflects the survey cycle. `is_census_year`
marks which is which.
}
+20
View File
@@ -107,3 +107,23 @@ call. Those summary rows are quantiles **within each category**, not
quantiles of each peer's total — see the `@return` section before summing
them.
}
\section{Reading `coverage`}{
`provenance$coverage` reports `n_units_reporting` against
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
it counts cohort members with rows for the category you asked for, not
cohort members collected that year.** A government that was surveyed and
genuinely spends nothing in that category is indistinguishable here from one
that was never surveyed.
The ratio is therefore **not a response rate** and must not be used as one.
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
`category = "Police"`; the 174-city gap is overwhelmingly cities that
contract policing to the county sheriff, not non-response.
The comparison that *is* valid is the same category across a census year
(ending in 2 or 7) and a sample year, where the real-zero component is
roughly constant and the difference reflects the survey cycle. `is_census_year`
marks which is which.
}
+22 -2
View File
@@ -13,7 +13,9 @@ cog_revenue(
basis = c("harmonized", "raw"),
recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE
complete = FALSE,
limit = NULL,
offset = NULL
)
}
\arguments{
@@ -22,7 +24,16 @@ cog_revenue(
\item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.}
`summary_categories.category`), or `NULL` for all categories broken out
one row each. The reserved value `"All Categories"` instead returns a
single summed row per `(year, canonical_govid, subtype)`, covering every
category inside the requested concept's subtype scope. It cannot be
combined with other category names, and it is not the same thing as
`revenue_concept = "total"`: the concept chooses which subtypes are in
scope, `"All Categories"` chooses whether rows inside that scope are
broken out or summed. Because the result keeps one row per
`revenue_subtype`, filtering the returned frame to
`revenue_subtype == "own_source"` gives an own-source revenue total.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
@@ -106,6 +117,15 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
result is materialized. `NULL` (the default) returns every matching row,
exactly as before this parameter existed. Mutually exclusive with
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
+33 -4
View File
@@ -13,7 +13,9 @@ cog_spending(
basis = c("harmonized", "raw"),
recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE
complete = FALSE,
limit = NULL,
offset = NULL
)
}
\arguments{
@@ -22,7 +24,16 @@ cog_spending(
\item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.}
`summary_categories.category`), or `NULL` for all categories broken out
one row each. The reserved value `"All Categories"` instead returns a
single summed row per `(year, canonical_govid, subtype)`, covering every
category inside the requested concept's subtype scope. It cannot be
combined with other category names, and it is not the same thing as
`expenditure_concept = "total"`: the concept chooses which subtypes are in
scope, `"All Categories"` chooses whether rows inside that scope are
broken out or summed. Because the result keeps one row per
`spend_subtype`, filtering the returned frame to
`spend_subtype == "operations"` gives an operating-expenditure total.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
@@ -97,7 +108,12 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG.}
IG. When `category = "All Categories"` is combined with
`expenditure_concept = "total"`, this detection cannot run (it keys on
per-category rows, which all-categories mode collapses to one literal
value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
possibly-false `FALSE`; query an explicit `category` to get a real
answer.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
@@ -121,6 +137,15 @@ possibly-misleading `"harmonized"`/`"raw"` value.}
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
result is materialized. `NULL` (the default) returns every matching row,
exactly as before this parameter existed. Mutually exclusive with
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
@@ -130,7 +155,11 @@ Tibble with columns `year`, `canonical_govid`, `gov_name`,
and `value_source` when `complete = TRUE`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
whose `completion` block reports `applied`, `rows_filled`, and the
per-year `absence_means` rule that was applied.
per-year `absence_means` rule that was applied. When `limit` is set,
also carries a `total_rows` attribute: the full unpaginated row count,
computed by the same query (`COUNT(*) OVER()`) rather than a second
round trip -- so a caller walking pages never has to ask "how many are
there" separately.
}
\description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
@@ -0,0 +1,58 @@
test_that('cog_geographic_rollup() accepts "All Categories" and agrees with per-category sums', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
expect_gt(nrow(govs), 1L)
ids <- list(city = utils::head(govs$canonical_govid, 25L))
by_cat <- cog_geographic_rollup(ids, category = NULL, years = 2019L)
total <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L)
expect_setequal(unique(total$category), "All Categories")
# one row per (govid, subtype) that appears in the per-category result
key_by_cat <- unique(paste(by_cat$canonical_govid, by_cat$spend_subtype))
key_total <- paste(total$canonical_govid, total$spend_subtype)
expect_setequal(key_total, key_by_cat)
lhs <- tapply(by_cat$amt_nominal, paste(by_cat$canonical_govid, by_cat$spend_subtype), sum)
rhs <- tapply(total$amt_nominal, key_total, sum)
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
})
test_that('"All Categories" survives per_capita and inflation adjustment through the rollup', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = utils::head(govs$canonical_govid, 10L))
r <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
per_capita = TRUE, adjust_to_year = 2020L)
expect_true(all(c("amt_per_capita_nominal", "amt_real", "amt_per_capita_real") %in% names(r)))
expect_setequal(unique(r$category), "All Categories")
expect_true(all(is.finite(r$amt_real)))
})
test_that('cog_geographic_rollup() still refuses expenditure_concept = "total" with "All Categories"', {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = utils::head(govs$canonical_govid, 5L))
expect_error(
cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
expenditure_concept = "total")
)
})
test_that("n_units_reporting is category-conditional, not a response rate", {
skip_if_no_corpus()
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
ids <- list(city = govs$canonical_govid)
police <- cog_geographic_rollup(ids, category = "Police", years = 2012L)
allcat <- cog_geographic_rollup(ids, category = "All Categories", years = 2012L)
cov_police <- cog_explain(police, format = "list")$coverage
cov_all <- cog_explain(allcat, format = "list")$coverage
# Same year, same requested govids, same collection -- yet a single category
# reports fewer units than the all-categories query. That gap is real zeros,
# not non-response, which is exactly why the ratio is not a response rate.
expect_lte(cov_police$n_units_reporting, cov_all$n_units_reporting)
expect_identical(cov_police$n_units_expected, cov_all$n_units_expected)
})
+267
View File
@@ -0,0 +1,267 @@
# Baseline at branch point: 843 PASS / 0 FAIL / 0 SKIP / 0 WARN (2026-08-05, origin/main 2fc9e75)
test_that(".build_verb_sql emits a literal category and no category filter in all-categories mode", {
sql <- uscogdata:::.build_verb_sql(
view = "spending_annotated",
subtype_col = "spend_subtype",
govid = "552025209777",
years = 2019L,
category = NULL,
subtype_scope = c("operations", "capital"),
all_categories = TRUE
)
expect_match(sql, "'All Categories' AS category", fixed = TRUE)
# no category filter of any kind
expect_false(grepl("AND category IN", sql, fixed = TRUE))
# category is not a grouping key
expect_false(grepl("GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
sql, fixed = TRUE))
# the subtype allowlist still applies -- this is what makes the sum a concept
expect_match(sql, "AND spend_subtype IN ('operations','capital')", fixed = TRUE)
})
test_that(".build_verb_sql is unchanged when all_categories is FALSE", {
args <- list(
view = "spending_annotated", subtype_col = "spend_subtype",
govid = "552025209777", years = 2019L, category = NULL,
subtype_scope = c("operations", "capital")
)
old <- do.call(uscogdata:::.build_verb_sql, args)
new <- do.call(uscogdata:::.build_verb_sql, c(args, list(all_categories = FALSE)))
expect_identical(old, new)
expect_match(new, "GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
fixed = TRUE)
})
test_that(".ALL_CATEGORIES is the exact reserved string", {
expect_identical(uscogdata:::.ALL_CATEGORIES, "All Categories")
})
test_that('cog_spending(category = "All Categories") sums to the per-category total', {
gov <- "552025209777"
by_cat <- cog_spending(gov, 2019L)
total <- cog_spending(gov, 2019L, category = "All Categories")
expect_true(nrow(total) > 0L)
expect_setequal(unique(total$category), "All Categories")
# one row per subtype present in the by-category result
expect_setequal(unique(total$spend_subtype), unique(by_cat$spend_subtype))
expect_equal(nrow(total), length(unique(by_cat$spend_subtype)))
# the dollars agree, per subtype
lhs <- tapply(by_cat$amt_nominal, by_cat$spend_subtype, sum)
rhs <- tapply(total$amt_nominal, total$spend_subtype, sum)
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
})
test_that('"All Categories" respects expenditure_concept', {
gov <- "552025209777"
prim <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "primary")
dir <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "direct")
# direct = primary plus interest and insurance benefits, so it is never smaller
expect_gte(sum(dir$amt_nominal), sum(prim$amt_nominal))
})
test_that('"All Categories" works on revenue and respects revenue_concept', {
gov <- "552025209777"
gen <- cog_revenue(gov, 2019L, category = "All Categories",
revenue_concept = "general")
tot <- cog_revenue(gov, 2019L, category = "All Categories",
revenue_concept = "total")
expect_setequal(unique(gen$category), "All Categories")
expect_gte(sum(tot$amt_nominal), sum(gen$amt_nominal))
})
test_that('"All Categories" cannot be combined with another category', {
expect_error(
cog_spending("552025209777", 2019L, category = c("All Categories", "Police")),
class = "uscogdata_all_categories_not_combinable"
)
})
test_that('"All Categories" is recorded in provenance', {
r <- cog_spending("552025209777", 2019L, category = "All Categories")
expect_identical(cog_explain(r, format = "list")$category, "All Categories")
})
test_that('"All Categories" combines with subtype to give operating totals', {
gov <- "552025209777"
ops_by_cat <- cog_spending(gov, 2019L)
ops_by_cat <- ops_by_cat[ops_by_cat$spend_subtype == "operations", ]
ops_total <- cog_spending(gov, 2019L, category = "All Categories")
ops_total <- ops_total[ops_total$spend_subtype == "operations", ]
expect_equal(sum(ops_total$amt_nominal), sum(ops_by_cat$amt_nominal),
tolerance = 1e-8)
})
test_that('cog_categories() advertises "All Categories" for both flows', {
all <- cog_categories()
rows <- all[all$category == "All Categories", ]
expect_setequal(rows$category_type, c("expenditure", "revenue"))
expect_true(all(is.na(rows$subtype)))
expect_true(all(is.na(rows$n_codes)))
})
test_that('cog_categories(type=) still scopes, including the pseudo-category', {
sp <- cog_categories(type = "spending")
expect_setequal(unique(sp$category_type), "expenditure")
expect_true("All Categories" %in% sp$category)
rev <- cog_categories(type = "revenue")
expect_setequal(unique(rev$category_type), "revenue")
expect_true("All Categories" %in% rev$category)
# balances have no concept vocabulary, so no pseudo-category
bal <- cog_categories(type = "balance")
expect_false("All Categories" %in% bal$category)
})
test_that('cog_categories(pattern=) matches the pseudo-category', {
hit <- cog_categories(pattern = "^All Categories$")
expect_equal(nrow(hit), 2L)
})
# --- final whole-branch review fixes ---------------------------------------
test_that('complete = TRUE is refused when combined with "All Categories"', {
# .completion_grid_sql() would emit `AND c.category IN ('All Categories')`,
# match zero crosswalk rows, and the early return in .complete_result()
# would stamp completion$applied = TRUE, rows_filled = 0 -- reading as "the
# grid was checked and nothing was missing" when nothing was actually
# checked. Filling a summed row has no defined semantics, so the verb must
# refuse the combination outright (finding 2).
expect_error(
cog_spending("552025209777", 2019L, category = "All Categories",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
expect_error(
cog_revenue("552025209777", 2019L, category = "All Categories",
complete = TRUE),
class = "uscogdata_complete_unsupported"
)
})
test_that('cog_balances() rejects "All Categories" instead of silently returning zero rows', {
# cog_balances() reuses .validate_verb_inputs() but did not pass
# allow_all_categories = TRUE, so "All Categories" used to become
# `AND category IN ('All Categories')` against balance_annotated -- 0
# matching crosswalk rows, 0 rows back, no error (finding 3). Holdings are
# a stock with no concept vocabulary to sum across, so the honest answer is
# to refuse, the same way cog_spending()/cog_revenue() refuse other
# nonsensical combinations.
expect_error(
cog_balances("552025209777", 2019L, category = "All Categories"),
class = "uscogdata_all_categories_unsupported"
)
# An ordinary category still works -- this is not a blanket regression.
r <- suppressMessages(
cog_balances("552025209777", 2019L, category = "Fund Balances")
)
expect_gt(nrow(r), 0L)
})
test_that('expenditure_concept_direct_suppressed is NA, not FALSE, when categories are collapsed', {
# .detect_direct_suppressed() keys on
# paste(year, canonical_govid, category, sep = "\r"). In all-categories
# mode every row carries the literal "All Categories" value, so an IG-only
# row's key collides with any ordinary Direct row for the same
# (year, govid) -- has_direct reads TRUE whenever the government has ANY
# direct spending at all, candidate is always empty, and the detector can
# never fire. Before the fix this silently reported FALSE, an affirmative
# claim the code did not actually compute (finding 1). NA is the honest
# answer: cog_explain(x, format = "list") is required here, since without
# format = "list" it returns the result tibble, not the provenance list.
gov <- "552025209777"
t <- cog_spending(gov, 2019L, category = "All Categories",
expenditure_concept = "total")
prov <- cog_explain(t, format = "list")
expect_true(is.na(prov$expenditure_concept_direct_suppressed))
expect_false(isTRUE(prov$expenditure_concept_direct_suppressed))
expect_match(prov$expenditure_concept_note, "unavailable", fixed = TRUE)
# A per-category "total" query on the same government/year is unaffected --
# the detector can still key correctly and reports a strict logical.
t_by_cat <- cog_spending(gov, 2019L, expenditure_concept = "total")
prov_by_cat <- cog_explain(t_by_cat, format = "list")
expect_false(is.na(prov_by_cat$expenditure_concept_direct_suppressed))
})
test_that('"All Categories" still signposts coverage gaps (finding 6, final whole-branch review)', {
# .build_suggestions()'s candidate sub-select used to be keyed on
# `category`, e.g. `WHERE category IN ('All Categories')`. Since
# .ALL_CATEGORIES is never itself a row in summary_categories.category,
# that sub-select always came back empty in all-categories mode, so
# `candidates` was empty and .build_suggestions() short-circuited to
# list() -- coverage signposting was structurally impossible for the one
# mode whose whole selling point is "you cannot sum the wrong scope"
# (uscogdata#9's entire point, silently defeated).
#
# AL state government, FY2011, category = "Corrections": this category has
# no legacy leaf rows in FY2011 (aggregate-flagged E04/E05 family), so the
# per-category query returns 0 rows and 3 recipe-hint suggestions fire
# (empty_year path). All-categories mode does not have an empty year --
# the government has other primary spending in FY2011 -- but the same
# suppressed Corrections dollars are still excluded from the summed total,
# so the fix (scoping the candidate sub-select by subtype_col/subtype_scope
# instead of by category, symmetric with .build_verb_sql()) must still
# surface them via the suppressed_component path.
gov <- "010000226085"
by_cat <- suppressMessages(cog_spending(gov, 2011L, category = "Corrections"))
sugg_by_cat <- cog_explain(by_cat, format = "list")$suggestions
expect_gt(length(sugg_by_cat), 0L)
all_cat <- suppressMessages(cog_spending(gov, 2011L, category = "All Categories"))
sugg_all_cat <- cog_explain(all_cat, format = "list")$suggestions
expect_gt(length(sugg_all_cat), 0L)
# The same Corrections recipe that fired per-category must also fire in
# all-categories mode -- not just some unrelated recipe.
ids_by_cat <- vapply(sugg_by_cat, function(s) s$recipe_id %||% "", character(1))
ids_all_cat <- vapply(sugg_all_cat, function(s) s$recipe_id %||% "", character(1))
expect_true("corrections_combined" %in% ids_by_cat)
expect_true("corrections_combined" %in% ids_all_cat)
# In all-categories mode the government DOES have other primary spending
# in FY2011 (the year itself is not a gap), so the suggestion can only have
# fired via the suppressed_component path, not empty_year.
corr_all <- sugg_all_cat[[which(ids_all_cat == "corrections_combined")]]
expect_identical(corr_all$trigger, "suppressed_component")
expect_gt(corr_all$suppressed_amount, 0)
})
test_that('"All Categories" candidate scoping is symmetric with .build_verb_sql() -- subtype, not category', {
# Direct assertion on the mechanism itself (finding 6): in all-categories
# mode .build_suggestions() must scope its candidate recipe sub-select by
# subtype_col/subtype_scope, not by the literal "All Categories" value.
# Passing all_categories = FALSE with the identical category value proves
# the branch -- not merely the subtype_col/subtype_scope arguments' mere
# presence -- is what changes the query.
con <- uscogdata:::.ensure_session()
none <- uscogdata:::.build_suggestions(
con, govid = "010000226085", years = 2011L,
category = "All Categories", result = NULL, basis = "harmonized",
flow_prefixes = c("E", "F", "G"),
long_view = "spending_long_harmonized",
all_categories = FALSE,
subtype_col = "spend_subtype",
subtype_scope = c("operations", "capital", "assistance")
)
expect_length(none, 0L)
scoped <- uscogdata:::.build_suggestions(
con, govid = "010000226085", years = 2011L,
category = "All Categories", result = NULL, basis = "harmonized",
flow_prefixes = c("E", "F", "G"),
long_view = "spending_long_harmonized",
all_categories = TRUE,
subtype_col = "spend_subtype",
subtype_scope = c("operations", "capital", "assistance")
)
expect_gt(length(scoped), 0L)
})
+10 -2
View File
@@ -29,7 +29,9 @@ test_that("cog_categories(type = 'spending') returns only expenditure rows", {
# joined with the I/Q/Y flow batch -- the last two characters of Census's
# expenditure taxonomy. `interest` is what makes the three-concept model
# computable: primary = direct minus debt service.
expect_true(all(r$subtype %in%
# Exclude pseudo-category which has NA for subtype
r_crosswalk <- r[r$category != "All Categories", ]
expect_true(all(r_crosswalk$subtype %in%
c("operations", "capital", "intergovernmental", "assistance",
"interest", "insurance_benefits")))
})
@@ -54,7 +56,9 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
# (A90) revenue by definition, which is what makes both of its published
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
expect_true(all(r$subtype %in%
# Exclude pseudo-category which has NA for subtype
r_crosswalk <- r[r$category != "All Categories", ]
expect_true(all(r_crosswalk$subtype %in%
c("own_source", "federal", "state", "local_aid",
"insurance_trust", "utility", "liquor_store")))
})
@@ -69,6 +73,8 @@ test_that("cog_categories(pattern = ...) filters case-insensitively", {
test_that("cog_categories has one row per (category, subtype)", {
skip_if_no_corpus()
r <- cog_categories()
# Exclude pseudo-category which is not a crosswalk entry
r <- r[r$category != "All Categories", ]
key <- paste(r$category, r$subtype, sep = "|")
expect_equal(length(key), length(unique(key)))
})
@@ -76,6 +82,8 @@ test_that("cog_categories has one row per (category, subtype)", {
test_that("cog_categories item_codes is non-empty comma-separated string", {
skip_if_no_corpus()
r <- cog_categories()
# Exclude pseudo-category which has NA for n_codes and item_codes
r <- r[r$category != "All Categories", ]
expect_true(all(nzchar(r$item_codes)))
expect_true(all(r$n_codes >= 1L))
# n_codes should equal count of commas + 1
+29
View File
@@ -0,0 +1,29 @@
# Mirror of test-spending-pagination.R for cog_revenue(), which shares the
# same .verb_spendrev()/.build_verb_sql() pushdown -- see that file for the
# incident this fixes.
test_that("cog_revenue limit/offset page correctly and report total_rows", {
skip_if_no_corpus()
full <- cog_revenue("121011212191", years = 2019:2020, category = NULL)
page <- cog_revenue("121011212191", years = 2019:2020, category = NULL,
limit = 5L, offset = 3L)
expect_equal(nrow(page), 5L)
expect_equal(page[c("year", "canonical_govid", "revenue_subtype", "category")],
full[4:8, c("year", "canonical_govid", "revenue_subtype", "category")],
ignore_attr = TRUE)
expect_equal(attr(page, "total_rows"), nrow(full))
})
test_that("cog_revenue limit unset by default leaves total_rows absent", {
skip_if_no_corpus()
r <- cog_revenue("121011212191", 2020L, "Property Tax")
expect_null(attr(r, "total_rows"))
})
test_that("cog_revenue complete + limit conflict aborts the same way as cog_spending", {
skip_if_no_corpus()
expect_error(
cog_revenue("121011212191", 2020L, "Property Tax", complete = TRUE, limit = 5L),
class = "uscogdata_complete_pagination_conflict"
)
})
+95
View File
@@ -0,0 +1,95 @@
# cog-api's paginate() used to slice an already-fully-materialized result:
# every page of a deep sweep re-ran the whole query and re-listified every
# row, just to keep 1000 and discard the rest. For a 193,105-row fleet-wide
# query walked 194 pages deep, that repeated the full cost 194 times and
# wedged the production server for hours (2026-08-06 incident). limit/offset
# here push the slice into the SQL itself, so a page costs O(limit), not
# O(full result).
test_that("limit without offset returns the first page, matching the unpaginated head", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
limit = 10L)
expect_equal(nrow(page), 10L)
expect_equal(page[c("year", "canonical_govid", "spend_subtype", "category")],
full[1:10, c("year", "canonical_govid", "spend_subtype", "category")],
ignore_attr = TRUE)
})
test_that("offset skips ahead without gaps or overlap", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
page2 <- cog_spending("121011212191", years = 2019:2020, category = NULL,
limit = 10L, offset = 10L)
expect_equal(nrow(page2), 10L)
expect_equal(page2[c("year", "canonical_govid", "spend_subtype", "category")],
full[11:20, c("year", "canonical_govid", "spend_subtype", "category")],
ignore_attr = TRUE)
})
test_that("walking every page reconstructs the unpaginated result exactly", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
n <- nrow(full)
limit <- 7L
pages <- list()
offset <- 0L
repeat {
p <- cog_spending("121011212191", years = 2019:2020, category = NULL,
limit = limit, offset = offset)
if (nrow(p) == 0L) break
pages[[length(pages) + 1L]] <- p
offset <- offset + limit
if (offset > n + limit) stop("test runaway: paging did not terminate")
}
walked <- dplyr::bind_rows(pages)
expect_equal(nrow(walked), n)
key_cols <- c("year", "canonical_govid", "spend_subtype", "category", "amt_nominal")
expect_equal(walked[key_cols], full[key_cols], ignore_attr = TRUE)
})
test_that("total_rows attribute reports the full unpaginated count", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
limit = 5L, offset = 0L)
expect_equal(attr(page, "total_rows"), nrow(full))
})
test_that("offset past the end returns zero rows, not an error", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
limit = 10L, offset = nrow(full) + 100L)
expect_equal(nrow(page), 0L)
expect_equal(attr(page, "total_rows"), nrow(full))
})
test_that("limit is unset by default -- unpaginated calls are unaffected", {
skip_if_no_corpus()
r <- cog_spending("121011212191", 2020L, "Corrections")
expect_null(attr(r, "total_rows"))
})
test_that("per_capita and adjust_to_year still apply correctly within a page", {
skip_if_no_corpus()
full <- cog_spending("121011212191", years = 2020L, category = NULL,
per_capita = TRUE, adjust_to_year = 2022L)
page <- cog_spending("121011212191", years = 2020L, category = NULL,
per_capita = TRUE, adjust_to_year = 2022L,
limit = 3L, offset = 2L)
expect_equal(page[c("amt_nominal", "amt_real", "amt_per_capita_nominal",
"amt_per_capita_real")],
full[3:5, c("amt_nominal", "amt_real", "amt_per_capita_nominal",
"amt_per_capita_real")],
ignore_attr = TRUE)
})
test_that("complete = TRUE with limit aborts -- pagination over a partial grid is undefined", {
skip_if_no_corpus()
expect_error(
cog_spending("121011212191", 2020L, "Corrections", complete = TRUE, limit = 5L),
class = "uscogdata_complete_pagination_conflict"
)
})