cog-api's paginate() sliced an already-fully-materialized result: every page of a deep sweep re-ran the whole cog_spending()/cog_revenue() query and re-listified every row, just to keep up to 1000 and discard the rest. A 193,105-row/194-page fleet-wide sweep (cog_explorer's Southern guide, corpus summary build) repeated that full cost 194 times and wedged the production server for hours on 2026-08-06 -- single request, CPU-bound, single-threaded plumber process, no other request could get through, not even /health. cog_spending()/cog_revenue() gain optional limit/offset, pushed into .build_verb_sql() as SQL LIMIT/OFFSET behind the existing (already deterministic) ORDER BY. The full unpaginated row count rides along via COUNT(*) OVER() in the same scan -- exposed as a total_rows attribute -- so a caller walking pages never needs a second round trip to ask how many there are. A page now costs O(limit), not O(full result). Mutually exclusive with complete = TRUE (which fills a grid over the FULL requested (year, category) space -- pagination over a partial slice of already-grouped rows has no defined meaning for the cells it would fill) and with recipe (whose result comes from a separate, not-yet-wired query path). Both abort with a clear classed condition rather than silently ignoring the parameter. limit/offset default to NULL; every existing call site is unaffected.
171 lines
8.5 KiB
R
171 lines
8.5 KiB
R
% Generated by roxygen2: do not edit by hand
|
|
% Please edit documentation in R/spending.R
|
|
\name{cog_spending}
|
|
\alias{cog_spending}
|
|
\title{Summarized spending by category}
|
|
\usage{
|
|
cog_spending(
|
|
govid,
|
|
years,
|
|
category = NULL,
|
|
per_capita = FALSE,
|
|
adjust_to_year = NULL,
|
|
basis = c("harmonized", "raw"),
|
|
recipe = NULL,
|
|
expenditure_concept = c("primary", "direct", "total"),
|
|
complete = FALSE,
|
|
limit = NULL,
|
|
offset = NULL
|
|
)
|
|
}
|
|
\arguments{
|
|
\item{govid}{Character vector of `canonical_govid` values.}
|
|
|
|
\item{years}{Integer vector of years.}
|
|
|
|
\item{category}{Character vector of category names (from
|
|
`summary_categories.category`), or `NULL` for all categories broken out
|
|
one row each. The reserved value `"All Categories"` instead returns a
|
|
single summed row per `(year, canonical_govid, subtype)`, covering every
|
|
category inside the requested concept's subtype scope. It cannot be
|
|
combined with other category names, and it is not the same thing as
|
|
`expenditure_concept = "total"`: the concept chooses which subtypes are in
|
|
scope, `"All Categories"` chooses whether rows inside that scope are
|
|
broken out or summed. Because the result keeps one row per
|
|
`spend_subtype`, filtering the returned frame to
|
|
`spend_subtype == "operations"` gives an operating-expenditure total.}
|
|
|
|
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
|
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
|
Census F-33 population from `gov_population_yearly`. Result also gains
|
|
a `pop_source` column with values `"census_f33"` or `"unavailable"`
|
|
(the latter for gov types 4/5 and any row whose population is missing
|
|
in that year).}
|
|
|
|
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
|
|
or `NULL` for nominal only.}
|
|
|
|
\item{basis}{`"harmonized"` (default) sums item codes through the
|
|
cross-vintage harmonization mapping (folding series-break-affected
|
|
codes onto a comparable target and excluding aggregate / discontinued
|
|
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
|
|
reproduces the pre-Phase-R2 behavior (published item codes, no
|
|
folding). On a corpus with `schema_version < 5` (no harmonization
|
|
tables), `basis` silently resolves to `"raw"` when left at its default
|
|
and the resolution is recorded in the provenance; explicitly passing
|
|
`basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
|
|
is set (see below).}
|
|
|
|
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
|
|
multi-code cross-vintage series that a 1:1 harmonized_code mapping
|
|
can't express (e.g. a wide-era aggregate that only splits into leaf
|
|
codes in the modern era). Mutually exclusive with `category`. The
|
|
result's subtype column reads `"recipe"` and `category` reads the
|
|
recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
|
|
`basis` entirely (it joins `long` directly rather than going through
|
|
the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
|
|
argument is ignored and the result's provenance reports
|
|
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
|
FALSE`, pointing at the `recipe` block instead) rather than a
|
|
possibly-misleading `"harmonized"`/`"raw"` value.}
|
|
|
|
\item{expenditure_concept}{Which spending concept to return. Concepts are
|
|
defined as sets of the crosswalk's `spend_subtype` values -- never as
|
|
item-code first letters, which cannot classify correctly (prefix `Y`
|
|
alone spans revenue, expenditure, and balance codes):
|
|
|
|
* `"primary"` (default) -- the government's own service provision:
|
|
`operations` + `capital` + `assistance` subtypes.
|
|
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
|
|
`interest` (interest on debt) and `insurance_benefits` (insurance
|
|
trust benefit payments, e.g. pensions -- Census manual section
|
|
5.2.2.1 includes payments to retirees in Direct).
|
|
* `"total"` -- `direct` plus the intergovernmental leg: payments to
|
|
local governments (`M` codes), to the state government (`L` codes,
|
|
excluding the `L--` family-total rollup), and state payments to
|
|
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
|
|
`spend_subtype == "intergovernmental"`. Requires the active corpus's
|
|
`summary_categories` to carry M/L rows (added by cog_pipeline PR
|
|
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
|
|
older corpus rather than silently under-reporting. Mutually
|
|
exclusive with `recipe` (a recipe already defines its own component
|
|
codes).
|
|
|
|
**Do not sum `"total"` results across levels of government** (e.g.
|
|
state + county + city): a state's `M12` payment to a school district is
|
|
the same dollar the district reports as its own direct `E12`, so
|
|
summing both double-counts it. This matters in particular with
|
|
[cog_geographic_rollup()], which sums across exactly that kind of
|
|
multi-layer government set.
|
|
|
|
In the legacy wide era (<= FY2011), some functions are published ONLY
|
|
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
|
split), which the Direct leg excludes by construction but the IG leg
|
|
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
|
query, any (year, category) where this leaves intergovernmental rows
|
|
with NO Direct counterpart is flagged: the affected rows' `notes`
|
|
name the harmonization recipe that recovers the missing Direct
|
|
component (when one exists), and
|
|
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
|
figure in those rows is the intergovernmental leg alone, not Direct +
|
|
IG. When `category = "All Categories"` is combined with
|
|
`expenditure_concept = "total"`, this detection cannot run (it keys on
|
|
per-category rows, which all-categories mode collapses to one literal
|
|
value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
|
|
possibly-false `FALSE`; query an explicit `category` to get a real
|
|
answer.}
|
|
|
|
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
|
corpus does not carry still appears, labelled with **why** it is
|
|
missing, and add a `value_source` column to every row:
|
|
|
|
* `"reported"` — the corpus carries this cell.
|
|
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
|
Census published `$0`. `amt_nominal` is `0`.
|
|
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
|
government did not report, and the value is unknown. `amt_nominal` is
|
|
`NA`, **not** `0` — writing a zero there would invent data.
|
|
|
|
The grid comes from the corpus's `code_set` table, scoped to each
|
|
government's own type, so a county is never filled with cells only a
|
|
state can report. Reported rows are passed through untouched.
|
|
|
|
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
|
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
|
when `representation`/`code_set` began shipping; aborts with class
|
|
`uscogdata_representation_unavailable` otherwise. Not available with
|
|
`recipe` or with `expenditure_concept = "total"` (class
|
|
`uscogdata_complete_unsupported`) — neither draws its cells from
|
|
`code_set`.}
|
|
|
|
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
|
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
|
result is materialized. `NULL` (the default) returns every matching row,
|
|
exactly as before this parameter existed. Mutually exclusive with
|
|
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
|
|
|
|
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
|
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
|
}
|
|
\value{
|
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
|
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
|
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
|
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
|
and `value_source` when `complete = TRUE`.
|
|
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
|
whose `completion` block reports `applied`, `rows_filled`, and the
|
|
per-year `absence_means` rule that was applied. When `limit` is set,
|
|
also carries a `total_rows` attribute: the full unpaginated row count,
|
|
computed by the same query (`COUNT(*) OVER()`) rather than a second
|
|
round trip -- so a caller walking pages never has to ask "how many are
|
|
there" separately.
|
|
}
|
|
\description{
|
|
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
|
returned in **full U.S. dollars** (the raw corpus stores them in $1,000s;
|
|
this verb multiplies by 1000 so downstream code can freely rescale to
|
|
millions/billions). The conversion is recorded in the provenance attribute
|
|
under `transformations$units_conversion`.
|
|
}
|