Files
uscogdata/man/cog_spending.Rd
T
jared 93300ae0c1
R-CMD-check / check (push) Successful in 3m5s
feat: three-concept expenditure model classified by crosswalk membership (#11)
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:

  primary = operations + capital + assistance
  direct  = primary + interest + insurance_benefits   (Census Direct)
  total   = direct + intergovernmental                (M/L/Q via ig views)

- inst/sql: flow views (20-25) select by crosswalk membership;
  summary_categories moves to 11- so it registers before them (DuckDB
  binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
  school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
  harmonization exclusion count, and the complete = TRUE grid;
  flow_prefixes survives only to scope recipe suggestions.
  cog_geographic_rollup/cog_peer_compare accept primary|direct, still
  refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
  (uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
  proof is asserted against the crosswalk, not the default
  cog_revenue() call (which stays General Revenue pending #12).

Suite: 696 pass / 0 fail / 1 skip (#12, expected).

Closes #11
2026-07-30 16:56:50 -04:00

142 lines
6.8 KiB
R

% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/spending.R
\name{cog_spending}
\alias{cog_spending}
\title{Summarized spending by category}
\usage{
cog_spending(
govid,
years,
category = NULL,
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
expenditure_concept = c("primary", "direct", "total"),
complete = FALSE
)
}
\arguments{
\item{govid}{Character vector of `canonical_govid` values.}
\item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
Census F-33 population from `gov_population_yearly`. Result also gains
a `pop_source` column with values `"census_f33"` or `"unavailable"`
(the latter for gov types 4/5 and any row whose population is missing
in that year).}
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
or `NULL` for nominal only.}
\item{basis}{`"harmonized"` (default) sums item codes through the
cross-vintage harmonization mapping (folding series-break-affected
codes onto a comparable target and excluding aggregate / discontinued
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
reproduces the pre-Phase-R2 behavior (published item codes, no
folding). On a corpus with `schema_version < 5` (no harmonization
tables), `basis` silently resolves to `"raw"` when left at its default
and the resolution is recorded in the provenance; explicitly passing
`basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
is set (see below).}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
multi-code cross-vintage series that a 1:1 harmonized_code mapping
can't express (e.g. a wide-era aggregate that only splits into leaf
codes in the modern era). Mutually exclusive with `category`. The
result's subtype column reads `"recipe"` and `category` reads the
recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
`basis` entirely (it joins `long` directly rather than going through
the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{expenditure_concept}{Which spending concept to return. Concepts are
defined as sets of the crosswalk's `spend_subtype` values -- never as
item-code first letters, which cannot classify correctly (prefix `Y`
alone spans revenue, expenditure, and balance codes):
* `"primary"` (default) -- the government's own service provision:
`operations` + `capital` + `assistance` subtypes.
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
`interest` (interest on debt) and `insurance_benefits` (insurance
trust benefit payments, e.g. pensions -- Census manual section
5.2.2.1 includes payments to retirees in Direct).
* `"total"` -- `direct` plus the intergovernmental leg: payments to
local governments (`M` codes), to the state government (`L` codes,
excluding the `L--` family-total rollup), and state payments to
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
`spend_subtype == "intergovernmental"`. Requires the active corpus's
`summary_categories` to carry M/L rows (added by cog_pipeline PR
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
older corpus rather than silently under-reporting. Mutually
exclusive with `recipe` (a recipe already defines its own component
codes).
**Do not sum `"total"` results across levels of government** (e.g.
state + county + city): a state's `M12` payment to a school district is
the same dollar the district reports as its own direct `E12`, so
summing both double-counts it. This matters in particular with
[cog_geographic_rollup()], which sums across exactly that kind of
multi-layer government set.
In the legacy wide era (<= FY2011), some functions are published ONLY
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
split), which the Direct leg excludes by construction but the IG leg
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
query, any (year, category) where this leaves intergovernmental rows
with NO Direct counterpart is flagged: the affected rows' `notes`
name the harmonization recipe that recovers the missing Direct
component (when one exists), and
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
figure in those rows is the intergovernmental leg alone, not Direct +
IG.}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
whose `completion` block reports `applied`, `rows_filled`, and the
per-year `absence_means` rule that was applied.
}
\description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
returned in **full U.S. dollars** (the raw corpus stores them in $1,000s;
this verb multiplies by 1000 so downstream code can freely rescale to
millions/billions). The conversion is recorded in the provenance attribute
under `transformations$units_conversion`.
}