Files
uscogdata/man/cog_spending.Rd
T
jared 4de915b557 feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24
recipes / schema_version >= 5), and a recipe= argument on cog_spending()/
cog_revenue() that runs a recipe's generic multi-code join instead of the
category view: SUM(amt * weight) across whichever component codes are
present for a (year, canonical_govid), scoped by gov_type_scope. The join
deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes
these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY
as aggregate rows, with leaf codes first appearing in 2012, so excluding
aggregates would zero out the wide-era half of every recipe. This is safe
by corpus construction: wide-era rows are aggregate-only, modern rows are
leaf-only, and every component is year-scoped, so there is no
double-counting. recipe= is mutually exclusive with category=; the result's
subtype column reads "recipe" and category reads the recipe's label.

Adds recipe-component-driven signposting: when a basis="harmonized" +
category query comes back with zero rows in a requested year, and a
harmonization recipe covering that category would actually produce rows
for this government in that year (via the same join .run_recipe() uses),
the recipe is surfaced in provenance$suggestions plus one
cli::cli_inform() message. This is deliberately keyed off recipe
components rather than harmonization_map's suggested_recipe_id column
(which is empty on every live row -- the wide era's split families are
NA-by-construction via aggregate exclusion, not an NA ruling to hang a
suggestion off of).

Also populates the previously-always-empty provenance$series_break_refs
(schema v5 only: series_breaks_pq rows whose fin_code is among the
observed codes and whose break_year falls in the requested span), and
extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series
breaks sections.
2026-07-18 23:33:02 -04:00

66 lines
2.7 KiB
R

% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/spending.R
\name{cog_spending}
\alias{cog_spending}
\title{Summarized spending by category}
\usage{
cog_spending(
govid,
years,
category = NULL,
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL
)
}
\arguments{
\item{govid}{Character vector of `canonical_govid` values.}
\item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
Census F-33 population from `gov_population_yearly`. Result also gains
a `pop_source` column with values `"census_f33"` or `"unavailable"`
(the latter for gov types 4/5 and any row whose population is missing
in that year).}
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
or `NULL` for nominal only.}
\item{basis}{`"harmonized"` (default) sums item codes through the
cross-vintage harmonization mapping (folding series-break-affected
codes onto a comparable target and excluding aggregate / discontinued
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
reproduces the pre-Phase-R2 behavior (published item codes, no
folding). On a corpus with `schema_version < 5` (no harmonization
tables), `basis` silently resolves to `"raw"` when left at its default
and the resolution is recorded in the provenance; explicitly passing
`basis = "harmonized"` on such a corpus aborts.}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
multi-code cross-vintage series that a 1:1 harmonized_code mapping
can't express (e.g. a wide-era aggregate that only splits into leaf
codes in the modern era). Mutually exclusive with `category`. The
result's subtype column reads `"recipe"` and `category` reads the
recipe's label. Requires `schema_version >= 5`.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
}
\description{
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
returned in **full U.S. dollars** (the raw corpus stores them in $1,000s;
this verb multiplies by 1000 so downstream code can freely rescale to
millions/billions). The conversion is recorded in the provenance attribute
under `transformations$units_conversion`.
}