Files
jared 0fbae00e27
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
feat: name a cohort by state/type predicate instead of a 40k-id IN list
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00

171 lines
8.2 KiB
R

% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/revenue.R
\name{cog_revenue}
\alias{cog_revenue}
\title{Summarized revenue by category}
\usage{
cog_revenue(
govid = NULL,
years,
category = NULL,
per_capita = FALSE,
adjust_to_year = NULL,
basis = c("harmonized", "raw"),
recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE,
limit = NULL,
offset = NULL,
state = NULL,
type = NULL
)
}
\arguments{
\item{govid}{Character vector of `canonical_govid` values, or `NULL` to name
the cohort by `state`/`type` instead. One of `govid`, `state`, or `type`
is required.}
\item{years}{Integer vector of years.}
\item{category}{Character vector of category names (from
`summary_categories.category`), or `NULL` for all categories broken out
one row each. The reserved value `"All Categories"` instead returns a
single summed row per `(year, canonical_govid, subtype)`, covering every
category inside the requested concept's subtype scope. It cannot be
combined with other category names, and it is not the same thing as
`revenue_concept = "total"`: the concept chooses which subtypes are in
scope, `"All Categories"` chooses whether rows inside that scope are
broken out or summed. Because the result keeps one row per
`revenue_subtype`, filtering the returned frame to
`revenue_subtype == "own_source"` gives an own-source revenue total.}
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
Census F-33 population from `gov_population_yearly`. Result also gains
a `pop_source` column with values `"census_f33"` or `"unavailable"`
(the latter for gov types 4/5 and any row whose population is missing
in that year).}
\item{adjust_to_year}{Integer base year for CPI-U real-dollar conversion,
or `NULL` for nominal only.}
\item{basis}{`"harmonized"` (default) sums item codes through the
cross-vintage harmonization mapping (folding series-break-affected
codes onto a comparable target and excluding aggregate / discontinued
rows -- see the `harmonization` block in `cog_explain()`); `"raw"`
reproduces the pre-Phase-R2 behavior (published item codes, no
folding). On a corpus with `schema_version < 5` (no harmonization
tables), `basis` silently resolves to `"raw"` when left at its default
and the resolution is recorded in the provenance; explicitly passing
`basis = "harmonized"` on such a corpus aborts. Ignored when `recipe`
is set (see below).}
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]) for
multi-code cross-vintage series that a 1:1 harmonized_code mapping
can't express (e.g. a wide-era aggregate that only splits into leaf
codes in the modern era). Mutually exclusive with `category`. The
result's subtype column reads `"recipe"` and `category` reads the
recipe's label. Requires `schema_version >= 5`. A recipe query bypasses
`basis` entirely (it joins `long` directly rather than going through
the `*_annotated`/`*_annotated_harmonized` views), so the `basis`
argument is ignored and the result's provenance reports
`basis = "recipe"` with an inert `harmonization` block (`applied =
FALSE`, pointing at the `recipe` block instead) rather than a
possibly-misleading `"harmonized"`/`"raw"` value.}
\item{revenue_concept}{Which of Census's two published revenue concepts to
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
values -- never as item-code first letters, which cannot classify
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
prefix `X` does the same):
* `"general"` (default) -- Census General Revenue: `own_source` +
`federal` + `state` + `local_aid`. The manual defines this concept by
subtraction (section 4.3: *"General revenue comprises all revenue
except that classified as liquor store, utility, or insurance trust
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
insurance trust revenue are all excluded.
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
`general` plus utility, liquor store, and insurance trust revenue
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
`X01`/`X02`/`X05`/`X08`).
The two are related by Census's own identity, `Total Revenue = General +
Utility + Liquor Store + Insurance Trust`.
Note that the employee-retirement (`X`) codes stop at FY2016, when those
systems moved out of the annual finance file into the separate Annual
Survey of Public Pensions, so a `"total"` series steps down at the
FY2016/FY2017 seam for reasons that are about collection scope rather
than revenue (series breaks `SB197`-`SB202`).}
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
corpus does not carry still appears, labelled with **why** it is
missing, and add a `value_source` column to every row:
* `"reported"` — the corpus carries this cell.
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
Census published `$0`. `amt_nominal` is `0`.
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
government did not report, and the value is unknown. `amt_nominal` is
`NA`, **not** `0` — writing a zero there would invent data.
The grid comes from the corpus's `code_set` table, scoped to each
government's own type, so a county is never filled with cells only a
state can report. Reported rows are passed through untouched.
Defaults to `FALSE` (the historical behaviour: absent cells simply do
not appear). Needs a corpus published from 2026-07-29 onward, which is
when `representation`/`code_set` began shipping; aborts with class
`uscogdata_representation_unavailable` otherwise. Not available with
`recipe` or with `expenditure_concept = "total"` (class
`uscogdata_complete_unsupported`) — neither draws its cells from
`code_set`.}
\item{limit}{Maximum number of result rows to return, pushed into the SQL
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
result is materialized. `NULL` (the default) returns every matching row,
exactly as before this parameter existed. Mutually exclusive with
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
\item{offset}{Rows to skip before `limit` starts counting (0-based).
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
\item{state, type}{Name the cohort by predicate instead of by id: `state` is
a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
`"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
the same vocabulary, and the same internal coercion, as
[cog_gov_search()]. Both default to `NULL`.
The cohort is then expressed as a subquery against `canonical_fips_xwalk`
inside each statement rather than round-tripped through R as a literal id
list. For a fleet-scale cohort that is the difference between a
301,591-character `IN` list re-parsed in 5--8 statements per call and a
constant-size predicate: measured at **94 ms versus 449 ms** for the same
FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
7% of the no-filter floor.
Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
governments in `govid` that also match the predicate -- rather than one
silently taking precedence. Naming no cohort at all (`govid`, `state` and
`type` all `NULL`) aborts with class `uscogdata_no_cohort`.
When the cohort is named by predicate, `provenance$scope$govids_found`
and `govids_missing` are empty -- there is no id list to report against --
and `provenance$scope$cohort` carries `state`, `type` and
`n_governments` instead. A `govid`-named cohort reports exactly as before.}
}
\value{
Tibble with columns `year`, `canonical_govid`, `gov_name`,
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
and `value_source` when `complete = TRUE`.
}
\description{
Mirror of [cog_spending()] for revenue categories. One row per
`(year, canonical_govid, revenue_subtype, category)`. Amounts are returned
in **full U.S. dollars** (raw Census values are in $1,000s; this verb
multiplies by 1000 and records the conversion in `provenance`).
}