Files
uscogdata/R/revenue.R
T
jared 0fbae00e27
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
feat: name a cohort by state/type predicate instead of a 40k-id IN list
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00

84 lines
4.2 KiB
R

# R/revenue.R
#' Summarized revenue by category
#'
#' Mirror of [cog_spending()] for revenue categories. One row per
#' `(year, canonical_govid, revenue_subtype, category)`. Amounts are returned
#' in **full U.S. dollars** (raw Census values are in $1,000s; this verb
#' multiplies by 1000 and records the conversion in `provenance`).
#'
#' @inheritParams cog_spending
#' @param category Character vector of category names (from
#' `summary_categories.category`), or `NULL` for all categories broken out
#' one row each. The reserved value `"All Categories"` instead returns a
#' single summed row per `(year, canonical_govid, subtype)`, covering every
#' category inside the requested concept's subtype scope. It cannot be
#' combined with other category names, and it is not the same thing as
#' `revenue_concept = "total"`: the concept chooses which subtypes are in
#' scope, `"All Categories"` chooses whether rows inside that scope are
#' broken out or summed. Because the result keeps one row per
#' `revenue_subtype`, filtering the returned frame to
#' `revenue_subtype == "own_source"` gives an own-source revenue total.
#' @param revenue_concept Which of Census's two published revenue concepts to
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
#' values -- never as item-code first letters, which cannot classify
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
#' prefix `X` does the same):
#'
#' * `"general"` (default) -- Census General Revenue: `own_source` +
#' `federal` + `state` + `local_aid`. The manual defines this concept by
#' subtraction (section 4.3: *"General revenue comprises all revenue
#' except that classified as liquor store, utility, or insurance trust
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
#' insurance trust revenue are all excluded.
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
#' `general` plus utility, liquor store, and insurance trust revenue
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
#' `X01`/`X02`/`X05`/`X08`).
#'
#' The two are related by Census's own identity, `Total Revenue = General +
#' Utility + Liquor Store + Insurance Trust`.
#'
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
#' systems moved out of the annual finance file into the separate Annual
#' Survey of Public Pensions, so a `"total"` series steps down at the
#' FY2016/FY2017 seam for reasons that are about collection scope rather
#' than revenue (series breaks `SB197`-`SB202`).
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
#' and `value_source` when `complete = TRUE`.
#' @export
cog_revenue <- function(govid = NULL, years, category = NULL,
per_capita = FALSE, adjust_to_year = NULL,
basis = c("harmonized", "raw"), recipe = NULL,
revenue_concept = c("general", "total"),
complete = FALSE, limit = NULL, offset = NULL,
state = NULL, type = NULL) {
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
# membership does -- General Revenue, i.e. everything except
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
# this verb's recipe families (see R/suggestions.R).
.verb_spendrev(
verb = "cog_revenue",
view_base = "revenue_annotated",
subtype_col = "revenue_subtype",
flow_prefixes = c("T", "A", "U", "B", "C", "D"),
call = match.call(),
govid = govid,
years = years,
category = category,
per_capita = per_capita,
adjust_to_year = adjust_to_year,
basis = basis,
recipe = recipe,
revenue_concept = revenue_concept,
complete = complete,
limit = limit,
offset = offset,
state = state,
type = type
)
}