verbs were left materializing everything and slicing in R -- the pattern behind the 2026-08-06 production incident. cog_gov_search() had no LIMIT at all, so an unfiltered call returns the entire 40,336-row crosswalk. Extracted the #39 machinery into R/pagination.R first (.validate_pagination(), .paginate_sql(), .take_pagination_total()) rather than growing a third inline copy: three definitions of what total_rows means is three places for it to drift. Conflict refusals stay at the call sites because each verb's conflict set differs. .verb_spendrev() now uses the shared helpers and is unchanged in behaviour. The empty-page fallback query is now passed as a thunk, so the unpaginated SQL is only BUILT when an offset actually lands past the end instead of on every paged call. Two things #57 did not anticipate: - cog_gov_search()'s ORDER BY was not a total order. population_acs DESC NULLS LAST leaves ties -- and the whole NULL block -- in scan order, so two requests can order them differently and a paged sweep duplicates one row while dropping another. Added canonical_govid as tiebreaker. Unpaginated output changes only in the relative order of already-tied rows. - Basket mode returns one resolved row per requested name plus a sidecar covering all of them, so a page of it is not a page of anything the caller asked for. Refused with uscogdata_basket_pagination_conflict rather than silently ignoring the arguments. Both default to NULL, so cog-api adopts them behind its existing formals() probe with no lockstep deploy. Suite: 1067 passed, 0 failed, 0 warnings (2 pre-existing live-corpus skips).
123 lines
5.4 KiB
R
123 lines
5.4 KiB
R
% Generated by roxygen2: do not edit by hand
|
|
% Please edit documentation in R/balances.R
|
|
\name{cog_balances}
|
|
\alias{cog_balances}
|
|
\title{Cash and security holdings for one or more governments}
|
|
\usage{
|
|
cog_balances(
|
|
govid = NULL,
|
|
years,
|
|
category = NULL,
|
|
per_capita = FALSE,
|
|
adjust_to_year = NULL,
|
|
basis = c("harmonized", "raw"),
|
|
recipe = NULL,
|
|
state = NULL,
|
|
type = NULL,
|
|
limit = NULL,
|
|
offset = NULL
|
|
)
|
|
}
|
|
\arguments{
|
|
\item{govid}{Canonical govid(s): a character vector, or a data frame with a
|
|
`canonical_govid` column (e.g. from [cog_gov_search()]). `NULL` to name
|
|
the cohort by `state`/`type` instead.}
|
|
|
|
\item{years}{Integer vector of fiscal years.}
|
|
|
|
\item{category}{Optional character vector of categories to keep. One of
|
|
`"Fund Balances"`, `"Insurance Trust Balances"`,
|
|
`"Retirement System Holdings"`. There is deliberately no `subtype`
|
|
argument: for holdings, `category` is a strict coarsening of
|
|
`balance_subtype` (unlike the money verbs, where the two axes cross), so
|
|
every combination would be either redundant or empty.
|
|
`category = "Fund Balances"` is exactly the `general` family
|
|
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
|
one `dplyr::filter()` away. The reserved pseudo-category
|
|
`"All Categories"` (see [cog_spending()]) is **not** supported here and
|
|
errors with class `uscogdata_all_categories_unsupported`: it sums a
|
|
concept's subtype scope, and holdings are a stock with no concept
|
|
vocabulary to sum across. Omit `category` to get every category broken
|
|
out instead.}
|
|
|
|
\item{per_capita}{Divide holdings by population. Note this is a **stock per
|
|
resident** (reserves per person), which is *not* comparable to
|
|
[cog_spending()]'s per-capita figures -- those are a flow per person.}
|
|
|
|
\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).}
|
|
|
|
\item{basis}{Accepted for uniformity with the money verbs, but currently a
|
|
**no-op**: `harmonization_map` carries no balance-code rows, so harmonized
|
|
and raw space are identical for holdings. Reported in
|
|
`provenance$basis_note`.}
|
|
|
|
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]).
|
|
`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
|
|
wide era to the modern one.}
|
|
|
|
\item{state, type}{Name the cohort by predicate instead of by id: `state` is
|
|
a 2-letter USPS abbreviation (or a FIPS code) and `type` is one of
|
|
`"state"`, `"county"`, `"city"`, `"township"` (or the integer `0:3`) --
|
|
the same vocabulary, and the same internal coercion, as
|
|
[cog_gov_search()]. Both default to `NULL`.
|
|
|
|
The cohort is then expressed as a subquery against `canonical_fips_xwalk`
|
|
inside each statement rather than round-tripped through R as a literal id
|
|
list. For a fleet-scale cohort that is the difference between a
|
|
301,591-character `IN` list re-parsed in 5--8 statements per call and a
|
|
constant-size predicate: measured at **94 ms versus 449 ms** for the same
|
|
FY2022 aggregate over the 20,106-government `type = "city"` cohort, within
|
|
7% of the no-filter floor.
|
|
|
|
Supplying `govid` **and** `state`/`type` INTERSECTS them -- the
|
|
governments in `govid` that also match the predicate -- rather than one
|
|
silently taking precedence. Naming no cohort at all (`govid`, `state` and
|
|
`type` all `NULL`) aborts with class `uscogdata_no_cohort`.
|
|
|
|
When the cohort is named by predicate, `provenance$scope$govids_found`
|
|
and `govids_missing` are empty -- there is no id list to report against --
|
|
and `provenance$scope$cohort` carries `state`, `type` and
|
|
`n_governments` instead. A `govid`-named cohort reports exactly as before.}
|
|
|
|
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
|
rather than applied after materializing every row. `NULL` (default)
|
|
returns everything. Cannot be combined with `recipe` -- see `offset` and
|
|
`total_rows`.}
|
|
|
|
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
|
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
|
}
|
|
\value{
|
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
|
`balance_subtype`, `category`, `amt_nominal`, `codes_included`,
|
|
`aggregate_fallback`, plus optional `amt_per_capita_nominal` and
|
|
`pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
|
|
`adjust_to_year` is set), and optional `amt_per_capita_real` (only when
|
|
**both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
|
|
nominal per-capita column to deflate otherwise). Amounts are full US
|
|
dollars.
|
|
|
|
Carries a `provenance` attribute matching
|
|
`inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
|
|
`not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
|
|
every balance subtype in the mounted corpus, not only the observed ones)
|
|
and `truncated` (the observed subtypes whose coverage falls short of the
|
|
requested years). `expenditure_concept`/`revenue_concept` are `NA` --
|
|
holdings are a stock, not a flow, so neither concept vocabulary applies.
|
|
|
|
When `limit` is set, also carries a `total_rows` attribute: the full
|
|
unpaginated row count, computed by the same query (`COUNT(*) OVER()`)
|
|
rather than a second scan.
|
|
}
|
|
\description{
|
|
Returns Census cash-and-security holdings (`category_type = "balance"`):
|
|
fund balances, retirement system holdings and insurance trust balances.
|
|
}
|
|
\section{Holdings are not GAAP fund balance}{
|
|
|
|
Census holdings are **gross** -- no liabilities are netted -- so a reserve
|
|
ratio built from them overstates what is actually available. They are not
|
|
comparable to a GAAP fund balance from an ACFR.
|
|
}
|
|
|