Files
uscogdata/vignettes/total-spending.Rmd
T
jared 3bd9b1f011 docs: total-spending vignette + README section on Direct vs Total
Task 7 (final) of the expenditure_concept plan. The vignette leads with the
two archetype questions -- a single government's own trend (either concept
works, held fixed across years) vs a cross-government rollup (direct only,
with the refusal error from cog_geographic_rollup() shown and explained) --
walked through with code that runs against the bundled fixture corpus
(years 2011/2012/2019/2020, substituting for "2017 vs today"). Explains the
double-counting mechanism (a state's M44 payment to a county is the same
dollar as the county's own E44/F44), why Total = Direct + M + L rather than
Direct + M, and the composition rules (expenditure_concept is orthogonal to
basis, mutually exclusive with recipe). README gets a short pointer section
with the one-line rule.
2026-07-27 11:23:31 -04:00

218 lines
8.8 KiB
Plaintext

---
title: "Total spending: Direct, Total, and when each is right"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
```{r setup, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```
# Two questions that sound the same but aren't
"Total spending" means two different things depending on whether the question
is about one government or several:
1. **"What did my county spend in total, a decade ago vs today?"** — one
government, tracked over time. Either `direct` or `total` spending answers
this correctly, as long as the same concept is used for both years.
2. **"How do all the counties in my state compare, a decade ago vs today,
against the neighboring state?"** — several governments, summed together.
Here only `direct` gives the right answer; summing `total` across
governments double-counts money that passes between them.
`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`)
controls which of these a query answers. This vignette walks through both
questions with code that actually runs against the package's bundled fixture
corpus, then explains why the second question refuses `"total"` outright.
```{r}
library(uscogdata)
# Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50
# states) so this vignette knits without network access. In real use,
# USCOGDATA_URL is instead set to the published corpus URL -- see README.md.
Sys.setenv(USCOGDATA_URL = paste0(
system.file("extdata/fixture_corpus", package = "uscogdata"), "/"
))
```
The fixture doesn't carry 2017 or the present year, so the examples below use
the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs
today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published
corpus and swap in real years; the mechanics are identical.
# Archetype 1: one government's own trend
For a single government, `total` is a legitimate way to describe "everything
this government spent, including money it handed to other governments to
spend on its behalf":
```{r}
al_total <- cog_spending(
"010000226085", # Alabama, the state government
years = c(2012, 2020),
category = "Highways",
expenditure_concept = "total"
)
al_total
```
The `intergovernmental` rows are what `"total"` adds on top of `"direct"`
(`capital` + `operations`): Alabama's own payments out to counties and
cities for highway work. Because this query only ever concerns Alabama,
including that piece is safe -- there's no other government's number it
could be double-counted against.
`"direct"` (the default) answers the same trend question just as validly:
```{r}
al_direct <- cog_spending(
"010000226085", years = c(2012, 2020), category = "Highways"
# expenditure_concept = "direct" is the default; shown here for contrast
)
al_direct
```
Both are internally consistent series. What breaks the comparison is
**switching concepts between the two years being compared** -- e.g. `direct`
for 2012 and `total` for 2020 -- which manufactures a trend that isn't
really there. Pick one concept for a given question and hold it fixed across
every year in the series.
# Archetype 2: a cross-government rollup
`cog_geographic_rollup()` sums spending across state/county/city layers for
a place. Its default -- and, as shown below, its *only* accepted value for
`expenditure_concept` -- is `"direct"`:
```{r}
fl_rollup <- cog_geographic_rollup(
govids = list(
state = "120000226351", # Florida
county = c("121011212191", "121099101897") # Broward + Palm Beach
),
category = "Highways",
years = c(2012, 2020)
)
fl_rollup
```
For the neighboring state, the comparison is a single government, so it's a
plain `cog_spending()` call rather than a rollup:
```{r}
ga_state <- cog_spending(
"130000226087", years = c(2012, 2020), category = "Highways" # Georgia
)
ga_state
```
Now the same rollup, but asking for `expenditure_concept = "total"`:
```{r, error = TRUE}
cog_geographic_rollup(
govids = list(state = "120000226351", county = "121011212191"),
category = "Highways",
years = 2020,
expenditure_concept = "total"
)
```
`cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason)
refuses `"total"` outright rather than silently returning an inflated
number. The next section is why.
# The mechanism
Suppose Alabama gives a county $10M toward a highway project. That $10M
shows up **twice** in the underlying corpus:
- Once on Alabama's own record, coded `M44` ("to local governments,
Highways") -- Alabama's intergovernmental leg.
- Again on the county's record, coded `E44` / `F44` ("Highways, current
operations" / "capital outlay") -- the county's direct spending, because
the county is the government that actually lets the contract and pays the
paving crew.
`direct` (item codes `E`/`F`/`G`) only ever counts the second of those --
the government that actually did the spending. `total` (Direct plus the
`M`/`L` intergovernmental legs) counts the first one *as well*, which is
exactly right for describing Alabama's own budget: Alabama's `total`
genuinely includes the $10M it committed to highways, whether it built the
road itself or paid the county to. But sum `total` across Alabama **and**
the county, and that $10M is counted twice -- once as Alabama's payment out,
once as the county's spending in -- reporting $20M of highway work for $10M
actually spent.
This is exactly the shape of query `cog_geographic_rollup()` exists to run
(summing across layers of government), so it refuses `"total"` rather than
silently overstating every multi-layer figure it produces.
# How big is the risk in practice
Intergovernmental transfers aren't evenly distributed by government type.
Measured on the full published corpus, intergovernmental spending as a
share of a government's own Direct spending is:
| Government type | Intergovernmental / Direct |
|---|---|
| State | 17.2% |
| County | 1.8% |
| City | 0.8% |
So the Direct/Total choice matters overwhelmingly for **state** governments
-- a state's Total genuinely differs from its Direct by a meaningful margin,
while for a county or city the two are close. That's also why the mistake
this vignette warns about is easy to make unnoticed at the county/city level
and costly at the state level: rolling up every government in a state using
`total` instead of `direct` overstates the true figure -- measured at 7.6%
for Alabama in FY2019, and 11.6% nationally.
# Why Total = Direct + M + L, not Direct + M
It's tempting to assume `total` only needs to add `M` (payments to local
governments). But not all of the money a county or city receives arrives
directly from its state as an `M` payment -- some flows through as `L`
(payments *to* the state government), which the state government then
redistributes as `M`. On the published corpus, `L` is 0 for state
governments (a state has no "payments to the state government" leg of its
own) but is 91.6% the size of `M` for counties and 188.3% the size of `M`
for cities -- so a `total` that omitted `L` would silently undercount Total
specifically for local governments. `cog_spending(expenditure_concept =
"total")` includes both legs (excluding the `L--` family-total rollup row,
which would double-count its own components).
# Composition rules
- `expenditure_concept` (whose spending counts -- Direct vs Direct plus
intergovernmental) is **orthogonal** to `basis` (which vintage of the
item-code space a query resolves against -- `"harmonized"` vs `"raw"`).
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
valid, meaningful query, and so is every other pairing.
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
recipe already defines its own component codes (some recipes have their
own matching intergovernmental counterpart recipe instead -- see
`cog_recipes()` and the "firing suggestion" notes surfaced in
`cog_spending()`'s provenance), so layering a second, generic `total`
union on top of a recipe query has no well-defined meaning. Passing both
together aborts with an error naming the conflict.
- `expenditure_concept` is a **spending-only** concept: `cog_revenue()`
doesn't expose it (revenue's own intergovernmental codes are a different
axis -- see `?cog_revenue`).
# Summary
- Comparing one government to itself over time: `"direct"` or `"total"`
both work -- pick one and hold it fixed across every year compared.
- Comparing or summing across governments -- counties within a state, a
state against its neighbor, cities against counties: use `"direct"`.
`cog_geographic_rollup()` and `cog_peer_compare()` enforce this by
refusing `"total"`.
- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local
governments + `L` to the state government, excluding the `L--`
family-total row).