Task 7 (final) of the expenditure_concept plan. The vignette leads with the two archetype questions -- a single government's own trend (either concept works, held fixed across years) vs a cross-government rollup (direct only, with the refusal error from cog_geographic_rollup() shown and explained) -- walked through with code that runs against the bundled fixture corpus (years 2011/2012/2019/2020, substituting for "2017 vs today"). Explains the double-counting mechanism (a state's M44 payment to a county is the same dollar as the county's own E44/F44), why Total = Direct + M + L rather than Direct + M, and the composition rules (expenditure_concept is orthogonal to basis, mutually exclusive with recipe). README gets a short pointer section with the one-line rule.
218 lines
8.8 KiB
Plaintext
218 lines
8.8 KiB
Plaintext
---
|
|
title: "Total spending: Direct, Total, and when each is right"
|
|
output: rmarkdown::html_vignette
|
|
vignette: >
|
|
%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}
|
|
%\VignetteEngine{knitr::rmarkdown}
|
|
%\VignetteEncoding{UTF-8}
|
|
---
|
|
|
|
```{r setup, include = FALSE}
|
|
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
|
|
```
|
|
|
|
# Two questions that sound the same but aren't
|
|
|
|
"Total spending" means two different things depending on whether the question
|
|
is about one government or several:
|
|
|
|
1. **"What did my county spend in total, a decade ago vs today?"** — one
|
|
government, tracked over time. Either `direct` or `total` spending answers
|
|
this correctly, as long as the same concept is used for both years.
|
|
2. **"How do all the counties in my state compare, a decade ago vs today,
|
|
against the neighboring state?"** — several governments, summed together.
|
|
Here only `direct` gives the right answer; summing `total` across
|
|
governments double-counts money that passes between them.
|
|
|
|
`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`)
|
|
controls which of these a query answers. This vignette walks through both
|
|
questions with code that actually runs against the package's bundled fixture
|
|
corpus, then explains why the second question refuses `"total"` outright.
|
|
|
|
```{r}
|
|
library(uscogdata)
|
|
|
|
# Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50
|
|
# states) so this vignette knits without network access. In real use,
|
|
# USCOGDATA_URL is instead set to the published corpus URL -- see README.md.
|
|
Sys.setenv(USCOGDATA_URL = paste0(
|
|
system.file("extdata/fixture_corpus", package = "uscogdata"), "/"
|
|
))
|
|
```
|
|
|
|
The fixture doesn't carry 2017 or the present year, so the examples below use
|
|
the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs
|
|
today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published
|
|
corpus and swap in real years; the mechanics are identical.
|
|
|
|
# Archetype 1: one government's own trend
|
|
|
|
For a single government, `total` is a legitimate way to describe "everything
|
|
this government spent, including money it handed to other governments to
|
|
spend on its behalf":
|
|
|
|
```{r}
|
|
al_total <- cog_spending(
|
|
"010000226085", # Alabama, the state government
|
|
years = c(2012, 2020),
|
|
category = "Highways",
|
|
expenditure_concept = "total"
|
|
)
|
|
al_total
|
|
```
|
|
|
|
The `intergovernmental` rows are what `"total"` adds on top of `"direct"`
|
|
(`capital` + `operations`): Alabama's own payments out to counties and
|
|
cities for highway work. Because this query only ever concerns Alabama,
|
|
including that piece is safe -- there's no other government's number it
|
|
could be double-counted against.
|
|
|
|
`"direct"` (the default) answers the same trend question just as validly:
|
|
|
|
```{r}
|
|
al_direct <- cog_spending(
|
|
"010000226085", years = c(2012, 2020), category = "Highways"
|
|
# expenditure_concept = "direct" is the default; shown here for contrast
|
|
)
|
|
al_direct
|
|
```
|
|
|
|
Both are internally consistent series. What breaks the comparison is
|
|
**switching concepts between the two years being compared** -- e.g. `direct`
|
|
for 2012 and `total` for 2020 -- which manufactures a trend that isn't
|
|
really there. Pick one concept for a given question and hold it fixed across
|
|
every year in the series.
|
|
|
|
# Archetype 2: a cross-government rollup
|
|
|
|
`cog_geographic_rollup()` sums spending across state/county/city layers for
|
|
a place. Its default -- and, as shown below, its *only* accepted value for
|
|
`expenditure_concept` -- is `"direct"`:
|
|
|
|
```{r}
|
|
fl_rollup <- cog_geographic_rollup(
|
|
govids = list(
|
|
state = "120000226351", # Florida
|
|
county = c("121011212191", "121099101897") # Broward + Palm Beach
|
|
),
|
|
category = "Highways",
|
|
years = c(2012, 2020)
|
|
)
|
|
fl_rollup
|
|
```
|
|
|
|
For the neighboring state, the comparison is a single government, so it's a
|
|
plain `cog_spending()` call rather than a rollup:
|
|
|
|
```{r}
|
|
ga_state <- cog_spending(
|
|
"130000226087", years = c(2012, 2020), category = "Highways" # Georgia
|
|
)
|
|
ga_state
|
|
```
|
|
|
|
Now the same rollup, but asking for `expenditure_concept = "total"`:
|
|
|
|
```{r, error = TRUE}
|
|
cog_geographic_rollup(
|
|
govids = list(state = "120000226351", county = "121011212191"),
|
|
category = "Highways",
|
|
years = 2020,
|
|
expenditure_concept = "total"
|
|
)
|
|
```
|
|
|
|
`cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason)
|
|
refuses `"total"` outright rather than silently returning an inflated
|
|
number. The next section is why.
|
|
|
|
# The mechanism
|
|
|
|
Suppose Alabama gives a county $10M toward a highway project. That $10M
|
|
shows up **twice** in the underlying corpus:
|
|
|
|
- Once on Alabama's own record, coded `M44` ("to local governments,
|
|
Highways") -- Alabama's intergovernmental leg.
|
|
- Again on the county's record, coded `E44` / `F44` ("Highways, current
|
|
operations" / "capital outlay") -- the county's direct spending, because
|
|
the county is the government that actually lets the contract and pays the
|
|
paving crew.
|
|
|
|
`direct` (item codes `E`/`F`/`G`) only ever counts the second of those --
|
|
the government that actually did the spending. `total` (Direct plus the
|
|
`M`/`L` intergovernmental legs) counts the first one *as well*, which is
|
|
exactly right for describing Alabama's own budget: Alabama's `total`
|
|
genuinely includes the $10M it committed to highways, whether it built the
|
|
road itself or paid the county to. But sum `total` across Alabama **and**
|
|
the county, and that $10M is counted twice -- once as Alabama's payment out,
|
|
once as the county's spending in -- reporting $20M of highway work for $10M
|
|
actually spent.
|
|
|
|
This is exactly the shape of query `cog_geographic_rollup()` exists to run
|
|
(summing across layers of government), so it refuses `"total"` rather than
|
|
silently overstating every multi-layer figure it produces.
|
|
|
|
# How big is the risk in practice
|
|
|
|
Intergovernmental transfers aren't evenly distributed by government type.
|
|
Measured on the full published corpus, intergovernmental spending as a
|
|
share of a government's own Direct spending is:
|
|
|
|
| Government type | Intergovernmental / Direct |
|
|
|---|---|
|
|
| State | 17.2% |
|
|
| County | 1.8% |
|
|
| City | 0.8% |
|
|
|
|
So the Direct/Total choice matters overwhelmingly for **state** governments
|
|
-- a state's Total genuinely differs from its Direct by a meaningful margin,
|
|
while for a county or city the two are close. That's also why the mistake
|
|
this vignette warns about is easy to make unnoticed at the county/city level
|
|
and costly at the state level: rolling up every government in a state using
|
|
`total` instead of `direct` overstates the true figure -- measured at 7.6%
|
|
for Alabama in FY2019, and 11.6% nationally.
|
|
|
|
# Why Total = Direct + M + L, not Direct + M
|
|
|
|
It's tempting to assume `total` only needs to add `M` (payments to local
|
|
governments). But not all of the money a county or city receives arrives
|
|
directly from its state as an `M` payment -- some flows through as `L`
|
|
(payments *to* the state government), which the state government then
|
|
redistributes as `M`. On the published corpus, `L` is 0 for state
|
|
governments (a state has no "payments to the state government" leg of its
|
|
own) but is 91.6% the size of `M` for counties and 188.3% the size of `M`
|
|
for cities -- so a `total` that omitted `L` would silently undercount Total
|
|
specifically for local governments. `cog_spending(expenditure_concept =
|
|
"total")` includes both legs (excluding the `L--` family-total rollup row,
|
|
which would double-count its own components).
|
|
|
|
# Composition rules
|
|
|
|
- `expenditure_concept` (whose spending counts -- Direct vs Direct plus
|
|
intergovernmental) is **orthogonal** to `basis` (which vintage of the
|
|
item-code space a query resolves against -- `"harmonized"` vs `"raw"`).
|
|
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
|
|
valid, meaningful query, and so is every other pairing.
|
|
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
|
|
recipe already defines its own component codes (some recipes have their
|
|
own matching intergovernmental counterpart recipe instead -- see
|
|
`cog_recipes()` and the "firing suggestion" notes surfaced in
|
|
`cog_spending()`'s provenance), so layering a second, generic `total`
|
|
union on top of a recipe query has no well-defined meaning. Passing both
|
|
together aborts with an error naming the conflict.
|
|
- `expenditure_concept` is a **spending-only** concept: `cog_revenue()`
|
|
doesn't expose it (revenue's own intergovernmental codes are a different
|
|
axis -- see `?cog_revenue`).
|
|
|
|
# Summary
|
|
|
|
- Comparing one government to itself over time: `"direct"` or `"total"`
|
|
both work -- pick one and hold it fixed across every year compared.
|
|
- Comparing or summing across governments -- counties within a state, a
|
|
state against its neighbor, cities against counties: use `"direct"`.
|
|
`cog_geographic_rollup()` and `cog_peer_compare()` enforce this by
|
|
refusing `"total"`.
|
|
- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local
|
|
governments + `L` to the state government, excluding the `L--`
|
|
family-total row).
|