--- title: "Total spending: Primary, Direct, Total, and when each is right" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` # Two questions that sound the same but aren't "Total spending" means two different things depending on whether the question is about one government or several: 1. **"What did my county spend in total, a decade ago vs today?"** — one government, tracked over time. Any concept answers this correctly, as long as the same concept is used for both years. 2. **"How do all the counties in my state compare, a decade ago vs today, against the neighboring state?"** — several governments, summed together. Here only a non-intergovernmental concept (`primary` or `direct`) gives the right answer; summing `total` across governments double-counts money that passes between them. `cog_spending()`'s `expenditure_concept` argument controls which of these a query answers, via three nested concepts defined as sets of the crosswalk's `spend_subtype` values (never item-code first letters — the letter `Y` alone spans revenue, expenditure, and balance codes): - `"primary"` (the default) — the government's own service provision: `operations` + `capital` + `assistance`. - `"direct"` — Census's published Direct Expenditure: `primary` plus `interest` on debt and `insurance_benefits` (e.g. pension payments). - `"total"` — `direct` plus the `intergovernmental` leg. This vignette walks through both questions with code that actually runs against the package's bundled fixture corpus, then explains why the second question refuses `"total"` outright. Before any of the numbers below: every amount column here — `amt_nominal`, `amt_real`, and their `amt_per_capita_*` counterparts — is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus preserves that in its own `amt` column, but the verbs multiply by 1000 on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317 billion. Do not scale it again. ```{r} library(uscogdata) # Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50 # states) so this vignette knits without network access. In real use, # USCOGDATA_URL is instead set to the published corpus URL -- see README.md. Sys.setenv(USCOGDATA_URL = paste0( system.file("extdata/fixture_corpus", package = "uscogdata"), "/" )) ``` The fixture doesn't carry 2017 or the present year, so the examples below use the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published corpus and swap in real years; the mechanics are identical. # Archetype 1: one government's own trend For a single government, `total` is a legitimate way to describe "everything this government spent, including money it handed to other governments to spend on its behalf": ```{r} al_total <- cog_spending( "010000226085", # Alabama, the state government years = c(2012, 2020), category = "Highways", expenditure_concept = "total" ) al_total ``` The `intergovernmental` rows are what `"total"` adds on top of the non-intergovernmental subtypes (here `capital` + `operations`): Alabama's own payments out to counties and cities for highway work. Because this query only ever concerns Alabama, including that piece is safe -- there's no other government's number it could be double-counted against. `"primary"` (the default) answers the same trend question just as validly (for Highways, which maps only to operations/capital codes, `"primary"` and `"direct"` coincide -- there is no highway-specific interest or insurance benefit to add): ```{r} al_primary <- cog_spending( "010000226085", years = c(2012, 2020), category = "Highways" # expenditure_concept = "primary" is the default; shown here for contrast ) al_primary ``` Both are internally consistent series. What breaks the comparison is **switching concepts between the two years being compared** -- e.g. `direct` for 2012 and `total` for 2020 -- which manufactures a trend that isn't really there. Pick one concept for a given question and hold it fixed across every year in the series. # Archetype 2: a cross-government rollup `cog_geographic_rollup()` sums spending across state/county/city layers for a place. Its default is `"primary"`, and (as shown below) it accepts only the non-intergovernmental concepts, `"primary"` and `"direct"`: ```{r} fl_rollup <- cog_geographic_rollup( govids = list( state = "120000226351", # Florida county = c("121011212191", "121099101897") # Broward + Palm Beach ), category = "Highways", years = c(2012, 2020) ) fl_rollup ``` For the neighboring state, the comparison is a single government, so it's a plain `cog_spending()` call rather than a rollup: ```{r} ga_state <- cog_spending( "130000226087", years = c(2012, 2020), category = "Highways" # Georgia ) ga_state ``` Now the same rollup, but asking for `expenditure_concept = "total"`: ```{r, error = TRUE} cog_geographic_rollup( govids = list(state = "120000226351", county = "121011212191"), category = "Highways", years = 2020, expenditure_concept = "total" ) ``` `cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason) refuses `"total"` outright rather than silently returning an inflated number. The next section is why. # The mechanism Suppose Alabama gives a county $10M toward a highway project. That $10M shows up **twice** in the underlying corpus: - Once on Alabama's own record, coded `M44` ("to local governments, Highways") -- Alabama's intergovernmental leg. - Again on the county's record, coded `E44` / `F44` ("Highways, current operations" / "capital outlay") -- the county's direct spending, because the county is the government that actually lets the contract and pays the paving crew. `primary` and `direct` (the crosswalk's non-intergovernmental expenditure subtypes) only ever count the second of those -- the government that actually did the spending. `total` (Direct plus the intergovernmental leg) counts the first one *as well*, which is exactly right for describing Alabama's own budget: Alabama's `total` genuinely includes the $10M it committed to highways, whether it built the road itself or paid the county to. But sum `total` across Alabama **and** the county, and that $10M is counted twice -- once as Alabama's payment out, once as the county's spending in -- reporting $20M of highway work for $10M actually spent. This is exactly the shape of query `cog_geographic_rollup()` exists to run (summing across layers of government), so it refuses `"total"` rather than silently overstating every multi-layer figure it produces. # How big is the risk in practice Intergovernmental transfers aren't evenly distributed by government type. Measured against the bundled fixture corpus (all 50 states, each of its four years -- 2011, 2012, 2019, 2020), intergovernmental spending as a share of a government's own Direct spending is: | Government type | Intergovernmental / Direct | |---|---| | State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) | | County | 3.3%-4.8% (varies by year) | | City | 2.4%-2.9% (varies by year) | So the Direct/Total choice matters overwhelmingly for **state** governments -- a state's Total genuinely differs from its Direct by more than a third, while for a county or city the two are close. (The state share is much larger than pre-#11 measurements suggested, because the intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state payments to school systems -- for most states the single largest transfer they make.) That's also why the mistake this vignette warns about is easy to make unnoticed at the county/city level and costly at the state level: rolling up every government using `total` instead of `primary`/`direct` overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally. # Why Total = Direct + M + L + Q, not Direct + M It's tempting to assume `total` only needs to add `M`. But the intergovernmental leg has three families, all money the queried government itself pays **out** -- they're not different accounts of a receiving government's revenue. `M` is what it pays to other **local** governments (e.g. a county paying a city for a shared paving contract); `L` is what it pays **up** to its **state** government (e.g. a county's contribution to a state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments to **school systems** (K-12 and higher-ed aid -- for most states the single largest transfer they make, and the piece the pre-#11 prefix allowlist silently dropped, finding F-017). A government's Total genuinely includes every leg it pays, because each is its own spending, just routed to a different kind of recipient. On the bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has no "payments to the state government" leg of its own) but is 43%-51% the size of `M` for counties (varies by year) and 144%-189% the size of `M` for cities (varies by year; 166% pooled across all four) -- so a `total` that omitted `L` would silently undercount Total specifically for local governments, and for cities `L` is often the *larger* of the two legs. `cog_spending(expenditure_concept = "total")` includes every leg (excluding the `L--` family-total rollup row, which would double-count its own components). # Composition rules - `expenditure_concept` (whose spending counts -- Primary, Direct, or Direct plus intergovernmental) is **orthogonal** to `basis` (which vintage of the item-code space a query resolves against -- `"harmonized"` vs `"raw"`). They combine freely: `expenditure_concept = "total", basis = "raw"` is a valid, meaningful query, and so is every other pairing. - `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a recipe already defines its own component codes (some recipes have their own matching intergovernmental counterpart recipe instead -- see `cog_recipes()` and the "firing suggestion" notes surfaced in `cog_spending()`'s provenance), so layering a second, generic `total` union on top of a recipe query has no well-defined meaning. Passing both together aborts with an error naming the conflict. - `expenditure_concept` is a **spending-only** concept: `cog_revenue()` doesn't expose it (revenue's own intergovernmental codes are a different axis -- see `?cog_revenue`). # Summary - Comparing one government to itself over time: any concept works -- pick one and hold it fixed across every year compared. - Comparing or summing across governments -- counties within a state, a state against its neighbor, cities against counties: use `"primary"` (the default) or `"direct"`. `cog_geographic_rollup()` and `cog_peer_compare()` enforce this by refusing `"total"`. - `"primary"` = operations + capital + assistance. `"direct"` = primary + interest on debt + insurance trust benefits (Census's published Direct Expenditure). `"total"` = direct + intergovernmental (`M` to local governments, `L` to the state government excluding the `L--` family-total row, and `Q11`/`Q12`/`Q18` state payments to school systems).