Compare commits
88
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6392a74013
|
||
|
|
de2ba0cfb9 | ||
|
|
e912a926c2 | ||
|
|
44e4953f94
|
||
|
|
a5500f0b6a
|
||
|
|
da2839f885
|
||
|
|
331399ab86
|
||
|
|
5582c6cb57
|
||
|
|
013af5b0d2
|
||
|
|
4a92f36d44
|
||
|
|
a67735f121
|
||
|
|
b9f7f7d8d3
|
||
|
|
4300b636b1
|
||
|
|
99e1e86e37
|
||
|
|
c042ee0b90
|
||
|
|
29dc8199e0
|
||
|
|
619b167ab1
|
||
|
|
dfda39051e
|
||
|
|
785f3af16d
|
||
|
|
c587c8ba87 | ||
|
|
6a06302036
|
||
|
|
e3ab26c3e6 | ||
|
|
498950afa6
|
||
|
|
a5f86d87b3
|
||
|
|
61b9c95731
|
||
|
|
44e9b40b86
|
||
|
|
f1e9aa383a
|
||
|
|
12a9be110f
|
||
|
|
503fa6562f
|
||
|
|
11ae99c382
|
||
|
|
5e22e940e7
|
||
|
|
2fc9e7585b | ||
|
|
8bf9c4ccc1
|
||
|
|
77074621d8
|
||
|
|
4b749205a5
|
||
|
|
f77adb6c83
|
||
|
|
230f3401c4
|
||
|
|
7522b48a08
|
||
|
|
693f8d81a6
|
||
|
|
db35fa9058
|
||
|
|
cabe2e2799
|
||
|
|
6cd219a291 | ||
|
|
e067a5930f
|
||
|
|
342debaefa
|
||
|
|
b59b79b2d5 | ||
|
|
5668d6b102
|
||
|
|
0a6d878a36 | ||
|
|
da726a61f6
|
||
|
|
03c313b46d | ||
|
|
2c532bde19
|
||
|
|
a9e80858d4
|
||
|
|
fde62eb6cc
|
||
|
|
22c2478634
|
||
|
|
225cd60968
|
||
|
|
b03f095e49
|
||
|
|
724b6bd58b
|
||
|
|
82e4face4e
|
||
|
|
6c5bdb3048
|
||
|
|
b8189aeb7f
|
||
|
|
90d2e6019e
|
||
|
|
769164c824
|
||
|
|
de3a58d105
|
||
|
|
cdb574d3d0
|
||
|
|
a281a9621f
|
||
|
|
825ac394f2
|
||
|
|
d09bfd6aef
|
||
|
|
a11e29a0e0
|
||
|
|
7ac4dc6882
|
||
|
|
57212e3399
|
||
|
|
9f9d40e1c3
|
||
|
|
d7e14156ff
|
||
|
|
de7ccbebc7 | ||
|
|
4b23dbd9f4
|
||
|
|
93300ae0c1
|
||
|
|
5d77d39711 | ||
|
|
7d798b9937
|
||
|
|
915a4d0678 | ||
|
|
6f98d061a9 | ||
|
|
d95c9032c5
|
||
|
|
af85a23ea7
|
||
|
|
8db944e4a0 | ||
|
|
2e8383b098
|
||
|
|
d006dea6e4
|
||
|
|
ebac39e6de | ||
|
|
47dc08c4b0 | ||
|
|
1d553a788f
|
||
|
|
c375c55da7
|
||
|
|
82acda6f93 |
+3
-3
@@ -3,17 +3,17 @@
|
||||
^\.Rproj\.user$
|
||||
^_pkgdown\.yml$
|
||||
^docs$
|
||||
^Meta$
|
||||
^doc$
|
||||
^pkgdown$
|
||||
^\.github$
|
||||
^LICENSE\.md$
|
||||
^\.git$
|
||||
^\.gitignore$
|
||||
\.gitkeep$
|
||||
^vignettes$
|
||||
^specs$
|
||||
^plans$
|
||||
^doc$
|
||||
^Meta$
|
||||
^\.gitea$
|
||||
^CLAUDE\.md$
|
||||
^\.superpowers$
|
||||
^CONTRIBUTING\.md$
|
||||
|
||||
@@ -11,6 +11,21 @@ jobs:
|
||||
steps:
|
||||
- name: Install system libraries and Node.js (required by actions/checkout)
|
||||
run: |
|
||||
# Switch apt to HTTPS mirrors. Measured from this runner on
|
||||
# 2026-08-04: the SAME index file takes 20.1s over http:// and 3.1s
|
||||
# over https://. apt fetches many indexes serially, so http:// does
|
||||
# not read as "slow" -- it reads as a hang (zero bytes in
|
||||
# /var/cache/apt/archives after 3+ minutes, apt's http workers parked
|
||||
# in S state). rocker/r-ver:4.4 already ships ca-certificates and
|
||||
# apt 2.8.3 has the https method built in, so nothing needs to be
|
||||
# installed over http first to bootstrap this.
|
||||
# `|| true` because the step runs under `sh -e`: on an image whose
|
||||
# sources live in the other location, the missing-file sed must not
|
||||
# kill the job.
|
||||
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
|
||||
/etc/apt/sources.list.d/ubuntu.sources 2>/dev/null || true
|
||||
sed -i -E 's#http://(archive|security)\.ubuntu\.com#https://\1.ubuntu.com#g' \
|
||||
/etc/apt/sources.list 2>/dev/null || true
|
||||
apt-get update -qq
|
||||
apt-get install -y --no-install-recommends \
|
||||
nodejs git \
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
# Multi-platform R CMD check, running on the GitHub mirror.
|
||||
#
|
||||
# This exists because the canonical Gitea runner is Linux-only, and this
|
||||
# package hard-depends on duckdb and httr2 -- both compiled, both with real
|
||||
# platform variance -- while having never been checked on Windows or macOS.
|
||||
# A large share of the audience is on Windows.
|
||||
#
|
||||
# Gitea reads .gitea/workflows and GitHub reads .github/workflows, so this
|
||||
# file is inert on the canonical repo and coexists with the Gitea CI that
|
||||
# remains authoritative for deploys.
|
||||
#
|
||||
# The suite needs NO credentials: tests/testthat/setup.R points USCOGDATA_URL
|
||||
# at the bundled fixture corpus. That is exactly why inst/extdata/fixture_corpus
|
||||
# must never be added to .Rbuildignore.
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
|
||||
name: R-CMD-check
|
||||
|
||||
permissions: read-all
|
||||
|
||||
jobs:
|
||||
R-CMD-check:
|
||||
runs-on: ${{ matrix.config.os }}
|
||||
name: ${{ matrix.config.os }} (${{ matrix.config.r }})
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
config:
|
||||
- {os: macos-latest, r: 'release'}
|
||||
- {os: windows-latest, r: 'release'}
|
||||
- {os: ubuntu-latest, r: 'devel', http-user-agent: 'release'}
|
||||
- {os: ubuntu-latest, r: 'release'}
|
||||
|
||||
env:
|
||||
GITHUB_PAT: ${{ secrets.GITHUB_TOKEN }}
|
||||
R_KEEP_PKG_SOURCE: yes
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: r-lib/actions/setup-pandoc@v2
|
||||
|
||||
- uses: r-lib/actions/setup-r@v2
|
||||
with:
|
||||
r-version: ${{ matrix.config.r }}
|
||||
http-user-agent: ${{ matrix.config.http-user-agent }}
|
||||
use-public-rspm: true
|
||||
|
||||
- uses: r-lib/actions/setup-r-dependencies@v2
|
||||
with:
|
||||
extra-packages: any::rcmdcheck
|
||||
needs: check
|
||||
|
||||
- uses: r-lib/actions/check-r-package@v2
|
||||
with:
|
||||
upload-snapshots: true
|
||||
build_args: 'c("--no-manual")'
|
||||
File diff suppressed because it is too large
Load Diff
@@ -28,8 +28,23 @@ USCOGDATA_URL (local path or https://)
|
||||
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
|
||||
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
|
||||
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
|
||||
- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated`
|
||||
- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
|
||||
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
|
||||
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
|
||||
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
|
||||
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
|
||||
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
|
||||
`revenue_annotated`/`ig_annotated`), and metadata views
|
||||
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
|
||||
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
|
||||
`representation`, `code_set`)
|
||||
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
|
||||
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
|
||||
**stock** (a balance at a point in time) rather than a **flow** (activity
|
||||
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
|
||||
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
|
||||
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
|
||||
disclosure + measured per-subtype coverage windows).
|
||||
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
|
||||
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
|
||||
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
|
||||
@@ -45,28 +60,31 @@ USCOGDATA_URL (local path or https://)
|
||||
Any value without `://` is treated as a local path by `.is_local_path()` and reads
|
||||
`manifest.json` directly from disk (no HTTP, no TTL cache).
|
||||
|
||||
## Current State (2026-04-27)
|
||||
## Current State (2026-08-03)
|
||||
|
||||
**Version:** 0.1.0 (pre-release)
|
||||
**Branch:** `main`, commit `d65e9fe`
|
||||
**Tests:** 181 PASS / 0 FAIL / 0 SKIP
|
||||
**Branch:** `feat/cog-balances-25`, commit `fde62eb`
|
||||
**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
|
||||
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
|
||||
|
||||
### Completed (Tasks 2.1–2.7)
|
||||
|
||||
All 8 exported verbs implemented and tested:
|
||||
`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`,
|
||||
`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`,
|
||||
plus `cog_categories`.
|
||||
All **14** exports implemented and tested (measured from `NAMESPACE`):
|
||||
`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
|
||||
`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
|
||||
`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
|
||||
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
|
||||
|
||||
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years
|
||||
2019+2020, all 50 states). Tests run fully offline — no credentials needed.
|
||||
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
|
||||
2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
|
||||
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
|
||||
|
||||
### Remaining to v0.1 release
|
||||
|
||||
1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports;
|
||||
full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`.
|
||||
Vignettes can be stubbed for v0.1.
|
||||
1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
|
||||
`README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
|
||||
`total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
|
||||
`pkgdown::build_site()` has never been run (no `docs/`).
|
||||
|
||||
2. **Phase 3 — cog_explorer bridge:** create
|
||||
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
|
||||
@@ -99,6 +117,17 @@ devtools::test()
|
||||
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
|
||||
- Return value is always a `tbl_df` with a `provenance` attribute
|
||||
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
|
||||
- SQL lives in `inst/sql/` — never inline SQL strings in R files
|
||||
- SQL has two layers. **View definitions** live in `inst/sql/` and are
|
||||
registered by `.register_views()`, which globs the directory in sorted order
|
||||
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
|
||||
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
|
||||
a numbered `.sql` file; build a query in R.
|
||||
- No arrow dependency — DuckDB reads parquet natively
|
||||
- `withr` is a Suggests-only dep; only used in tests
|
||||
|
||||
## Domain context — read this first
|
||||
|
||||
**Before doing any work in this repo, read `~/.claude/memory/values/civilytics.md`.**
|
||||
It carries the purpose, direction, and constraints for this domain. It is not optional
|
||||
context — read it before planning or writing code, not after. (An `@` import will not
|
||||
work here; project-level imports don't preload. The read is the mechanism.)
|
||||
|
||||
+105
@@ -0,0 +1,105 @@
|
||||
# Contributing to uscogdata
|
||||
|
||||
Thanks for reading this — a package like this gets better mostly through people
|
||||
noticing that a number looks wrong.
|
||||
|
||||
## Where the code lives
|
||||
|
||||
Development happens on **Gitea**, at
|
||||
`gitea.civilytics.org/Civilytics/uscogdata`. The repository at
|
||||
`github.com/civilytics/uscogdata` is a **mirror** that accepts issues and pull
|
||||
requests.
|
||||
|
||||
## What happens to a GitHub pull request
|
||||
|
||||
Open it normally. Behind the scenes it is fetched and landed on the canonical
|
||||
Gitea repository, then syncs back:
|
||||
|
||||
```sh
|
||||
git fetch github refs/pull/42/head:pr-42
|
||||
git switch main && git merge --no-ff pr-42
|
||||
git push origin main # Gitea -> mirror -> GitHub
|
||||
```
|
||||
|
||||
Because the merge preserves your commits at their original SHAs, **GitHub marks
|
||||
your PR merged on its own** as soon as the mirror syncs. So:
|
||||
|
||||
> If your pull request closes as "Merged" without anyone visibly clicking
|
||||
> Merge, that is the normal, successful outcome — not a rejection.
|
||||
|
||||
Substantial contributions get a `ctb` entry in `DESCRIPTION`, which surfaces in
|
||||
`citation("uscogdata")`.
|
||||
|
||||
There is no CLA and no DCO sign-off requirement.
|
||||
|
||||
## Running the tests
|
||||
|
||||
```r
|
||||
devtools::test() # bundled fixture; no network, no credentials
|
||||
```
|
||||
|
||||
`tests/testthat/setup.R` points `USCOGDATA_URL` at
|
||||
`inst/extdata/fixture_corpus/` automatically — a four-year slice (2011, 2012,
|
||||
2019, 2020) covering all 50 states. That is the whole data setup.
|
||||
|
||||
## Testing against the live corpus
|
||||
|
||||
```sh
|
||||
USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test(filter = "live-corpus")'
|
||||
```
|
||||
|
||||
This is worth understanding rather than skipping. Until 0.3.0 the package
|
||||
**could not read a remote corpus at all** — the partitioned view used a glob,
|
||||
and DuckDB cannot expand a glob over generic HTTP. It went unnoticed for months
|
||||
because every test path used a local corpus (the bundled fixture), and so did
|
||||
the production API (a host mount). Nothing exercised the package the way a new
|
||||
user does.
|
||||
|
||||
`test-live-corpus.R` is the only test that runs with no `USCOGDATA_URL`, no
|
||||
option, and no fixture. If you change anything touching view registration,
|
||||
manifest handling, or configuration, run it.
|
||||
|
||||
## Do not exclude the fixture from the build
|
||||
|
||||
There is a temptation to add `^inst/extdata/fixture_corpus$` to
|
||||
`.Rbuildignore` because 15 MB feels large for a package. Don't:
|
||||
|
||||
- `vignette("total-spending")` reads from it and would fail to build.
|
||||
- `R CMD check` on r-universe and GitHub Actions would have no corpus, so the
|
||||
suite could not run without credentials.
|
||||
|
||||
This package is not going to CRAN, so its 5 MB guidance does not apply. A
|
||||
package-size NOTE in `R CMD check` is expected and acceptable.
|
||||
|
||||
## Downstream consumers
|
||||
|
||||
`cog-api` depends on this package and its CI clones uscogdata at
|
||||
`USCOGDATA_REF`, **defaulting to `main`**. There is no pin. Anything merged
|
||||
here reaches the API's next build, so before merging a change to the reader,
|
||||
run the API suite against your branch:
|
||||
|
||||
```sh
|
||||
Rscript -e "remotes::install_local('/path/to/uscogdata', upgrade = 'never')"
|
||||
cd /path/to/cog-api/api/tests/testthat
|
||||
Rscript -e 'testthat::test_dir(".", stop_on_failure = TRUE)'
|
||||
```
|
||||
|
||||
The API calls only exported verbs, so internal refactors are usually safe —
|
||||
but "usually" is not a release gate.
|
||||
|
||||
## Release checklist
|
||||
|
||||
1. `devtools::test()` — green against the bundled fixture, offline.
|
||||
2. `USCOGDATA_LIVE_TEST=true devtools::test()` — green against the live corpus.
|
||||
3. cog-api suite green against this branch (above).
|
||||
4. `devtools::check(args = "--as-cran")` — 0 errors, 0 warnings.
|
||||
5. `pkgdown::build_site()` completes.
|
||||
6. Vignettes resolve from an installed copy:
|
||||
`vignette("total-spending", package = "uscogdata")`.
|
||||
7. **Cold-start check**: on a machine that has never had this package,
|
||||
install it and run the README quickstart verbatim with no environment
|
||||
variables set. This is the only check that catches a
|
||||
corpus-unreachable defect, and its absence is why 0.3.0 needed fixing.
|
||||
8. Bump `Version` and add a `NEWS.md` section.
|
||||
9. Tag, then update the r-universe registry pin at
|
||||
`github.com/civilytics/civilytics.r-universe.dev`.
|
||||
+11
-4
@@ -1,14 +1,21 @@
|
||||
Package: uscogdata
|
||||
Type: Package
|
||||
Title: Curated Reader for the Civilytics US Census of Governments Finance Corpus
|
||||
Version: 0.1.0
|
||||
Authors@R:
|
||||
person("Civilytics", , , "jknowles@gmail.com", role = c("aut", "cre"))
|
||||
Version: 0.3.0
|
||||
Authors@R: c(
|
||||
person(c("Jared", "E."), "Knowles",
|
||||
email = "jared@civilytics.com",
|
||||
role = c("aut", "cre"),
|
||||
comment = c(ORCID = "0000-0003-0005-9478")),
|
||||
person("Civilytics Consulting LLC", role = c("cph", "fnd")))
|
||||
Description: Curated R verbs over the Civilytics US Census of Governments
|
||||
finance corpus. Provides unit-level financial profiles, geographic
|
||||
rollups, and peer comparisons with auditable provenance and built-in
|
||||
cross-vintage correctness.
|
||||
License: MIT + file LICENSE
|
||||
URL: https://github.com/civilytics/uscogdata,
|
||||
https://civilytics.r-universe.dev/uscogdata
|
||||
BugReports: https://github.com/civilytics/uscogdata/issues
|
||||
Encoding: UTF-8
|
||||
LazyData: false
|
||||
Depends: R (>= 4.1)
|
||||
@@ -32,4 +39,4 @@ Config/testthat/edition: 3
|
||||
VignetteBuilder: knitr
|
||||
RoxygenNote: 7.3.3
|
||||
MinCorpusSchema: 4
|
||||
MaxCorpusSchema: 5
|
||||
MaxCorpusSchema: 7
|
||||
|
||||
@@ -1,2 +1,2 @@
|
||||
YEAR: 2026
|
||||
COPYRIGHT HOLDER: Civilytics
|
||||
COPYRIGHT HOLDER: Civilytics Consulting LLC
|
||||
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
# MIT License
|
||||
|
||||
Copyright (c) 2026 Civilytics Consulting LLC
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
@@ -1,5 +1,6 @@
|
||||
# Generated by roxygen2: do not edit by hand
|
||||
|
||||
export(cog_balances)
|
||||
export(cog_basket_resolution)
|
||||
export(cog_basket_unresolved)
|
||||
export(cog_categories)
|
||||
|
||||
@@ -1,100 +1,101 @@
|
||||
# uscogdata 0.1.0 (development)
|
||||
# uscogdata 0.3.0
|
||||
|
||||
## Breaking: corpus schema_version 4 (Phase P canonical ids)
|
||||
First public release.
|
||||
|
||||
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
|
||||
`MaxCorpusSchema` in `DESCRIPTION` are both `4`); older corpora built
|
||||
against schema 3 are rejected by `cog_open()` with a clear version-mismatch
|
||||
error. `canonical_govid` is now uniformly 12 characters across every
|
||||
vintage the corpus covers (previously a mix of 9-char legacy ids and
|
||||
12-char FIPS ids depending on source year) — **every hardcoded
|
||||
`canonical_govid` literal from a pre-Phase-P corpus is now invalid** and
|
||||
must be re-resolved via `cog_gov_search()` or the new `canonical_alias`
|
||||
lookup table. `canonical_fips_xwalk` gains four columns
|
||||
(`legacy_govs_id`, `census_geoid`, `id_source`; `confidence` is renamed to
|
||||
`pop_confidence`) and a companion `canonical_alias` table ships in the
|
||||
corpus for mapping legacy/alternate ids onto the current canonical
|
||||
namespace. The bundled fixture corpus (`inst/extdata/fixture_corpus/`) has
|
||||
been regenerated against the Phase P publish tree, now ships the full
|
||||
`canonical_fips_xwalk` and `canonical_alias` master tables alongside the
|
||||
2019-2020 long partitions, and is reproducible via
|
||||
`data-raw/regenerate_fixture_corpus.R`.
|
||||
`uscogdata` provides curated R verbs over the Civilytics US Census of
|
||||
Governments finance corpus: unit-level financial profiles, geographic rollups
|
||||
and peer comparisons, with auditable provenance on every result.
|
||||
|
||||
## Clearer errors when `USCOGDATA_URL` is unconfigured or returns non-JSON
|
||||
## What it covers
|
||||
|
||||
* `cog_open()` now aborts with the `uscogdata_url_not_configured` error
|
||||
class when the resolved corpus URL still contains the placeholder
|
||||
`REPLACE_WITH_SHARE_TOKEN` sentinel (or is empty). The message lists both
|
||||
remediation paths (`Sys.setenv(USCOGDATA_URL = ...)` and
|
||||
`options(uscogdata.url = ...)`) and points at the bundled fixture for
|
||||
offline testing. Previously the package proceeded to fetch the placeholder
|
||||
URL, cached the resulting HTML welcome page, and failed downstream with a
|
||||
cryptic `jsonlite` lexical-error.
|
||||
* `.fetch_or_cache_manifest()` now parses the HTTP response body before
|
||||
persisting it. Non-JSON responses (login pages, 404 HTML) raise
|
||||
`uscogdata_invalid_manifest` with the URL, Content-Type, and underlying
|
||||
parse error — and never write to the on-disk cache.
|
||||
* Manifest cache writes are now atomic (write to `manifest.json.tmp.<pid>`
|
||||
in `cache_dir`, then `file.rename` over the target), so an interrupted
|
||||
fetch cannot replace a previously-good cache.
|
||||
* Existing caches with non-JSON content (poisoned by the prior code path)
|
||||
are silently refetched instead of returning a parse error to the caller.
|
||||
* Local `USCOGDATA_URL` paths whose `manifest.json` is not valid JSON now
|
||||
surface the same `uscogdata_invalid_manifest` class with file context.
|
||||
Government types 0-3 (state, county, municipality, township), FY1967-FY2024 --
|
||||
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
|
||||
or FY1969. Special districts (type 4) and school districts (type 5) are out of
|
||||
scope pending validation.
|
||||
|
||||
## Per-capita denominators now use per-year Census F-33 population
|
||||
## The verbs
|
||||
|
||||
* `cog_spending()` and `cog_revenue()` previously divided all years' amounts
|
||||
by a single ACS 2018-2022 estimate (`canonical_fips_xwalk.population_acs`),
|
||||
producing biased per-capita values for time-series analysis. They now
|
||||
divide by the F-33 `population` recorded on each gov-year via the new
|
||||
`gov_population_yearly` view. Result tibbles gain a `pop_source` column
|
||||
with values `"census_f33"` or `"unavailable"`. `notes` is updated to
|
||||
concatenate multiple notes with `"; "`.
|
||||
`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and holdings;
|
||||
`cog_gov_search()` to resolve place names (including basket mode for many at
|
||||
once); `cog_find_peers()` and `cog_peer_compare()` for cohorts;
|
||||
`cog_geographic_rollup()` for aggregates; `cog_categories()`, `cog_recipes()`,
|
||||
`cog_manifest()` and `cog_explain()` for metadata and provenance; and
|
||||
`cog_mirror()` for a local copy of the corpus.
|
||||
|
||||
## Peer cohorts can be set to a chosen year
|
||||
## Reading the corpus now works out of the box
|
||||
|
||||
* `cog_find_peers()` adds a `year` argument (default: most recent year for
|
||||
which the target has an observed population in `gov_population_yearly`).
|
||||
The returned column previously named `population_acs` is now `population`
|
||||
and reflects the cohort year's vintage. The cohort year is attached to the
|
||||
returned tibble as `attr(x, "cohort_year")`.
|
||||
* `cog_peer_compare()` now stamps a `cohort_year` column on its result (read
|
||||
from the peers tibble's attribute) and records `cohort_year` plus
|
||||
`cohort_govids` in provenance. When the caller supplies a bare character
|
||||
vector instead of a `cog_find_peers()` result, `cohort_year` is `NA`.
|
||||
* The package reads the published corpus over HTTPS **with no configuration**.
|
||||
Previously the default was a placeholder sentinel and no document in the
|
||||
package supplied a working URL, so a new user had no path to a session.
|
||||
* Remote reads work at all. The partitioned view used a glob, and DuckDB
|
||||
cannot expand a glob over generic HTTP -- there is no directory listing to
|
||||
expand against. Partition paths are now enumerated from the corpus manifest,
|
||||
which is host-agnostic: an HTTPS mirror, a Nextcloud share and a local
|
||||
`cog_mirror()` copy all take the same path.
|
||||
* Nothing is written to disk in remote mode; DuckDB fetches only the row
|
||||
groups a query needs.
|
||||
|
||||
## Rollups exclude govs missing population
|
||||
## Four things to know before your first query
|
||||
|
||||
* `cog_geographic_rollup(per_capita = TRUE)` drops rows whose government has
|
||||
`pop_source == "unavailable"` and records the dropped govids in
|
||||
`provenance$rollup$excluded_govids`. This excludes special districts
|
||||
(type 4) and school districts (type 5) from per-capita rollups by design.
|
||||
* **Amounts are in full US dollars.** The raw Census files report thousands;
|
||||
the verbs multiply by 1000 on the way out. Do not multiply again.
|
||||
* **Multi-government aggregates disclose their coverage.** The Census is a
|
||||
complete enumeration only in years ending in 2 and 7; every other year is a
|
||||
sample. Every such result carries `provenance$coverage` with per-year
|
||||
`n_units_reporting`.
|
||||
* **Absence means two different things.** Before FY2012 an absent cell means
|
||||
Census published $0; from FY2012 it means not reported. `complete = TRUE`
|
||||
labels which.
|
||||
* **Series breaks reach you unasked.** Catalogued breaks intersecting your
|
||||
query appear in provenance and in `cog_explain()`.
|
||||
|
||||
## New: vignette and provenance metadata
|
||||
## Known limits
|
||||
|
||||
* New vignette `population-denominators` covers the four population sources,
|
||||
the type-4/5 coverage gap, the popyear quirk, and how to build moving-window
|
||||
peer cohorts manually.
|
||||
* Provenance gains `transformations$per_capita$popyear_range` and
|
||||
`pop_source_counts`. `cog_explain()` renders both.
|
||||
* Special districts (type 4) and school districts (type 5) are out of scope.
|
||||
* Per-capita rollups exclude governments with no F-33 population, which is by
|
||||
design but does silently narrow a rollup.
|
||||
* `n_units_reporting` is category-conditional and is not a response rate.
|
||||
* Employee-retirement (`X`) codes stop at FY2016, when those systems moved to
|
||||
the Annual Survey of Public Pensions.
|
||||
|
||||
# uscogdata 0.2.0
|
||||
|
||||
## New features
|
||||
|
||||
* `cog_gov_search()` gains a **basket mode**: passing vector `name`
|
||||
/ `state` / `type` arguments resolves multiple place names in one
|
||||
call and returns a tibble of canonical rows in input order, ready
|
||||
to pipe into `cog_spending()` / `cog_revenue()`. Per-row resolution
|
||||
follows an exact-then-substring matching algorithm with deterministic
|
||||
disambiguation; ambiguous and missing entries are surfaced via a
|
||||
sidecar audit tibble plus a single console summary message.
|
||||
* New exports `cog_basket_resolution()` and `cog_basket_unresolved()`
|
||||
expose the basket sidecar for iterative query refinement.
|
||||
* `cog_spending()` and `cog_revenue()` accept the reserved category
|
||||
`"All Categories"`, returning one summed row per
|
||||
`(year, canonical_govid, subtype)` across every category inside the
|
||||
requested concept's subtype scope. Filtering the result to
|
||||
`spend_subtype == "operations"` gives an operating-expenditure total.
|
||||
`cog_geographic_rollup()` inherits it,
|
||||
which is the efficient way to build a geographic total — previously a
|
||||
caller had to issue one rollup per category and sum the results
|
||||
(cog-api#37).
|
||||
|
||||
## Breaking changes
|
||||
`"All Categories"` is not the same thing as `expenditure_concept = "total"`.
|
||||
The concept chooses which subtypes are in scope; `"All Categories"` chooses
|
||||
whether the rows inside that scope are broken out or summed.
|
||||
|
||||
* The first formal of `cog_gov_search()` was renamed from `pattern`
|
||||
to `name`. All existing call sites in `cog_explorer/` and the
|
||||
package itself use positional first-arg, so this rename is
|
||||
non-breaking in practice. Callers that pass `pattern = ...` by name
|
||||
must update to `name = ...`.
|
||||
* `cog_categories()` advertises `"All Categories"` for the expenditure and
|
||||
revenue vocabularies, so the reserved value is discoverable.
|
||||
|
||||
* Coverage signposting (see "Signposting now catches partially-suppressed
|
||||
categories" below) now also works in `category = "All Categories"` mode.
|
||||
The recipe-suggestion candidate query used to be scoped by `category`,
|
||||
which is never a match for the reserved `"All Categories"` value, so
|
||||
`provenance$suggestions` always came back empty there — the one mode whose
|
||||
whole point is "you cannot sum the wrong scope" was silently unable to
|
||||
signal a wrong scope. The candidate query is now scoped by the concept's
|
||||
subtype allowlist instead, symmetric with how `.build_verb_sql()` itself
|
||||
scopes the summed total: Los Angeles County FY2011, `category = "All
|
||||
Categories"` still excludes $271,589,000 of aggregate-published Public
|
||||
Welfare (`E68`), but now names `recipe = "welfare_cash_e68_wide"` to
|
||||
recover it instead of reporting zero suggestions.
|
||||
|
||||
## Documentation
|
||||
|
||||
* `cog_geographic_rollup()` and `cog_peer_compare()` now document that
|
||||
`provenance$coverage`'s `n_units_reporting` is **category-conditional** and
|
||||
is not a response rate: a government that was surveyed and genuinely spends
|
||||
nothing in the requested category is indistinguishable from one never
|
||||
surveyed (uscogdata#36).
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
# R/balance_caveats.R
|
||||
#
|
||||
# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and
|
||||
# security holdings. Each one silently invalidates an obvious analysis, so
|
||||
# they travel in provenance (machine-readable, for cog-api#26) rather than
|
||||
# living only in prose.
|
||||
#
|
||||
# Two of the four are already carried by the code-driven series-break
|
||||
# builders and are deliberately NOT duplicated here:
|
||||
# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via
|
||||
# series_break_refs on the recipe path, the only path that observes those
|
||||
# codes.
|
||||
# What remains is the GAAP distinction (a constant) and the coverage windows
|
||||
# (measured, never hardcoded, so they stay correct as the corpus grows).
|
||||
|
||||
#' Per-subtype observed year extents, plus which requested families are
|
||||
#' truncated relative to the requested span.
|
||||
#' @noRd
|
||||
.balance_caveats <- function(con, codes_observed, years) {
|
||||
cw <- .balance_coverage_windows(con)
|
||||
|
||||
observed_subtypes <- if (length(codes_observed) == 0L) {
|
||||
character(0)
|
||||
} else {
|
||||
DBI::dbGetQuery(con, sprintf(
|
||||
"SELECT DISTINCT balance_subtype FROM summary_categories
|
||||
WHERE item_code IN (%s) AND balance_subtype IS NOT NULL",
|
||||
.sql_lit_chr(codes_observed)
|
||||
))$balance_subtype
|
||||
}
|
||||
|
||||
# A family is "truncated" when the caller asked for years outside the span
|
||||
# that family actually covers -- the FY2016 employee-retirement termination
|
||||
# and the FY2021 end of the W family are both this shape.
|
||||
truncated <- character(0)
|
||||
if (length(years) > 0L) {
|
||||
for (s in observed_subtypes) {
|
||||
w <- cw[[s]]
|
||||
if (is.null(w)) next
|
||||
if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s)
|
||||
}
|
||||
}
|
||||
|
||||
list(
|
||||
not_gaap = TRUE,
|
||||
not_gaap_note = paste0(
|
||||
"Census holdings are gross -- no liabilities are netted -- and are NOT ",
|
||||
"GAAP fund balance. A reserve ratio built from them overstates what is ",
|
||||
"actually available."
|
||||
),
|
||||
coverage_window = cw,
|
||||
truncated = sort(unique(truncated))
|
||||
)
|
||||
}
|
||||
|
||||
#' Per-subtype [min year, max year] extents for EVERY balance subtype in the
|
||||
#' mounted corpus, memoised for the session.
|
||||
#'
|
||||
#' The query carries no govid and no year predicate -- its answer is a property
|
||||
#' of the mounted corpus alone and cannot change between calls -- but it scans
|
||||
#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime
|
||||
#' on the bundled fixture and would be a per-request throughput ceiling once
|
||||
#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and
|
||||
#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`.
|
||||
#'
|
||||
#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking
|
||||
#' "is there a family I missed?" needs every window. The observed-scoped field
|
||||
#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json.
|
||||
#' @noRd
|
||||
.balance_coverage_windows <- function(con) {
|
||||
cached <- .uscogdata_env$balance_coverage_windows
|
||||
if (!is.null(cached)) return(cached)
|
||||
|
||||
windows <- DBI::dbGetQuery(con,
|
||||
"SELECT c.balance_subtype AS subtype,
|
||||
MIN(l.year) AS year_min,
|
||||
MAX(l.year) AS year_max
|
||||
FROM balance_long l
|
||||
JOIN summary_categories c USING (item_code)
|
||||
WHERE c.balance_subtype IS NOT NULL
|
||||
GROUP BY 1
|
||||
ORDER BY 1"
|
||||
)
|
||||
|
||||
cw <- stats::setNames(
|
||||
lapply(seq_len(nrow(windows)),
|
||||
function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))),
|
||||
windows$subtype
|
||||
)
|
||||
.uscogdata_env$balance_coverage_windows <- cw
|
||||
cw
|
||||
}
|
||||
|
||||
#' TRUE the first time `key` is seen this session, FALSE thereafter.
|
||||
#' Reset by cog_close().
|
||||
#' @noRd
|
||||
.balance_caveat_once <- function(key) {
|
||||
seen <- .uscogdata_env$balance_caveats_shown
|
||||
if (is.null(seen)) seen <- character(0)
|
||||
if (key %in% seen) return(FALSE)
|
||||
.uscogdata_env$balance_caveats_shown <- c(seen, key)
|
||||
TRUE
|
||||
}
|
||||
|
||||
#' Emit at most one message per caveat class per session.
|
||||
#' @noRd
|
||||
.emit_balance_caveats <- function(caveats) {
|
||||
if (.balance_caveat_once("not_gaap")) {
|
||||
cli::cli_inform(c(
|
||||
"!" = "Census holdings are gross and are {.strong not} GAAP fund balance.",
|
||||
"i" = "No liabilities are netted; a reserve ratio built from them overstates available funds."
|
||||
))
|
||||
}
|
||||
if (length(caveats$truncated) > 0L &&
|
||||
.balance_caveat_once("coverage_window")) {
|
||||
cli::cli_inform(c(
|
||||
"!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.",
|
||||
"i" = "See {.code provenance$balance_caveats$coverage_window}."
|
||||
))
|
||||
}
|
||||
invisible(NULL)
|
||||
}
|
||||
+171
@@ -0,0 +1,171 @@
|
||||
# R/balances.R
|
||||
#
|
||||
# Cash and security holdings. A third verb rather than an argument on a money
|
||||
# verb because holdings are a STOCK -- a balance at a point in time -- while
|
||||
# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money
|
||||
# verbs' whole argument vocabulary (expenditure_concept, revenue_concept,
|
||||
# complete=) describes flows and is meaningless here, so this deliberately
|
||||
# does NOT route through .verb_spendrev().
|
||||
|
||||
#' Cash and security holdings for one or more governments
|
||||
#'
|
||||
#' Returns Census cash-and-security holdings (`category_type = "balance"`):
|
||||
#' fund balances, retirement system holdings and insurance trust balances.
|
||||
#'
|
||||
#' @section Holdings are not GAAP fund balance:
|
||||
#' Census holdings are **gross** -- no liabilities are netted -- so a reserve
|
||||
#' ratio built from them overstates what is actually available. They are not
|
||||
#' comparable to a GAAP fund balance from an ACFR.
|
||||
#'
|
||||
#' @param govid Canonical govid(s): a character vector, or a data frame with a
|
||||
#' `canonical_govid` column (e.g. from [cog_gov_search()]).
|
||||
#' @param years Integer vector of fiscal years.
|
||||
#' @param category Optional character vector of categories to keep. One of
|
||||
#' `"Fund Balances"`, `"Insurance Trust Balances"`,
|
||||
#' `"Retirement System Holdings"`. There is deliberately no `subtype`
|
||||
#' argument: for holdings, `category` is a strict coarsening of
|
||||
#' `balance_subtype` (unlike the money verbs, where the two axes cross), so
|
||||
#' every combination would be either redundant or empty.
|
||||
#' `category = "Fund Balances"` is exactly the `general` family
|
||||
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
||||
#' one `dplyr::filter()` away. The reserved pseudo-category
|
||||
#' `"All Categories"` (see [cog_spending()]) is **not** supported here and
|
||||
#' errors with class `uscogdata_all_categories_unsupported`: it sums a
|
||||
#' concept's subtype scope, and holdings are a stock with no concept
|
||||
#' vocabulary to sum across. Omit `category` to get every category broken
|
||||
#' out instead.
|
||||
#' @param per_capita Divide holdings by population. Note this is a **stock per
|
||||
#' resident** (reserves per person), which is *not* comparable to
|
||||
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
|
||||
#' @param adjust_to_year Deflate to this year's dollars (CPI-U).
|
||||
#' @param basis Accepted for uniformity with the money verbs, but currently a
|
||||
#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized
|
||||
#' and raw space are identical for holdings. Reported in
|
||||
#' `provenance$basis_note`.
|
||||
#' @param recipe Optional harmonization recipe id (see [cog_recipes()]).
|
||||
#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
|
||||
#' wide era to the modern one.
|
||||
#'
|
||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`,
|
||||
#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and
|
||||
#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
|
||||
#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when
|
||||
#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
|
||||
#' nominal per-capita column to deflate otherwise). Amounts are full US
|
||||
#' dollars.
|
||||
#'
|
||||
#' Carries a `provenance` attribute matching
|
||||
#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
|
||||
#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
|
||||
#' every balance subtype in the mounted corpus, not only the observed ones)
|
||||
#' and `truncated` (the observed subtypes whose coverage falls short of the
|
||||
#' requested years). `expenditure_concept`/`revenue_concept` are `NA` --
|
||||
#' holdings are a stock, not a flow, so neither concept vocabulary applies.
|
||||
#' @export
|
||||
cog_balances <- function(govid, years, category = NULL,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||
call <- match.call()
|
||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||
# Coerce FIRST, validate second: .validate_verb_inputs() asserts
|
||||
# is.character(govid), and a data-frame govid (cog_gov_search() output) has
|
||||
# not been unwrapped yet at this point.
|
||||
govid <- .coerce_govid_input(govid)
|
||||
# The money verbs' validator, reused rather than re-implemented (R/spending.R).
|
||||
# It covers the exact superset cog_balances() needs -- including the
|
||||
# recipe/category mutual-exclusivity guard -- so a second local copy would
|
||||
# only be a place for the two to drift apart. This is the same kind of
|
||||
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
|
||||
# route the verb through .verb_spendrev(), which stays deliberately unused
|
||||
# here because its flow vocabulary is meaningless for a stock.
|
||||
#
|
||||
# allow_all_categories is left at its FALSE default (contrast
|
||||
# .verb_spendrev(), which passes TRUE): the all-categories mode's "sum"
|
||||
# only means something in terms of a concept's subtype scope, and holdings
|
||||
# have no concept vocabulary. The reuse above is exactly why this can be a
|
||||
# one-line default rather than a second bespoke check -- see the
|
||||
# validator's own doc comment for the incident that made that matter.
|
||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
||||
recipe)
|
||||
years <- as.integer(years)
|
||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||
|
||||
con <- .ensure_session()
|
||||
.require_balance_support(con)
|
||||
scope <- .check_govids_in_scope(govid)
|
||||
|
||||
basis_note <- paste0(
|
||||
"`basis` has no effect on holdings: harmonization_map carries no ",
|
||||
"balance-code rows, so harmonized and raw space are identical here."
|
||||
)
|
||||
|
||||
manifest <- .uscogdata_env$manifest
|
||||
recipe_block <- NULL
|
||||
category_for_prov <- category
|
||||
|
||||
if (!is.null(recipe)) {
|
||||
.require_schema_v5(con, manifest, "recipe =")
|
||||
.validate_recipe_id(con, recipe)
|
||||
comps <- .recipe_components(con, recipe)
|
||||
recipe_label <- comps$label[[1]]
|
||||
result <- .run_recipe(con, recipe, govid, years)
|
||||
sql <- attr(result, "sql_query")
|
||||
result <- .shape_recipe_result(result, "balance_subtype", recipe_label)
|
||||
recipe_block <- list(
|
||||
recipe_id = recipe, label = recipe_label,
|
||||
components = .df_to_row_list(comps)
|
||||
)
|
||||
category_for_prov <- recipe_label
|
||||
} else {
|
||||
sql <- .build_verb_sql("balance_annotated", "balance_subtype",
|
||||
govid, years, category,
|
||||
ig_view = NULL, subtype_scope = NULL)
|
||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
}
|
||||
|
||||
# Order matters (matches .verb_spendrev()): per-capita first, so
|
||||
# .attach_real_dollars() deflates the nominal per-capita column into
|
||||
# amt_per_capita_real rather than needing amt_per_capita_nominal recomputed.
|
||||
if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid)
|
||||
if (!is.null(adjust_to_year)) {
|
||||
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
||||
}
|
||||
|
||||
prov <- .build_provenance(
|
||||
verb = "cog_balances", call = call, govid = govid, years = years,
|
||||
category = category_for_prov, per_capita = per_capita,
|
||||
adjust_to_year = adjust_to_year, result = result, sql = sql,
|
||||
subtype_col = "balance_subtype",
|
||||
basis = basis, basis_note = basis_note,
|
||||
# Neither concept vocabulary applies to a stock.
|
||||
expenditure_concept = NA_character_,
|
||||
revenue_concept = NA_character_,
|
||||
recipe = recipe_block
|
||||
)
|
||||
prov$scope$govids_found <- scope$found
|
||||
prov$scope$govids_missing <- scope$missing
|
||||
|
||||
prov$balance_caveats <- .balance_caveats(
|
||||
con, prov$codes_summed$observed, years
|
||||
)
|
||||
.emit_balance_caveats(prov$balance_caveats)
|
||||
|
||||
attr(result, "provenance") <- prov
|
||||
result
|
||||
}
|
||||
|
||||
#' Abort unless the mounted corpus classifies balance codes.
|
||||
#'
|
||||
#' `balance_subtype` arrived with cog_pipeline #76/#77 without a
|
||||
#' schema_version bump, so the check is on the column, not the version.
|
||||
#' @noRd
|
||||
.require_balance_support <- function(con) {
|
||||
if (.corpus_has_balance_subtype(con)) return(invisible(TRUE))
|
||||
cli::cli_abort(
|
||||
c("This corpus does not classify cash and security holdings.",
|
||||
i = "`summary_categories` has no {.field balance_subtype} column.",
|
||||
i = "Republish from cog_pipeline at #76/#77 or later."),
|
||||
class = "uscogdata_no_balance_support"
|
||||
)
|
||||
}
|
||||
@@ -44,11 +44,18 @@
|
||||
|
||||
#' Count + sum item-level rows that basis="harmonized" excludes because they
|
||||
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
|
||||
#' requested flow type (spending or revenue), govids, and years. Only
|
||||
#' meaningful when the resolved basis is "harmonized"; returns an
|
||||
#' applied = FALSE stub otherwise (raw basis never excludes rows this way).
|
||||
#' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
|
||||
#' the same subtype-membership classification the verb SQL uses, never
|
||||
#' item-code prefixes), govids, and years. Only meaningful when the resolved
|
||||
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
|
||||
#' basis never excludes rows this way).
|
||||
#'
|
||||
#' The intergovernmental leg is deliberately outside this count even for
|
||||
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
|
||||
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
|
||||
#' @noRd
|
||||
.build_harmonization_block <- function(con, govid, years, resolved, flow_prefixes) {
|
||||
.build_harmonization_block <- function(con, govid, years, resolved,
|
||||
subtype_col, subtype_scope) {
|
||||
if (!identical(resolved$basis, "harmonized")) {
|
||||
return(list(
|
||||
applied = FALSE,
|
||||
@@ -63,9 +70,11 @@
|
||||
FROM long
|
||||
WHERE canonical_govid IN (%s) AND year IN (%s)
|
||||
AND NOT is_aggregate AND harmonized_code IS NULL
|
||||
AND LEFT(item_code, 1) IN (%s)",
|
||||
AND item_code IN (
|
||||
SELECT item_code FROM summary_categories WHERE %s IN (%s)
|
||||
)",
|
||||
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
|
||||
.sql_lit_chr(flow_prefixes)
|
||||
subtype_col, .sql_lit_chr(subtype_scope)
|
||||
)
|
||||
na <- DBI::dbGetQuery(con, sql)
|
||||
|
||||
|
||||
+41
-8
@@ -5,23 +5,33 @@
|
||||
#' Returns the category taxonomy exposed by the corpus's
|
||||
#' `summary_categories` view, grouped to one row per
|
||||
#' `(category, subtype)` pair. Use this to discover valid `category`
|
||||
#' values for [cog_spending()] / [cog_revenue()] /
|
||||
#' values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
|
||||
#' [cog_geographic_rollup()] and to audit which Census item codes feed
|
||||
#' each category.
|
||||
#'
|
||||
#' @param type Either `NULL` (default, return both spending and revenue
|
||||
#' rows), `"spending"`, or `"revenue"`.
|
||||
#' `subtype` COALESCEs the crosswalk's three subtype columns, so it carries
|
||||
#' `spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
|
||||
#' `balance_subtype` on balance rows. Note that [cog_balances()] itself takes
|
||||
#' no `subtype` argument — for holdings, `category` is a strict coarsening of
|
||||
#' `balance_subtype` — but the value is surfaced here because it is the
|
||||
#' discovery surface downstream consumers build their vocabulary from.
|
||||
#'
|
||||
#' @param type Either `NULL` (default, every row: expenditure, revenue and
|
||||
#' balance), `"spending"`, `"revenue"`, or `"balance"`.
|
||||
#' @param pattern Optional regex matched case-insensitively against the
|
||||
#' `category` column (e.g. `"Police"` or `"Tax"`).
|
||||
#' @return Tibble with columns `category`, `category_type`, `subtype`,
|
||||
#' `n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
||||
#' `category_type`, `category`, `subtype`.
|
||||
#' `category_type`, `category`, `subtype`. Includes one row per flow for the
|
||||
#' reserved pseudo-category `"All Categories"`, which carries `NA` for
|
||||
#' `subtype`, `n_codes` and `item_codes` because it is a query mode rather
|
||||
#' than a crosswalk entry — see [cog_spending()]'s `category` argument.
|
||||
#' @export
|
||||
cog_categories <- function(type = NULL, pattern = NULL) {
|
||||
if (!is.null(type)) {
|
||||
if (!is.character(type) || length(type) != 1L ||
|
||||
!type %in% c("spending", "revenue")) {
|
||||
cli::cli_abort('`type` must be NULL, "spending", or "revenue".')
|
||||
!type %in% c("spending", "revenue", "balance")) {
|
||||
cli::cli_abort('`type` must be NULL, "spending", "revenue", or "balance".')
|
||||
}
|
||||
}
|
||||
if (!is.null(pattern) &&
|
||||
@@ -48,7 +58,7 @@ cog_categories <- function(type = NULL, pattern = NULL) {
|
||||
|
||||
sql <- paste(
|
||||
"SELECT category, category_type,
|
||||
COALESCE(spend_subtype, revenue_subtype) AS subtype,
|
||||
COALESCE(spend_subtype, revenue_subtype, balance_subtype) AS subtype,
|
||||
COUNT(DISTINCT item_code) AS n_codes,
|
||||
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS item_codes
|
||||
FROM summary_categories",
|
||||
@@ -56,5 +66,28 @@ cog_categories <- function(type = NULL, pattern = NULL) {
|
||||
"GROUP BY category, category_type, subtype
|
||||
ORDER BY category_type, category, subtype"
|
||||
)
|
||||
tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
out <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
|
||||
# The reserved pseudo-category is a query mode, not a crosswalk row, so it
|
||||
# has no item codes to report -- hence NA rather than 0 for n_codes. It is
|
||||
# emitted for the two FLOW vocabularies only: cog_balances() returns a stock
|
||||
# and has no concept argument to sum within.
|
||||
pseudo <- tibble::tibble(
|
||||
category = .ALL_CATEGORIES,
|
||||
category_type = c("expenditure", "revenue"),
|
||||
subtype = NA_character_,
|
||||
n_codes = NA_integer_,
|
||||
item_codes = NA_character_
|
||||
)
|
||||
if (!is.null(type)) {
|
||||
db_type <- if (type == "spending") "expenditure" else type
|
||||
pseudo <- pseudo[pseudo$category_type == db_type, , drop = FALSE]
|
||||
}
|
||||
if (!is.null(pattern) && nrow(pseudo) > 0L) {
|
||||
keep <- grepl(pattern, pseudo$category, ignore.case = TRUE)
|
||||
pseudo <- pseudo[keep, , drop = FALSE]
|
||||
}
|
||||
if (nrow(pseudo) == 0L) return(out)
|
||||
out <- rbind(out, pseudo)
|
||||
out[order(out$category_type, out$category, out$subtype), , drop = FALSE]
|
||||
}
|
||||
|
||||
+149
@@ -0,0 +1,149 @@
|
||||
# R/complete.R
|
||||
#
|
||||
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
|
||||
# cell the corpus does not carry still appears, labelled with WHY it is
|
||||
# missing.
|
||||
#
|
||||
# The corpus stopped storing the wide era's explicit zeros
|
||||
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
|
||||
#
|
||||
# <= FY2011 dense_source absent => Census published $0 (census_zero)
|
||||
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
|
||||
#
|
||||
# Before sparsification a wide-era query whose cells were all $0 came back as
|
||||
# explicit $0 rows; afterwards it came back empty, with nothing to say which
|
||||
# of the two meanings applied. This restores that -- and improves on it,
|
||||
# because the pre-sparsification corpus could not distinguish the two either.
|
||||
#
|
||||
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
|
||||
# That difference is the entire point: writing 0 into a modern absence would
|
||||
# invent data, which is the error the representation contract exists to stop.
|
||||
|
||||
#' @noRd
|
||||
.abort_complete_unsupported <- function(reason, alternative) {
|
||||
cli::cli_abort(c(
|
||||
"{.code complete = TRUE} is not supported for this query.",
|
||||
x = reason,
|
||||
i = alternative
|
||||
), class = "uscogdata_complete_unsupported")
|
||||
}
|
||||
|
||||
#' @noRd
|
||||
.require_representation <- function(con, manifest) {
|
||||
needed <- c("representation.parquet", "code_set.parquet")
|
||||
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
|
||||
logical(1))]
|
||||
if (length(missing) == 0L) return(invisible(TRUE))
|
||||
cli::cli_abort(c(
|
||||
"This corpus does not publish the representation contract.",
|
||||
x = "Missing: {.file {missing}}.",
|
||||
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
|
||||
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
|
||||
), class = "uscogdata_representation_unavailable")
|
||||
}
|
||||
|
||||
#' The cells a government-year COULD carry: every code in force for that
|
||||
#' government's own type, mapped through `summary_categories`, restricted to
|
||||
#' the calling verb's crosswalk subtype scope (the same subtype-membership
|
||||
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
|
||||
#' operations/capital/assistance) and (when given) its category filter.
|
||||
#'
|
||||
#' Scoped by `govs_type` deliberately. Filling against the union of all types
|
||||
#' would invent cells that the government can never report -- a county row for
|
||||
#' "state IG transfer to school districts" -- and those inventions would then
|
||||
#' be indistinguishable from real census zeros.
|
||||
#'
|
||||
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
|
||||
#' aggregate rows. Without it the grid would offer cells the verb structurally
|
||||
#' never returns, so every one of them would fill as a phantom $0.
|
||||
#' @noRd
|
||||
.completion_grid_sql <- function(subtype_col, govid, years, category,
|
||||
subtype_scope) {
|
||||
category_pred <- if (is.null(category)) {
|
||||
""
|
||||
} else {
|
||||
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
|
||||
}
|
||||
sprintf(
|
||||
"SELECT DISTINCT
|
||||
cs.year,
|
||||
x.canonical_govid,
|
||||
x.gov_name,
|
||||
c.%1$s AS subtype_value,
|
||||
c.category,
|
||||
r.absence_means
|
||||
FROM code_set cs
|
||||
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
|
||||
JOIN summary_categories c ON c.item_code = cs.item_code
|
||||
JOIN representation r ON r.year = cs.year
|
||||
WHERE x.canonical_govid IN (%2$s)
|
||||
AND cs.year IN (%3$s)
|
||||
AND NOT cs.is_aggregate
|
||||
AND c.category IS NOT NULL
|
||||
AND c.%1$s IN (%4$s)
|
||||
%5$s",
|
||||
subtype_col, .sql_lit_chr(govid),
|
||||
paste(as.integer(years), collapse = ","),
|
||||
.sql_lit_chr(subtype_scope), category_pred
|
||||
)
|
||||
}
|
||||
|
||||
#' Fill `result` out to the full grid, stamping `value_source` on every row.
|
||||
#'
|
||||
#' Returns the completed tibble with a `.completion` attribute carrying the
|
||||
#' provenance block. Reported rows are passed through untouched -- filling
|
||||
#' must never alter or drop what the corpus actually published.
|
||||
#' @noRd
|
||||
.complete_result <- function(result, con, subtype_col, govid, years, category,
|
||||
subtype_scope) {
|
||||
grid <- tibble::as_tibble(DBI::dbGetQuery(
|
||||
con, .completion_grid_sql(subtype_col, govid, years, category, subtype_scope)
|
||||
))
|
||||
|
||||
result$value_source <- rep("reported", nrow(result))
|
||||
if (nrow(grid) == 0L) {
|
||||
attr(result, ".completion") <- list(
|
||||
applied = TRUE, rows_filled = 0L, absence_means = list()
|
||||
)
|
||||
return(result)
|
||||
}
|
||||
|
||||
names(grid)[names(grid) == "subtype_value"] <- subtype_col
|
||||
key <- function(d) {
|
||||
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
|
||||
}
|
||||
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
|
||||
|
||||
if (nrow(missing) > 0L) {
|
||||
filled <- tibble::tibble(
|
||||
year = as.integer(missing$year),
|
||||
canonical_govid = as.character(missing$canonical_govid),
|
||||
gov_name = as.character(missing$gov_name),
|
||||
category = as.character(missing$category),
|
||||
# census_zero is a value Census published; not_reported is unknown and
|
||||
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
|
||||
amt_nominal = ifelse(missing$absence_means == "census_zero",
|
||||
0, NA_real_),
|
||||
codes_included = NA_character_,
|
||||
aggregate_fallback = NA,
|
||||
value_source = as.character(missing$absence_means)
|
||||
)
|
||||
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
|
||||
if ("notes" %in% names(result)) filled$notes <- NA_character_
|
||||
|
||||
result <- dplyr::bind_rows(result, filled)
|
||||
result <- result[order(result$year, result$canonical_govid,
|
||||
result[[subtype_col]], result$category), ,
|
||||
drop = FALSE]
|
||||
}
|
||||
|
||||
rules <- unique(grid[, c("year", "absence_means")])
|
||||
attr(result, ".completion") <- list(
|
||||
applied = TRUE,
|
||||
rows_filled = nrow(missing),
|
||||
absence_means = stats::setNames(
|
||||
as.list(as.character(rules$absence_means)), as.character(rules$year)
|
||||
)
|
||||
)
|
||||
result
|
||||
}
|
||||
+12
-1
@@ -5,7 +5,18 @@
|
||||
.uscogdata_env <- new.env(parent = emptyenv())
|
||||
|
||||
.uscogdata_defaults <- list(
|
||||
url = "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/",
|
||||
# Public HuggingFace mirror of the published corpus: CC-BY-4.0, no
|
||||
# credential, CDN-backed. This is the default so `library(uscogdata)`
|
||||
# followed by a verb works with zero configuration -- previously the
|
||||
# default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in the
|
||||
# package supplied a working URL, so a new user had no path to a session.
|
||||
#
|
||||
# The trailing slash is required: every consumer concatenates onto this
|
||||
# (see .resolve_url(), which enforces it anyway).
|
||||
#
|
||||
# Override with USCOGDATA_URL or options(uscogdata.url=) to read a
|
||||
# Nextcloud share or a local copy made by cog_mirror().
|
||||
url = "https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/",
|
||||
cache_dir = NULL,
|
||||
manifest_ttl_secs = 3600L
|
||||
)
|
||||
|
||||
+107
@@ -0,0 +1,107 @@
|
||||
# R/coverage.R
|
||||
#
|
||||
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
|
||||
# findings F-020 and F-023).
|
||||
#
|
||||
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
|
||||
# 7. Every other year is a sample, and the sample varies enormously: on the
|
||||
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
|
||||
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
|
||||
# what the verbs have always done -- correctly -- but the return value said
|
||||
# nothing about it, so a statewide total resting on 18% of the universe looked
|
||||
# exactly like one resting on 98%.
|
||||
#
|
||||
# Owner's settled design: a `coverage` argument selecting WHICH units to
|
||||
# include, plus always-on metadata saying how many there were either way. The
|
||||
# principle behind it: using these verbs correctly must not require the caller
|
||||
# to know the survey calendar.
|
||||
|
||||
# Years ending in 2 or 7 are full censuses of every government; all others are
|
||||
# samples.
|
||||
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
|
||||
|
||||
#' @noRd
|
||||
.is_census_year <- function(years) {
|
||||
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
|
||||
}
|
||||
|
||||
#' @noRd
|
||||
.validate_coverage <- function(coverage) {
|
||||
tryCatch(
|
||||
match.arg(coverage, c("all", "census", "consistent")),
|
||||
error = function(e) {
|
||||
cli::cli_abort(
|
||||
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
|
||||
class = "uscogdata_invalid_coverage", parent = e
|
||||
)
|
||||
}
|
||||
)
|
||||
}
|
||||
|
||||
#' Restrict `years` to census years for `coverage = "census"`.
|
||||
#'
|
||||
#' Aborts rather than returning an empty result when the requested range holds
|
||||
#' no census year: silently handing back zero rows for a query the caller
|
||||
#' believes they made is the failure mode this whole issue is about.
|
||||
#' @noRd
|
||||
.apply_census_years <- function(years, coverage, verb) {
|
||||
if (!identical(coverage, "census")) return(as.integer(years))
|
||||
keep <- as.integer(years)[.is_census_year(years)]
|
||||
if (length(keep) == 0L) {
|
||||
cli::cli_abort(c(
|
||||
"{.code coverage = \"census\"} leaves no years to query.",
|
||||
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
|
||||
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
|
||||
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
|
||||
), class = "uscogdata_no_census_years")
|
||||
}
|
||||
sort(keep)
|
||||
}
|
||||
|
||||
#' Keep only units that report in EVERY requested year (a balanced panel).
|
||||
#'
|
||||
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
|
||||
#' filter (the peer-comparison target, which is the subject of the comparison
|
||||
#' rather than a member of the cohort being balanced).
|
||||
#' @noRd
|
||||
.filter_consistent <- function(result, years, id_col = "canonical_govid",
|
||||
keep_ids = character(0)) {
|
||||
years <- unique(as.integer(years))
|
||||
if (nrow(result) == 0L || length(years) <= 1L) return(result)
|
||||
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
|
||||
present <- vapply(ids, function(g) {
|
||||
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
|
||||
}, logical(1))
|
||||
consistent <- c(ids[present], keep_ids)
|
||||
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
|
||||
drop = FALSE]
|
||||
}
|
||||
|
||||
#' Per-year coverage metadata, always attached regardless of mode.
|
||||
#'
|
||||
#' Built from the REQUESTED years rather than the years present in the result,
|
||||
#' so a year in which nothing reported still appears -- with
|
||||
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
|
||||
#' missing year fails to make.
|
||||
#'
|
||||
#' `n_units_reporting` describes the result the caller actually received, so
|
||||
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
|
||||
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
|
||||
#' `n_units_reporting` is the number that tells the truth.
|
||||
#' @noRd
|
||||
.coverage_table <- function(result, years, n_expected,
|
||||
id_col = "canonical_govid", rows = NULL) {
|
||||
years <- sort(unique(as.integer(years)))
|
||||
src <- if (is.null(rows)) result else rows
|
||||
reporting <- vapply(years, function(y) {
|
||||
ids <- src[[id_col]][as.integer(src$year) == y]
|
||||
length(unique(ids[!is.na(ids)]))
|
||||
}, integer(1))
|
||||
tibble::tibble(
|
||||
year = years,
|
||||
n_units_reporting = as.integer(reporting),
|
||||
n_units_expected = rep(as.integer(n_expected), length(years)),
|
||||
is_census_year = .is_census_year(years)
|
||||
)
|
||||
}
|
||||
+114
-2
@@ -11,6 +11,30 @@
|
||||
#' returns `result` invisibly for chaining. `"list"` returns the raw
|
||||
#' provenance list (identical to `attr(result, "provenance")`).
|
||||
#' @return Either `result` (invisibly) or the provenance list.
|
||||
#' @section Two kinds of series break:
|
||||
#' Catalogued breaks reach you without being asked for, in two disjoint
|
||||
#' fields, because a caveat about one series and a caveat about the whole
|
||||
#' corpus are different claims:
|
||||
#'
|
||||
#' * **`series_break_refs`** — breaks matched against the item codes actually
|
||||
#' present in this result. A break in one code you queried.
|
||||
#' * **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
|
||||
#' which are statements about the corpus rather than about any one code:
|
||||
#' dollar precision across the 1976/1977 boundary (`SB085`), imputation
|
||||
#' exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
|
||||
#' representation change (`SB194`), and the FY2017 government-identifier
|
||||
#' change (`SB086`). These are selected on the break-year window alone.
|
||||
#'
|
||||
#' `SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
|
||||
#' crosses the boundary where an absent cell stops meaning "Census published
|
||||
#' $0" and starts meaning "not reported".
|
||||
#' @section Other provenance blocks:
|
||||
#' `transformations$units_conversion` records the `$1,000s`-to-dollars
|
||||
#' multiply that every amount column has already had applied.
|
||||
#' `transformations$per_capita` records the population denominator and its
|
||||
#' year range. `coverage` and `coverage_mode` appear on multi-government
|
||||
#' results (see [cog_geographic_rollup()]). `completion` appears when
|
||||
#' `complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
|
||||
#' @export
|
||||
cog_explain <- function(result, format = c("print", "list")) {
|
||||
format <- match.arg(format)
|
||||
@@ -60,7 +84,20 @@ cog_explain <- function(result, format = c("print", "list")) {
|
||||
cli::cli_text("Basis: {prov$basis}{note}")
|
||||
}
|
||||
|
||||
if (!is.null(prov$expenditure_concept)) {
|
||||
# Each verb reports its OWN concept. Both fields are always present (each
|
||||
# defaults to its concept's default), so printing `expenditure_concept`
|
||||
# unconditionally would tell a cog_revenue() caller "Concept: primary",
|
||||
# which names a spending concept their result has nothing to do with.
|
||||
if (identical(prov$verb, "cog_revenue")) {
|
||||
if (!is.null(prov$revenue_concept)) {
|
||||
cli::cli_text("Concept: {prov$revenue_concept} revenue")
|
||||
}
|
||||
} else if (identical(prov$verb, "cog_balances")) {
|
||||
# Both concept fields are deliberately NA here (a stock has no flow
|
||||
# concept). Printing the raw NA reads as a missing value rather than an
|
||||
# intentional one, so say what it means instead.
|
||||
cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)")
|
||||
} else if (!is.null(prov$expenditure_concept)) {
|
||||
concept_note <- if (!is.null(prov$expenditure_concept_note) &&
|
||||
!is.na(prov$expenditure_concept_note)) {
|
||||
sprintf(" (%s)", prov$expenditure_concept_note)
|
||||
@@ -112,17 +149,92 @@ cog_explain <- function(result, format = c("print", "list")) {
|
||||
if (length(prov$suggestions) > 0L) {
|
||||
cli::cli_h2("Suggestions")
|
||||
sugg_lines <- vapply(prov$suggestions, function(s) {
|
||||
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
|
||||
line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
|
||||
s$available_years[1], s$available_years[2], s$hint)
|
||||
if (isTRUE(s$suppressed_amount > 0)) {
|
||||
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
|
||||
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
|
||||
paste0("FY", s$suppressed_years, collapse = ", "),
|
||||
paste(s$suppressed_codes, collapse = ", ")))
|
||||
}
|
||||
line
|
||||
}, character(1))
|
||||
cli::cli_ul(sugg_lines)
|
||||
}
|
||||
|
||||
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
|
||||
cli::cli_h2("Reporting coverage")
|
||||
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
|
||||
cov <- prov$coverage
|
||||
cli::cli_ul(sprintf(
|
||||
"%d: %d of %d units reporting (%.0f%%) -- %s year",
|
||||
cov$year, cov$n_units_reporting, cov$n_units_expected,
|
||||
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
|
||||
ifelse(cov$is_census_year, "census", "sample")
|
||||
))
|
||||
if (any(!cov$is_census_year)) {
|
||||
cli::cli_text(
|
||||
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
if (isTRUE(prov$completion$applied)) {
|
||||
cli::cli_h2("Completion")
|
||||
cli::cli_text(
|
||||
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
|
||||
)
|
||||
rules <- prov$completion$absence_means
|
||||
if (length(rules) > 0L) {
|
||||
cli::cli_ul(vapply(names(rules), function(y) {
|
||||
sprintf("%s: an absent cell means %s", y,
|
||||
if (identical(rules[[y]], "census_zero")) {
|
||||
"Census published $0 (filled as 0)"
|
||||
} else {
|
||||
"the government did not report (filled as NA, not 0)"
|
||||
})
|
||||
}, character(1)))
|
||||
}
|
||||
}
|
||||
|
||||
if (length(prov$series_break_refs) > 0L) {
|
||||
cli::cli_h2("Series breaks")
|
||||
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
||||
}
|
||||
|
||||
# Kept in a section of its own: these qualify the whole result, so folding
|
||||
# them in with the per-code breaks above would invite reading them as a
|
||||
# caveat about one series.
|
||||
if (length(prov$corpus_break_refs) > 0L) {
|
||||
cli::cli_h2("Corpus-wide caveats")
|
||||
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
|
||||
}
|
||||
|
||||
# Balance results only (NULL on money-verb provenance, so they are
|
||||
# unaffected). This is the ONLY on-demand surface for the GAAP disclosure:
|
||||
# .emit_balance_caveats() fires at most once per session, and is routinely
|
||||
# consumed by a suppressMessages() call or by a knitted chunk nobody reads,
|
||||
# so a caller who deliberately audits a result with cog_explain() must still
|
||||
# be told.
|
||||
bc <- prov$balance_caveats
|
||||
if (!is.null(bc)) {
|
||||
cli::cli_h2("Holdings caveats")
|
||||
if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note)
|
||||
if (length(bc$truncated) > 0L) {
|
||||
cli::cli_text(
|
||||
"Requested years extend beyond what these families actually cover:"
|
||||
)
|
||||
cli::cli_ul(vapply(bc$truncated, function(s) {
|
||||
w <- bc$coverage_window[[s]]
|
||||
if (length(w) == 2L) {
|
||||
sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2])
|
||||
} else {
|
||||
s
|
||||
}
|
||||
}, character(1)))
|
||||
}
|
||||
}
|
||||
|
||||
cli::cli_h2("Transformations")
|
||||
uc <- prov$transformations$units_conversion
|
||||
if (isTRUE(uc$applied)) {
|
||||
|
||||
+2
-2
@@ -28,7 +28,7 @@
|
||||
"*" = "{.code Sys.setenv(USCOGDATA_URL = \"<url-or-local-path>/\")}",
|
||||
"*" = "{.code options(uscogdata.url = \"<url-or-local-path>/\")}",
|
||||
i = "For an offline smoke test, use the bundled fixture: {.code system.file(\"extdata/fixture_corpus\", package = \"uscogdata\")}.",
|
||||
i = "For the live Civilytics corpus, request the Nextcloud share URL from the package maintainer."
|
||||
i = "The public corpus is the default: unset USCOGDATA_URL to use it, or point it at a local copy made by {.code cog_mirror()}."
|
||||
), class = "uscogdata_url_not_configured")
|
||||
}
|
||||
invisible(url)
|
||||
@@ -138,7 +138,7 @@
|
||||
#' year, matching canonical_fips_xwalk) rather than as-of-year; as-of-year
|
||||
#' moved to the *_asof columns. This package's own geography always came from
|
||||
#' the xwalk (already present-based), so behaviour is unchanged.
|
||||
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L)) {
|
||||
.validate_schema <- function(manifest, supported = c(4L, 5L, 6L, 7L)) {
|
||||
if (!manifest$schema_version %in% supported) {
|
||||
cli::cli_abort(c(
|
||||
"Corpus schema version mismatch.",
|
||||
|
||||
@@ -19,6 +19,13 @@
|
||||
#' target's population at `year` to produce absolute bounds. If `FALSE`,
|
||||
#' `pop_range` is interpreted as absolute population counts.
|
||||
#' @param max_peers Integer cap on the number of peers returned.
|
||||
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
|
||||
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
||||
#' most recent census year with an observed population, so a cohort is not
|
||||
#' built from a sample year in which most of the candidate universe is
|
||||
#' absent. `"consistent"` needs a year range, which cohort selection does not
|
||||
#' have, so it selects like `"all"` and is carried on the result as
|
||||
#' `attr(x, "coverage")` for [cog_peer_compare()].
|
||||
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
|
||||
#' `attr(x, "cohort_year")`.
|
||||
@@ -29,7 +36,9 @@ cog_find_peers <- function(target_govid,
|
||||
same_state = FALSE,
|
||||
pop_range = c(0.7, 1.3),
|
||||
is_ratio = TRUE,
|
||||
max_peers = 10L) {
|
||||
max_peers = 10L,
|
||||
coverage = c("all", "census", "consistent")) {
|
||||
coverage <- .validate_coverage(coverage)
|
||||
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
||||
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
||||
}
|
||||
@@ -59,7 +68,7 @@ cog_find_peers <- function(target_govid,
|
||||
))
|
||||
}
|
||||
|
||||
cohort_year <- .resolve_cohort_year(con, target_govid, year)
|
||||
cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
|
||||
|
||||
pop_sql <- sprintf(
|
||||
"SELECT population FROM gov_population_yearly
|
||||
@@ -107,12 +116,34 @@ cog_find_peers <- function(target_govid,
|
||||
attr(peers, "cohort_year") <- as.integer(cohort_year)
|
||||
attr(peers, "pop_range") <- as.numeric(pop_range)
|
||||
attr(peers, "is_ratio") <- isTRUE(is_ratio)
|
||||
attr(peers, "coverage") <- coverage
|
||||
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
|
||||
peers
|
||||
}
|
||||
|
||||
# `coverage` picks the cohort vintage when the caller did not name one.
|
||||
# "census" snaps to the most recent CENSUS year with an observed population,
|
||||
# so a cohort is not silently built from a sample year in which most of the
|
||||
# candidate universe is absent. "consistent" is a comparison-time concept --
|
||||
# it needs a year RANGE, which cohort selection does not have -- so it selects
|
||||
# like "all" here and is carried on the result for cog_peer_compare().
|
||||
#' @noRd
|
||||
.resolve_cohort_year <- function(con, target_govid, year) {
|
||||
.resolve_cohort_year <- function(con, target_govid, year,
|
||||
coverage = "all") {
|
||||
if (!is.null(year)) return(as.integer(year))
|
||||
if (identical(coverage, "census")) {
|
||||
sql <- sprintf(
|
||||
"SELECT MAX(year) AS y FROM gov_population_yearly
|
||||
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
|
||||
.sql_lit_chr(target_govid)
|
||||
)
|
||||
y <- DBI::dbGetQuery(con, sql)$y
|
||||
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
|
||||
cli::cli_abort(c(
|
||||
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
|
||||
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
|
||||
), class = "uscogdata_no_census_years")
|
||||
}
|
||||
sql <- sprintf(
|
||||
"SELECT MAX(year) AS y FROM gov_population_yearly
|
||||
WHERE canonical_govid = %s",
|
||||
@@ -133,7 +164,9 @@ cog_find_peers <- function(target_govid,
|
||||
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||
#' call.
|
||||
#' call. Those summary rows are quantiles **within each category**, not
|
||||
#' quantiles of each peer's total — see the `@return` section before summing
|
||||
#' them.
|
||||
#'
|
||||
#' @param target_govid Character scalar.
|
||||
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
||||
@@ -143,10 +176,36 @@ cog_find_peers <- function(target_govid,
|
||||
#' @param per_capita Default `TRUE` — peer compare usually normalizes by
|
||||
#' population.
|
||||
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
|
||||
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
|
||||
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||
#' single-government queries but cannot be used here because combining Total
|
||||
#' across peer sets counts intergovernmental transfers twice.
|
||||
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
|
||||
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
||||
#' refused here because combining Total across peer sets counts
|
||||
#' intergovernmental transfers twice; `"primary"` and `"direct"` combine
|
||||
#' safely.
|
||||
#' @param coverage How to handle the Census of Governments survey cycle,
|
||||
#' which is a **complete census only in years ending in 2 and 7** -- every
|
||||
#' other year is a sample, and the sample varies enormously (on the bundled
|
||||
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||
#' and 112 in FY2019).
|
||||
#'
|
||||
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
||||
#' behaviour, so existing code keeps working.
|
||||
#' * `"census"` -- census years only. Aborts if the requested range holds
|
||||
#' none, rather than silently returning nothing.
|
||||
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
||||
#' a balanced panel.
|
||||
#'
|
||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
#' statement about the **survey calendar**, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
#' report. `n_units_reporting` is the number that tells the truth.
|
||||
#'
|
||||
#' The comparison target is exempt from `"consistent"` balancing -- it is the
|
||||
#' subject of the comparison, not a member of the cohort -- and the
|
||||
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
|
||||
#' cohort actually returned. `n_units_reporting` counts peers only, against
|
||||
#' the cohort size: "3 of your 15 peers reported in FY2019".
|
||||
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
||||
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
||||
@@ -155,12 +214,57 @@ cog_find_peers <- function(target_govid,
|
||||
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||
#' `cohort_year`, and `cohort_govids`.
|
||||
#'
|
||||
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||
#' Each one is computed **within each `(year, spend_subtype,
|
||||
#' category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||
#' median peer's value in that one category*, not *the value of the median
|
||||
#' peer's total*. The median peer for Police and the median peer for Fire
|
||||
#' are usually different governments, so summing `summary_*` rows across
|
||||
#' categories does not give any peer's total and misstates the band it
|
||||
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||
#' one cohort, with a sign flip at FY2012.
|
||||
#'
|
||||
#' Facet by `role` **and** `category` (the documented use, and what the
|
||||
#' rows are built for). For a genuine "median peer's total spending" line,
|
||||
#' sum each peer's own categories first and take the quantile of those
|
||||
#' per-government totals:
|
||||
#'
|
||||
#' ```r
|
||||
#' library(dplyr)
|
||||
#' cmp |>
|
||||
#' filter(role %in% c("target", "peer")) |>
|
||||
#' group_by(year, role, canonical_govid) |>
|
||||
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||
#' filter(role == "peer") |>
|
||||
#' group_by(year) |>
|
||||
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||
#' ```
|
||||
#' @section Reading `coverage`:
|
||||
#' `provenance$coverage` reports `n_units_reporting` against
|
||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
#' it counts cohort members with rows for the category you asked for, not
|
||||
#' cohort members collected that year.** A government that was surveyed and
|
||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
||||
#' that was never surveyed.
|
||||
#'
|
||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
#' contract policing to the county sheriff, not non-response.
|
||||
#'
|
||||
#' The comparison that *is* valid is the same category across a census year
|
||||
#' (ending in 2 or 7) and a sample year, where the real-zero component is
|
||||
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
|
||||
#' marks which is which.
|
||||
#' @export
|
||||
cog_peer_compare <- function(target_govid, peers, category, years,
|
||||
per_capita = TRUE, adjust_to_year = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
coverage = c("all", "census", "consistent")) {
|
||||
call <- match.call()
|
||||
expenditure_concept <- match.arg(expenditure_concept)
|
||||
coverage <- .validate_coverage(coverage)
|
||||
if (identical(expenditure_concept, "total")) {
|
||||
.abort_concept_not_aggregatable("cog_peer_compare")
|
||||
}
|
||||
@@ -183,9 +287,21 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
||||
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
|
||||
all_govids <- unique(c(target_govid, peer_govids))
|
||||
|
||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||
years <- .apply_census_years(years, coverage, "cog_peer_compare")
|
||||
|
||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
|
||||
expenditure_concept = expenditure_concept)
|
||||
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
|
||||
|
||||
# The target is exempt from balancing: it is the subject of the comparison,
|
||||
# not a member of the cohort being balanced, and dropping it would leave a
|
||||
# peer comparison with nothing to compare. Filtering happens BEFORE the
|
||||
# quantiles below, so a "consistent" cohort's summary rows describe that
|
||||
# cohort rather than the unbalanced one.
|
||||
if (identical(coverage, "consistent")) {
|
||||
r <- .filter_consistent(r, years, keep_ids = target_govid)
|
||||
}
|
||||
|
||||
value_col <- .peer_value_col(per_capita, adjust_to_year)
|
||||
|
||||
summary_rows <- .peer_summary_rows(r, value_col)
|
||||
@@ -206,6 +322,14 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
||||
canonical_govid = target_govid,
|
||||
gov_name = unique(r$gov_name[r$role == "target"])
|
||||
)
|
||||
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
|
||||
# reported in FY2019". Including the target would inflate every count by one
|
||||
# and make a cohort that has entirely stopped reporting look non-empty.
|
||||
prov$coverage_mode <- coverage
|
||||
prov$coverage <- .coverage_table(
|
||||
out, years, length(peer_govids),
|
||||
rows = r[r$role == "peer", , drop = FALSE]
|
||||
)
|
||||
attr(out, "provenance") <- prov
|
||||
out
|
||||
}
|
||||
|
||||
+32
-4
@@ -6,11 +6,13 @@
|
||||
per_capita, adjust_to_year, result, sql,
|
||||
subtype_col, basis = NA_character_,
|
||||
basis_note = NA_character_,
|
||||
expenditure_concept = "direct",
|
||||
expenditure_concept = "primary",
|
||||
expenditure_concept_note = NA_character_,
|
||||
expenditure_concept_direct_suppressed = FALSE,
|
||||
revenue_concept = "general",
|
||||
harmonization = NULL, recipe = NULL,
|
||||
suggestions = list()) {
|
||||
suggestions = list(),
|
||||
completion = NULL) {
|
||||
manifest <- .uscogdata_env$manifest
|
||||
|
||||
codes <- result[["codes_included"]]
|
||||
@@ -37,11 +39,20 @@
|
||||
|
||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||
con <- .uscogdata_env$con
|
||||
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) {
|
||||
have_con <- !is.null(con) && DBI::dbIsValid(con)
|
||||
break_refs <- if (have_con) {
|
||||
.build_series_break_refs(con, codes_observed, years, schema_version)
|
||||
} else {
|
||||
character(0)
|
||||
}
|
||||
# Corpus-wide caveats travel separately: they qualify the whole result
|
||||
# rather than one series, and they do not depend on codes_observed (see
|
||||
# .build_corpus_break_refs()).
|
||||
corpus_refs <- if (have_con) {
|
||||
.build_corpus_break_refs(con, years, schema_version)
|
||||
} else {
|
||||
character(0)
|
||||
}
|
||||
|
||||
list(
|
||||
verb = verb,
|
||||
@@ -56,7 +67,16 @@
|
||||
basis_note = basis_note,
|
||||
expenditure_concept = expenditure_concept,
|
||||
expenditure_concept_note = expenditure_concept_note,
|
||||
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
|
||||
# isTRUE() alone would collapse a deliberate NA (all-categories mode,
|
||||
# where suppression detection cannot run -- see .verb_spendrev()) down to
|
||||
# FALSE, turning "we don't know" back into the false claim this field
|
||||
# exists to avoid. Preserve NA; otherwise normalize to a strict logical.
|
||||
expenditure_concept_direct_suppressed = if (isTRUE(is.na(expenditure_concept_direct_suppressed))) {
|
||||
NA
|
||||
} else {
|
||||
isTRUE(expenditure_concept_direct_suppressed)
|
||||
},
|
||||
revenue_concept = revenue_concept,
|
||||
harmonization = harmonization %||% list(
|
||||
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
|
||||
note = NA_character_
|
||||
@@ -116,6 +136,14 @@
|
||||
)
|
||||
),
|
||||
series_break_refs = break_refs,
|
||||
corpus_break_refs = corpus_refs,
|
||||
# What `complete = TRUE` filled, and the rule it filled by. Always
|
||||
# present so a consumer can read `completion$applied` without testing
|
||||
# for the key -- an absent block and applied = FALSE would otherwise be
|
||||
# indistinguishable from an older reader version.
|
||||
completion = completion %||% list(
|
||||
applied = FALSE, rows_filled = 0L, absence_means = list()
|
||||
),
|
||||
manifest = list(
|
||||
schema_version = as.integer(manifest$schema_version),
|
||||
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
||||
|
||||
+50
-3
@@ -8,14 +8,57 @@
|
||||
#' multiplies by 1000 and records the conversion in `provenance`).
|
||||
#'
|
||||
#' @inheritParams cog_spending
|
||||
#' @param category Character vector of category names (from
|
||||
#' `summary_categories.category`), or `NULL` for all categories broken out
|
||||
#' one row each. The reserved value `"All Categories"` instead returns a
|
||||
#' single summed row per `(year, canonical_govid, subtype)`, covering every
|
||||
#' category inside the requested concept's subtype scope. It cannot be
|
||||
#' combined with other category names, and it is not the same thing as
|
||||
#' `revenue_concept = "total"`: the concept chooses which subtypes are in
|
||||
#' scope, `"All Categories"` chooses whether rows inside that scope are
|
||||
#' broken out or summed. Because the result keeps one row per
|
||||
#' `revenue_subtype`, filtering the returned frame to
|
||||
#' `revenue_subtype == "own_source"` gives an own-source revenue total.
|
||||
#' @param revenue_concept Which of Census's two published revenue concepts to
|
||||
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
||||
#' values -- never as item-code first letters, which cannot classify
|
||||
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
|
||||
#' prefix `X` does the same):
|
||||
#'
|
||||
#' * `"general"` (default) -- Census General Revenue: `own_source` +
|
||||
#' `federal` + `state` + `local_aid`. The manual defines this concept by
|
||||
#' subtraction (section 4.3: *"General revenue comprises all revenue
|
||||
#' except that classified as liquor store, utility, or insurance trust
|
||||
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
|
||||
#' insurance trust revenue are all excluded.
|
||||
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
|
||||
#' `general` plus utility, liquor store, and insurance trust revenue
|
||||
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
|
||||
#' `X01`/`X02`/`X05`/`X08`).
|
||||
#'
|
||||
#' The two are related by Census's own identity, `Total Revenue = General +
|
||||
#' Utility + Liquor Store + Insurance Trust`.
|
||||
#'
|
||||
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
|
||||
#' systems moved out of the annual finance file into the separate Annual
|
||||
#' Survey of Public Pensions, so a `"total"` series steps down at the
|
||||
#' FY2016/FY2017 seam for reasons that are about collection scope rather
|
||||
#' than revenue (series breaks `SB197`-`SB202`).
|
||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
#' and `value_source` when `complete = TRUE`.
|
||||
#' @export
|
||||
cog_revenue <- function(govid, years, category = NULL,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
revenue_concept = c("general", "total"),
|
||||
complete = FALSE, limit = NULL, offset = NULL) {
|
||||
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
|
||||
# membership does -- General Revenue, i.e. everything except
|
||||
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
|
||||
# this verb's recipe families (see R/suggestions.R).
|
||||
.verb_spendrev(
|
||||
verb = "cog_revenue",
|
||||
view_base = "revenue_annotated",
|
||||
@@ -28,6 +71,10 @@ cog_revenue <- function(govid, years, category = NULL,
|
||||
per_capita = per_capita,
|
||||
adjust_to_year = adjust_to_year,
|
||||
basis = basis,
|
||||
recipe = recipe
|
||||
recipe = recipe,
|
||||
revenue_concept = revenue_concept,
|
||||
complete = complete,
|
||||
limit = limit,
|
||||
offset = offset
|
||||
)
|
||||
}
|
||||
|
||||
+66
-9
@@ -19,30 +19,72 @@
|
||||
#' `state`, `county`, `city`. Each element is a character vector of
|
||||
#' `canonical_govid` values. At least one layer required.
|
||||
#' @param category Single category name or character vector (passed through
|
||||
#' to [cog_spending()]).
|
||||
#' to [cog_spending()]), or the reserved `"All Categories"` for one summed
|
||||
#' row per `(year, canonical_govid, subtype)` covering every category in the
|
||||
#' concept's scope. `"All Categories"` is the efficient way to build a
|
||||
#' geographic total: without it a caller must issue one rollup per category
|
||||
#' and sum the results themselves.
|
||||
#' @param years Integer vector of years.
|
||||
#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year
|
||||
#' population from `gov_population_yearly`. Govs with missing population
|
||||
#' are excluded from the result.
|
||||
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
|
||||
#' @param expenditure_concept `"direct"` (default) or `"total"`. Currently only
|
||||
#' `"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||
#' single-government queries but cannot be used here because combining Total
|
||||
#' across multiple layers of government double-counts intergovernmental
|
||||
#' transfers (a state's payment to a school district is the same dollar the
|
||||
#' district reports as its own Direct spending).
|
||||
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
|
||||
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
||||
#' refused here because combining Total across multiple layers of
|
||||
#' government double-counts intergovernmental transfers (a state's payment
|
||||
#' to a school district is the same dollar the district reports as its own
|
||||
#' Direct spending); `"primary"` and `"direct"` combine safely.
|
||||
#' @param coverage How to handle the Census of Governments survey cycle,
|
||||
#' which is a **complete census only in years ending in 2 and 7** -- every
|
||||
#' other year is a sample, and the sample varies enormously (on the bundled
|
||||
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||
#' and 112 in FY2019).
|
||||
#'
|
||||
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
||||
#' behaviour, so existing code keeps working.
|
||||
#' * `"census"` -- census years only. Aborts if the requested range holds
|
||||
#' none, rather than silently returning nothing.
|
||||
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
||||
#' a balanced panel.
|
||||
#'
|
||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
#' statement about the **survey calendar**, never a claim of completeness:
|
||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
#' report. `n_units_reporting` is the number that tells the truth.
|
||||
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
||||
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
||||
#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a
|
||||
#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`,
|
||||
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
||||
#' @section Reading `coverage`:
|
||||
#' `provenance$coverage` reports `n_units_reporting` against
|
||||
#' `n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
#' it counts governments with rows for the category you asked for, not
|
||||
#' governments collected that year.** A government that was surveyed and
|
||||
#' genuinely spends nothing in that category is indistinguishable here from one
|
||||
#' that was never surveyed.
|
||||
#'
|
||||
#' The ratio is therefore **not a response rate** and must not be used as one.
|
||||
#' In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
#' `category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
#' contract policing to the county sheriff, not non-response.
|
||||
#'
|
||||
#' The comparison that *is* valid is the same category across a census year
|
||||
#' (ending in 2 or 7) and a sample year, where the real-zero component is
|
||||
#' roughly constant and the difference reflects the survey cycle. `is_census_year`
|
||||
#' marks which is which.
|
||||
#' @export
|
||||
cog_geographic_rollup <- function(govids, category, years,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
coverage = c("all", "census", "consistent")) {
|
||||
call <- match.call()
|
||||
expenditure_concept <- match.arg(expenditure_concept)
|
||||
coverage <- .validate_coverage(coverage)
|
||||
if (identical(expenditure_concept, "total")) {
|
||||
.abort_concept_not_aggregatable("cog_geographic_rollup")
|
||||
}
|
||||
@@ -59,11 +101,21 @@ cog_geographic_rollup <- function(govids, category, years,
|
||||
layer = rep(layer_names, lengths(govids))
|
||||
)
|
||||
|
||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||
# coverage = "census" drops non-census years BEFORE the query rather than
|
||||
# after: a sample year's rows are not wanted at all, and fetching them only
|
||||
# to discard them would also let them into the coverage table.
|
||||
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
|
||||
|
||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
|
||||
expenditure_concept = expenditure_concept)
|
||||
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
|
||||
relationship = "many-to-many")
|
||||
r$scope_note <- .rollup_scope_note(r$layer)
|
||||
|
||||
if (identical(coverage, "consistent")) {
|
||||
r <- .filter_consistent(r, years)
|
||||
}
|
||||
|
||||
excluded <- character(0)
|
||||
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
|
||||
drop <- r$pop_source == "unavailable"
|
||||
@@ -82,6 +134,11 @@ cog_geographic_rollup <- function(govids, category, years,
|
||||
included_govids = included,
|
||||
excluded_govids = excluded
|
||||
)
|
||||
# n_units_expected is the universe the CALLER named -- the govids passed in
|
||||
# -- not the national universe. That is what makes the ratio meaningful:
|
||||
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
|
||||
prov$coverage_mode <- coverage
|
||||
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
|
||||
attr(r, "provenance") <- prov
|
||||
|
||||
r
|
||||
|
||||
+19
-7
@@ -6,8 +6,11 @@
|
||||
#' the cross-vintage canonical-government registry. Operates in two modes:
|
||||
#'
|
||||
#' * **Utility mode** (single `name`, the original behavior): returns all
|
||||
#' rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||
#' `population_acs` descending. Useful for exploratory lookups.
|
||||
#' rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||
#' substring**, sorted by `population_acs` descending. Useful for
|
||||
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||
#' government is findable by its own complete name even when that name
|
||||
#' contains parentheses or a period.
|
||||
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||
#' single canonical govid and returns a tibble in input order, suitable
|
||||
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||
@@ -19,7 +22,8 @@
|
||||
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||
#' Single hit -> resolved. Multiple -> step 4.
|
||||
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||
#' 3. **Substring fallback:** case-insensitive literal substring against
|
||||
#' `gov_name` (metacharacters escaped).
|
||||
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||
#' `status = "no_match"`. Multiple hits -> step 4.
|
||||
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||
@@ -48,7 +52,7 @@
|
||||
#' [cog_spending()], [cog_revenue()].
|
||||
#' @examples
|
||||
#' \dontrun{
|
||||
#' # Utility mode — exploratory regex lookup
|
||||
#' # Utility mode — exploratory substring lookup
|
||||
#' cog_gov_search("broward", state = "FL")
|
||||
#'
|
||||
#' # Basket mode — resolve a known cohort
|
||||
@@ -98,9 +102,16 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
||||
if (!is.character(name) || length(name) != 1L) {
|
||||
cli::cli_abort("`name` must be a length-1 character string.")
|
||||
}
|
||||
# Escaped, so `name` is a literal case-insensitive substring -- the same
|
||||
# treatment basket mode has always given it. Interpolating it raw made a
|
||||
# government unfindable by its own name whenever that name contains a
|
||||
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
|
||||
# match-everything wildcard, and let malformed pattern text reach the
|
||||
# engine as an error -- which cog-api surfaced as a 500, reachable by
|
||||
# typing a real name one character at a time (uscogdata#16, F-025).
|
||||
preds <- c(preds,
|
||||
sprintf("regexp_matches(gov_name, %s, 'i')",
|
||||
.sql_lit_chr(name)))
|
||||
.sql_lit_chr(.escape_regex(name))))
|
||||
}
|
||||
if (!is.null(state)) {
|
||||
st_fips <- .coerce_state_to_fips(state)
|
||||
@@ -136,8 +147,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
||||
#' @noRd
|
||||
.escape_regex <- function(x) {
|
||||
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
||||
# literal substring in the DuckDB regexp_matches call (substring fallback
|
||||
# only; utility-mode intentionally preserves regex behavior).
|
||||
# literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
|
||||
# utility mode used to interpolate raw, which was a defect rather than a
|
||||
# feature -- see the call site and uscogdata#16.
|
||||
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
||||
}
|
||||
|
||||
|
||||
+33
-1
@@ -14,9 +14,41 @@
|
||||
sql <- sprintf(
|
||||
"SELECT DISTINCT break_id
|
||||
FROM series_breaks_pq
|
||||
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d
|
||||
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
|
||||
AND break_year BETWEEN %d AND %d
|
||||
ORDER BY break_id",
|
||||
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
||||
)
|
||||
DBI::dbGetQuery(con, sql)$break_id
|
||||
}
|
||||
|
||||
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
|
||||
#' `"ALL"` rather than an item code. They qualify the whole result, so they
|
||||
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
|
||||
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
|
||||
#' before uscogdata#19. Selection is on the break_year window alone: which
|
||||
#' codes a result happens to contain is irrelevant to a caveat about the
|
||||
#' corpus.
|
||||
#'
|
||||
#' All four catalogued entries are *boundary* caveats (dollar precision
|
||||
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
|
||||
#' representation change at 2012, the id scheme change at 2017), so the same
|
||||
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
|
||||
#' path uses is the right one -- a request that never crosses the boundary
|
||||
#' is not affected by it.
|
||||
#'
|
||||
#' Returned separately from `series_break_refs` so a consumer can tell a
|
||||
#' whole-result caveat from a break in one series; the two are disjoint by
|
||||
#' construction.
|
||||
#' @noRd
|
||||
.build_corpus_break_refs <- function(con, years, schema_version) {
|
||||
if (schema_version < 5L || length(years) == 0L) return(character(0))
|
||||
sql <- sprintf(
|
||||
"SELECT DISTINCT break_id
|
||||
FROM series_breaks_pq
|
||||
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
|
||||
ORDER BY break_id",
|
||||
min(as.integer(years)), max(as.integer(years))
|
||||
)
|
||||
DBI::dbGetQuery(con, sql)$break_id
|
||||
}
|
||||
|
||||
+4
-1
@@ -12,7 +12,7 @@ cog_open <- function(url = .resolve_url(),
|
||||
DBI::dbExecute(con, "INSTALL httpfs; LOAD httpfs;")
|
||||
|
||||
manifest <- .fetch_or_cache_manifest(url, cache_dir)
|
||||
.validate_schema(manifest, supported = c(4L, 5L, 6L))
|
||||
.validate_schema(manifest, supported = c(4L, 5L, 6L, 7L))
|
||||
.validate_scope(manifest)
|
||||
|
||||
.register_views(con, url, manifest)
|
||||
@@ -95,4 +95,7 @@ cog_close <- function() {
|
||||
}
|
||||
.uscogdata_env$con <- NULL
|
||||
.uscogdata_env$manifest <- NULL
|
||||
.uscogdata_env$balance_caveats_shown <- NULL
|
||||
# Memoised corpus-constant; a different corpus may be mounted next.
|
||||
.uscogdata_env$balance_coverage_windows <- NULL
|
||||
}
|
||||
|
||||
+438
-48
@@ -1,5 +1,65 @@
|
||||
# R/spending.R
|
||||
|
||||
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
|
||||
# `spend_subtype` values. Classification is crosswalk membership, never
|
||||
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
|
||||
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
|
||||
# route it (finding F-018).
|
||||
#
|
||||
# primary = operations + capital + assistance (the default)
|
||||
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
|
||||
# total = direct + intergovernmental (via the ig_* views)
|
||||
#
|
||||
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
|
||||
# than intergovernmental -- including payments to retirees, i.e. insurance
|
||||
# trust benefits. Verified against Census's own published FY2020 state
|
||||
# aggregates (20statetypepu.txt): `total` reproduces the published
|
||||
# expenditure sum to the dollar; omitting insurance benefits understates
|
||||
# California's Direct by 10.9%.
|
||||
.spend_subtypes_primary <- c("operations", "capital", "assistance")
|
||||
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
|
||||
|
||||
# The reserved pseudo-category. Deliberately NOT "Total": `category = "Total"`
|
||||
# would sit one argument away from `expenditure_concept = "total"` and mean
|
||||
# something different -- the concept selects WHICH SUBTYPES are in scope, this
|
||||
# selects whether the rows inside that scope are broken out by category or
|
||||
# summed. "All Categories" states the operation and cannot be misread as the
|
||||
# concept.
|
||||
.ALL_CATEGORIES <- "All Categories"
|
||||
|
||||
#' @noRd
|
||||
.expenditure_concept_subtypes <- function(concept) {
|
||||
switch(concept,
|
||||
primary = .spend_subtypes_primary,
|
||||
# "total" = the direct subtypes here PLUS the intergovernmental leg,
|
||||
# which travels through the ig_* views rather than this scope (see
|
||||
# .build_verb_sql()).
|
||||
direct = ,
|
||||
total = .spend_subtypes_direct
|
||||
)
|
||||
}
|
||||
|
||||
# The two revenue concepts (uscogdata#12), again as crosswalk subtype sets.
|
||||
# Census's manual section 4.3 defines the first by SUBTRACTING from the second
|
||||
# -- "General revenue comprises all revenue except that classified as liquor
|
||||
# store, utility, or insurance trust revenue" -- giving the identity
|
||||
#
|
||||
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
||||
#
|
||||
# Verified against Census's own computed concept fields (IndFin FY2012,
|
||||
# Wisconsin state): 31,410,686 + 0 + 0 + 4,469,906 = 35,880,592, exact.
|
||||
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
|
||||
.revenue_subtypes_total <- c(.revenue_subtypes_general, "utility",
|
||||
"liquor_store", "insurance_trust")
|
||||
|
||||
#' @noRd
|
||||
.revenue_concept_subtypes <- function(concept) {
|
||||
switch(concept,
|
||||
general = .revenue_subtypes_general,
|
||||
total = .revenue_subtypes_total
|
||||
)
|
||||
}
|
||||
|
||||
#' Summarized spending by category
|
||||
#'
|
||||
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||
@@ -11,7 +71,16 @@
|
||||
#' @param govid Character vector of `canonical_govid` values.
|
||||
#' @param years Integer vector of years.
|
||||
#' @param category Character vector of category names (from
|
||||
#' `summary_categories.category`), or `NULL` for all categories.
|
||||
#' `summary_categories.category`), or `NULL` for all categories broken out
|
||||
#' one row each. The reserved value `"All Categories"` instead returns a
|
||||
#' single summed row per `(year, canonical_govid, subtype)`, covering every
|
||||
#' category inside the requested concept's subtype scope. It cannot be
|
||||
#' combined with other category names, and it is not the same thing as
|
||||
#' `expenditure_concept = "total"`: the concept chooses which subtypes are in
|
||||
#' scope, `"All Categories"` chooses whether rows inside that scope are
|
||||
#' broken out or summed. Because the result keeps one row per
|
||||
#' `spend_subtype`, filtering the returned frame to
|
||||
#' `spend_subtype == "operations"` gives an operating-expenditure total.
|
||||
#' @param per_capita If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||
#' `amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||
#' Census F-33 population from `gov_population_yearly`. Result also gains
|
||||
@@ -42,22 +111,34 @@
|
||||
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||
#' FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
#' possibly-misleading `"harmonized"`/`"raw"` value.
|
||||
#' @param expenditure_concept `"direct"` (default) returns only the
|
||||
#' government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
#' from prior releases. `"total"` additionally UNIONs in the
|
||||
#' intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
#' to the state government (`L` codes, excluding the `L--` family-total
|
||||
#' rollup) -- so results gain rows with `spend_subtype ==
|
||||
#' "intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
#' to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
#' `uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
#' silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
#' already defines its own component codes). **Do not sum `"total"`
|
||||
#' results across levels of government** (e.g. state + county + city):
|
||||
#' a state's `M12` payment to a school district is the same dollar the
|
||||
#' district reports as its own direct `E12`, so summing both double-counts
|
||||
#' it. This matters in particular with [cog_geographic_rollup()], which
|
||||
#' sums across exactly that kind of multi-layer government set.
|
||||
#' @param expenditure_concept Which spending concept to return. Concepts are
|
||||
#' defined as sets of the crosswalk's `spend_subtype` values -- never as
|
||||
#' item-code first letters, which cannot classify correctly (prefix `Y`
|
||||
#' alone spans revenue, expenditure, and balance codes):
|
||||
#'
|
||||
#' * `"primary"` (default) -- the government's own service provision:
|
||||
#' `operations` + `capital` + `assistance` subtypes.
|
||||
#' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
|
||||
#' `interest` (interest on debt) and `insurance_benefits` (insurance
|
||||
#' trust benefit payments, e.g. pensions -- Census manual section
|
||||
#' 5.2.2.1 includes payments to retirees in Direct).
|
||||
#' * `"total"` -- `direct` plus the intergovernmental leg: payments to
|
||||
#' local governments (`M` codes), to the state government (`L` codes,
|
||||
#' excluding the `L--` family-total rollup), and state payments to
|
||||
#' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
|
||||
#' `spend_subtype == "intergovernmental"`. Requires the active corpus's
|
||||
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
|
||||
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
|
||||
#' older corpus rather than silently under-reporting. Mutually
|
||||
#' exclusive with `recipe` (a recipe already defines its own component
|
||||
#' codes).
|
||||
#'
|
||||
#' **Do not sum `"total"` results across levels of government** (e.g.
|
||||
#' state + county + city): a state's `M12` payment to a school district is
|
||||
#' the same dollar the district reports as its own direct `E12`, so
|
||||
#' summing both double-counts it. This matters in particular with
|
||||
#' [cog_geographic_rollup()], which sums across exactly that kind of
|
||||
#' multi-layer government set.
|
||||
#'
|
||||
#' In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
@@ -69,17 +150,63 @@
|
||||
#' component (when one exists), and
|
||||
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
#' IG.
|
||||
#' IG. When `category = "All Categories"` is combined with
|
||||
#' `expenditure_concept = "total"`, this detection cannot run (it keys on
|
||||
#' per-category rows, which all-categories mode collapses to one literal
|
||||
#' value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
|
||||
#' possibly-false `FALSE`; query an explicit `category` to get a real
|
||||
#' answer.
|
||||
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
||||
#' corpus does not carry still appears, labelled with **why** it is
|
||||
#' missing, and add a `value_source` column to every row:
|
||||
#'
|
||||
#' * `"reported"` — the corpus carries this cell.
|
||||
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
#' Census published `$0`. `amt_nominal` is `0`.
|
||||
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
#' government did not report, and the value is unknown. `amt_nominal` is
|
||||
#' `NA`, **not** `0` — writing a zero there would invent data.
|
||||
#'
|
||||
#' The grid comes from the corpus's `code_set` table, scoped to each
|
||||
#' government's own type, so a county is never filled with cells only a
|
||||
#' state can report. Reported rows are passed through untouched.
|
||||
#'
|
||||
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
#' when `representation`/`code_set` began shipping; aborts with class
|
||||
#' `uscogdata_representation_unavailable` otherwise. Not available with
|
||||
#' `recipe` or with `expenditure_concept = "total"` (class
|
||||
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
#' `code_set`.
|
||||
#' @param limit Maximum number of result rows to return, pushed into the SQL
|
||||
#' query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
||||
#' result is materialized. `NULL` (the default) returns every matching row,
|
||||
#' exactly as before this parameter existed. Mutually exclusive with
|
||||
#' `recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.
|
||||
#' @param offset Rows to skip before `limit` starts counting (0-based).
|
||||
#' Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.
|
||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
#' and `value_source` when `complete = TRUE`.
|
||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
||||
#' per-year `absence_means` rule that was applied. When `limit` is set,
|
||||
#' also carries a `total_rows` attribute: the full unpaginated row count,
|
||||
#' computed by the same query (`COUNT(*) OVER()`) rather than a second
|
||||
#' round trip -- so a caller walking pages never has to ask "how many are
|
||||
#' there" separately.
|
||||
#' @export
|
||||
cog_spending <- function(govid, years, category = NULL,
|
||||
per_capita = FALSE, adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
complete = FALSE, limit = NULL, offset = NULL) {
|
||||
# flow_prefixes no longer classifies rows (crosswalk subtype membership
|
||||
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
|
||||
# machinery to this verb's recipe families (see R/suggestions.R; the
|
||||
# catalog only has E/F/G-component direct-expenditure recipes).
|
||||
.verb_spendrev(
|
||||
verb = "cog_spending",
|
||||
view_base = "spending_annotated",
|
||||
@@ -93,7 +220,10 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
adjust_to_year = adjust_to_year,
|
||||
basis = basis,
|
||||
recipe = recipe,
|
||||
expenditure_concept = expenditure_concept
|
||||
expenditure_concept = expenditure_concept,
|
||||
complete = complete,
|
||||
limit = limit,
|
||||
offset = offset
|
||||
)
|
||||
}
|
||||
|
||||
@@ -101,8 +231,9 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
.abort_concept_not_aggregatable <- function(verb) {
|
||||
cli::cli_abort(c(
|
||||
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
|
||||
"*" = "Use {.code expenditure_concept = \"direct\"} (the default) for any \\
|
||||
comparison or sum that spans more than one government.",
|
||||
"*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
|
||||
{.code \"direct\"} for any comparison or sum that spans more than \\
|
||||
one government.",
|
||||
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
|
||||
money it hands to other governments. The receiving government reports \\
|
||||
that same dollar again as its own Direct when it actually spends it, \\
|
||||
@@ -118,26 +249,67 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
govid, years, category,
|
||||
per_capita, adjust_to_year,
|
||||
basis = c("harmonized", "raw"), recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")) {
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
revenue_concept = c("general", "total"),
|
||||
complete = FALSE, limit = NULL, offset = NULL) {
|
||||
basis_explicit <- length(basis) == 1L
|
||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
||||
# condition; wrap it so an invalid expenditure_concept aborts consistently
|
||||
# with the rest of this package's validation (cli::cli_abort -> rlang_error).
|
||||
expenditure_concept <- tryCatch(
|
||||
match.arg(expenditure_concept, c("direct", "total")),
|
||||
match.arg(expenditure_concept, c("primary", "direct", "total")),
|
||||
error = function(e) {
|
||||
cli::cli_abort(
|
||||
"`expenditure_concept` must be one of {.val direct} or {.val total}.",
|
||||
"`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
|
||||
class = "uscogdata_invalid_expenditure_concept",
|
||||
parent = e
|
||||
)
|
||||
}
|
||||
)
|
||||
|
||||
revenue_concept <- tryCatch(
|
||||
match.arg(revenue_concept, c("general", "total")),
|
||||
error = function(e) {
|
||||
cli::cli_abort(
|
||||
"`revenue_concept` must be one of {.val general} or {.val total}.",
|
||||
class = "uscogdata_invalid_revenue_concept",
|
||||
parent = e
|
||||
)
|
||||
}
|
||||
)
|
||||
|
||||
# The concept's subtype scope. Every code path below -- the verb SQL, the
|
||||
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
|
||||
# by crosswalk subtype membership, never by item-code prefix. The
|
||||
# expenditure "total" concept's extra intergovernmental leg is the one
|
||||
# exception: it travels through the ig_* views rather than this scope,
|
||||
# because its legacy rows are aggregate-flagged.
|
||||
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
|
||||
.expenditure_concept_subtypes(expenditure_concept)
|
||||
} else {
|
||||
.revenue_concept_subtypes(revenue_concept)
|
||||
}
|
||||
|
||||
govid <- .coerce_govid_input(govid, arg = "govid")
|
||||
# allow_all_categories = TRUE: cog_spending()/cog_revenue() are the two
|
||||
# verbs the reserved pseudo-category is defined for. cog_balances() shares
|
||||
# this validator but leaves the argument at its FALSE default, so it
|
||||
# rejects "All Categories" instead of silently returning zero rows
|
||||
# (finding 3, all-categories review).
|
||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
||||
recipe)
|
||||
recipe, allow_all_categories = TRUE)
|
||||
|
||||
# Recognize the reserved pseudo-category. Detected after type validation so a
|
||||
# non-character `category` still fails with the ordinary type error.
|
||||
all_categories <- !is.null(category) && .ALL_CATEGORIES %in% category
|
||||
if (all_categories && length(category) > 1L) {
|
||||
cli::cli_abort(c(
|
||||
"{.val {(.ALL_CATEGORIES)}} cannot be combined with other categories.",
|
||||
"i" = "It already sums every category in the requested concept's scope.",
|
||||
"*" = "Ask for it alone, or list the specific categories you want."
|
||||
), class = "uscogdata_all_categories_not_combinable")
|
||||
}
|
||||
|
||||
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
|
||||
cli::cli_abort(c(
|
||||
@@ -148,10 +320,10 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
}
|
||||
|
||||
# .verb_spendrev() is shared with cog_revenue(), which never exposes
|
||||
# expenditure_concept and always resolves it to "direct" -- so nothing on
|
||||
# the public API can reach this today. But it's a cheap guard against a
|
||||
# expenditure_concept and always resolves it to the default -- so nothing
|
||||
# on the public API can reach this today. But it's a cheap guard against a
|
||||
# future call (direct or via a modified cog_revenue()) that would UNION
|
||||
# the IG leg's expenditure M/L rows into a revenue result, which has no
|
||||
# the IG leg's expenditure M/L/Q rows into a revenue result, which has no
|
||||
# matching IG view and no sensible meaning.
|
||||
if (identical(expenditure_concept, "total") &&
|
||||
!identical(view_base, "spending_annotated")) {
|
||||
@@ -165,17 +337,71 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
)
|
||||
}
|
||||
|
||||
complete <- isTRUE(complete)
|
||||
if (complete && !is.null(recipe)) {
|
||||
.abort_complete_unsupported(
|
||||
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
|
||||
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
|
||||
)
|
||||
}
|
||||
if (complete && identical(expenditure_concept, "total")) {
|
||||
.abort_complete_unsupported(
|
||||
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
|
||||
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
||||
)
|
||||
}
|
||||
if (complete && all_categories) {
|
||||
.abort_complete_unsupported(
|
||||
"`category = \"All Categories\"` collapses the category dimension that `code_set` grids over (see `.completion_grid_sql()`), so there is no per-category grid left to fill -- filling a summed row has no defined semantics.",
|
||||
"Drop `complete`, or use `complete = TRUE` with an explicit `category` (or `category = NULL` for every category)."
|
||||
)
|
||||
}
|
||||
|
||||
# limit/offset push the page into the SQL itself (see .build_verb_sql()),
|
||||
# so the two things that would make "a page of what" ambiguous are refused
|
||||
# up front rather than silently ignored: complete = TRUE fills a grid over
|
||||
# the FULL requested (year, category) space, and a recipe's result comes
|
||||
# from .run_recipe()'s own query, which this function does not touch.
|
||||
if (!is.null(limit)) {
|
||||
limit <- as.integer(limit)
|
||||
if (length(limit) != 1L || is.na(limit) || limit < 0L) {
|
||||
cli::cli_abort("`limit` must be a single non-negative integer.",
|
||||
class = "uscogdata_invalid_pagination")
|
||||
}
|
||||
offset <- if (is.null(offset)) 0L else as.integer(offset)
|
||||
if (length(offset) != 1L || is.na(offset) || offset < 0L) {
|
||||
cli::cli_abort("`offset` must be a single non-negative integer.",
|
||||
class = "uscogdata_invalid_pagination")
|
||||
}
|
||||
if (complete) {
|
||||
cli::cli_abort(c(
|
||||
"`limit`/`offset` cannot be combined with `complete = TRUE`.",
|
||||
"i" = "`complete` fills a grid over the FULL requested (year, category) space; paginating a slice of already-grouped rows has no defined meaning for the cells it would fill.",
|
||||
"*" = "Drop `limit`/`offset`, or drop `complete`."
|
||||
), class = "uscogdata_complete_pagination_conflict")
|
||||
}
|
||||
if (!is.null(recipe)) {
|
||||
cli::cli_abort(c(
|
||||
"`limit`/`offset` cannot be combined with `recipe`.",
|
||||
"i" = "A recipe's result comes from a separate query (`.run_recipe()`) that pagination is not wired into yet.",
|
||||
"*" = "Drop `limit`/`offset`, or drop `recipe`."
|
||||
), class = "uscogdata_recipe_pagination_conflict")
|
||||
}
|
||||
}
|
||||
|
||||
years <- as.integer(years)
|
||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||
|
||||
con <- .ensure_session()
|
||||
manifest <- .uscogdata_env$manifest
|
||||
scope <- .check_govids_in_scope(govid)
|
||||
if (complete) .require_representation(con, manifest)
|
||||
|
||||
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
||||
|
||||
recipe_block <- NULL
|
||||
category_for_prov <- category
|
||||
total_rows <- NULL # set below only when limit is non-NULL (non-recipe path)
|
||||
if (!is.null(recipe)) {
|
||||
.require_schema_v5(con, manifest, "recipe =")
|
||||
.validate_recipe_id(con, recipe)
|
||||
@@ -197,8 +423,44 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
} else {
|
||||
NULL
|
||||
}
|
||||
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view)
|
||||
sql <- .build_verb_sql(view, subtype_col, govid, years,
|
||||
if (all_categories) NULL else category,
|
||||
ig_view, subtype_scope,
|
||||
all_categories = all_categories,
|
||||
limit = limit, offset = offset)
|
||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
if (!is.null(limit)) {
|
||||
# COUNT(*) OVER() rides along as an ordinary column so the total comes
|
||||
# from the same scan when this page has any rows -- see
|
||||
# .build_verb_sql(). An empty page (offset past the end) carries no
|
||||
# such row to read it from, so that one case falls back to a second,
|
||||
# unpaginated COUNT(*) query rather than reporting a wrong zero.
|
||||
if (nrow(result) > 0L) {
|
||||
total_rows <- result$pagination_total_rows[[1]]
|
||||
result$pagination_total_rows <- NULL
|
||||
} else {
|
||||
count_sql <- sprintf(
|
||||
"SELECT COUNT(*) AS n FROM (%s) AS _uncounted",
|
||||
.build_verb_sql(view, subtype_col, govid, years,
|
||||
if (all_categories) NULL else category,
|
||||
ig_view, subtype_scope,
|
||||
all_categories = all_categories)
|
||||
)
|
||||
total_rows <- as.integer(DBI::dbGetQuery(con, count_sql)$n[[1]])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# Fill BEFORE per_capita / inflation so the added cells get the same
|
||||
# treatment as reported ones: a census_zero stays $0 per capita and in real
|
||||
# dollars, and a not_reported stays NA through both rather than becoming a
|
||||
# spurious 0.
|
||||
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
|
||||
if (complete) {
|
||||
result <- .complete_result(result, con, subtype_col, govid, years,
|
||||
category, subtype_scope)
|
||||
completion <- attr(result, ".completion")
|
||||
attr(result, ".completion") <- NULL
|
||||
}
|
||||
|
||||
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
||||
@@ -229,7 +491,7 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
basis_for_prov <- resolved$basis
|
||||
basis_note_for_prov <- resolved$note
|
||||
harmonization <- .build_harmonization_block(
|
||||
con, govid, years, resolved, flow_prefixes
|
||||
con, govid, years, resolved, subtype_col, subtype_scope
|
||||
)
|
||||
# C1(a): gap detection must run against the Direct leg alone. `result`
|
||||
# can also carry UNION'd intergovernmental rows (expenditure_concept =
|
||||
@@ -246,7 +508,11 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
}
|
||||
suggestions <- .build_suggestions(con, govid, years, category,
|
||||
direct_leg_result,
|
||||
resolved$basis, flow_prefixes)
|
||||
resolved$basis, flow_prefixes,
|
||||
.select_long_view(view_base, resolved$basis),
|
||||
all_categories = all_categories,
|
||||
subtype_col = subtype_col,
|
||||
subtype_scope = subtype_scope)
|
||||
}
|
||||
|
||||
# C1(b): when expenditure_concept = "total", flag any row where the IG
|
||||
@@ -258,13 +524,32 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
# direct spending in that category, which is correct, ordinary data). When
|
||||
# a covering recipe is found, both the row-level notes and the provenance
|
||||
# say so rather than pass silently as a plausible Total.
|
||||
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
|
||||
#
|
||||
# In all-categories mode this cannot run at all: .detect_direct_suppressed()
|
||||
# keys on (year, canonical_govid, category), and every row shares the same
|
||||
# literal "All Categories" value, so the key collides across every real
|
||||
# category for that (year, govid) -- an IG-only row for a suppressed
|
||||
# category becomes indistinguishable from one sharing a key with an
|
||||
# unrelated category's ordinary Direct row. `has_direct` would then read
|
||||
# TRUE whenever the government has ANY direct spending at all, and the
|
||||
# detector could never fire. Rather than run it and report a false FALSE,
|
||||
# skip it and record NA -- the provenance must stop making a claim it
|
||||
# cannot support (finding 1, all-categories review).
|
||||
suppression_unavailable <- all_categories &&
|
||||
identical(expenditure_concept, "total")
|
||||
direct_suppressed_info <- if (suppression_unavailable) {
|
||||
list(flag = rep(NA, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
||||
} else if (identical(expenditure_concept, "total")) {
|
||||
.detect_direct_suppressed(con, result, subtype_col)
|
||||
} else {
|
||||
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
||||
}
|
||||
direct_suppressed <- direct_suppressed_info$flag
|
||||
direct_suppressed_flag <- isTRUE(any(direct_suppressed))
|
||||
direct_suppressed_flag <- if (suppression_unavailable) {
|
||||
NA
|
||||
} else {
|
||||
isTRUE(any(direct_suppressed))
|
||||
}
|
||||
|
||||
result$notes <- .notes_column(result, direct_suppressed_info$notes)
|
||||
|
||||
@@ -273,9 +558,20 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
# leg is suppressed for at least one requested (year, category), append an
|
||||
# explicit warning rather than let the base note's "Total = Direct + IG"
|
||||
# framing stand unqualified for rows where that arithmetic didn't happen.
|
||||
# When suppression detection itself is unavailable (all-categories mode),
|
||||
# say so instead of silently reusing the unqualified base note.
|
||||
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
|
||||
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
|
||||
if (direct_suppressed_flag) {
|
||||
if (suppression_unavailable) {
|
||||
paste0(
|
||||
base_note,
|
||||
" NOTE: direct-leg-suppression detection is unavailable when ",
|
||||
"`category = \"All Categories\"` -- it keys on per-category rows, ",
|
||||
"which this mode collapses. `expenditure_concept_direct_suppressed` ",
|
||||
"is NA here rather than a possibly-false FALSE; query an explicit ",
|
||||
"`category` (or `category = NULL`) to get a real answer."
|
||||
)
|
||||
} else if (isTRUE(direct_suppressed_flag)) {
|
||||
paste0(
|
||||
base_note,
|
||||
" NOTE: for at least one requested (year, category) the Direct leg ",
|
||||
@@ -307,23 +603,42 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
expenditure_concept = expenditure_concept,
|
||||
expenditure_concept_note = expenditure_concept_note_for_prov,
|
||||
expenditure_concept_direct_suppressed = direct_suppressed_flag,
|
||||
revenue_concept = revenue_concept,
|
||||
harmonization = harmonization,
|
||||
recipe = recipe_block,
|
||||
suggestions = suggestions
|
||||
suggestions = suggestions,
|
||||
completion = completion
|
||||
)
|
||||
prov$scope$govids_found <- scope$found
|
||||
prov$scope$govids_missing <- scope$missing
|
||||
attr(result, "provenance") <- prov
|
||||
attr(result, ".popyear_range") <- NULL
|
||||
# Attached here, after every downstream transform (per_capita/real-dollar
|
||||
# joins, notes, subtype filtering), the same way provenance is -- an
|
||||
# attribute set before those runs is not guaranteed to survive them.
|
||||
if (!is.null(limit)) attr(result, "total_rows") <- total_rows
|
||||
|
||||
if (length(suggestions) > 0L) .inform_suggestions(suggestions)
|
||||
|
||||
result
|
||||
}
|
||||
|
||||
#' Shared input validation for the money/holdings verbs.
|
||||
#'
|
||||
#' `allow_all_categories` gates the reserved pseudo-category
|
||||
#' `.ALL_CATEGORIES` ("All Categories"). It is meaningful only where a
|
||||
#' concept's subtype scope defines what "all" sums over --
|
||||
#' `cog_spending()`/`cog_revenue()`, via `.verb_spendrev()`, pass `TRUE`.
|
||||
#' `cog_balances()` leaves it at the `FALSE` default: holdings are a stock
|
||||
#' with no concept vocabulary to sum across (see R/balances.R), and before
|
||||
#' this guard existed `cog_balances(category = "All Categories")` silently
|
||||
#' matched zero crosswalk rows and returned an empty result with no error
|
||||
#' (finding 3, all-categories review). This validator is shared specifically
|
||||
#' so the three verbs cannot drift apart on this again.
|
||||
#' @noRd
|
||||
.validate_verb_inputs <- function(govid, years, category,
|
||||
per_capita, adjust_to_year, recipe = NULL) {
|
||||
per_capita, adjust_to_year, recipe = NULL,
|
||||
allow_all_categories = FALSE) {
|
||||
if (!is.character(govid) || length(govid) == 0L) {
|
||||
cli::cli_abort("`govid` must be a non-empty character vector.")
|
||||
}
|
||||
@@ -333,6 +648,14 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
if (!is.null(category) && !is.character(category)) {
|
||||
cli::cli_abort("`category` must be character or NULL.")
|
||||
}
|
||||
if (!allow_all_categories && !is.null(category) &&
|
||||
.ALL_CATEGORIES %in% category) {
|
||||
cli::cli_abort(c(
|
||||
"{.val {(.ALL_CATEGORIES)}} is not supported here.",
|
||||
i = "It sums a spending or revenue concept's subtype scope; this verb has no concept vocabulary to sum across.",
|
||||
i = "Use {.fn cog_spending} or {.fn cog_revenue} for an all-categories total."
|
||||
), class = "uscogdata_all_categories_unsupported")
|
||||
}
|
||||
if (!is.logical(per_capita) || length(per_capita) != 1L) {
|
||||
cli::cli_abort("`per_capita` must be a length-1 logical.")
|
||||
}
|
||||
@@ -361,6 +684,18 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
|
||||
}
|
||||
|
||||
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
|
||||
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
|
||||
#' anti-joins the LONG view rather than the annotated one: they have identical
|
||||
#' row membership (the annotated views are the long views plus LEFT JOINs, see
|
||||
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
|
||||
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
|
||||
#' the suppression test is asking about.
|
||||
#' @noRd
|
||||
.select_long_view <- function(view_base, basis) {
|
||||
.select_view(sub("_annotated$", "_long", view_base), basis)
|
||||
}
|
||||
|
||||
#' @noRd
|
||||
.select_ig_view <- function(basis) {
|
||||
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
|
||||
@@ -406,18 +741,37 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
|
||||
#' @noRd
|
||||
.build_verb_sql <- function(view, subtype_col, govid, years, category,
|
||||
ig_view = NULL) {
|
||||
ig_view = NULL, subtype_scope = NULL,
|
||||
all_categories = FALSE, limit = NULL, offset = NULL) {
|
||||
govid_lit <- .sql_lit_chr(govid)
|
||||
years_lit <- paste(as.integer(years), collapse = ",")
|
||||
category_pred <- if (is.null(category)) {
|
||||
# In all-categories mode there is no category filter: the sum is defined by
|
||||
# the concept's SUBTYPE allowlist (subtype_pred below), which is the real
|
||||
# concept boundary. Filtering by category as well would be a no-op at best
|
||||
# and, if the crosswalk ever gained an uncategorized code, a silent
|
||||
# under-count of the very total this mode exists to guarantee.
|
||||
category_pred <- if (all_categories || is.null(category)) {
|
||||
""
|
||||
} else {
|
||||
sprintf("AND category IN (%s)", .sql_lit_chr(category))
|
||||
}
|
||||
|
||||
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
|
||||
# The base views carry every subtype of their flow (spending_annotated has
|
||||
# all five non-IG expenditure subtypes); the concept narrows here. For
|
||||
# "total", the IG leg's rows are 'intergovernmental', so that value joins
|
||||
# the allowlist exactly when ig_view is present.
|
||||
subtype_pred <- if (is.null(subtype_scope)) {
|
||||
""
|
||||
} else {
|
||||
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
|
||||
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
|
||||
}
|
||||
|
||||
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
|
||||
# never UNION: the two legs are disjoint by item_code prefix (E/F/G vs M/L),
|
||||
# so de-duplication would be pure cost, and a silent row-drop if two
|
||||
# never UNION: the two legs are disjoint by crosswalk subtype (the direct
|
||||
# view excludes 'intergovernmental'; the IG view is only that), so
|
||||
# de-duplication would be pure cost, and a silent row-drop if two
|
||||
# governments ever reported identical values.
|
||||
source_expr <- if (is.null(ig_view)) {
|
||||
view
|
||||
@@ -435,13 +789,26 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
# though its dollars came entirely from an aggregate row, silently
|
||||
# suppressing the "Aggregate fallback applied" note on exactly the rows
|
||||
# this feature exists to surface.
|
||||
sprintf(
|
||||
|
||||
# Collapse the category dimension. subtype is deliberately KEPT: it is what
|
||||
# lets a caller filter the result to `spend_subtype == "operations"` and
|
||||
# get an operating-expenditure total, the measure a fiscal comparison
|
||||
# actually wants. (There is no `subtype` argument -- this is a post-hoc
|
||||
# filter on the returned column, not a query parameter.)
|
||||
category_select <- if (all_categories) {
|
||||
sprintf("%s AS category", .sql_lit_chr(.ALL_CATEGORIES))
|
||||
} else {
|
||||
"category"
|
||||
}
|
||||
category_group <- if (all_categories) "" else ", category"
|
||||
|
||||
base_sql <- sprintf(
|
||||
"SELECT
|
||||
year,
|
||||
canonical_govid,
|
||||
COALESCE(xwalk_gov_name, gov_name) AS gov_name,
|
||||
%1$s,
|
||||
category,
|
||||
%7$s,
|
||||
SUM(amt) * 1000.0 AS amt_nominal,
|
||||
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
|
||||
bool_or(is_aggregate) AS aggregate_fallback
|
||||
@@ -449,10 +816,33 @@ cog_spending <- function(govid, years, category = NULL,
|
||||
WHERE canonical_govid IN (%3$s)
|
||||
AND year IN (%4$s)
|
||||
%5$s
|
||||
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
|
||||
ORDER BY year, canonical_govid, %1$s, category",
|
||||
subtype_col, source_expr, govid_lit, years_lit, category_pred
|
||||
%6$s
|
||||
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s%8$s
|
||||
ORDER BY year, canonical_govid, %1$s%8$s",
|
||||
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred,
|
||||
category_select, category_group
|
||||
)
|
||||
|
||||
# limit/offset push the page into the query itself instead of pulling every
|
||||
# matching row across the network only to slice and discard most of it
|
||||
# afterward (the pattern behind the 2026-08-06 production incident: a
|
||||
# 193,105-row/194-page sweep re-ran the full query and re-listified every
|
||||
# row on EVERY page). COUNT(*) OVER() rides along as an ordinary column so
|
||||
# the caller gets the true total from this same scan -- see the call site
|
||||
# in .verb_spendrev(), which reads it off row 1 and strips it back out.
|
||||
# The outer SELECT * wrapping (rather than appending LIMIT/OFFSET directly
|
||||
# to base_sql) is what makes COUNT(*) OVER() see the post-GROUP-BY row
|
||||
# count, not the pre-aggregation one.
|
||||
if (is.null(limit)) {
|
||||
base_sql
|
||||
} else {
|
||||
sprintf(
|
||||
"SELECT *, COUNT(*) OVER() AS pagination_total_rows
|
||||
FROM (%s) AS _paged
|
||||
LIMIT %d OFFSET %d",
|
||||
base_sql, limit, offset
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
#' @noRd
|
||||
|
||||
+127
-18
@@ -1,8 +1,16 @@
|
||||
# R/suggestions.R
|
||||
# Recipe-component-driven signposting: when a basis = "harmonized" query for
|
||||
# a category comes back with a coverage gap in some requested years (the
|
||||
# result has no rows at all in that year) that a harmonization recipe would
|
||||
# actually fill for this government, surface that recipe as a suggestion.
|
||||
# Recipe-component-driven signposting. When a basis = "harmonized" query for
|
||||
# a category comes back incomplete in some requested year -- and a
|
||||
# harmonization recipe would actually fill it for this government -- surface
|
||||
# that recipe as a suggestion. "Incomplete" has two forms, and a recipe
|
||||
# qualifies on either:
|
||||
# 1. empty_year -- the result has no rows at all in that year.
|
||||
# 2. suppressed_component -- the result HAS rows, but a component code
|
||||
# carries dollars the verb's own long view structurally excludes
|
||||
# (aggregate-published, or absent from summary_categories). This is
|
||||
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
|
||||
# aggregate-only E67/E68, so form 1 never fired and the caller got a
|
||||
# number a third too low with no signpost at all.
|
||||
#
|
||||
# This is deliberately keyed off the recipe catalog's component codes, not
|
||||
# off harmonization_map rows: no live map row carries a non-blank
|
||||
@@ -48,11 +56,39 @@
|
||||
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
|
||||
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
|
||||
#' lookup scoped to the calling verb's own flow family.
|
||||
#' @param long_view Name of the verb's own long view (from
|
||||
#' `.select_long_view()`), passed through to `.suppressed_components()` to
|
||||
#' measure the second qualifying path (uscogdata#9).
|
||||
#' @param all_categories `TRUE` when the caller's `category` is the reserved
|
||||
#' pseudo-category (`.ALL_CATEGORIES`). Defaults to `FALSE` so no other
|
||||
#' caller's behaviour changes. When `TRUE`, the candidate-recipe sub-select
|
||||
#' is scoped by `subtype_col`/`subtype_scope` instead of by `category` --
|
||||
#' symmetric with `.build_verb_sql()`'s own all-categories branch (see
|
||||
#' R/spending.R): the concept's subtype allowlist is the real scope
|
||||
#' boundary, not any literal category value, and
|
||||
#' `.ALL_CATEGORIES` ("All Categories") is never itself a row in
|
||||
#' `summary_categories.category`, so leaving the category-keyed sub-select
|
||||
#' in place here always returned zero candidates and silently disabled
|
||||
#' signposting in all-categories mode (final whole-branch review, finding
|
||||
#' 6).
|
||||
#' @param subtype_col Name of the `summary_categories` subtype column to
|
||||
#' scope by when `all_categories = TRUE` (`"spend_subtype"` or
|
||||
#' `"revenue_subtype"` -- the same value `.build_verb_sql()` already
|
||||
#' receives as its own `subtype_col`). Ignored when `all_categories =
|
||||
#' FALSE`. `NULL` by default.
|
||||
#' @param subtype_scope Character vector of subtype values to scope by when
|
||||
#' `all_categories = TRUE` (the same value `.build_verb_sql()` already
|
||||
#' receives as its own `subtype_scope` -- the concept's subtype allowlist,
|
||||
#' e.g. `.expenditure_concept_subtypes(expenditure_concept)`). Ignored when
|
||||
#' `all_categories = FALSE`. `NULL` by default.
|
||||
#' @return List of `list(recipe_id, label, available_years, hint,
|
||||
#' ig_recipe_id)`, possibly empty.
|
||||
#' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
|
||||
#' suppressed_codes)`, possibly empty.
|
||||
#' @noRd
|
||||
.build_suggestions <- function(con, govid, years, category, result, basis,
|
||||
flow_prefixes) {
|
||||
flow_prefixes, long_view,
|
||||
all_categories = FALSE,
|
||||
subtype_col = NULL, subtype_scope = NULL) {
|
||||
if (!identical(basis, "harmonized") || is.null(category)) return(list())
|
||||
|
||||
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
|
||||
@@ -67,16 +103,35 @@
|
||||
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
|
||||
# should never be suggested as a coverage-gap filler for EITHER verb, not
|
||||
# just kept from being named as the *counterpart* of another suggestion.
|
||||
#
|
||||
# The inner sub-select is the concept boundary (finding 6, final
|
||||
# whole-branch review): in all-categories mode it is scoped by
|
||||
# `subtype_col`/`subtype_scope` -- the same allowlist `.build_verb_sql()`
|
||||
# applies as a WHERE predicate to make the summed result a *concept*, not
|
||||
# by `category` (`.ALL_CATEGORIES` is never a row in
|
||||
# `summary_categories.category`, so a category-keyed sub-select always
|
||||
# came back empty here). The M/L exclusion below is unchanged either way.
|
||||
candidate_scope_sql <- if (isTRUE(all_categories)) {
|
||||
sprintf(
|
||||
"SELECT DISTINCT item_code FROM summary_categories WHERE %s IN (%s)",
|
||||
subtype_col, .sql_lit_chr(subtype_scope)
|
||||
)
|
||||
} else {
|
||||
sprintf(
|
||||
"SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)",
|
||||
.sql_lit_chr(category)
|
||||
)
|
||||
}
|
||||
candidates <- DBI::dbGetQuery(con, sprintf(
|
||||
"SELECT DISTINCT recipe_id FROM harmonization_recipes
|
||||
WHERE component_code IN (
|
||||
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)
|
||||
%s
|
||||
)
|
||||
AND recipe_id NOT IN (
|
||||
SELECT DISTINCT recipe_id FROM harmonization_recipes
|
||||
WHERE LEFT(component_code, 1) IN ('M', 'L')
|
||||
)",
|
||||
.sql_lit_chr(category)
|
||||
candidate_scope_sql
|
||||
))$recipe_id
|
||||
if (length(candidates) == 0L) return(list())
|
||||
|
||||
@@ -86,7 +141,28 @@
|
||||
unique(as.integer(result$year))
|
||||
}
|
||||
gap_years <- setdiff(as.integer(years), result_years)
|
||||
if (length(gap_years) == 0L) return(list())
|
||||
|
||||
# Path 2 (uscogdata#9): component dollars this government holds that the
|
||||
# verb's own view structurally excludes. Measured across ALL requested
|
||||
# years, not just gap years -- the whole point is that a year with rows can
|
||||
# still be missing dollars. Scoped to the calling verb's own flow_prefixes
|
||||
# (I1) -- see `.suppressed_components()`'s own roxygen for why.
|
||||
#
|
||||
# This runs unconditionally whenever there are candidates -- an earlier
|
||||
# revision of this fix wave tried a free, in-memory pre-check
|
||||
# (`.needs_suppression_query()`) to skip the round trip on an already-
|
||||
# covered path, but a scoped re-review measured it against the fixture and
|
||||
# found it didn't pay for itself (it skipped ~3% of healthy calls, ~0% of
|
||||
# the multi-govid batch shape it was meant to help, at a net cost increase
|
||||
# once its own always-run metadata query was counted) while adding an
|
||||
# untested exactness invariant -- that `result$codes_included` and this
|
||||
# anti-join share the harmonized `item_code` space -- whose silent
|
||||
# violation would kill signposting, the exact failure class uscogdata#9
|
||||
# exists to prevent. Owner's call: keep this simple; a batch-aware
|
||||
# optimization, if one is worth building, is a separate issue.
|
||||
supp <- .suppressed_components(con, candidates, govid, years, long_view, flow_prefixes)
|
||||
|
||||
if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
|
||||
|
||||
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
|
||||
"SELECT recipe_id, any_value(label) AS label,
|
||||
@@ -97,11 +173,12 @@
|
||||
.sql_lit_chr(candidates)
|
||||
)))
|
||||
|
||||
# Which (recipe_id, year) pairs the recipe's own generic join actually
|
||||
# covers for this government, restricted to the gap years -- the same
|
||||
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
|
||||
# no is_aggregate filter), just checking existence instead of summing.
|
||||
covered <- DBI::dbGetQuery(con, sprintf(
|
||||
# Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
|
||||
# covers for this government, restricted to the gap years.
|
||||
covered <- if (length(gap_years) == 0L) {
|
||||
data.frame(recipe_id = character(0), year = integer(0))
|
||||
} else {
|
||||
DBI::dbGetQuery(con, sprintf(
|
||||
"SELECT DISTINCT r.recipe_id, l.year
|
||||
FROM long l
|
||||
JOIN harmonization_recipes r
|
||||
@@ -116,16 +193,35 @@
|
||||
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
||||
paste(gap_years, collapse = ",")
|
||||
))
|
||||
}
|
||||
|
||||
suggestions <- list()
|
||||
for (rid in candidates) {
|
||||
if (!rid %in% covered$recipe_id) next
|
||||
empty_hit <- rid %in% covered$recipe_id
|
||||
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
|
||||
supp_hit <- nrow(s_rows) > 0L
|
||||
if (!empty_hit && !supp_hit) next
|
||||
m <- meta[meta$recipe_id == rid, ]
|
||||
suggestions[[length(suggestions) + 1L]] <- list(
|
||||
recipe_id = rid,
|
||||
label = m$label[[1]],
|
||||
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
|
||||
hint = sprintf("re-run with recipe = '%s'", rid)
|
||||
hint = sprintf("re-run with recipe = '%s'", rid),
|
||||
# An empty year is the stronger claim -- the category returned nothing
|
||||
# at all -- so it wins when both paths qualify. The suppressed_* fields
|
||||
# are still populated, so an empty_year fire also reports its dollars.
|
||||
trigger = if (empty_hit) "empty_year" else "suppressed_component",
|
||||
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
|
||||
suppressed_years = if (supp_hit) {
|
||||
sort(unique(as.integer(s_rows$year)))
|
||||
} else {
|
||||
integer(0)
|
||||
},
|
||||
suppressed_codes = if (supp_hit) {
|
||||
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
|
||||
} else {
|
||||
character(0)
|
||||
}
|
||||
)
|
||||
}
|
||||
.attach_ig_counterparts(con, suggestions, flow_prefixes)
|
||||
@@ -232,12 +328,25 @@
|
||||
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
|
||||
#' continuation line is appended naming the intergovernmental counterpart
|
||||
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
|
||||
#' same bullet under cli, not a new bullet).
|
||||
#' same bullet under cli, not a new bullet). Same treatment for
|
||||
#' `suppressed_amount` (uscogdata#9): only present when dollars were
|
||||
#' actually measured as excluded (an `empty_year` fire can carry them too --
|
||||
#' see `.build_suggestions()` -- so this keys off the amount, not `trigger`).
|
||||
#' @noRd
|
||||
.inform_suggestions <- function(suggestions) {
|
||||
bullets <- vapply(suggestions, function(s) {
|
||||
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
|
||||
s$available_years[1], s$available_years[2], s$hint)
|
||||
# Only present when dollars were actually measured as excluded. An
|
||||
# empty_year fire can carry them too -- the year had no rows AND the
|
||||
# component was suppressed -- which is strictly more informative.
|
||||
if (isTRUE(s$suppressed_amount > 0)) {
|
||||
bullet <- paste0(bullet, sprintf(
|
||||
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
|
||||
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
|
||||
paste0("FY", s$suppressed_years, collapse = ", "),
|
||||
paste(s$suppressed_codes, collapse = ", ")))
|
||||
}
|
||||
if (!is.null(s$ig_recipe_id)) {
|
||||
bullet <- paste0(bullet, sprintf(
|
||||
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
|
||||
@@ -245,7 +354,7 @@
|
||||
bullet
|
||||
}, character(1))
|
||||
cli::cli_inform(c(
|
||||
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
|
||||
i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
|
||||
stats::setNames(bullets, rep("*", length(bullets)))
|
||||
))
|
||||
}
|
||||
|
||||
+115
@@ -0,0 +1,115 @@
|
||||
# R/suppression.R
|
||||
# Split out of R/suggestions.R (2026-08-05) to keep files under the project's
|
||||
# 400-line limit. Owns the second qualifying path for coverage signposting
|
||||
# (uscogdata#9): measuring, per government, the component dollars the
|
||||
# calling verb's own long view structurally excludes (aggregate-published,
|
||||
# or absent from summary_categories). See R/suggestions.R for the
|
||||
# orchestrator (`.build_suggestions()`) that calls this and the full
|
||||
# uscogdata#9 background.
|
||||
|
||||
#' Measure, per (recipe, year), the component dollars this government holds
|
||||
#' that the calling verb's own long view structurally excludes.
|
||||
#'
|
||||
#' This is the second qualifying path for a suggestion (uscogdata#9). The
|
||||
#' first -- row absence -- only fires when a category returns NOTHING in a
|
||||
#' requested year, which is how Corrections behaves in the wide era. Public
|
||||
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
|
||||
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
|
||||
#' 2011, and absent from `summary_categories` entirely) are dropped. The
|
||||
#' caller gets a plausible number a third too low, silently.
|
||||
#'
|
||||
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
|
||||
#' view rather than restating its WHERE clause, so this stays correct if
|
||||
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
|
||||
#' anti-join is keyed on `item_code`, which is sound only because
|
||||
#' harmonization never renames a recipe component -- asserted by the "no
|
||||
#' recipe component is ever renamed by harmonization" test in
|
||||
#' tests/testthat/test-recipes.R.
|
||||
#'
|
||||
#' Note what this deliberately does NOT count as suppressed: a component
|
||||
#' excluded from the RESULT for scoping reasons -- because it belongs to a
|
||||
#' different `category`, or because `expenditure_concept` narrowed the
|
||||
#' subtypes -- is still present in the view, so it never fires. Suggesting a
|
||||
#' recipe is a coverage fix, not a category redefinition.
|
||||
#'
|
||||
#' `flow_prefixes` (uscogdata#9 review, finding I1) restricts the measured
|
||||
#' components to the CALLING VERB's own flow family (`c("E","F","G")` for
|
||||
#' spending, `c("T","A","U","B","C","D")` for revenue). Without this, a
|
||||
#' candidate recipe belonging to the OTHER flow family is always absent from
|
||||
#' this verb's view (by construction -- `cog_revenue()`'s view never carries
|
||||
#' an E-coded row) and so was always reported as "suppressed", fabricating a
|
||||
#' dollar claim across flow families (`cog_revenue(category = "Corrections")`
|
||||
#' claimed $3.63B excluded that `cog_spending()` reports and fully accounts
|
||||
#' for). Filtering on `LEFT(r.component_code, 1)` also drops M/L-prefixed
|
||||
#' components from measurement under `cog_spending()` (`flow_prefixes` never
|
||||
#' includes "M"/"L") -- harmless today, because a recipe's own M/L components
|
||||
#' (e.g. `corrections_ig_local_combined`'s M04/M05) are present in the view
|
||||
#' in every year they exist and so never fired as suppressed anyway, but
|
||||
#' worth recording since this filter is now the thing relied on to prevent
|
||||
#' it.
|
||||
#'
|
||||
#' @param con Active DuckDB connection.
|
||||
#' @param candidates Character vector of recipe ids to measure.
|
||||
#' @param govid Character vector of canonical_govid values.
|
||||
#' @param years Integer vector of requested years.
|
||||
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
|
||||
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
|
||||
#' `.build_suggestions()`). Only recipe components whose first character is
|
||||
#' in this set are measured.
|
||||
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
|
||||
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
|
||||
#' nothing is suppressed.
|
||||
#' @noRd
|
||||
.suppressed_components <- function(con, candidates, govid, years, long_view,
|
||||
flow_prefixes) {
|
||||
empty <- tibble::tibble(
|
||||
recipe_id = character(0), year = numeric(0),
|
||||
suppressed_amount = numeric(0), suppressed_codes = character(0)
|
||||
)
|
||||
if (length(candidates) == 0L) return(empty)
|
||||
|
||||
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
|
||||
# never be quoted safely. It is always internally derived from a fixed
|
||||
# view_base, so an off-allowlist value is a programming error, not input.
|
||||
if (!long_view %in% c("spending_long", "spending_long_harmonized",
|
||||
"revenue_long", "revenue_long_harmonized")) {
|
||||
cli::cli_abort(
|
||||
"Internal error: unexpected `long_view` {.val {long_view}}.",
|
||||
class = "uscogdata_internal_error"
|
||||
)
|
||||
}
|
||||
|
||||
sql <- sprintf(
|
||||
"SELECT r.recipe_id,
|
||||
l.year,
|
||||
SUM(l.amt) * 1000.0 AS suppressed_amount,
|
||||
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
|
||||
AS suppressed_codes
|
||||
FROM long l
|
||||
JOIN harmonization_recipes r
|
||||
ON l.item_code = r.component_code
|
||||
AND l.year BETWEEN r.year_min AND r.year_max
|
||||
AND (r.gov_type_scope = 'all'
|
||||
OR (r.gov_type_scope = 'state' AND l.type = 0)
|
||||
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
|
||||
WHERE r.recipe_id IN (%1$s)
|
||||
AND l.canonical_govid IN (%2$s)
|
||||
AND l.year IN (%3$s)
|
||||
AND l.amt <> 0
|
||||
AND LEFT(r.component_code, 1) IN (%5$s)
|
||||
AND NOT EXISTS (
|
||||
SELECT 1 FROM %4$s v
|
||||
WHERE v.canonical_govid = l.canonical_govid
|
||||
AND v.year = l.year
|
||||
AND v.item_code = l.item_code
|
||||
AND v.year IN (%3$s) -- restated: enables partition pruning (I3a)
|
||||
AND v.canonical_govid IN (%2$s) -- restated: pushes the govid filter (I3a)
|
||||
)
|
||||
GROUP BY 1, 2
|
||||
ORDER BY 1, 2",
|
||||
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
||||
paste(as.integer(years), collapse = ","), long_view,
|
||||
.sql_lit_chr(flow_prefixes)
|
||||
)
|
||||
tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||
}
|
||||
@@ -32,6 +32,101 @@
|
||||
"45-ig_annotated_harmonized.sql"
|
||||
)
|
||||
|
||||
# The representation contract (cog_pipeline#64): two parquet tables that say
|
||||
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
|
||||
# on schema_version, because the sparsification that introduced them did not
|
||||
# bump the version -- the pre-sparsification corpus this package shipped
|
||||
# against until 2026-07-30 was already schema v6 and carried neither table.
|
||||
# Keying off the version number would therefore register a view over a file
|
||||
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
|
||||
# check exists to tolerate.
|
||||
.representation_view_files <- c(
|
||||
"36-representation.sql" = "representation.parquet",
|
||||
"37-code_set.sql" = "code_set.parquet"
|
||||
)
|
||||
|
||||
# Cash and security holdings (uscogdata#25). 46- selects
|
||||
# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and
|
||||
# WITHOUT a schema_version bump -- so neither existing gate applies:
|
||||
# .harmonization_view_files keys on schema_version, .representation_view_files
|
||||
# on the presence of a FILE. Here the discriminator is a COLUMN on a table
|
||||
# that exists either way. CREATE VIEW resolves its source schema eagerly, so
|
||||
# on an older corpus 46- would fail at registration with "Binder Error:
|
||||
# Referenced column balance_subtype not found" rather than at query time.
|
||||
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
|
||||
|
||||
#' Does the mounted corpus's `summary_categories` carry `balance_subtype`?
|
||||
#' Probed against the live connection rather than the manifest, because the
|
||||
#' manifest describes files, not columns.
|
||||
#' @noRd
|
||||
.corpus_has_balance_subtype <- function(con) {
|
||||
n <- DBI::dbGetQuery(con,
|
||||
"SELECT COUNT(*) AS n FROM information_schema.columns
|
||||
WHERE table_name = 'summary_categories'
|
||||
AND column_name = 'balance_subtype'"
|
||||
)$n
|
||||
isTRUE(as.integer(n) > 0L)
|
||||
}
|
||||
|
||||
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
|
||||
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
|
||||
#' works identically for a local fixture and a remote share.
|
||||
#' @noRd
|
||||
.corpus_has_table <- function(manifest, file) {
|
||||
paths <- vapply(manifest$files$metadata %||% list(),
|
||||
function(f) as.character(f$path %||% ""), character(1))
|
||||
file %in% basename(paths)
|
||||
}
|
||||
|
||||
#' Build the SQL path expression for the partitioned `long` table.
|
||||
#'
|
||||
#' DuckDB cannot expand a glob over generic HTTP: there is no directory
|
||||
#' listing to expand against, and `allow_asterisks_in_http_paths` only
|
||||
#' forwards the literal `**/*` as a filename, which 404s. Measured against
|
||||
#' the published corpus on 2026-08-08, an explicit file list returns the
|
||||
#' same 46,148,034 rows the (working) `hf://` glob does, and
|
||||
#' `hive_partitioning = true` still recovers `year` from the paths.
|
||||
#'
|
||||
#' The manifest already enumerates every partition, so we build the list
|
||||
#' from it. This is host-agnostic -- Nextcloud, HuggingFace and a local
|
||||
#' fixture take the same path -- where an `hf://` URL would tie the reader
|
||||
#' to one vendor's protocol and still need special-casing, since manifest
|
||||
#' fetching goes through httr2, which cannot speak `hf://`.
|
||||
#'
|
||||
#' Falls back to the glob when the manifest carries no partition list: a
|
||||
#' hand-built manifest in a test (see test-views.R) or a corpus predating
|
||||
#' the field. Both are local, where globbing works.
|
||||
#' @noRd
|
||||
.long_files_sql <- function(url, manifest) {
|
||||
parts <- manifest$files$long_partitions %||% list()
|
||||
if (length(parts) == 0L) {
|
||||
return(.sql_lit_chr(paste0(url, "data/long/**/*.parquet")))
|
||||
}
|
||||
paths <- vapply(parts, function(p) as.character(p$path), character(1))
|
||||
paste0("[", .sql_lit_chr(paste0(url, paths)), "]")
|
||||
}
|
||||
|
||||
#' Substitute the corpus-location tokens in a view's SQL text.
|
||||
#'
|
||||
#' One place knows the token vocabulary. `.register_views()` and the tests
|
||||
#' that execute a view file directly both route through here. This exists
|
||||
#' because four test sites had hand-rolled the `{url}` substitution -- one
|
||||
#' of them commented as doing it "exactly as .register_views() does" -- and
|
||||
#' every one of them broke the moment a second token was introduced.
|
||||
#'
|
||||
#' `{long_files}` must be substituted BEFORE `{url}`: it expands to a string
|
||||
#' that itself contains the url, so the reverse order leaves the token in
|
||||
#' place and DuckDB's parser fails on the brace.
|
||||
#'
|
||||
#' `manifest` defaults to empty, which routes `.long_files_sql()` to its glob
|
||||
#' fallback -- correct for the local temp corpora the direct-execution tests
|
||||
#' build.
|
||||
#' @noRd
|
||||
.render_view_sql <- function(sql, url, manifest = list()) {
|
||||
sql <- gsub("\\{long_files\\}", .long_files_sql(url, manifest), sql, fixed = FALSE)
|
||||
gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||||
}
|
||||
|
||||
#' Register DuckDB views from inst/sql/ SQL files
|
||||
#' @noRd
|
||||
.register_views <- function(con, url, manifest) {
|
||||
@@ -39,9 +134,13 @@
|
||||
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||
for (f in files) {
|
||||
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next
|
||||
base <- basename(f)
|
||||
if (base %in% .harmonization_view_files && schema_version < 5L) next
|
||||
if (base %in% names(.representation_view_files) &&
|
||||
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
|
||||
if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next
|
||||
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
||||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||||
sql <- .render_view_sql(sql, url, manifest)
|
||||
DBI::dbExecute(con, sql)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,78 +1,247 @@
|
||||
# uscogdata
|
||||
|
||||
Curated R reader for the Civilytics US Census of Governments finance corpus.
|
||||
<!-- badges: start -->
|
||||
[](https://civilytics.r-universe.dev/uscogdata)
|
||||
[](LICENSE.md)
|
||||
<!-- badges: end -->
|
||||
|
||||
Provides unit-level financial profiles, geographic rollups, and peer comparisons
|
||||
with auditable provenance and built-in cross-vintage correctness. Reads the
|
||||
published corpus (Hive-partitioned parquet + manifest.json) directly from
|
||||
Nextcloud via DuckDB httpfs — no local bulk downloads required.
|
||||
A curated R reader for the Civilytics US Census of Governments finance corpus —
|
||||
every dollar that US state, county, municipal and township governments reported
|
||||
raising and spending, from **FY1967 to FY2024**, in one queryable place.
|
||||
|
||||
## Status
|
||||
The Census of Governments is the only nationwide source for local government
|
||||
finance, and it is hard to use: item codes change meaning across vintages,
|
||||
government identifiers were renumbered in 2017, and an absent value means
|
||||
"published zero" in one era and "not reported" in the next. This package
|
||||
handles each of those problems, and it tells you when it has — every result
|
||||
carries provenance describing what was converted, what was aggregated, and
|
||||
which known series breaks intersect your query.
|
||||
|
||||
Under active development (Phase 2 of the cog_pipeline project). See
|
||||
`../cog_pipeline/docs/reader-specification.md` for the reader contract this
|
||||
package implements.
|
||||
**Scope:** government types 0–3 (state, county, municipality, township).
|
||||
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
|
||||
or FY1969. Special districts (type 4) and school districts (type 5) are
|
||||
excluded pending validation.
|
||||
|
||||
## Installation
|
||||
## Where the data comes from
|
||||
|
||||
The corpus is published and documented at the **[US Census of Governments
|
||||
Finance API](https://pages.civilytics.org/cog-api/)**. Start there for how the
|
||||
data was built, how the identifier and item-code reconciliation works, and what
|
||||
the corpus does and does not cover.
|
||||
|
||||
- **[API documentation and walkthroughs](https://pages.civilytics.org/cog-api/)**
|
||||
— reference, data dictionary, and worked examples such as the
|
||||
[Southern states guide](https://pages.civilytics.org/cog-api/cog-api-south-guide.html)
|
||||
- **[Live API](https://cog-api.civilytics.org/api/v1/)** — the same corpus over
|
||||
HTTP, for Tableau, Python, or anything that isn't R
|
||||
- **[Bulk corpus on Hugging Face](https://huggingface.co/datasets/civilytics/us-cog-finance)**
|
||||
— CC-BY-4.0; the same parquet files this package reads
|
||||
- **[Census Bureau source data](https://www.census.gov/programs-surveys/gov-finances.html)**
|
||||
— the underlying public files
|
||||
|
||||
## Install
|
||||
|
||||
```r
|
||||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
||||
install.packages("uscogdata",
|
||||
repos = c("https://civilytics.r-universe.dev",
|
||||
"https://cloud.r-project.org"))
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
||||
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
||||
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
||||
|
||||
## Direct vs Total spending
|
||||
|
||||
`cog_spending(..., expenditure_concept = c("direct", "total"))` controls
|
||||
whose spending a result counts. `"direct"` (the default) is a government's
|
||||
own current operations, capital outlay, and other direct spending. `"total"`
|
||||
additionally adds in the intergovernmental legs — money it hands to other
|
||||
governments to spend on its behalf — which is meaningful for describing one
|
||||
government's own budget over time, but double-counts when summed across
|
||||
governments (a state's payment to a county is the same dollar the county
|
||||
reports as its own direct spending).
|
||||
|
||||
**Rule of thumb: any figure that spans more than one government uses
|
||||
`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this
|
||||
by refusing `expenditure_concept = "total"`. See
|
||||
`vignette("total-spending", package = "uscogdata")` for the full
|
||||
explanation with worked examples.
|
||||
|
||||
## Developer notes
|
||||
|
||||
### Testing
|
||||
|
||||
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
||||
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
|
||||
all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
|
||||
at this fixture, so the full test suite runs offline with no network
|
||||
dependency:
|
||||
Or from source:
|
||||
|
||||
```r
|
||||
devtools::test() # uses bundled fixture, no credentials required
|
||||
pak::pkg_install("git::https://gitea.civilytics.org/Civilytics/uscogdata.git")
|
||||
```
|
||||
|
||||
### Releasing against the live corpus
|
||||
## Quickstart
|
||||
|
||||
Before cutting a release, run the test suite against the published corpus to
|
||||
catch any drift between the fixture and the real data:
|
||||
No configuration, no credentials, no download. The package reads the published
|
||||
corpus over HTTPS by default.
|
||||
|
||||
```r
|
||||
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
|
||||
devtools::test()
|
||||
library(uscogdata)
|
||||
|
||||
# Resolve a place name to a canonical government id
|
||||
madison <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||||
madison$canonical_govid
|
||||
#> [1] "552025209777"
|
||||
|
||||
# Police spending, inflation-adjusted and per capita
|
||||
spend <- cog_spending(
|
||||
madison$canonical_govid,
|
||||
years = 2012:2022,
|
||||
category = "Police",
|
||||
per_capita = TRUE,
|
||||
adjust_to_year = 2023
|
||||
)
|
||||
|
||||
# What did that result do to the numbers, and what should you know about them?
|
||||
cog_explain(spend)
|
||||
```
|
||||
|
||||
When the live-corpus run is clean, strip the fixture from the built package by
|
||||
adding this line to `.Rbuildignore`:
|
||||
`years` is required — there is no implicit full-history default.
|
||||
|
||||
```
|
||||
^inst/extdata/fixture_corpus$
|
||||
## Two ways to read the corpus
|
||||
|
||||
| | Remote (default) | Mirrored |
|
||||
|---|---|---|
|
||||
| Setup | none | `cog_mirror(dest)`, 190.6 MB once |
|
||||
| Disk used | **0 MB** — HTTP range requests only | 190.6 MB |
|
||||
| Per query | ~4 s (one government, one year)<br>~6 s (one government, 23 years) | local speed |
|
||||
| Good for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility |
|
||||
|
||||
Nothing is written to disk in remote mode: DuckDB fetches the parquet footer,
|
||||
works out which row groups it needs, and reads only those. Nothing is cached
|
||||
between sessions either, so every query goes back to the network.
|
||||
|
||||
The default points at a public HuggingFace mirror of the corpus. If you would
|
||||
rather not depend on a third party — for reproducibility, for an air-gapped
|
||||
environment, or on principle — **the escape hatch is one function call**:
|
||||
|
||||
```r
|
||||
cog_mirror("~/cog-corpus")
|
||||
Sys.setenv(USCOGDATA_URL = "~/cog-corpus/")
|
||||
```
|
||||
|
||||
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when
|
||||
the bundled fixture is absent, so no test code changes are needed for the
|
||||
release run or after stripping the fixture.
|
||||
After that, nothing in your analysis touches an external service.
|
||||
|
||||
### Configuration
|
||||
|
||||
- `USCOGDATA_URL` — corpus root: an HTTPS URL or a local path, **trailing slash required**
|
||||
- `USCOGDATA_CACHE_DIR` — where the manifest is cached (default: user cache dir)
|
||||
- `USCOGDATA_MANIFEST_TTL_SECS` — manifest re-fetch interval (default 3600)
|
||||
|
||||
## Amounts are in full US dollars
|
||||
|
||||
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
||||
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
|
||||
|
||||
The raw Census source files report **thousands of dollars**, and the corpus's
|
||||
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
||||
you never have to. The conversion is recorded in every result:
|
||||
|
||||
```r
|
||||
attr(spend, "provenance")$transformations$units_conversion
|
||||
#> $applied TRUE
|
||||
#> $source_unit "$1,000s (raw Census)"
|
||||
#> $target_unit "$USD"
|
||||
#> $multiplier 1000
|
||||
```
|
||||
|
||||
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
||||
`$1,000s` — which is true of the raw Census files and of the corpus's own `amt`
|
||||
column — that rule does not apply to anything a `cog_*()` verb hands you.
|
||||
Applying it twice overstates every figure by 1000x, and the result looks
|
||||
plausible rather than obviously wrong.
|
||||
|
||||
## Concepts worth understanding before you publish a number
|
||||
|
||||
### Primary vs Direct vs Total spending
|
||||
|
||||
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
|
||||
controls *whose* spending a result counts. Concepts are defined as sets of the
|
||||
crosswalk's `spend_subtype` values, never item-code first letters — the letter
|
||||
`Y` alone spans revenue, expenditure and balance codes.
|
||||
|
||||
- **`"primary"`** (default) — the government's own service provision: current
|
||||
operations, capital outlay, assistance payments.
|
||||
- **`"direct"`** — Census's published Direct Expenditure: `primary` plus
|
||||
interest on debt and insurance trust benefits (e.g. pensions).
|
||||
- **`"total"`** — adds the intergovernmental leg, money handed to other
|
||||
governments to spend. Meaningful for one government's own budget over time,
|
||||
but it double-counts when summed across governments: a state's payment to a
|
||||
county is the same dollar the county reports as its own direct spending.
|
||||
|
||||
**Rule of thumb: any figure spanning more than one government uses `primary`
|
||||
or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce that
|
||||
by refusing `"total"` outright. Worked examples in
|
||||
`vignette("total-spending", package = "uscogdata")`.
|
||||
|
||||
### General vs Total revenue
|
||||
|
||||
`cog_revenue(..., revenue_concept = c("general", "total"))`:
|
||||
|
||||
- **`"general"`** (default) — Census General Revenue: own-source taxes,
|
||||
charges and miscellaneous, plus federal, state and local aid.
|
||||
- **`"total"`** — General plus utility revenue (`A91`–`A94`), liquor store
|
||||
revenue (`A90`), and insurance trust revenue.
|
||||
|
||||
Census defines these by its own identity:
|
||||
|
||||
```
|
||||
Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
||||
```
|
||||
|
||||
Two things to know before switching to `"total"`. **Utility revenue is large
|
||||
for cities** — measured on the bundled fixture, utility plus liquor store is
|
||||
15.9% of city revenue, against 1.2% for states and 1.7% for counties. And the
|
||||
**employee-retirement (`X`) codes stop at FY2016**, when those systems moved to
|
||||
the separate Annual Survey of Public Pensions, so a `"total"` series steps down
|
||||
at the FY2016/FY2017 boundary for reasons of collection scope, not revenue
|
||||
(series breaks `SB197`–`SB209`).
|
||||
|
||||
### Reporting coverage: the Census is only sometimes a census
|
||||
|
||||
**The Census of Governments is a complete enumeration only in years ending in
|
||||
2 and 7.** Every other year is a sample, and the sample varies enormously —
|
||||
measured on the bundled fixture, Wisconsin's 608-city universe rolls up 597
|
||||
governments in FY2012 and 112 in FY2019.
|
||||
|
||||
A statewide total resting on a fifth of the universe looks exactly like one
|
||||
resting on all of it, so every multi-government result now says which it is:
|
||||
|
||||
```r
|
||||
attr(rollup, "provenance")$coverage # per-year n_units_reporting, is_census_year
|
||||
```
|
||||
|
||||
`cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` take a
|
||||
`coverage` argument — `"all"` (default), `"census"` (census years only), or
|
||||
`"consistent"` (only units reporting in every requested year, a balanced
|
||||
panel).
|
||||
|
||||
`n_units_reporting` is **category-conditional**, and it is not a response rate. A government that was surveyed and genuinely spends
|
||||
nothing in the requested category is indistinguishable from one never surveyed.
|
||||
|
||||
### Absent cells mean two different things
|
||||
|
||||
Before FY2012, an absent cell means Census published `$0`. From FY2012 on, it
|
||||
means not reported. `cog_spending(..., complete = TRUE)` fills the requested
|
||||
grid and labels every row with which it is, via `value_source`:
|
||||
|
||||
| `value_source` | meaning | `amt_nominal` |
|
||||
|---|---|---|
|
||||
| `reported` | the corpus carries this cell | as published |
|
||||
| `census_zero` | dense-source year (≤ FY2011), absent — Census published `$0` | `0` |
|
||||
| `not_reported` | sparse-source year (≥ FY2012), absent — unknown | `NA` |
|
||||
|
||||
That `NA` is deliberate. Filling a modern absence with `0` would invent data.
|
||||
|
||||
### Series breaks surface on their own
|
||||
|
||||
Catalogued breaks that intersect your query appear in provenance whether or not
|
||||
you went looking for them — `series_break_refs` for breaks in a specific item code, and
|
||||
`corpus_break_refs` for caveats about the corpus as a whole (dollar precision
|
||||
across the 1976/1977 boundary, the FY2017 identifier change, the FY2012
|
||||
dense→sparse representation change). `cog_explain()` prints both.
|
||||
|
||||
## How to cite
|
||||
|
||||
```r
|
||||
citation("uscogdata")
|
||||
```
|
||||
|
||||
The corpus itself is published under CC-BY-4.0. Cite it as:
|
||||
|
||||
> Civilytics Consulting. US Census of Governments finance corpus.
|
||||
> https://huggingface.co/datasets/civilytics/us-cog-finance
|
||||
|
||||
## Contributing
|
||||
|
||||
Development happens on [Gitea](https://gitea.civilytics.org/Civilytics/uscogdata);
|
||||
[GitHub](https://github.com/civilytics/uscogdata) is a mirror that accepts
|
||||
issues and pull requests. See [CONTRIBUTING.md](CONTRIBUTING.md) for how a
|
||||
patch gets from there to here.
|
||||
|
||||
## License
|
||||
|
||||
MIT © Civilytics Consulting LLC. See [LICENSE.md](LICENSE.md).
|
||||
|
||||
+27
-5
@@ -1,19 +1,41 @@
|
||||
url: ~
|
||||
url: https://civilytics.r-universe.dev/uscogdata
|
||||
|
||||
template:
|
||||
bootstrap: 5
|
||||
|
||||
reference:
|
||||
- title: Financial data
|
||||
desc: Spending, revenue and balance-sheet holdings for one or more governments.
|
||||
contents:
|
||||
- cog_spending
|
||||
- cog_revenue
|
||||
- cog_balances
|
||||
- title: Search & basket
|
||||
desc: Resolve place names into canonical govids.
|
||||
contents:
|
||||
- cog_gov_search
|
||||
- cog_basket_resolution
|
||||
- cog_basket_unresolved
|
||||
- title: Session
|
||||
- title: Comparison & aggregation
|
||||
desc: Peer cohorts and geographic aggregates.
|
||||
contents:
|
||||
- has_keyword("internal")
|
||||
- cog_find_peers
|
||||
- cog_peer_compare
|
||||
- cog_geographic_rollup
|
||||
- title: Corpus metadata
|
||||
desc: >
|
||||
What the corpus contains, where a given result came from, and how to
|
||||
hold a local copy of it.
|
||||
contents:
|
||||
- cog_categories
|
||||
- cog_recipes
|
||||
- cog_manifest
|
||||
- cog_explain
|
||||
- cog_mirror
|
||||
|
||||
articles:
|
||||
- title: Getting started
|
||||
- title: Concepts
|
||||
navbar: ~
|
||||
contents: []
|
||||
contents:
|
||||
- total-spending
|
||||
- population-denominators
|
||||
|
||||
@@ -11,12 +11,14 @@
|
||||
# Each partition is a full year (all states/govs) as published, so
|
||||
# Broward County FL and every other previously-pinned government stay
|
||||
# covered without any per-gov slicing logic.
|
||||
# 2. Copies the full canonical_fips_xwalk.parquet, canonical_alias.parquet,
|
||||
# summary_categories.parquet, harmonization_map.parquet,
|
||||
# harmonization_recipes.parquet, and series_breaks.parquet metadata
|
||||
# tables as-is (these are small cross-vintage registries, not
|
||||
# partitioned by year, so the fixture ships the complete tables rather
|
||||
# than a year-scoped subset).
|
||||
# 2. Copies every metadata parquet the publish tree ships (see
|
||||
# .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
|
||||
# registries, not partitioned by year, so the fixture ships the complete
|
||||
# tables rather than a year-scoped subset. representation.parquet and
|
||||
# code_set.parquet are what make the sparse wide era interpretable --
|
||||
# absence means "$0" in a dense_source year and "not reported" in a
|
||||
# sparse_source one -- so a fixture without them cannot represent the
|
||||
# published corpus.
|
||||
# 3. Resyncs the four reference docs (data_dictionary.md,
|
||||
# reader-specification.md, README.md, series_breaks.md) from the
|
||||
# publish tree's docs/.
|
||||
@@ -38,6 +40,22 @@
|
||||
# source("data-raw/regenerate_fixture_corpus.R")
|
||||
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
|
||||
|
||||
# Every metadata parquet the publish tree ships, in the order they appear in
|
||||
# the corpus manifest. Single source of truth for both the copy step and the
|
||||
# fixture manifest, so the two can never drift apart.
|
||||
.FIXTURE_METADATA_FILES <- c(
|
||||
"canonical_alias.parquet",
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"census_collection_coverage.parquet",
|
||||
"code_set.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"lineage_events.parquet",
|
||||
"representation.parquet",
|
||||
"series_breaks.parquet",
|
||||
"summary_categories.parquet"
|
||||
)
|
||||
|
||||
regenerate_fixture_corpus <- function(
|
||||
publish_cache_dir = file.path(
|
||||
"..", "cog_pipeline", "_targets", "publish_cache"
|
||||
@@ -100,20 +118,11 @@ regenerate_fixture_corpus <- function(
|
||||
invisible(NULL)
|
||||
}
|
||||
|
||||
# Copy the full (not year-scoped) canonical_fips_xwalk, canonical_alias,
|
||||
# summary_categories, and (schema v5+) harmonization_map/
|
||||
# harmonization_recipes/series_breaks parquet tables.
|
||||
# Copy the full (not year-scoped) metadata tables listed in
|
||||
# .FIXTURE_METADATA_FILES.
|
||||
#' @noRd
|
||||
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
|
||||
files <- c(
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"canonical_alias.parquet",
|
||||
"summary_categories.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"series_breaks.parquet"
|
||||
)
|
||||
for (f in files) {
|
||||
for (f in .FIXTURE_METADATA_FILES) {
|
||||
src <- file.path(publish_cache_dir, "data", f)
|
||||
dst <- file.path(fixture_dir, "data", f)
|
||||
if (!file.exists(src)) {
|
||||
@@ -179,15 +188,7 @@ regenerate_fixture_corpus <- function(
|
||||
)
|
||||
})
|
||||
|
||||
metadata_files <- c(
|
||||
"canonical_alias.parquet",
|
||||
"canonical_fips_xwalk.parquet",
|
||||
"summary_categories.parquet",
|
||||
"harmonization_map.parquet",
|
||||
"harmonization_recipes.parquet",
|
||||
"series_breaks.parquet"
|
||||
)
|
||||
metadata <- lapply(metadata_files, function(f) {
|
||||
metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
|
||||
rel <- file.path("data", f)
|
||||
path <- file.path(fixture_dir, rel)
|
||||
list(
|
||||
@@ -203,13 +204,16 @@ regenerate_fixture_corpus <- function(
|
||||
pipeline_commit = source_manifest$pipeline_commit,
|
||||
fixture_note = paste(
|
||||
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
|
||||
"corpus available via USCOGDATA_URL. Regenerated for Phase R2",
|
||||
"(schema_version 5, harmonization_map/harmonization_recipes/",
|
||||
"series_breaks parquet tables added). 2011/2012 straddle the",
|
||||
"wide-aggregate -> modern-leaf format boundary exercised by basis=",
|
||||
"\"harmonized\" and recipe= queries; 2019/2020 retain the prior",
|
||||
"per-capita/CPI regression anchors. Full canonical_fips_xwalk master",
|
||||
"and canonical_alias lookup table included via",
|
||||
"corpus available via USCOGDATA_URL. Regenerated from the sparsified",
|
||||
"schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
|
||||
"zeros, so FY2011 absence means Census published $0 while FY2012+",
|
||||
"absence means not reported. representation.parquet and",
|
||||
"code_set.parquet carry that rule and ship in full, as do every other",
|
||||
"metadata table in the publish tree. 2011/2012 straddle both the",
|
||||
"wide-aggregate -> modern-leaf format boundary (exercised by",
|
||||
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
|
||||
"representation boundary (SB194); 2019/2020 retain the prior",
|
||||
"per-capita/CPI regression anchors. Regenerated via",
|
||||
"data-raw/regenerate_fixture_corpus.R."
|
||||
),
|
||||
data_vintage = source_manifest$data_vintage,
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+30
-10
@@ -1,8 +1,8 @@
|
||||
{
|
||||
"schema_version": 6,
|
||||
"built_at": "2026-07-27T13:04:05Z",
|
||||
"pipeline_commit": "6098baf",
|
||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.",
|
||||
"built_at": "2026-08-03T16:51:32Z",
|
||||
"pipeline_commit": "e7394a4",
|
||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
|
||||
"data_vintage": {
|
||||
"source_vintages": {
|
||||
"2012": "10162019",
|
||||
@@ -36,9 +36,9 @@
|
||||
{
|
||||
"year": 2011,
|
||||
"path": "data/long/year=2011/part-0.parquet",
|
||||
"sha256": "84302ab364dc9fc3b3fbbc3c3f8b826e3508b4d73ff7c42d094d3863cd1e37b5",
|
||||
"row_count": 2864212,
|
||||
"size_bytes": 3845911
|
||||
"sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
|
||||
"row_count": 496004,
|
||||
"size_bytes": 2202455
|
||||
},
|
||||
{
|
||||
"year": 2012,
|
||||
@@ -74,9 +74,14 @@
|
||||
"description": "canonical_fips_xwalk.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/summary_categories.parquet",
|
||||
"sha256": "0985b607f3f35a8dff62c0561261ab6922423b81d11c07b03bcb3e3461f85e33",
|
||||
"description": "summary_categories.parquet"
|
||||
"path": "data/census_collection_coverage.parquet",
|
||||
"sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
|
||||
"description": "census_collection_coverage.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/code_set.parquet",
|
||||
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
|
||||
"description": "code_set.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/harmonization_map.parquet",
|
||||
@@ -88,10 +93,25 @@
|
||||
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
|
||||
"description": "harmonization_recipes.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/lineage_events.parquet",
|
||||
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
|
||||
"description": "lineage_events.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/representation.parquet",
|
||||
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
|
||||
"description": "representation.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/series_breaks.parquet",
|
||||
"sha256": "b0b6794b6887a4f300079adfa10029c2a77109faa4952fbff1c5a270793cc02b",
|
||||
"sha256": "731998516cd802f63fcf7fb66053c7a62b7be955ab0794cad4a4979cb7628b87",
|
||||
"description": "series_breaks.parquet"
|
||||
},
|
||||
{
|
||||
"path": "data/summary_categories.parquet",
|
||||
"sha256": "e3b0efa00ce713b8f45829b89cfde24b55333f26101f0495df82d85997d18d8e",
|
||||
"description": "summary_categories.parquet"
|
||||
}
|
||||
]
|
||||
},
|
||||
|
||||
@@ -14,25 +14,96 @@
|
||||
"basis_note": { "type": ["string", "null"] },
|
||||
"expenditure_concept": {
|
||||
"type": "string",
|
||||
"enum": ["direct", "total"],
|
||||
"description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments."
|
||||
"enum": ["primary", "direct", "total"],
|
||||
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
|
||||
},
|
||||
"expenditure_concept_note": {
|
||||
"type": ["string", "null"],
|
||||
"description": "How the intergovernmental leg was assembled; null for 'direct'."
|
||||
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
|
||||
},
|
||||
"expenditure_concept_direct_suppressed": {
|
||||
"type": "boolean",
|
||||
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
|
||||
"type": ["boolean", "null"],
|
||||
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. null (NA) when expenditure_concept = 'total' AND category = 'All Categories': the detector keys on per-category rows, which that mode collapses, so suppression cannot be computed -- see `expenditure_concept_note`. See the affected rows' `notes` for the recovering recipe, if any."
|
||||
},
|
||||
"revenue_concept": {
|
||||
"type": "string",
|
||||
"enum": ["general", "total"],
|
||||
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
|
||||
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
|
||||
},
|
||||
"harmonization": { "type": "object" },
|
||||
"recipe": { "type": ["object", "null"] },
|
||||
"suggestions": { "type": "array" },
|
||||
"suggestions": {
|
||||
"type": "array",
|
||||
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["recipe_id", "label", "available_years", "hint", "ig_recipe_id",
|
||||
"trigger", "suppressed_amount", "suppressed_years", "suppressed_codes"],
|
||||
"properties": {
|
||||
"recipe_id": { "type": "string" },
|
||||
"label": { "type": "string" },
|
||||
"available_years": {
|
||||
"type": "array",
|
||||
"items": { "type": "integer" },
|
||||
"description": "[year_min, year_max] of the recipe's component coverage."
|
||||
},
|
||||
"hint": { "type": "string" },
|
||||
"ig_recipe_id": {
|
||||
"type": ["string", "null"],
|
||||
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
|
||||
},
|
||||
"trigger": {
|
||||
"type": "string",
|
||||
"enum": ["empty_year", "suppressed_component"],
|
||||
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars this government reports in the requested years that the verb's underlying long view structurally excludes -- aggregate-published, carrying no harmonized code, or absent from summary_categories. This is NOT the same thing as 'excluded from the result': a component present in the view under a different category (a scoping choice, e.g. a different `category` or a narrower `expenditure_concept`) contributes 0 and never fires. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way, using the same underlying-view measurement, and can be 0 even on an 'empty_year' fire."
|
||||
},
|
||||
"suppressed_amount": {
|
||||
"type": "number",
|
||||
"description": "Full US dollars this government reports, in the recipe's component codes, in the requested years, that the verb's underlying long view structurally excludes (aggregate-published, carrying no harmonized code, or absent from summary_categories) -- summed across those years. This is NOT the same quantity as 'what the result excludes': a component present in the view under a different category or a narrower `expenditure_concept` is scoped out on purpose, counts as 0 here, and is not suppression. 0 does not always mean full coverage -- see 'trigger' and 'empty_year'. May be negative where Census publishes a negative `amt` for the excluded rows."
|
||||
},
|
||||
"suppressed_years": {
|
||||
"type": "array",
|
||||
"items": { "type": "integer" },
|
||||
"description": "The requested years contributing to suppressed_amount."
|
||||
},
|
||||
"suppressed_codes": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"description": "The excluded component item codes, sorted."
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"scope": { "type": "object" },
|
||||
"codes_summed": { "type": "object" },
|
||||
"aggregate_fallback": { "type": ["object", "null"] },
|
||||
"transformations":{ "type": "object" },
|
||||
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
||||
"completion": {
|
||||
"type": "object",
|
||||
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
|
||||
"properties": {
|
||||
"applied": { "type": "boolean" },
|
||||
"rows_filled": { "type": "integer" },
|
||||
"absence_means": { "type": "object" }
|
||||
}
|
||||
},
|
||||
"corpus_break_refs": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
||||
},
|
||||
"balance_caveats": {
|
||||
"type": ["object", "null"],
|
||||
"description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.",
|
||||
"properties": {
|
||||
"not_gaap": { "type": "boolean" },
|
||||
"not_gaap_note": { "type": "string" },
|
||||
"coverage_window": { "type": "object" },
|
||||
"truncated": { "type": "array", "items": { "type": "string" } }
|
||||
}
|
||||
},
|
||||
"manifest": { "type": "object" },
|
||||
"sql_query": { "type": "string" }
|
||||
}
|
||||
|
||||
@@ -1,3 +1,7 @@
|
||||
CREATE OR REPLACE VIEW long AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true);
|
||||
-- {long_files} carries its own quoting: a bracketed list of every partition
|
||||
-- the manifest enumerates, or a single quoted glob on fallback. Do NOT wrap
|
||||
-- it in quotes. See .long_files_sql() in R/views.R for why a glob alone
|
||||
-- cannot work over HTTP.
|
||||
FROM read_parquet({long_files}, hive_partitioning = true);
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
-- Category crosswalk. Numbered 11 (not with the other reference tables at
|
||||
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
|
||||
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
|
||||
-- already exist when they register.
|
||||
CREATE OR REPLACE VIEW summary_categories AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/summary_categories.parquet');
|
||||
@@ -1,5 +1,22 @@
|
||||
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
|
||||
-- (summary_categories.category_type = 'expenditure'), never by item-code
|
||||
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
|
||||
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
|
||||
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
|
||||
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
|
||||
-- every non-intergovernmental expenditure subtype: operations, capital,
|
||||
-- assistance, interest, insurance_benefits.
|
||||
--
|
||||
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
|
||||
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
|
||||
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
|
||||
-- 24-ig_long.sql).
|
||||
CREATE OR REPLACE VIEW spending_long AS
|
||||
SELECT *
|
||||
FROM long
|
||||
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G')
|
||||
WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE category_type = 'expenditure'
|
||||
AND spend_subtype <> 'intergovernmental'
|
||||
)
|
||||
AND NOT is_aggregate;
|
||||
|
||||
@@ -1,5 +1,18 @@
|
||||
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
|
||||
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
|
||||
--
|
||||
-- Carries EVERY revenue subtype. Which of Census's two published concepts a
|
||||
-- query actually returns is decided per revenue_concept in R
|
||||
-- (.verb_spendrev), exactly as expenditure_concept narrows spending_long:
|
||||
-- general = own_source + federal + state + local_aid (the default)
|
||||
-- total = general + utility + liquor_store + insurance_trust
|
||||
-- Census defines the first by subtracting the other three from the second
|
||||
-- (manual section 4.3), so both concepts need all four families present here.
|
||||
CREATE OR REPLACE VIEW revenue_long AS
|
||||
SELECT *
|
||||
FROM long
|
||||
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D')
|
||||
WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE category_type = 'revenue'
|
||||
)
|
||||
AND NOT is_aggregate;
|
||||
|
||||
@@ -1,6 +1,16 @@
|
||||
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
|
||||
-- classification, applied to harmonized_code (the code the row is folded
|
||||
-- onto) rather than the published item_code. Safe because the harmonized
|
||||
-- space is leaf-only and every harmonized_code in the corpus is a
|
||||
-- summary_categories member (verified at fixture regen; a code the
|
||||
-- crosswalk cannot classify would be silently dropped here).
|
||||
CREATE OR REPLACE VIEW spending_long_harmonized AS
|
||||
SELECT * REPLACE (harmonized_code AS item_code)
|
||||
FROM long
|
||||
WHERE NOT is_aggregate
|
||||
AND harmonized_code IS NOT NULL
|
||||
AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G');
|
||||
AND harmonized_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE category_type = 'expenditure'
|
||||
AND spend_subtype <> 'intergovernmental'
|
||||
);
|
||||
|
||||
@@ -1,6 +1,12 @@
|
||||
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
|
||||
-- classification (every revenue subtype; the concept narrows in R), applied
|
||||
-- to harmonized_code rather than the published item_code.
|
||||
CREATE OR REPLACE VIEW revenue_long_harmonized AS
|
||||
SELECT * REPLACE (harmonized_code AS item_code)
|
||||
FROM long
|
||||
WHERE NOT is_aggregate
|
||||
AND harmonized_code IS NOT NULL
|
||||
AND LEFT(harmonized_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D');
|
||||
AND harmonized_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE category_type = 'revenue'
|
||||
);
|
||||
|
||||
+12
-5
@@ -1,4 +1,6 @@
|
||||
-- Intergovernmental expenditure rows (M = to local govts, L = to state govts).
|
||||
-- Intergovernmental expenditure rows: crosswalk spend_subtype =
|
||||
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
|
||||
-- = state payments to school systems -- uscogdata#11, finding F-017).
|
||||
--
|
||||
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
|
||||
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
|
||||
@@ -9,10 +11,15 @@
|
||||
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
|
||||
-- the pipeline's recipe joins use.
|
||||
--
|
||||
-- `L--` IS excluded: it is the IG-to-state FAMILY TOTAL and genuinely rolls up
|
||||
-- the L-NN codes, so including it would double-count.
|
||||
-- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
|
||||
-- rolls up the L-NN codes, so including it would double-count. The crosswalk
|
||||
-- deliberately carries no `--` family-total codes, so membership excludes it
|
||||
-- (guarded by "the IG leg never includes the L-- family total" in
|
||||
-- tests/testthat/test-expenditure-concept.R).
|
||||
CREATE OR REPLACE VIEW ig_long AS
|
||||
SELECT *
|
||||
FROM long
|
||||
WHERE LEFT(item_code, 1) IN ('M', 'L')
|
||||
AND item_code NOT LIKE '%--';
|
||||
WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE spend_subtype = 'intergovernmental'
|
||||
);
|
||||
|
||||
@@ -8,8 +8,15 @@
|
||||
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
|
||||
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
|
||||
-- while never dropping a row.
|
||||
--
|
||||
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
|
||||
-- rather than the COALESCEd code: every IG harmonization target (M36) is
|
||||
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
|
||||
-- the column that exists on every row.
|
||||
CREATE OR REPLACE VIEW ig_long_harmonized AS
|
||||
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
|
||||
FROM long
|
||||
WHERE LEFT(item_code, 1) IN ('M', 'L')
|
||||
AND item_code NOT LIKE '%--';
|
||||
WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE spend_subtype = 'intergovernmental'
|
||||
);
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
-- Cash and security holdings, classified by crosswalk MEMBERSHIP on
|
||||
-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot
|
||||
-- do this job -- the X and Y families each span revenue, expenditure AND
|
||||
-- balance).
|
||||
--
|
||||
-- These rows are STOCKS: a balance at a point in time, not a flow over a
|
||||
-- fiscal year. Summing a stock with a flow is meaningless, which is why they
|
||||
-- live behind a third view rather than as a subtype of either money view, and
|
||||
-- why neither spending_long nor revenue_long can reach them.
|
||||
--
|
||||
-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era
|
||||
-- aggregate-only holdings codes (X40/X41) are deliberately outside this view;
|
||||
-- they are reachable only through the recipe path, which bypasses this filter
|
||||
-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2).
|
||||
CREATE OR REPLACE VIEW balance_long AS
|
||||
SELECT *
|
||||
FROM long
|
||||
WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories
|
||||
WHERE category_type = 'balance'
|
||||
)
|
||||
AND NOT is_aggregate;
|
||||
@@ -1,3 +0,0 @@
|
||||
CREATE OR REPLACE VIEW summary_categories AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/summary_categories.parquet');
|
||||
@@ -0,0 +1,3 @@
|
||||
CREATE OR REPLACE VIEW representation AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/representation.parquet');
|
||||
@@ -0,0 +1,3 @@
|
||||
CREATE OR REPLACE VIEW code_set AS
|
||||
SELECT *
|
||||
FROM read_parquet('{url}data/code_set.parquet');
|
||||
@@ -0,0 +1,16 @@
|
||||
CREATE OR REPLACE VIEW balance_annotated AS
|
||||
SELECT
|
||||
s.*,
|
||||
x.gov_name AS xwalk_gov_name,
|
||||
x.govs_type,
|
||||
x.type_label,
|
||||
x.fips_state AS xwalk_fips_state,
|
||||
x.fips_county AS xwalk_fips_county,
|
||||
x.fips_place,
|
||||
x.population_acs,
|
||||
c.category,
|
||||
c.category_type,
|
||||
c.balance_subtype
|
||||
FROM balance_long s
|
||||
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
|
||||
LEFT JOIN summary_categories c USING (item_code);
|
||||
@@ -0,0 +1,81 @@
|
||||
% Generated by roxygen2: do not edit by hand
|
||||
% Please edit documentation in R/balances.R
|
||||
\name{cog_balances}
|
||||
\alias{cog_balances}
|
||||
\title{Cash and security holdings for one or more governments}
|
||||
\usage{
|
||||
cog_balances(
|
||||
govid,
|
||||
years,
|
||||
category = NULL,
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
\item{govid}{Canonical govid(s): a character vector, or a data frame with a
|
||||
`canonical_govid` column (e.g. from [cog_gov_search()]).}
|
||||
|
||||
\item{years}{Integer vector of fiscal years.}
|
||||
|
||||
\item{category}{Optional character vector of categories to keep. One of
|
||||
`"Fund Balances"`, `"Insurance Trust Balances"`,
|
||||
`"Retirement System Holdings"`. There is deliberately no `subtype`
|
||||
argument: for holdings, `category` is a strict coarsening of
|
||||
`balance_subtype` (unlike the money verbs, where the two axes cross), so
|
||||
every combination would be either redundant or empty.
|
||||
`category = "Fund Balances"` is exactly the `general` family
|
||||
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
||||
one `dplyr::filter()` away. The reserved pseudo-category
|
||||
`"All Categories"` (see [cog_spending()]) is **not** supported here and
|
||||
errors with class `uscogdata_all_categories_unsupported`: it sums a
|
||||
concept's subtype scope, and holdings are a stock with no concept
|
||||
vocabulary to sum across. Omit `category` to get every category broken
|
||||
out instead.}
|
||||
|
||||
\item{per_capita}{Divide holdings by population. Note this is a **stock per
|
||||
resident** (reserves per person), which is *not* comparable to
|
||||
[cog_spending()]'s per-capita figures -- those are a flow per person.}
|
||||
|
||||
\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).}
|
||||
|
||||
\item{basis}{Accepted for uniformity with the money verbs, but currently a
|
||||
**no-op**: `harmonization_map` carries no balance-code rows, so harmonized
|
||||
and raw space are identical for holdings. Reported in
|
||||
`provenance$basis_note`.}
|
||||
|
||||
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]).
|
||||
`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
|
||||
wide era to the modern one.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`balance_subtype`, `category`, `amt_nominal`, `codes_included`,
|
||||
`aggregate_fallback`, plus optional `amt_per_capita_nominal` and
|
||||
`pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
|
||||
`adjust_to_year` is set), and optional `amt_per_capita_real` (only when
|
||||
**both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
|
||||
nominal per-capita column to deflate otherwise). Amounts are full US
|
||||
dollars.
|
||||
|
||||
Carries a `provenance` attribute matching
|
||||
`inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
|
||||
`not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
|
||||
every balance subtype in the mounted corpus, not only the observed ones)
|
||||
and `truncated` (the observed subtypes whose coverage falls short of the
|
||||
requested years). `expenditure_concept`/`revenue_concept` are `NA` --
|
||||
holdings are a stock, not a flow, so neither concept vocabulary applies.
|
||||
}
|
||||
\description{
|
||||
Returns Census cash-and-security holdings (`category_type = "balance"`):
|
||||
fund balances, retirement system holdings and insurance trust balances.
|
||||
}
|
||||
\section{Holdings are not GAAP fund balance}{
|
||||
|
||||
Census holdings are **gross** -- no liabilities are netted -- so a reserve
|
||||
ratio built from them overstates what is actually available. They are not
|
||||
comparable to a GAAP fund balance from an ACFR.
|
||||
}
|
||||
|
||||
+15
-4
@@ -7,8 +7,8 @@
|
||||
cog_categories(type = NULL, pattern = NULL)
|
||||
}
|
||||
\arguments{
|
||||
\item{type}{Either `NULL` (default, return both spending and revenue
|
||||
rows), `"spending"`, or `"revenue"`.}
|
||||
\item{type}{Either `NULL` (default, every row: expenditure, revenue and
|
||||
balance), `"spending"`, `"revenue"`, or `"balance"`.}
|
||||
|
||||
\item{pattern}{Optional regex matched case-insensitively against the
|
||||
`category` column (e.g. `"Police"` or `"Tax"`).}
|
||||
@@ -16,13 +16,24 @@ rows), `"spending"`, or `"revenue"`.}
|
||||
\value{
|
||||
Tibble with columns `category`, `category_type`, `subtype`,
|
||||
`n_codes`, `item_codes` (comma-separated, alphabetical). Sorted by
|
||||
`category_type`, `category`, `subtype`.
|
||||
`category_type`, `category`, `subtype`. Includes one row per flow for the
|
||||
reserved pseudo-category `"All Categories"`, which carries `NA` for
|
||||
`subtype`, `n_codes` and `item_codes` because it is a query mode rather
|
||||
than a crosswalk entry — see [cog_spending()]'s `category` argument.
|
||||
}
|
||||
\description{
|
||||
Returns the category taxonomy exposed by the corpus's
|
||||
`summary_categories` view, grouped to one row per
|
||||
`(category, subtype)` pair. Use this to discover valid `category`
|
||||
values for [cog_spending()] / [cog_revenue()] /
|
||||
values for [cog_spending()] / [cog_revenue()] / [cog_balances()] /
|
||||
[cog_geographic_rollup()] and to audit which Census item codes feed
|
||||
each category.
|
||||
}
|
||||
\details{
|
||||
`subtype` COALESCEs the crosswalk's three subtype columns, so it carries
|
||||
`spend_subtype` on expenditure rows, `revenue_subtype` on revenue rows and
|
||||
`balance_subtype` on balance rows. Note that [cog_balances()] itself takes
|
||||
no `subtype` argument — for holdings, `category` is a strict coarsening of
|
||||
`balance_subtype` — but the value is surfaced here because it is the
|
||||
discovery surface downstream consumers build their vocabulary from.
|
||||
}
|
||||
|
||||
@@ -21,3 +21,33 @@ Prints the structured provenance attached to a tibble returned by any
|
||||
`cog_*` verb, or returns it as a list for downstream use (MCP tools,
|
||||
dashboards, JSON export).
|
||||
}
|
||||
\section{Two kinds of series break}{
|
||||
|
||||
Catalogued breaks reach you without being asked for, in two disjoint
|
||||
fields, because a caveat about one series and a caveat about the whole
|
||||
corpus are different claims:
|
||||
|
||||
* **`series_break_refs`** — breaks matched against the item codes actually
|
||||
present in this result. A break in one code you queried.
|
||||
* **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
|
||||
which are statements about the corpus rather than about any one code:
|
||||
dollar precision across the 1976/1977 boundary (`SB085`), imputation
|
||||
exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
|
||||
representation change (`SB194`), and the FY2017 government-identifier
|
||||
change (`SB086`). These are selected on the break-year window alone.
|
||||
|
||||
`SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
|
||||
crosses the boundary where an absent cell stops meaning "Census published
|
||||
$0" and starts meaning "not reported".
|
||||
}
|
||||
|
||||
\section{Other provenance blocks}{
|
||||
|
||||
`transformations$units_conversion` records the `$1,000s`-to-dollars
|
||||
multiply that every amount column has already had applied.
|
||||
`transformations$per_capita` records the population denominator and its
|
||||
year range. `coverage` and `coverage_mode` appear on multi-government
|
||||
results (see [cog_geographic_rollup()]). `completion` appears when
|
||||
`complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
|
||||
}
|
||||
|
||||
|
||||
+10
-1
@@ -11,7 +11,8 @@ cog_find_peers(
|
||||
same_state = FALSE,
|
||||
pop_range = c(0.7, 1.3),
|
||||
is_ratio = TRUE,
|
||||
max_peers = 10L
|
||||
max_peers = 10L,
|
||||
coverage = c("all", "census", "consistent")
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -34,6 +35,14 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
|
||||
`pop_range` is interpreted as absolute population counts.}
|
||||
|
||||
\item{max_peers}{Integer cap on the number of peers returned.}
|
||||
|
||||
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
|
||||
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
||||
most recent census year with an observed population, so a cohort is not
|
||||
built from a sample year in which most of the candidate universe is
|
||||
absent. `"consistent"` needs a year range, which cohort selection does not
|
||||
have, so it selects like `"all"` and is carried on the result as
|
||||
`attr(x, "coverage")` for [cog_peer_compare()].}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||
|
||||
@@ -10,7 +10,8 @@ cog_geographic_rollup(
|
||||
years,
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
expenditure_concept = c("direct", "total")
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
coverage = c("all", "census", "consistent")
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -19,7 +20,11 @@ cog_geographic_rollup(
|
||||
`canonical_govid` values. At least one layer required.}
|
||||
|
||||
\item{category}{Single category name or character vector (passed through
|
||||
to [cog_spending()]).}
|
||||
to [cog_spending()]), or the reserved `"All Categories"` for one summed
|
||||
row per `(year, canonical_govid, subtype)` covering every category in the
|
||||
concept's scope. `"All Categories"` is the efficient way to build a
|
||||
geographic total: without it a caller must issue one rollup per category
|
||||
and sum the results themselves.}
|
||||
|
||||
\item{years}{Integer vector of years.}
|
||||
|
||||
@@ -29,12 +34,32 @@ are excluded from the result.}
|
||||
|
||||
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
|
||||
|
||||
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
|
||||
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||
single-government queries but cannot be used here because combining Total
|
||||
across multiple layers of government double-counts intergovernmental
|
||||
transfers (a state's payment to a school district is the same dollar the
|
||||
district reports as its own Direct spending).}
|
||||
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
|
||||
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
||||
refused here because combining Total across multiple layers of
|
||||
government double-counts intergovernmental transfers (a state's payment
|
||||
to a school district is the same dollar the district reports as its own
|
||||
Direct spending); `"primary"` and `"direct"` combine safely.}
|
||||
|
||||
\item{coverage}{How to handle the Census of Governments survey cycle,
|
||||
which is a **complete census only in years ending in 2 and 7** -- every
|
||||
other year is a sample, and the sample varies enormously (on the bundled
|
||||
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||
and 112 in FY2019).
|
||||
|
||||
* `"all"` (default) -- every unit that reported that year. Unchanged
|
||||
behaviour, so existing code keeps working.
|
||||
* `"census"` -- census years only. Aborts if the requested range holds
|
||||
none, rather than silently returning nothing.
|
||||
* `"consistent"` -- only units reporting in *every* requested year, giving
|
||||
a balanced panel.
|
||||
|
||||
Regardless of mode, `provenance$coverage` always carries per-year
|
||||
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
statement about the **survey calendar**, never a claim of completeness:
|
||||
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
report. `n_units_reporting` is the number that tells the truth.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||
@@ -59,3 +84,23 @@ the result. The dropped govids are recorded in
|
||||
(gov type 4) and school districts (gov type 5) from per-capita rollups
|
||||
by design — see `vignette('population-denominators')`.
|
||||
}
|
||||
\section{Reading `coverage`}{
|
||||
|
||||
`provenance$coverage` reports `n_units_reporting` against
|
||||
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
it counts governments with rows for the category you asked for, not
|
||||
governments collected that year.** A government that was surveyed and
|
||||
genuinely spends nothing in that category is indistinguishable here from one
|
||||
that was never surveyed.
|
||||
|
||||
The ratio is therefore **not a response rate** and must not be used as one.
|
||||
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
`category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
contract policing to the county sheriff, not non-response.
|
||||
|
||||
The comparison that *is* valid is the same category across a census year
|
||||
(ending in 2 or 7) and a sample year, where the real-zero component is
|
||||
roughly constant and the difference reflects the survey cycle. `is_census_year`
|
||||
marks which is which.
|
||||
}
|
||||
|
||||
|
||||
@@ -32,8 +32,11 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
||||
}
|
||||
\details{
|
||||
* **Utility mode** (single `name`, the original behavior): returns all
|
||||
rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||
`population_acs` descending. Useful for exploratory lookups.
|
||||
rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
||||
substring**, sorted by `population_acs` descending. Useful for
|
||||
exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
||||
government is findable by its own complete name even when that name
|
||||
contains parentheses or a period.
|
||||
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||
single canonical govid and returns a tibble in input order, suitable
|
||||
for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||
@@ -45,7 +48,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
||||
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||
Single hit -> resolved. Multiple -> step 4.
|
||||
3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||
3. **Substring fallback:** case-insensitive literal substring against
|
||||
`gov_name` (metacharacters escaped).
|
||||
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||
`status = "no_match"`. Multiple hits -> step 4.
|
||||
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||
@@ -58,7 +62,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
|
||||
}
|
||||
\examples{
|
||||
\dontrun{
|
||||
# Utility mode — exploratory regex lookup
|
||||
# Utility mode — exploratory substring lookup
|
||||
cog_gov_search("broward", state = "FL")
|
||||
|
||||
# Basket mode — resolve a known cohort
|
||||
|
||||
+82
-6
@@ -11,7 +11,8 @@ cog_peer_compare(
|
||||
years,
|
||||
per_capita = TRUE,
|
||||
adjust_to_year = NULL,
|
||||
expenditure_concept = c("direct", "total")
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
coverage = c("all", "census", "consistent")
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -29,10 +30,37 @@ population.}
|
||||
|
||||
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
|
||||
|
||||
\item{expenditure_concept}{`"direct"` (default) or `"total"`. Currently only
|
||||
`"direct"` is accepted; the `"total"` option exists in [cog_spending()] for
|
||||
single-government queries but cannot be used here because combining Total
|
||||
across peer sets counts intergovernmental transfers twice.}
|
||||
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
|
||||
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
||||
refused here because combining Total across peer sets counts
|
||||
intergovernmental transfers twice; `"primary"` and `"direct"` combine
|
||||
safely.}
|
||||
|
||||
\item{coverage}{How to handle the Census of Governments survey cycle,
|
||||
which is a **complete census only in years ending in 2 and 7** -- every
|
||||
other year is a sample, and the sample varies enormously (on the bundled
|
||||
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
||||
and 112 in FY2019).
|
||||
|
||||
* `"all"` (default) -- every unit that reported that year. Unchanged
|
||||
behaviour, so existing code keeps working.
|
||||
* `"census"` -- census years only. Aborts if the requested range holds
|
||||
none, rather than silently returning nothing.
|
||||
* `"consistent"` -- only units reporting in *every* requested year, giving
|
||||
a balanced panel.
|
||||
|
||||
Regardless of mode, `provenance$coverage` always carries per-year
|
||||
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
||||
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
||||
statement about the **survey calendar**, never a claim of completeness:
|
||||
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
||||
report. `n_units_reporting` is the number that tells the truth.
|
||||
|
||||
The comparison target is exempt from `"consistent"` balancing -- it is the
|
||||
subject of the comparison, not a member of the cohort -- and the
|
||||
`summary_*` quantiles are computed AFTER the filter, so they describe the
|
||||
cohort actually returned. `n_units_reporting` counts peers only, against
|
||||
the cohort size: "3 of your 15 peers reported in FY2019".}
|
||||
}
|
||||
\value{
|
||||
Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||
@@ -43,11 +71,59 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||
`cohort_year`, and `cohort_govids`.
|
||||
|
||||
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
||||
Each one is computed **within each `(year, spend_subtype,
|
||||
category)` cell** across the peer set, so a `summary_p50` row is *the
|
||||
median peer's value in that one category*, not *the value of the median
|
||||
peer's total*. The median peer for Police and the median peer for Fire
|
||||
are usually different governments, so summing `summary_*` rows across
|
||||
categories does not give any peer's total and misstates the band it
|
||||
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
||||
one cohort, with a sign flip at FY2012.
|
||||
|
||||
Facet by `role` **and** `category` (the documented use, and what the
|
||||
rows are built for). For a genuine "median peer's total spending" line,
|
||||
sum each peer's own categories first and take the quantile of those
|
||||
per-government totals:
|
||||
|
||||
```r
|
||||
library(dplyr)
|
||||
cmp |>
|
||||
filter(role %in% c("target", "peer")) |>
|
||||
group_by(year, role, canonical_govid) |>
|
||||
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
||||
filter(role == "peer") |>
|
||||
group_by(year) |>
|
||||
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
||||
```
|
||||
}
|
||||
\description{
|
||||
Pulls spending for the target plus a peer set (either a
|
||||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||
call.
|
||||
call. Those summary rows are quantiles **within each category**, not
|
||||
quantiles of each peer's total — see the `@return` section before summing
|
||||
them.
|
||||
}
|
||||
\section{Reading `coverage`}{
|
||||
|
||||
`provenance$coverage` reports `n_units_reporting` against
|
||||
`n_units_expected` per year. **`n_units_reporting` is category-conditional:
|
||||
it counts cohort members with rows for the category you asked for, not
|
||||
cohort members collected that year.** A government that was surveyed and
|
||||
genuinely spends nothing in that category is indistinguishable here from one
|
||||
that was never surveyed.
|
||||
|
||||
The ratio is therefore **not a response rate** and must not be used as one.
|
||||
In FY2022 — a complete census year — Georgia reports 393 of 567 cities for
|
||||
`category = "Police"`; the 174-city gap is overwhelmingly cities that
|
||||
contract policing to the county sheriff, not non-response.
|
||||
|
||||
The comparison that *is* valid is the same category across a census year
|
||||
(ending in 2 or 7) and a sample year, where the real-zero component is
|
||||
roughly constant and the difference reflects the survey cycle. `is_census_year`
|
||||
marks which is which.
|
||||
}
|
||||
|
||||
|
||||
+75
-3
@@ -11,7 +11,11 @@ cog_revenue(
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL
|
||||
recipe = NULL,
|
||||
revenue_concept = c("general", "total"),
|
||||
complete = FALSE,
|
||||
limit = NULL,
|
||||
offset = NULL
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -20,7 +24,16 @@ cog_revenue(
|
||||
\item{years}{Integer vector of years.}
|
||||
|
||||
\item{category}{Character vector of category names (from
|
||||
`summary_categories.category`), or `NULL` for all categories.}
|
||||
`summary_categories.category`), or `NULL` for all categories broken out
|
||||
one row each. The reserved value `"All Categories"` instead returns a
|
||||
single summed row per `(year, canonical_govid, subtype)`, covering every
|
||||
category inside the requested concept's subtype scope. It cannot be
|
||||
combined with other category names, and it is not the same thing as
|
||||
`revenue_concept = "total"`: the concept chooses which subtypes are in
|
||||
scope, `"All Categories"` chooses whether rows inside that scope are
|
||||
broken out or summed. Because the result keeps one row per
|
||||
`revenue_subtype`, filtering the returned frame to
|
||||
`revenue_subtype == "own_source"` gives an own-source revenue total.}
|
||||
|
||||
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||
@@ -55,12 +68,71 @@ argument is ignored and the result's provenance reports
|
||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{revenue_concept}{Which of Census's two published revenue concepts to
|
||||
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
||||
values -- never as item-code first letters, which cannot classify
|
||||
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
|
||||
prefix `X` does the same):
|
||||
|
||||
* `"general"` (default) -- Census General Revenue: `own_source` +
|
||||
`federal` + `state` + `local_aid`. The manual defines this concept by
|
||||
subtraction (section 4.3: *"General revenue comprises all revenue
|
||||
except that classified as liquor store, utility, or insurance trust
|
||||
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
|
||||
insurance trust revenue are all excluded.
|
||||
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
|
||||
`general` plus utility, liquor store, and insurance trust revenue
|
||||
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
|
||||
`X01`/`X02`/`X05`/`X08`).
|
||||
|
||||
The two are related by Census's own identity, `Total Revenue = General +
|
||||
Utility + Liquor Store + Insurance Trust`.
|
||||
|
||||
Note that the employee-retirement (`X`) codes stop at FY2016, when those
|
||||
systems moved out of the annual finance file into the separate Annual
|
||||
Survey of Public Pensions, so a `"total"` series steps down at the
|
||||
FY2016/FY2017 seam for reasons that are about collection scope rather
|
||||
than revenue (series breaks `SB197`-`SB202`).}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
|
||||
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
||||
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
||||
result is materialized. `NULL` (the default) returns every matching row,
|
||||
exactly as before this parameter existed. Mutually exclusive with
|
||||
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
|
||||
|
||||
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
||||
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
}
|
||||
\description{
|
||||
Mirror of [cog_spending()] for revenue categories. One row per
|
||||
|
||||
+99
-31
@@ -12,7 +12,10 @@ cog_spending(
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL,
|
||||
expenditure_concept = c("direct", "total")
|
||||
expenditure_concept = c("primary", "direct", "total"),
|
||||
complete = FALSE,
|
||||
limit = NULL,
|
||||
offset = NULL
|
||||
)
|
||||
}
|
||||
\arguments{
|
||||
@@ -21,7 +24,16 @@ cog_spending(
|
||||
\item{years}{Integer vector of years.}
|
||||
|
||||
\item{category}{Character vector of category names (from
|
||||
`summary_categories.category`), or `NULL` for all categories.}
|
||||
`summary_categories.category`), or `NULL` for all categories broken out
|
||||
one row each. The reserved value `"All Categories"` instead returns a
|
||||
single summed row per `(year, canonical_govid, subtype)`, covering every
|
||||
category inside the requested concept's subtype scope. It cannot be
|
||||
combined with other category names, and it is not the same thing as
|
||||
`expenditure_concept = "total"`: the concept chooses which subtypes are in
|
||||
scope, `"All Categories"` chooses whether rows inside that scope are
|
||||
broken out or summed. Because the result keeps one row per
|
||||
`spend_subtype`, filtering the returned frame to
|
||||
`spend_subtype == "operations"` gives an operating-expenditure total.}
|
||||
|
||||
\item{per_capita}{If `TRUE`, adds `amt_per_capita_nominal` (and
|
||||
`amt_per_capita_real` when `adjust_to_year` is set) using the per-year
|
||||
@@ -57,41 +69,97 @@ argument is ignored and the result's provenance reports
|
||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||
|
||||
\item{expenditure_concept}{`"direct"` (default) returns only the
|
||||
government's own direct spending (item codes `E`/`F`/`G`), unchanged
|
||||
from prior releases. `"total"` additionally UNIONs in the
|
||||
intergovernmental leg -- payments to local governments (`M` codes) and
|
||||
to the state government (`L` codes, excluding the `L--` family-total
|
||||
rollup) -- so results gain rows with `spend_subtype ==
|
||||
"intergovernmental"`. Requires the active corpus's `summary_categories`
|
||||
to carry M/L rows (added by cog_pipeline PR #59); aborts with class
|
||||
`uscogdata_ig_categories_unsupported` on an older corpus rather than
|
||||
silently under-reporting. Mutually exclusive with `recipe` (a recipe
|
||||
already defines its own component codes). **Do not sum `"total"`
|
||||
results across levels of government** (e.g. state + county + city):
|
||||
a state's `M12` payment to a school district is the same dollar the
|
||||
district reports as its own direct `E12`, so summing both double-counts
|
||||
it. This matters in particular with [cog_geographic_rollup()], which
|
||||
sums across exactly that kind of multi-layer government set.
|
||||
\item{expenditure_concept}{Which spending concept to return. Concepts are
|
||||
defined as sets of the crosswalk's `spend_subtype` values -- never as
|
||||
item-code first letters, which cannot classify correctly (prefix `Y`
|
||||
alone spans revenue, expenditure, and balance codes):
|
||||
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG.}
|
||||
* `"primary"` (default) -- the government's own service provision:
|
||||
`operations` + `capital` + `assistance` subtypes.
|
||||
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
|
||||
`interest` (interest on debt) and `insurance_benefits` (insurance
|
||||
trust benefit payments, e.g. pensions -- Census manual section
|
||||
5.2.2.1 includes payments to retirees in Direct).
|
||||
* `"total"` -- `direct` plus the intergovernmental leg: payments to
|
||||
local governments (`M` codes), to the state government (`L` codes,
|
||||
excluding the `L--` family-total rollup), and state payments to
|
||||
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
|
||||
`spend_subtype == "intergovernmental"`. Requires the active corpus's
|
||||
`summary_categories` to carry M/L rows (added by cog_pipeline PR
|
||||
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
|
||||
older corpus rather than silently under-reporting. Mutually
|
||||
exclusive with `recipe` (a recipe already defines its own component
|
||||
codes).
|
||||
|
||||
**Do not sum `"total"` results across levels of government** (e.g.
|
||||
state + county + city): a state's `M12` payment to a school district is
|
||||
the same dollar the district reports as its own direct `E12`, so
|
||||
summing both double-counts it. This matters in particular with
|
||||
[cog_geographic_rollup()], which sums across exactly that kind of
|
||||
multi-layer government set.
|
||||
|
||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
||||
split), which the Direct leg excludes by construction but the IG leg
|
||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
||||
query, any (year, category) where this leaves intergovernmental rows
|
||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
||||
name the harmonization recipe that recovers the missing Direct
|
||||
component (when one exists), and
|
||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
||||
IG. When `category = "All Categories"` is combined with
|
||||
`expenditure_concept = "total"`, this detection cannot run (it keys on
|
||||
per-category rows, which all-categories mode collapses to one literal
|
||||
value), so `expenditure_concept_direct_suppressed` is `NA` rather than a
|
||||
possibly-false `FALSE`; query an explicit `category` to get a real
|
||||
answer.}
|
||||
|
||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
||||
corpus does not carry still appears, labelled with **why** it is
|
||||
missing, and add a `value_source` column to every row:
|
||||
|
||||
* `"reported"` — the corpus carries this cell.
|
||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
||||
Census published `$0`. `amt_nominal` is `0`.
|
||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
||||
government did not report, and the value is unknown. `amt_nominal` is
|
||||
`NA`, **not** `0` — writing a zero there would invent data.
|
||||
|
||||
The grid comes from the corpus's `code_set` table, scoped to each
|
||||
government's own type, so a county is never filled with cells only a
|
||||
state can report. Reported rows are passed through untouched.
|
||||
|
||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
||||
when `representation`/`code_set` began shipping; aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
||||
`recipe` or with `expenditure_concept = "total"` (class
|
||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
||||
`code_set`.}
|
||||
|
||||
\item{limit}{Maximum number of result rows to return, pushed into the SQL
|
||||
query itself (`LIMIT`/`OFFSET`) rather than applied after the full
|
||||
result is materialized. `NULL` (the default) returns every matching row,
|
||||
exactly as before this parameter existed. Mutually exclusive with
|
||||
`recipe` and with `complete = TRUE` -- see `offset` and `total_rows`.}
|
||||
|
||||
\item{offset}{Rows to skip before `limit` starts counting (0-based).
|
||||
Ignored if `limit` is `NULL`; defaults to `0L` when `limit` is set.}
|
||||
}
|
||||
\value{
|
||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
||||
and `value_source` when `complete = TRUE`.
|
||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
||||
whose `completion` block reports `applied`, `rows_filled`, and the
|
||||
per-year `absence_means` rule that was applied. When `limit` is set,
|
||||
also carries a `total_rows` attribute: the full unpaginated row count,
|
||||
computed by the same query (`COUNT(*) OVER()`) rather than a second
|
||||
round trip -- so a caller walking pages never has to ask "how many are
|
||||
there" separately.
|
||||
}
|
||||
\description{
|
||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,281 @@
|
||||
# `cog_balances()` — a reader surface for cash and security holdings
|
||||
|
||||
**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26`
|
||||
**Date:** 2026-08-03 · **Status:** design, awaiting approval
|
||||
|
||||
Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped
|
||||
with `#11`/`#12` and is asserted at both view and verb level. This spec covers
|
||||
requirement 2 only: a way to query holdings.
|
||||
|
||||
## Decision: a verb, not an argument
|
||||
|
||||
`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`.
|
||||
|
||||
Holdings are a **stock** — a balance at a point in time — while the money verbs
|
||||
return **flows** over a fiscal year. The flow verbs' whole argument vocabulary
|
||||
is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe
|
||||
which flows Census aggregates into a published total, and `complete=` fills a
|
||||
grid of fiscal-year cells. Overloading a money verb would put a stock behind
|
||||
arguments that all assume a flow.
|
||||
|
||||
## The 14 codes
|
||||
|
||||
Measured against the published corpus 2026-08-03, not transcribed from the
|
||||
issue. `year_min`/`year_max` are observed row extents.
|
||||
|
||||
| `balance_subtype` | `category` | codes | observed years |
|
||||
|---|---|---|---|
|
||||
| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 |
|
||||
| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 |
|
||||
| | | `X47` | 1988–2016 |
|
||||
| | | `X30`, `Z77`, `Z78` | 2012–2016 |
|
||||
| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 |
|
||||
| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 |
|
||||
| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 |
|
||||
|
||||
## Architecture
|
||||
|
||||
### Two new views
|
||||
|
||||
Mirroring the `revenue_long` / `revenue_annotated` pair exactly:
|
||||
|
||||
- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate`
|
||||
- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and
|
||||
`summary_categories`, exposing `category`, `category_type`, `balance_subtype`
|
||||
|
||||
`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register
|
||||
with no new registration code.
|
||||
|
||||
### A third gate list in `R/views.R`
|
||||
|
||||
`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails
|
||||
at registration time, not at query time. `46-balance_annotated.sql` selects
|
||||
`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`.
|
||||
That arrived without a `schema_version` bump, so neither existing gate applies:
|
||||
`.harmonization_view_files` keys on `schema_version`, `.representation_view_files`
|
||||
on the presence of a *file*. The discriminator here is a **column on an existing
|
||||
table**.
|
||||
|
||||
```r
|
||||
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
|
||||
```
|
||||
|
||||
gated by probing `summary_categories` for `balance_subtype`, with
|
||||
`cog_balances()` erroring cleanly via `.require_balance_support()` on an older
|
||||
corpus — mirroring how `.require_schema_v5()` gates the harmonized views.
|
||||
|
||||
### `R/balances.R` — a dedicated path, not `.verb_spendrev()`
|
||||
|
||||
`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and
|
||||
`complete=` grid are all flow-specific, and four verbs depend on it. Threading a
|
||||
third mode through it adds branching to shared code for no reuse benefit.
|
||||
|
||||
Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`,
|
||||
`.build_corpus_break_refs()`, the population join, `.inflate()`, and
|
||||
`.coerce_govid_input()`.
|
||||
|
||||
Following the package's real two-layer convention: **view definitions** live in
|
||||
`inst/sql/`; **query construction** is inline `sprintf()` in R, as in
|
||||
`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R
|
||||
files", which the verb layer has never obeyed. Corrected in a separate commit —
|
||||
see Out of scope.)
|
||||
|
||||
## Signature
|
||||
|
||||
```r
|
||||
cog_balances(govid, years,
|
||||
category = NULL, # Fund Balances | Insurance Trust Balances |
|
||||
# Retirement System Holdings
|
||||
per_capita = FALSE,
|
||||
adjust_to_year = NULL,
|
||||
basis = c("harmonized", "raw"),
|
||||
recipe = NULL)
|
||||
```
|
||||
|
||||
Returns a `tbl_df` with a `provenance` attribute, like every other verb.
|
||||
|
||||
**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`,
|
||||
and `subtype` — see below.
|
||||
|
||||
**`per_capita` is offered.** Holdings per resident is a real measure (pension
|
||||
assets per capita, fund balance per resident). The roxygen `@param` states
|
||||
plainly that this is a *stock per resident* and is **not** comparable to
|
||||
`cog_spending()`'s per-capita figures.
|
||||
|
||||
**`basis` is currently a no-op** — `harmonization_map` has zero balance-code
|
||||
rows, so harmonized and raw are identical for holdings. Kept for uniformity
|
||||
with the money verbs (the API would otherwise special-case), and
|
||||
`provenance$basis_note` says so outright rather than letting it look meaningful.
|
||||
|
||||
**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
|
||||
to the modern one:
|
||||
|
||||
```
|
||||
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
|
||||
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
|
||||
```
|
||||
|
||||
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
|
||||
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
|
||||
other basis view. That is by design, not a defect:
|
||||
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
|
||||
era exposes these split families *only* as aggregates, and that the recipe join
|
||||
must therefore **not** filter `is_aggregate` — safe by construction, because
|
||||
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
|
||||
every component is year-scoped, so no double-count is possible. § 1 records the
|
||||
matching decision that the planned `X40→Z77` harmonization *map* rows were
|
||||
dropped and the continuity ships as recipes instead, which is why
|
||||
`harmonization_map` has no balance-code rows.
|
||||
|
||||
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
|
||||
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
|
||||
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
|
||||
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
|
||||
as intended.
|
||||
|
||||
### No `subtype` argument: `category` is a strict coarsening
|
||||
|
||||
`balance` is the only `category_type` in which `category` and the subtype column
|
||||
are **not** orthogonal. Measured against the published crosswalk:
|
||||
|
||||
| `category_type` | subtypes spanning more than one category |
|
||||
|---|---|
|
||||
| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) |
|
||||
| revenue | 1 of 7 (`own_source`) |
|
||||
| **balance** | **0 of 5** |
|
||||
|
||||
For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire)
|
||||
× *economic character* (operations, capital) — so both earn their place. For
|
||||
balance the relation is a strict tree:
|
||||
|
||||
```
|
||||
Fund Balances = {general} W01 W31 W61
|
||||
Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78
|
||||
Insurance Trust Balances = {unemployment_trust,
|
||||
workers_comp_trust,
|
||||
other_insurance_trust} Y07 Y08 Y21 Y61
|
||||
```
|
||||
|
||||
Exposing both would therefore admit no useful combination. Of the 15 possible
|
||||
pairs, 3 are redundant (the subtype already implies its category) and **12 are
|
||||
guaranteed empty for every government in every year** — and an impossible query
|
||||
would fail by returning an empty tibble, which reads as "this government holds
|
||||
none" rather than "you asked a contradiction."
|
||||
|
||||
Dropping `subtype` also keeps the verb aligned with the rest of the package: no
|
||||
uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in
|
||||
`.verb_spendrev()`, and the API layers its own `subtype` row filter on top
|
||||
(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for
|
||||
`/balances`.
|
||||
|
||||
`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the
|
||||
`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is
|
||||
isolating one of the three insurance funds in a single argument;
|
||||
`balance_subtype` remains a returned column, so that is one `dplyr::filter()`
|
||||
away.
|
||||
|
||||
## Caveat surfacing
|
||||
|
||||
`provenance$balance_caveats`, always present, plus one `cli_inform()` per
|
||||
session per caveat class when a query actually touches an affected family or
|
||||
year. Structured so `cog-api#26` can forward the fields verbatim.
|
||||
|
||||
Verified against `series_breaks.csv`, not assumed:
|
||||
|
||||
| # | Caveat | Covered by existing machinery? |
|
||||
|---|---|---|
|
||||
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
|
||||
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
|
||||
| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
|
||||
| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed |
|
||||
|
||||
On caveat 4's second condition: `.build_series_break_refs()` matches
|
||||
`break_year BETWEEN min(years) AND max(years)`, so a request spanning only
|
||||
2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a
|
||||
series sits entirely after the change, on one consistent basis, and flagging a
|
||||
break it never crosses would be noise. The same rule is applied deliberately in
|
||||
`.build_corpus_break_refs()`. An earlier draft of this row omitted the span
|
||||
condition and overclaimed.
|
||||
|
||||
`coverage_window` is derived per observed subtype family from the corpus, never
|
||||
hardcoded, so it stays correct as the corpus grows.
|
||||
|
||||
`series_break_refs` and `corpus_break_refs` are otherwise populated by the
|
||||
existing code-driven builders and need no change.
|
||||
|
||||
## Testing
|
||||
|
||||
New `tests/testthat/test-balances.R`. The bundled fixture covers all four
|
||||
fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y`
|
||||
throughout — so every test below runs offline.
|
||||
|
||||
- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing
|
||||
the already-asserted forward guard. Absence is verified against the raw corpus
|
||||
via `read_parquet` on `data/long`, never through the verb that creates it.
|
||||
- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019;
|
||||
`coverage_window` reports the termination and the console message fires once.
|
||||
- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the
|
||||
measured table above; the FY2002 valuation caveat fires only when the year
|
||||
range crosses 2002 *and* touches `employee_retirement`.
|
||||
- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`.
|
||||
- **`category = "Fund Balances"` is the `general` family.** Returns exactly
|
||||
`W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted
|
||||
rather than assumed.
|
||||
- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to
|
||||
exactly one `category`. Asserted against the crosswalk so that an upstream
|
||||
change breaking the tree — which would silently make `category` lossy —
|
||||
fails here rather than in a user's analysis.
|
||||
- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
|
||||
`X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
|
||||
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
|
||||
regression here would silently truncate a 45-year series to five.
|
||||
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
|
||||
FY2002 carries both in `provenance$series_break_refs`, so the book → market
|
||||
basis change is disclosed wherever `X40`/`X41` are actually observed.
|
||||
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
|
||||
`summary_categories` lacks `balance_subtype`.
|
||||
|
||||
## Out of scope, tracked separately
|
||||
|
||||
1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
|
||||
the seven holdings codes in `series_breaks.csv`. There is currently **no**
|
||||
entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
|
||||
`docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
|
||||
exactly this — *"candidate new `series_breaks.csv` entries (recommend
|
||||
`with_caution` documentation rows, no map action)"*. The follow-through never
|
||||
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
|
||||
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
|
||||
what keeps this out of the `joinable = "no"` identity-change rule, which
|
||||
would otherwise oblige a harmonization-map row.
|
||||
|
||||
Verify the break corpus-wide and census-to-census before writing the rows.
|
||||
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
|
||||
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
|
||||
signpost when it lands.
|
||||
|
||||
**Superseded:** an earlier draft of this spec proposed adding
|
||||
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
|
||||
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
|
||||
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
|
||||
are the documented § 1 decision, and the recipe path reaches them by design.
|
||||
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
|
||||
`param_contract`, and the `plumber.R` route signature. Lands after this.
|
||||
|
||||
**Two contract facts the API must carry forward**, both settled during
|
||||
implementation and easy to get wrong from the outside:
|
||||
|
||||
- `provenance$balance_caveats$coverage_window` is **corpus-scoped, not
|
||||
result-scoped**. It reports the observed year extent of *every* balance
|
||||
subtype in the corpus, not only the subtypes a given query returned — so a
|
||||
`category = "Fund Balances"` query still returns all five windows. That is
|
||||
deliberate: the windows describe what the corpus holds, which is what a
|
||||
consumer needs in order to know what it did *not* ask for. The sibling
|
||||
field `truncated` is the result-scoped one. Documented in
|
||||
`inst/schemas/provenance-v1.json` and mutation-guarded against silent
|
||||
inversion.
|
||||
- `balance_caveats` appears **only** on `cog_balances()` results. It is
|
||||
absent from `cog_spending()`/`cog_revenue()` provenance, and the schema
|
||||
says so — an API layer that assumes it is universal will read `NULL`.
|
||||
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7
|
||||
SQL views (there are 21), 181 tests (716), a two-year fixture (four years),
|
||||
and a "never inline SQL" rule the verb layer does not follow.
|
||||
@@ -0,0 +1,296 @@
|
||||
# `uscogdata` 0.3.0 — public release
|
||||
|
||||
**Date:** 2026-08-08 · **Status:** design, awaiting approval
|
||||
**Scope:** release-readiness, README, NEWS. Distribution mechanics recorded here as
|
||||
decided, sequenced after the package is clean.
|
||||
|
||||
`uscogdata` is feature-complete and the corpus it reads has been public on
|
||||
HuggingFace since 2026-08-07 (294 downloads as of this writing). The API built on
|
||||
it is live. What does not exist is a public *package*: the repo is private, there
|
||||
is no install path, and — measured, not assumed — **a stranger who installed it
|
||||
today could not read the corpus at all.**
|
||||
|
||||
This spec covers making that untrue.
|
||||
|
||||
## Decisions locked
|
||||
|
||||
| Decision | Choice |
|
||||
|---|---|
|
||||
| Canonical source | `gitea.civilytics.org/Civilytics/uscogdata`, flipped public |
|
||||
| Public mirror | `github.com/civilytics/uscogdata` — issues, PRs, multi-OS check, CDN |
|
||||
| Mirror mechanism | Gitea Actions non-force `git push` (not a push mirror) |
|
||||
| Binaries | `civilytics.r-universe.dev`, registry pinned to a release tag |
|
||||
| Author of record | Jared E. Knowles `<jared@civilytics.com>`, ORCID `0000-0003-0005-9478` |
|
||||
| Copyright | Civilytics Consulting LLC (`cph`, `fnd`) |
|
||||
| License | MIT (package) · CC-BY-4.0 (corpus) |
|
||||
| Corrections intake | Deferred — see *Out of scope* |
|
||||
| Other packages | Parked until this one walks the path end to end |
|
||||
|
||||
## P0 — the corpus is unreachable
|
||||
|
||||
Two independent faults, either of which alone is fatal.
|
||||
|
||||
**No corpus URL exists.** `R/config.R` defaults to the literal
|
||||
`REPLACE_WITH_SHARE_TOKEN` sentinel, and no file in the repo supplies a working
|
||||
one. A new user calling any verb gets `uscogdata_url_not_configured` with no path
|
||||
to resolution.
|
||||
|
||||
**Remote reads are broken regardless.** Every partitioned view globs:
|
||||
|
||||
```sql
|
||||
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true)
|
||||
```
|
||||
|
||||
DuckDB 1.5.5 refuses globs over generic HTTP. Its suggested
|
||||
`allow_asterisks_in_http_paths` escape hatch does not help — it forwards the
|
||||
literal `**/*` as a filename and 404s, because plain HTTP exposes no directory
|
||||
listing to expand against.
|
||||
|
||||
The package therefore works only against a **local path**. That is how the API
|
||||
runs it (`CORPUS_HOST_PATH` is a host mount on maxwell) and how the tests run
|
||||
(bundled fixture), which is why the fault went unnoticed. The README's headline
|
||||
claim — *"Reads the published corpus directly from Nextcloud via DuckDB httpfs —
|
||||
no local bulk downloads required"* — is currently false.
|
||||
|
||||
### Fix: enumerate from the manifest, do not glob
|
||||
|
||||
`manifest.json` already lists every partition under `files.long_partitions[]`
|
||||
with `path`, `year`, `sha256`, `row_count` and `size_bytes` — 56 of them.
|
||||
Substituting an explicit file list for the glob was measured against the
|
||||
published corpus on 2026-08-08:
|
||||
|
||||
| Path | Result |
|
||||
|---|---|
|
||||
| `https://…/data/long/**/*.parquet` (default) | error — globs unsupported over HTTP |
|
||||
| same, `allow_asterisks_in_http_paths = true` | error — literal `**/*` 404s |
|
||||
| `hf://datasets/civilytics/us-cog-finance/…` glob | 46,148,034 rows |
|
||||
| **explicit list over plain https** | **46,148,034 rows** |
|
||||
|
||||
`hive_partitioning = true` still recovers `year` from the paths under
|
||||
enumeration, so no downstream view or verb changes.
|
||||
|
||||
Enumeration is preferred over `hf://` deliberately. It is **host-agnostic** —
|
||||
Nextcloud, HuggingFace, or any static server take the same code path — where
|
||||
`hf://` would tie the default to one vendor's protocol and still need
|
||||
special-casing, since manifest fetching goes through `httr2`, which cannot speak
|
||||
`hf://`. Enumeration also *removes* a dependency (globbing) rather than adding
|
||||
one, and the manifest's per-file `sha256` becomes available for integrity
|
||||
checking later.
|
||||
|
||||
Views are registered from `inst/sql/` with `{url}` substitution in
|
||||
`R/views.R:.register_views()`. The list must be built once per session from the
|
||||
already-fetched manifest and substituted the same way, so the change is confined
|
||||
to view registration and does not touch verb code.
|
||||
|
||||
### Fix: ship a working default
|
||||
|
||||
`R/config.R`'s default becomes the public HuggingFace `resolve/main/` URL:
|
||||
CC-BY-4.0, no token to publish, CDN-backed, and it keeps maxwell's uplink out of
|
||||
the path — the same reasoning behind the GitHub mirror and r-universe.
|
||||
|
||||
This means `library(uscogdata)` followed by a verb works with **zero
|
||||
configuration**, which is what makes the package demonstrable in a README and
|
||||
later in a post. `USCOGDATA_URL` and `options(uscogdata.url=)` continue to
|
||||
override, so the Nextcloud copy and local mirrors are unaffected.
|
||||
|
||||
The `uscogdata_url_not_configured` error class stays — it still fires for an
|
||||
explicitly-set empty or placeholder URL — but ceases to be the default
|
||||
experience.
|
||||
|
||||
### Consequence: `cog_mirror()` is promoted
|
||||
|
||||
Measured cost of the remote default, from efron on a good connection:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Whole corpus | **190.6 MB**, 56 partitions, 46,148,034 rows, FY1967–FY2024 |
|
||||
| One government, one year | 1.5 s |
|
||||
| One government, all 56 years | 2.8 s |
|
||||
| Disk written | **0.00 MB** — range requests only; `external_file_cache` is in-memory |
|
||||
|
||||
Nothing persists locally beyond the shared `httpfs` extension in `~/.duckdb` (a
|
||||
few MB, once per machine, across all DuckDB use). Costs are RAM and per-query
|
||||
bandwidth, since nothing caches between sessions.
|
||||
|
||||
Those timings are raw scans. Real verbs additionally join crosswalks, resolve
|
||||
categories and assemble provenance, so end-to-end verb latency will be higher and
|
||||
**must be re-measured once the fix lands** — it cannot be measured today.
|
||||
|
||||
The corpus being only 190.6 MB makes `cog_mirror()` a first-class option rather
|
||||
than a developer footnote. The README presents **both paths**:
|
||||
|
||||
- **Remote (default, zero setup)** — trying it out, teaching, one-off questions.
|
||||
- **Mirrored (`cog_mirror()`, 190 MB once)** — repeated or heavy analysis,
|
||||
offline work, reproducibility, or preferring not to depend on HuggingFace.
|
||||
|
||||
The second is also the honest answer to the vendor-dependency question raised by
|
||||
defaulting to HuggingFace: **the escape hatch is one function call and 190 MB**,
|
||||
after which no analysis touches an external service. The README says so
|
||||
explicitly. That is the difference between a convenience default and lock-in.
|
||||
|
||||
## Release-readiness fixes
|
||||
|
||||
| # | Issue | Fix |
|
||||
|---|---|---|
|
||||
| 1 | `MaxCorpusSchema: 5` in DESCRIPTION; `.validate_schema()` accepts `4,5,6,7`; published corpus is **7** | `MaxCorpusSchema: 7` |
|
||||
| 2 | `^vignettes$` in `.Rbuildignore` — both vignettes absent from the installed package, while README tells users to run `vignette("total-spending")` | Remove `^vignettes$`, `^doc$`, `^Meta$`. Both vignettes build offline (`total-spending` reads the bundled fixture; `population-denominators` is `eval = FALSE`) |
|
||||
| 3 | `_pkgdown.yml` reference index covers 6 of 14 exports — pkgdown errors on missing topics | Add `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`; set `url:` |
|
||||
| 4 | No `URL:` / `BugReports:` in DESCRIPTION | Add both, pointing at the GitHub mirror |
|
||||
| 5 | No `LICENSE.md`; `LICENSE` holder reads `Civilytics` | `usethis::use_mit_license("Civilytics Consulting LLC")` |
|
||||
| 6 | README instructs stripping the fixture at release | Delete that section — see below |
|
||||
| 7 | `Authors@R` is an org with no human | Jared E. Knowles `aut`/`cre` + ORCID; Civilytics Consulting LLC `cph`/`fnd` |
|
||||
|
||||
**On #6.** The advice to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`
|
||||
is CRAN-sized thinking (5 MB limit) and this package is not going to CRAN.
|
||||
Stripping the 15 MB fixture would break `total-spending.Rmd`, which reads from
|
||||
it, and would leave r-universe and GitHub Actions unable to run the 28 test files
|
||||
without a corpus credential. **The fixture is what lets `R CMD check` pass
|
||||
anywhere with zero secrets** — precisely what public CI needs. It ships.
|
||||
|
||||
## README
|
||||
|
||||
The current README addresses someone standing inside the repo tree: status reads
|
||||
"Under active development (Phase 2 of the cog_pipeline project)", it points at
|
||||
`../cog_pipeline/docs/reader-specification.md`, the install line is commented
|
||||
out, and developer, testing and release sections sit above anything a user needs.
|
||||
|
||||
Restructured around a stranger, in this order:
|
||||
|
||||
1. **What this is** — one paragraph, and what the corpus covers (types 0–3,
|
||||
FY1967–FY2024, 46M rows, 190.6 MB).
|
||||
2. **Install** — r-universe first (binaries), git second.
|
||||
3. **Quickstart that actually runs** — resolve a government, get its history,
|
||||
print provenance. No configuration step.
|
||||
4. **Two ways to read the corpus** — remote default vs `cog_mirror()`, with the
|
||||
measured numbers and the independence note.
|
||||
5. **Amounts are in full US dollars** — kept near the top. This is the errata
|
||||
most likely to produce a wrong answer that looks plausible.
|
||||
6. **Concepts** — primary/direct/total spending, general/total revenue,
|
||||
coverage. Condensed, linking to the vignettes for the full treatment.
|
||||
7. **How to cite** — `citation("uscogdata")`, corpus CC-BY-4.0 attribution.
|
||||
8. **Contributing** — canonical-on-Gitea PR flow.
|
||||
|
||||
Developer notes, testing instructions and release procedure move to
|
||||
`CONTRIBUTING.md`. Every path reference to a sibling repo is removed or replaced
|
||||
with a URL that resolves for someone who has only this repo.
|
||||
|
||||
## NEWS.md
|
||||
|
||||
`NEWS.md` currently holds two sections. `0.2.0` is a legitimate changelog — the
|
||||
`"All Categories"` reserved value, the coverage-signposting fix, the
|
||||
`n_units_reporting` documentation — and it stays. Beneath it,
|
||||
`0.1.0 (development)` is a pre-release churn log: changes described relative to
|
||||
states no user has ever seen ("Breaking: corpus schema_version 4", "the package
|
||||
now requires…"), spanning the package's entire pre-release development. To a
|
||||
newcomer deciding whether to depend on this, that section reads as instability.
|
||||
|
||||
**A new `0.3.0` section is added at the top, framed as the first public
|
||||
release**: what the package does, what the corpus covers, and the caveats that
|
||||
are genuinely load-bearing. **`0.2.0` is kept verbatim.** **`0.1.0 (development)`
|
||||
is dropped** — that history stays in git, where it belongs.
|
||||
|
||||
The version is `0.3.0` rather than `0.2.0` because this release changes
|
||||
user-visible behaviour: remote corpus reads go from broken to working, and the
|
||||
default URL from a dead placeholder to a live corpus. It is also not `1.0.0` —
|
||||
the corpus still excludes government types 4 and 5 pending validation, so a
|
||||
stability promise would overclaim. No git tag exists for any prior version;
|
||||
`chore: release 0.2.0` bumped `DESCRIPTION` and `NEWS` only.
|
||||
|
||||
The substantive content is migrated, not deleted. These are hard-won and belong
|
||||
in documentation rather than buried in a changelog:
|
||||
|
||||
| Content | Destination |
|
||||
|---|---|
|
||||
| Coverage disclosure on multi-government aggregates (census vs sample years) | README concepts + `cog_geographic_rollup()` docs |
|
||||
| `complete = TRUE` three-way absence semantics (`reported` / `census_zero` / `not_reported`) | `cog_spending()` / `cog_revenue()` docs |
|
||||
| Series-break and corpus-break surfacing | README + `cog_explain()` docs |
|
||||
| $1,000s → full dollars conversion | README, already prominent |
|
||||
| Per-year F-33 population denominators | `population-denominators` vignette, already there |
|
||||
|
||||
This also makes NEWS reusable as raw material for the release announcement,
|
||||
which is the stated downstream purpose.
|
||||
|
||||
## Distribution mechanics
|
||||
|
||||
Recorded as decided; executed after the package is clean and checks are green.
|
||||
|
||||
**Sequence matters.** r-universe publishes check results the moment a package is
|
||||
registered. Registering before the fixes above land means a red badge on day one,
|
||||
which is a worse first impression than a week's delay.
|
||||
|
||||
1. `gitleaks` over full history. A coarse grep found nothing across 140 commits
|
||||
and the default corpus URL is still the placeholder sentinel, but a proper
|
||||
scan is the gate on an irreversible action.
|
||||
2. Flip the Gitea repo public. Disable Gitea issues on it, so there is exactly
|
||||
one inbox.
|
||||
3. Create `github.com/civilytics/uscogdata`. Add `.github/workflows/` for the
|
||||
Windows/macOS/Linux `R CMD check` matrix — the platforms the Gitea runner
|
||||
cannot provide, and which this package has never been tested on despite
|
||||
depending on duckdb and httr2. Gitea reads `.gitea/workflows`, GitHub reads
|
||||
`.github/workflows`; both live in one tree without colliding.
|
||||
4. Gitea Actions workflow pushing to GitHub **without `--force`**, so divergence
|
||||
fails loudly in CI rather than silently overwriting.
|
||||
5. Add `jared@civilytics.com` as a verified secondary email on the GitHub
|
||||
account — r-universe links maintainer identity by matching DESCRIPTION's email
|
||||
against registered GitHub emails, and the association only takes effect on the
|
||||
next build.
|
||||
6. Tag `v0.3.0`. Create `github.com/civilytics/civilytics.r-universe.dev` with a
|
||||
`packages.json` pinned to the tag, pointing at the GitHub mirror rather than
|
||||
Gitea so clone traffic stays off maxwell. Install the r-universe app.
|
||||
|
||||
### PR flow
|
||||
|
||||
Never press Merge on GitHub. A merge there is overwritten by the next sync, the
|
||||
PR still displays "Merged", and nothing says otherwise.
|
||||
|
||||
```sh
|
||||
git remote add github https://github.com/civilytics/uscogdata.git
|
||||
git config --add remote.github.fetch '+refs/pull/*/head:refs/remotes/github/pr/*'
|
||||
git fetch github
|
||||
git switch -c pr-42 github/pr/42 # test
|
||||
git switch main && git merge --no-ff pr-42
|
||||
git push origin main # Gitea -> mirror -> GitHub
|
||||
```
|
||||
|
||||
GitHub auto-closes a PR as merged once its head commit becomes an ancestor of the
|
||||
base branch, so `--no-ff` — which preserves the contributor's SHAs — makes the PR
|
||||
close itself when the mirror pushes. **For external PRs, merge; do not squash or
|
||||
rebase.** Squashing rewrites the SHAs, the auto-close never fires, and closing by
|
||||
hand reads to a first-time contributor as rejection.
|
||||
|
||||
`CONTRIBUTING.md` states this, and a GitHub Action comments it on incoming PRs.
|
||||
No CLA; no DCO.
|
||||
|
||||
## Verification
|
||||
|
||||
The release is not done until all of these pass:
|
||||
|
||||
1. `R CMD check --as-cran` clean on Linux, and on Windows and macOS via the
|
||||
GitHub matrix. This package has never been checked on the latter two.
|
||||
2. Full test suite (28 files) green against the **bundled fixture**, offline,
|
||||
with no credentials — the property public CI depends on.
|
||||
3. Full test suite green against the **live corpus**, which additionally
|
||||
exercises the enumeration fix that the fixture's local path cannot.
|
||||
4. `pkgdown::build_site()` completes.
|
||||
5. Both vignettes present in the built tarball and
|
||||
`vignette("total-spending", package = "uscogdata")` resolves from an
|
||||
installed copy.
|
||||
6. **Cold-start check on a machine that has never seen this package:** install
|
||||
from r-universe, `library(uscogdata)`, run the README quickstart verbatim with
|
||||
no environment variables set. This is the only test that catches the P0 class
|
||||
of fault, and its absence is why the fault survived.
|
||||
7. End-to-end verb latency re-measured against the live corpus and the README's
|
||||
numbers updated if they moved.
|
||||
|
||||
## Out of scope
|
||||
|
||||
- **Corrections intake.** Deferred by decision. Consequence: the release cannot
|
||||
invite data-error reports or make the "traceable and correctable" claim that
|
||||
most distinguishes this corpus from Census's own files. `BugReports:` points at
|
||||
package issues only. A verified correction should eventually terminate as a
|
||||
`lineage_event` or `series_break` row so it propagates through provenance to
|
||||
every consumer — that design is unstarted.
|
||||
- **Announcement posts.** Deferred. The API announcement is gated on corrections
|
||||
landing and merits a Civic Pulse edition.
|
||||
- **The rest of the R package backlog.** Parked until this one completes the path.
|
||||
- **`cog_pipeline` publication.** Stays private.
|
||||
@@ -6,6 +6,33 @@ fixture_corpus_path <- function() {
|
||||
if (nzchar(p)) paste0(p, "/") else ""
|
||||
}
|
||||
|
||||
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
|
||||
# or "" when it isn't there.
|
||||
#
|
||||
# Tests that assert on documentation content have to read the sources, and the
|
||||
# sources only exist when the suite runs from a checkout. Under R CMD check the
|
||||
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
|
||||
# shipped and `../../README.md` does not resolve -- so those tests must skip
|
||||
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
|
||||
# rcmdcheck, so the assertions are still enforced on every push; this only
|
||||
# stops them from failing a context that structurally cannot satisfy them.
|
||||
source_tree_path <- function(...) {
|
||||
p <- testthat::test_path("..", "..", ...)
|
||||
if (file.exists(p)) p else ""
|
||||
}
|
||||
|
||||
# Skip unless every named source file is present (see source_tree_path()).
|
||||
skip_if_no_source_tree <- function(...) {
|
||||
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
|
||||
character(1))
|
||||
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
|
||||
testthat::skip_if(
|
||||
any(missing),
|
||||
"package source tree not available (running against the installed package)"
|
||||
)
|
||||
invisible(paths)
|
||||
}
|
||||
|
||||
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
||||
skip_if_no_corpus <- function() {
|
||||
p <- fixture_corpus_path()
|
||||
@@ -58,6 +85,42 @@ with_doctored_schema_version <- function(version, code) {
|
||||
force(code)
|
||||
}
|
||||
|
||||
# Copy the bundled fixture to a temp dir with representation.parquet and
|
||||
# code_set.parquet removed (and dropped from the manifest's metadata list),
|
||||
# then run `code` against it. Models a corpus published BEFORE sparsification:
|
||||
# schema_version is left alone deliberately, because it was never bumped for
|
||||
# that change -- the pre-sparsification fixture this package shipped until
|
||||
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
|
||||
# is therefore the only honest signal, and this helper is what proves the
|
||||
# package keys off it rather than off the version number.
|
||||
with_corpus_missing_representation <- function(code) {
|
||||
src <- fixture_corpus_path()
|
||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
||||
|
||||
dropped <- c("representation.parquet", "code_set.parquet")
|
||||
file.remove(file.path(tmp, "data", dropped))
|
||||
|
||||
manifest_path <- file.path(tmp, "manifest.json")
|
||||
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
|
||||
m$files$metadata <- Filter(
|
||||
function(f) !basename(f$path) %in% dropped, m$files$metadata
|
||||
)
|
||||
writeLines(
|
||||
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
|
||||
manifest_path
|
||||
)
|
||||
|
||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
||||
uscogdata:::cog_close()
|
||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
||||
on.exit({
|
||||
uscogdata:::cog_close()
|
||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
||||
}, add = TRUE)
|
||||
force(code)
|
||||
}
|
||||
|
||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
||||
# rewritten to drop every M/L (intergovernmental) row, then run `code`
|
||||
# against it with a clean session (mirrors with_fixture_corpus()/
|
||||
@@ -91,3 +154,36 @@ with_corpus_missing_ig_categories <- function(code) {
|
||||
}, add = TRUE)
|
||||
force(code)
|
||||
}
|
||||
|
||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
||||
# rewritten to DROP the balance_subtype column, then run `code` against it.
|
||||
# Models a corpus published before cog_pipeline #76/#77. schema_version is
|
||||
# left untouched deliberately: that change shipped without a version bump, so
|
||||
# column presence is the only honest signal -- this helper is what proves the
|
||||
# package keys off it. Mirrors with_corpus_missing_ig_categories().
|
||||
with_corpus_missing_balance_subtype <- function(code) {
|
||||
src <- fixture_corpus_path()
|
||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
||||
|
||||
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
|
||||
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
|
||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbExecute(write_con, sprintf(
|
||||
"COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s))
|
||||
TO %s (FORMAT PARQUET)",
|
||||
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
|
||||
))
|
||||
file.remove(cats_path)
|
||||
file.rename(filtered_path, cats_path)
|
||||
|
||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
||||
uscogdata:::cog_close()
|
||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
||||
on.exit({
|
||||
uscogdata:::cog_close()
|
||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
||||
}, add = TRUE)
|
||||
force(code)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
test_that('cog_geographic_rollup() accepts "All Categories" and agrees with per-category sums', {
|
||||
skip_if_no_corpus()
|
||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
||||
expect_gt(nrow(govs), 1L)
|
||||
ids <- list(city = utils::head(govs$canonical_govid, 25L))
|
||||
|
||||
by_cat <- cog_geographic_rollup(ids, category = NULL, years = 2019L)
|
||||
total <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L)
|
||||
|
||||
expect_setequal(unique(total$category), "All Categories")
|
||||
# one row per (govid, subtype) that appears in the per-category result
|
||||
key_by_cat <- unique(paste(by_cat$canonical_govid, by_cat$spend_subtype))
|
||||
key_total <- paste(total$canonical_govid, total$spend_subtype)
|
||||
expect_setequal(key_total, key_by_cat)
|
||||
|
||||
lhs <- tapply(by_cat$amt_nominal, paste(by_cat$canonical_govid, by_cat$spend_subtype), sum)
|
||||
rhs <- tapply(total$amt_nominal, key_total, sum)
|
||||
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
|
||||
})
|
||||
|
||||
test_that('"All Categories" survives per_capita and inflation adjustment through the rollup', {
|
||||
skip_if_no_corpus()
|
||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
||||
ids <- list(city = utils::head(govs$canonical_govid, 10L))
|
||||
r <- cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
|
||||
per_capita = TRUE, adjust_to_year = 2020L)
|
||||
expect_true(all(c("amt_per_capita_nominal", "amt_real", "amt_per_capita_real") %in% names(r)))
|
||||
expect_setequal(unique(r$category), "All Categories")
|
||||
expect_true(all(is.finite(r$amt_real)))
|
||||
})
|
||||
|
||||
test_that('cog_geographic_rollup() still refuses expenditure_concept = "total" with "All Categories"', {
|
||||
skip_if_no_corpus()
|
||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
||||
ids <- list(city = utils::head(govs$canonical_govid, 5L))
|
||||
expect_error(
|
||||
cog_geographic_rollup(ids, category = "All Categories", years = 2019L,
|
||||
expenditure_concept = "total")
|
||||
)
|
||||
})
|
||||
|
||||
test_that("n_units_reporting is category-conditional, not a response rate", {
|
||||
skip_if_no_corpus()
|
||||
govs <- cog_gov_search(name = NULL, state = "WI", type = 2L)
|
||||
ids <- list(city = govs$canonical_govid)
|
||||
|
||||
police <- cog_geographic_rollup(ids, category = "Police", years = 2012L)
|
||||
allcat <- cog_geographic_rollup(ids, category = "All Categories", years = 2012L)
|
||||
|
||||
cov_police <- cog_explain(police, format = "list")$coverage
|
||||
cov_all <- cog_explain(allcat, format = "list")$coverage
|
||||
|
||||
# Same year, same requested govids, same collection -- yet a single category
|
||||
# reports fewer units than the all-categories query. That gap is real zeros,
|
||||
# not non-response, which is exactly why the ratio is not a response rate.
|
||||
expect_lte(cov_police$n_units_reporting, cov_all$n_units_reporting)
|
||||
expect_identical(cov_police$n_units_expected, cov_all$n_units_expected)
|
||||
})
|
||||
@@ -0,0 +1,267 @@
|
||||
# Baseline at branch point: 843 PASS / 0 FAIL / 0 SKIP / 0 WARN (2026-08-05, origin/main 2fc9e75)
|
||||
|
||||
test_that(".build_verb_sql emits a literal category and no category filter in all-categories mode", {
|
||||
sql <- uscogdata:::.build_verb_sql(
|
||||
view = "spending_annotated",
|
||||
subtype_col = "spend_subtype",
|
||||
govid = "552025209777",
|
||||
years = 2019L,
|
||||
category = NULL,
|
||||
subtype_scope = c("operations", "capital"),
|
||||
all_categories = TRUE
|
||||
)
|
||||
|
||||
expect_match(sql, "'All Categories' AS category", fixed = TRUE)
|
||||
# no category filter of any kind
|
||||
expect_false(grepl("AND category IN", sql, fixed = TRUE))
|
||||
# category is not a grouping key
|
||||
expect_false(grepl("GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
|
||||
sql, fixed = TRUE))
|
||||
# the subtype allowlist still applies -- this is what makes the sum a concept
|
||||
expect_match(sql, "AND spend_subtype IN ('operations','capital')", fixed = TRUE)
|
||||
})
|
||||
|
||||
test_that(".build_verb_sql is unchanged when all_categories is FALSE", {
|
||||
args <- list(
|
||||
view = "spending_annotated", subtype_col = "spend_subtype",
|
||||
govid = "552025209777", years = 2019L, category = NULL,
|
||||
subtype_scope = c("operations", "capital")
|
||||
)
|
||||
old <- do.call(uscogdata:::.build_verb_sql, args)
|
||||
new <- do.call(uscogdata:::.build_verb_sql, c(args, list(all_categories = FALSE)))
|
||||
expect_identical(old, new)
|
||||
expect_match(new, "GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, spend_subtype, category",
|
||||
fixed = TRUE)
|
||||
})
|
||||
|
||||
test_that(".ALL_CATEGORIES is the exact reserved string", {
|
||||
expect_identical(uscogdata:::.ALL_CATEGORIES, "All Categories")
|
||||
})
|
||||
|
||||
test_that('cog_spending(category = "All Categories") sums to the per-category total', {
|
||||
gov <- "552025209777"
|
||||
by_cat <- cog_spending(gov, 2019L)
|
||||
total <- cog_spending(gov, 2019L, category = "All Categories")
|
||||
|
||||
expect_true(nrow(total) > 0L)
|
||||
expect_setequal(unique(total$category), "All Categories")
|
||||
# one row per subtype present in the by-category result
|
||||
expect_setequal(unique(total$spend_subtype), unique(by_cat$spend_subtype))
|
||||
expect_equal(nrow(total), length(unique(by_cat$spend_subtype)))
|
||||
|
||||
# the dollars agree, per subtype
|
||||
lhs <- tapply(by_cat$amt_nominal, by_cat$spend_subtype, sum)
|
||||
rhs <- tapply(total$amt_nominal, total$spend_subtype, sum)
|
||||
expect_equal(as.numeric(rhs[names(lhs)]), as.numeric(lhs), tolerance = 1e-8)
|
||||
})
|
||||
|
||||
test_that('"All Categories" respects expenditure_concept', {
|
||||
gov <- "552025209777"
|
||||
prim <- cog_spending(gov, 2019L, category = "All Categories",
|
||||
expenditure_concept = "primary")
|
||||
dir <- cog_spending(gov, 2019L, category = "All Categories",
|
||||
expenditure_concept = "direct")
|
||||
# direct = primary plus interest and insurance benefits, so it is never smaller
|
||||
expect_gte(sum(dir$amt_nominal), sum(prim$amt_nominal))
|
||||
})
|
||||
|
||||
test_that('"All Categories" works on revenue and respects revenue_concept', {
|
||||
gov <- "552025209777"
|
||||
gen <- cog_revenue(gov, 2019L, category = "All Categories",
|
||||
revenue_concept = "general")
|
||||
tot <- cog_revenue(gov, 2019L, category = "All Categories",
|
||||
revenue_concept = "total")
|
||||
expect_setequal(unique(gen$category), "All Categories")
|
||||
expect_gte(sum(tot$amt_nominal), sum(gen$amt_nominal))
|
||||
})
|
||||
|
||||
test_that('"All Categories" cannot be combined with another category', {
|
||||
expect_error(
|
||||
cog_spending("552025209777", 2019L, category = c("All Categories", "Police")),
|
||||
class = "uscogdata_all_categories_not_combinable"
|
||||
)
|
||||
})
|
||||
|
||||
test_that('"All Categories" is recorded in provenance', {
|
||||
r <- cog_spending("552025209777", 2019L, category = "All Categories")
|
||||
expect_identical(cog_explain(r, format = "list")$category, "All Categories")
|
||||
})
|
||||
|
||||
test_that('"All Categories" combines with subtype to give operating totals', {
|
||||
gov <- "552025209777"
|
||||
ops_by_cat <- cog_spending(gov, 2019L)
|
||||
ops_by_cat <- ops_by_cat[ops_by_cat$spend_subtype == "operations", ]
|
||||
ops_total <- cog_spending(gov, 2019L, category = "All Categories")
|
||||
ops_total <- ops_total[ops_total$spend_subtype == "operations", ]
|
||||
expect_equal(sum(ops_total$amt_nominal), sum(ops_by_cat$amt_nominal),
|
||||
tolerance = 1e-8)
|
||||
})
|
||||
|
||||
test_that('cog_categories() advertises "All Categories" for both flows', {
|
||||
all <- cog_categories()
|
||||
rows <- all[all$category == "All Categories", ]
|
||||
expect_setequal(rows$category_type, c("expenditure", "revenue"))
|
||||
expect_true(all(is.na(rows$subtype)))
|
||||
expect_true(all(is.na(rows$n_codes)))
|
||||
})
|
||||
|
||||
test_that('cog_categories(type=) still scopes, including the pseudo-category', {
|
||||
sp <- cog_categories(type = "spending")
|
||||
expect_setequal(unique(sp$category_type), "expenditure")
|
||||
expect_true("All Categories" %in% sp$category)
|
||||
|
||||
rev <- cog_categories(type = "revenue")
|
||||
expect_setequal(unique(rev$category_type), "revenue")
|
||||
expect_true("All Categories" %in% rev$category)
|
||||
|
||||
# balances have no concept vocabulary, so no pseudo-category
|
||||
bal <- cog_categories(type = "balance")
|
||||
expect_false("All Categories" %in% bal$category)
|
||||
})
|
||||
|
||||
test_that('cog_categories(pattern=) matches the pseudo-category', {
|
||||
hit <- cog_categories(pattern = "^All Categories$")
|
||||
expect_equal(nrow(hit), 2L)
|
||||
})
|
||||
|
||||
# --- final whole-branch review fixes ---------------------------------------
|
||||
|
||||
test_that('complete = TRUE is refused when combined with "All Categories"', {
|
||||
# .completion_grid_sql() would emit `AND c.category IN ('All Categories')`,
|
||||
# match zero crosswalk rows, and the early return in .complete_result()
|
||||
# would stamp completion$applied = TRUE, rows_filled = 0 -- reading as "the
|
||||
# grid was checked and nothing was missing" when nothing was actually
|
||||
# checked. Filling a summed row has no defined semantics, so the verb must
|
||||
# refuse the combination outright (finding 2).
|
||||
expect_error(
|
||||
cog_spending("552025209777", 2019L, category = "All Categories",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
expect_error(
|
||||
cog_revenue("552025209777", 2019L, category = "All Categories",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
})
|
||||
|
||||
test_that('cog_balances() rejects "All Categories" instead of silently returning zero rows', {
|
||||
# cog_balances() reuses .validate_verb_inputs() but did not pass
|
||||
# allow_all_categories = TRUE, so "All Categories" used to become
|
||||
# `AND category IN ('All Categories')` against balance_annotated -- 0
|
||||
# matching crosswalk rows, 0 rows back, no error (finding 3). Holdings are
|
||||
# a stock with no concept vocabulary to sum across, so the honest answer is
|
||||
# to refuse, the same way cog_spending()/cog_revenue() refuse other
|
||||
# nonsensical combinations.
|
||||
expect_error(
|
||||
cog_balances("552025209777", 2019L, category = "All Categories"),
|
||||
class = "uscogdata_all_categories_unsupported"
|
||||
)
|
||||
# An ordinary category still works -- this is not a blanket regression.
|
||||
r <- suppressMessages(
|
||||
cog_balances("552025209777", 2019L, category = "Fund Balances")
|
||||
)
|
||||
expect_gt(nrow(r), 0L)
|
||||
})
|
||||
|
||||
test_that('expenditure_concept_direct_suppressed is NA, not FALSE, when categories are collapsed', {
|
||||
# .detect_direct_suppressed() keys on
|
||||
# paste(year, canonical_govid, category, sep = "\r"). In all-categories
|
||||
# mode every row carries the literal "All Categories" value, so an IG-only
|
||||
# row's key collides with any ordinary Direct row for the same
|
||||
# (year, govid) -- has_direct reads TRUE whenever the government has ANY
|
||||
# direct spending at all, candidate is always empty, and the detector can
|
||||
# never fire. Before the fix this silently reported FALSE, an affirmative
|
||||
# claim the code did not actually compute (finding 1). NA is the honest
|
||||
# answer: cog_explain(x, format = "list") is required here, since without
|
||||
# format = "list" it returns the result tibble, not the provenance list.
|
||||
gov <- "552025209777"
|
||||
t <- cog_spending(gov, 2019L, category = "All Categories",
|
||||
expenditure_concept = "total")
|
||||
prov <- cog_explain(t, format = "list")
|
||||
expect_true(is.na(prov$expenditure_concept_direct_suppressed))
|
||||
expect_false(isTRUE(prov$expenditure_concept_direct_suppressed))
|
||||
expect_match(prov$expenditure_concept_note, "unavailable", fixed = TRUE)
|
||||
|
||||
# A per-category "total" query on the same government/year is unaffected --
|
||||
# the detector can still key correctly and reports a strict logical.
|
||||
t_by_cat <- cog_spending(gov, 2019L, expenditure_concept = "total")
|
||||
prov_by_cat <- cog_explain(t_by_cat, format = "list")
|
||||
expect_false(is.na(prov_by_cat$expenditure_concept_direct_suppressed))
|
||||
})
|
||||
|
||||
test_that('"All Categories" still signposts coverage gaps (finding 6, final whole-branch review)', {
|
||||
# .build_suggestions()'s candidate sub-select used to be keyed on
|
||||
# `category`, e.g. `WHERE category IN ('All Categories')`. Since
|
||||
# .ALL_CATEGORIES is never itself a row in summary_categories.category,
|
||||
# that sub-select always came back empty in all-categories mode, so
|
||||
# `candidates` was empty and .build_suggestions() short-circuited to
|
||||
# list() -- coverage signposting was structurally impossible for the one
|
||||
# mode whose whole selling point is "you cannot sum the wrong scope"
|
||||
# (uscogdata#9's entire point, silently defeated).
|
||||
#
|
||||
# AL state government, FY2011, category = "Corrections": this category has
|
||||
# no legacy leaf rows in FY2011 (aggregate-flagged E04/E05 family), so the
|
||||
# per-category query returns 0 rows and 3 recipe-hint suggestions fire
|
||||
# (empty_year path). All-categories mode does not have an empty year --
|
||||
# the government has other primary spending in FY2011 -- but the same
|
||||
# suppressed Corrections dollars are still excluded from the summed total,
|
||||
# so the fix (scoping the candidate sub-select by subtype_col/subtype_scope
|
||||
# instead of by category, symmetric with .build_verb_sql()) must still
|
||||
# surface them via the suppressed_component path.
|
||||
gov <- "010000226085"
|
||||
|
||||
by_cat <- suppressMessages(cog_spending(gov, 2011L, category = "Corrections"))
|
||||
sugg_by_cat <- cog_explain(by_cat, format = "list")$suggestions
|
||||
expect_gt(length(sugg_by_cat), 0L)
|
||||
|
||||
all_cat <- suppressMessages(cog_spending(gov, 2011L, category = "All Categories"))
|
||||
sugg_all_cat <- cog_explain(all_cat, format = "list")$suggestions
|
||||
expect_gt(length(sugg_all_cat), 0L)
|
||||
|
||||
# The same Corrections recipe that fired per-category must also fire in
|
||||
# all-categories mode -- not just some unrelated recipe.
|
||||
ids_by_cat <- vapply(sugg_by_cat, function(s) s$recipe_id %||% "", character(1))
|
||||
ids_all_cat <- vapply(sugg_all_cat, function(s) s$recipe_id %||% "", character(1))
|
||||
expect_true("corrections_combined" %in% ids_by_cat)
|
||||
expect_true("corrections_combined" %in% ids_all_cat)
|
||||
|
||||
# In all-categories mode the government DOES have other primary spending
|
||||
# in FY2011 (the year itself is not a gap), so the suggestion can only have
|
||||
# fired via the suppressed_component path, not empty_year.
|
||||
corr_all <- sugg_all_cat[[which(ids_all_cat == "corrections_combined")]]
|
||||
expect_identical(corr_all$trigger, "suppressed_component")
|
||||
expect_gt(corr_all$suppressed_amount, 0)
|
||||
})
|
||||
|
||||
test_that('"All Categories" candidate scoping is symmetric with .build_verb_sql() -- subtype, not category', {
|
||||
# Direct assertion on the mechanism itself (finding 6): in all-categories
|
||||
# mode .build_suggestions() must scope its candidate recipe sub-select by
|
||||
# subtype_col/subtype_scope, not by the literal "All Categories" value.
|
||||
# Passing all_categories = FALSE with the identical category value proves
|
||||
# the branch -- not merely the subtype_col/subtype_scope arguments' mere
|
||||
# presence -- is what changes the query.
|
||||
con <- uscogdata:::.ensure_session()
|
||||
|
||||
none <- uscogdata:::.build_suggestions(
|
||||
con, govid = "010000226085", years = 2011L,
|
||||
category = "All Categories", result = NULL, basis = "harmonized",
|
||||
flow_prefixes = c("E", "F", "G"),
|
||||
long_view = "spending_long_harmonized",
|
||||
all_categories = FALSE,
|
||||
subtype_col = "spend_subtype",
|
||||
subtype_scope = c("operations", "capital", "assistance")
|
||||
)
|
||||
expect_length(none, 0L)
|
||||
|
||||
scoped <- uscogdata:::.build_suggestions(
|
||||
con, govid = "010000226085", years = 2011L,
|
||||
category = "All Categories", result = NULL, basis = "harmonized",
|
||||
flow_prefixes = c("E", "F", "G"),
|
||||
long_view = "spending_long_harmonized",
|
||||
all_categories = TRUE,
|
||||
subtype_col = "spend_subtype",
|
||||
subtype_scope = c("operations", "capital", "assistance")
|
||||
)
|
||||
expect_gt(length(scoped), 0L)
|
||||
})
|
||||
@@ -17,7 +17,16 @@
|
||||
# cog-api's llms.txt, which is silent on units).
|
||||
|
||||
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
||||
testthat::skip("Blocked on uscogdata#15 (finding F-004)")
|
||||
|
||||
# README and vignettes ship only in the source tree, not in the installed
|
||||
# package, so these assertions cannot run under R CMD check -- CI's earlier
|
||||
# testthat::test_local() step is what enforces them. See
|
||||
# skip_if_no_source_tree() in helper-fixture.R.
|
||||
docs <- skip_if_no_source_tree(
|
||||
"README.md",
|
||||
c("vignettes", "total-spending.Rmd"),
|
||||
c("vignettes", "population-denominators.Rmd")
|
||||
)
|
||||
|
||||
says_units <- function(path) {
|
||||
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
||||
@@ -25,10 +34,7 @@ test_that("returned amounts are documented as full US dollars where readers meet
|
||||
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
||||
}
|
||||
|
||||
expect_true(says_units(testthat::test_path("..", "..", "README.md")))
|
||||
expect_true(says_units(testthat::test_path("..", "..", "vignettes", "total-spending.Rmd")))
|
||||
expect_true(says_units(testthat::test_path("..", "..", "vignettes",
|
||||
"population-denominators.Rmd")))
|
||||
for (path in docs) expect_true(says_units(path))
|
||||
|
||||
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
||||
# The expected raw amount is read straight from the corpus's parquet
|
||||
|
||||
@@ -0,0 +1,539 @@
|
||||
test_that("balance views register and carry only balance codes", {
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
|
||||
views <- DBI::dbGetQuery(con,
|
||||
"SELECT table_name FROM information_schema.tables
|
||||
WHERE table_schema = 'main' AND table_type = 'VIEW'"
|
||||
)$table_name
|
||||
expect_true(all(c("balance_long", "balance_annotated") %in% views))
|
||||
|
||||
# Every item_code in balance_long is a category_type = 'balance' member.
|
||||
leak <- DBI::dbGetQuery(con,
|
||||
"SELECT COUNT(*) AS n FROM balance_long
|
||||
WHERE item_code NOT IN (
|
||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance')"
|
||||
)$n
|
||||
expect_identical(as.integer(leak), 0L)
|
||||
|
||||
# balance_annotated exposes the subtype column the verb groups on.
|
||||
cols <- DBI::dbGetQuery(con,
|
||||
"SELECT column_name FROM information_schema.columns
|
||||
WHERE table_name = 'balance_annotated'"
|
||||
)$column_name
|
||||
expect_true(all(c("category", "category_type", "balance_subtype") %in% cols))
|
||||
})
|
||||
|
||||
test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", {
|
||||
# Every category_type = 'balance' item_code in the bundled fixture has
|
||||
# is_aggregate = FALSE for every row of every year -- there is no real row
|
||||
# that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An
|
||||
# assertion against the live fixture (`WHERE is_aggregate` returns 0) is
|
||||
# therefore vacuous: it passes identically whether or not the view's
|
||||
# predicate is present. As with the 22-/23- and 24-/25- tests above, this
|
||||
# reads the real inst/sql/26-balance_long.sql text off disk and executes it
|
||||
# -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against
|
||||
# a synthetic hive-partitioned parquet tree that DOES contain an aggregate
|
||||
# row under a real balance item_code (W01), so a regression that drops the
|
||||
# predicate changes which rows survive.
|
||||
skip_if_no_corpus()
|
||||
|
||||
tmp <- withr::local_tempdir()
|
||||
part_dir <- file.path(tmp, "data", "long", "year=2004")
|
||||
dir.create(part_dir, recursive = TRUE)
|
||||
part_path <- file.path(part_dir, "part-0.parquet")
|
||||
|
||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbExecute(write_con, sprintf("
|
||||
COPY (
|
||||
SELECT * FROM (VALUES
|
||||
('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives
|
||||
('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate`
|
||||
) AS t(canonical_govid, item_code, amt, is_aggregate)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(part_path)))
|
||||
|
||||
DBI::dbExecute(write_con, sprintf("
|
||||
COPY (
|
||||
SELECT * FROM (VALUES
|
||||
('W01', 'Fund Balances', 'balance', NULL, NULL, 'general')
|
||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
||||
|
||||
sql_dir <- system.file("sql", package = "uscogdata")
|
||||
.read_view_sql <- function(filename) {
|
||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
||||
uscogdata:::.render_view_sql(txt, paste0(tmp, "/"))
|
||||
}
|
||||
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("26-balance_long.sql"))
|
||||
|
||||
rows <- DBI::dbGetQuery(con,
|
||||
"SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid"
|
||||
)
|
||||
expect_equal(nrow(rows), 1L)
|
||||
expect_equal(rows$canonical_govid, "bal-A")
|
||||
expect_equal(rows$amt, 100)
|
||||
})
|
||||
|
||||
test_that("balance views are skipped on a corpus without balance_subtype", {
|
||||
skip_if_no_corpus()
|
||||
with_corpus_missing_balance_subtype({
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
views <- DBI::dbGetQuery(con,
|
||||
"SELECT table_name FROM information_schema.tables
|
||||
WHERE table_schema = 'main' AND table_type = 'VIEW'"
|
||||
)$table_name
|
||||
# Registration must SKIP them, not error -- an older corpus stays usable.
|
||||
expect_false(any(c("balance_long", "balance_annotated") %in% views))
|
||||
expect_true("revenue_long" %in% views)
|
||||
|
||||
# ...and calling the verb on such a corpus must hit
|
||||
# .require_balance_support()'s curated abort (spec § Testing: "Gating"),
|
||||
# not a DuckDB binder error naming a view that was never registered.
|
||||
# Asserted on the CLASS: removing the guard still errors, so a bare
|
||||
# expect_error() would pass on the regression.
|
||||
expect_error(
|
||||
cog_balances("550000227544", 2019),
|
||||
class = "uscogdata_no_balance_support"
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("cog_balances returns holdings for a government that has them", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", 2019)
|
||||
expect_s3_class(r, "tbl_df")
|
||||
expect_true(nrow(r) > 0L)
|
||||
expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype",
|
||||
"category", "amt_nominal") %in% names(r)))
|
||||
expect_identical(sort(unique(r$category)),
|
||||
c("Fund Balances", "Insurance Trust Balances"))
|
||||
expect_false(is.null(attr(r, "provenance")))
|
||||
expect_identical(attr(r, "provenance")$verb, "cog_balances")
|
||||
})
|
||||
})
|
||||
|
||||
test_that('category = "Fund Balances" is exactly the general family', {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
||||
expect_identical(unique(r$balance_subtype), "general")
|
||||
codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
|
||||
expect_identical(codes, c("W01", "W31", "W61"))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("no flow code can reach cog_balances", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
|
||||
got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
|
||||
|
||||
# The expected set is read from the RAW corpus, never from the verb --
|
||||
# verifying an absence through the filter that creates it proves nothing.
|
||||
# A fresh, direct DuckDB connection against the raw parquet files (never
|
||||
# cog_open()'s session, never balance_long/balance_annotated) reads
|
||||
# parquet natively -- no arrow dependency needed (see CLAUDE.md).
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
||||
balance_codes <- DBI::dbGetQuery(con2, sprintf(
|
||||
"SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'",
|
||||
uscogdata:::.sql_lit_chr(cats_path)
|
||||
))$item_code
|
||||
|
||||
expect_true(length(got) > 0L)
|
||||
expect_true(all(got %in% balance_codes))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("every balance_subtype maps to exactly one category", {
|
||||
skip_if_no_corpus()
|
||||
# Dropping the `subtype` argument is only safe while this tree holds. If the
|
||||
# pipeline ever gives a balance subtype a second category, `category` becomes
|
||||
# a lossy filter -- fail HERE rather than in a user's analysis. Read via a
|
||||
# fresh direct DuckDB connection against the raw parquet file, not through
|
||||
# any registered view.
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
||||
b <- DBI::dbGetQuery(con2, sprintf(
|
||||
"SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'",
|
||||
uscogdata:::.sql_lit_chr(cats_path)
|
||||
))
|
||||
per_subtype <- tapply(b$category, b$balance_subtype,
|
||||
function(x) length(unique(x)))
|
||||
expect_true(all(per_subtype == 1L))
|
||||
})
|
||||
|
||||
test_that("cog_balances records found + missing govids in provenance", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
suppressMessages(
|
||||
r <- cog_balances(c("550000227544", "XXXINVALID"), 2019)
|
||||
)
|
||||
prov <- attr(r, "provenance")
|
||||
expect_equal(sort(prov$scope$govids_found), "550000227544")
|
||||
expect_equal(sort(prov$scope$govids_missing), "XXXINVALID")
|
||||
})
|
||||
})
|
||||
|
||||
test_that("per_capita divides holdings by population", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
plain <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
||||
pc <- cog_balances("550000227544", 2019, category = "Fund Balances",
|
||||
per_capita = TRUE)
|
||||
expect_true("amt_per_capita_nominal" %in% names(pc))
|
||||
expect_true("pop_source" %in% names(pc))
|
||||
expect_identical(pc$amt_nominal, plain$amt_nominal)
|
||||
|
||||
# Assert against the denominator read from the corpus, NOT against a
|
||||
# quantity derived from amt_per_capita_nominal itself -- dividing the
|
||||
# column back out would be tautological and would pass on any value.
|
||||
pop <- DBI::dbGetQuery(cog_open(), sprintf(
|
||||
"SELECT population FROM gov_population_yearly
|
||||
WHERE canonical_govid = %s AND year = 2019",
|
||||
uscogdata:::.sql_lit_chr("550000227544")
|
||||
))$population
|
||||
expect_length(pop, 1L)
|
||||
expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop,
|
||||
tolerance = 1e-8)
|
||||
|
||||
prov <- attr(pc, "provenance")
|
||||
expect_true(prov$transformations$per_capita$applied)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("adjust_to_year adds real dollars", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
||||
adjust_to_year = 2020)
|
||||
expect_true("amt_real" %in% names(r))
|
||||
# 2012 dollars inflated to 2020 must exceed nominal.
|
||||
expect_true(all(r$amt_real > r$amt_nominal))
|
||||
prov <- attr(r, "provenance")
|
||||
expect_true(prov$transformations$inflation$applied)
|
||||
expect_identical(prov$transformations$inflation$base_year, 2020L)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("per_capita and adjust_to_year compose", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
||||
per_capita = TRUE, adjust_to_year = 2020)
|
||||
expect_true("amt_per_capita_real" %in% names(r))
|
||||
# The per-capita column must be deflated by the SAME factor as the level
|
||||
# column -- this is what the ordering at R/balances.R:101-103 guarantees.
|
||||
# .attach_real_dollars() silently no-ops on the per-capita leg when
|
||||
# amt_per_capita_nominal does not exist yet (R/spending.R:664), so
|
||||
# reversing those two calls drops this column with no error at all.
|
||||
expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal,
|
||||
r$amt_real / r$amt_nominal, tolerance = 1e-8)
|
||||
|
||||
# And the documented condition is a conjunction: adjust_to_year ALONE
|
||||
# must not produce amt_per_capita_real (pins the @return wording).
|
||||
r2 <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
||||
adjust_to_year = 2020)
|
||||
expect_true("amt_real" %in% names(r2))
|
||||
expect_false("amt_per_capita_real" %in% names(r2))
|
||||
})
|
||||
})
|
||||
|
||||
# --- input validation ------------------------------------------------------
|
||||
|
||||
test_that("cog_balances validates its inputs like the money verbs", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
G <- "550000227544"
|
||||
# Pinned to the message, not bare expect_error(): every one of these
|
||||
# already produces *some* error or *some* quiet wrong answer today --
|
||||
# years = integer(0) leaks `Parser Error ... AND year IN ()` with the
|
||||
# generated SQL, recipe = c("a","b") throws "the condition has length > 1",
|
||||
# and the govid/category cases return 0 rows with no error at all.
|
||||
expect_error(cog_balances(G, integer(0)), "non-empty integer vector")
|
||||
expect_error(cog_balances(character(0), 2019), "non-empty character vector")
|
||||
expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL")
|
||||
expect_error(cog_balances(G, 2019, recipe = c("a", "b")),
|
||||
"length-1 character string")
|
||||
})
|
||||
})
|
||||
|
||||
test_that("validation runs after govid coercion, so a data frame still works", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# .validate_verb_inputs() asserts is.character(govid); it must therefore
|
||||
# run AFTER .coerce_govid_input(), never before, or the documented
|
||||
# data-frame input (cog_gov_search() output) would abort.
|
||||
df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE)
|
||||
r <- suppressMessages(cog_balances(df, 2019))
|
||||
expect_true(nrow(r) > 0L)
|
||||
expect_identical(unique(r$canonical_govid), "550000227544")
|
||||
})
|
||||
})
|
||||
|
||||
test_that("recipe and category are mutually exclusive", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
expect_error(
|
||||
cog_balances("550000227544", c(2011, 2012),
|
||||
category = "Fund Balances",
|
||||
recipe = "cash_securities_z77_wide"),
|
||||
class = "uscogdata_recipe_category_conflict"
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
# --- recipe = : the wide-era holdings bridge -------------------------------
|
||||
|
||||
test_that("recipe bridges the wide era into the modern one", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", c(2011, 2012),
|
||||
recipe = "cash_securities_z77_wide")
|
||||
# .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait,
|
||||
# not specific to this recipe -- see the money-verb recipe tests, which
|
||||
# only ever assert on it with expect_equal), so compare numerically rather
|
||||
# than with expect_identical()'s type-strict comparison.
|
||||
expect_equal(sort(r$year), c(2011, 2012))
|
||||
|
||||
# The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE
|
||||
# and therefore invisible to balance_long. If the recipe path ever starts
|
||||
# filtering aggregates, a 45-year series silently truncates to five --
|
||||
# this is the regression guard for phase_r_harmonization_review.md § 0.2.
|
||||
codes <- attr(r, "provenance")$codes_summed$observed
|
||||
expect_true("X40" %in% codes)
|
||||
expect_true("Z77" %in% codes)
|
||||
expect_true(all(r$amt_nominal > 0))
|
||||
|
||||
prov <- attr(r, "provenance")
|
||||
expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide")
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# 2002 is in the year vector deliberately, and must stay -- do not
|
||||
# "simplify" this back to c(2011, 2012).
|
||||
#
|
||||
# .build_series_break_refs() (R/series_breaks.R, shared with every verb)
|
||||
# matches breaks with `break_year BETWEEN min(years) AND max(years)`, and
|
||||
# SB195's break_year is 2002. A c(2011, 2012) span never crosses the
|
||||
# FY2002 book -> market change -- that whole span sits after it, on one
|
||||
# consistent basis -- so NOT disclosing SB195 there is correct behaviour,
|
||||
# not a gap (same reasoning as the "a request that never crosses the
|
||||
# boundary is not affected by it" comment on .build_corpus_break_refs()).
|
||||
#
|
||||
# The property actually worth testing is: a recipe query that observes
|
||||
# X40 AND spans FY2002 discloses SB195. This fixture has no 2002
|
||||
# partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020
|
||||
# partitions exist), so including 2002 in `years` widens the
|
||||
# break-matching window without changing which rows the recipe join
|
||||
# returns -- verified empirically: r$year below is exactly {2011, 2012}
|
||||
# whether or not 2002 is in the request (see task-4-report.md).
|
||||
# Removing 2002 would silently turn this back into the non-crossing case
|
||||
# above and destroy the test's purpose.
|
||||
r <- cog_balances("550000227544", c(2002, 2011, 2012),
|
||||
recipe = "cash_securities_z77_wide")
|
||||
expect_equal(sort(r$year), c(2011, 2012))
|
||||
refs <- attr(r, "provenance")$series_break_refs
|
||||
# SB195 sits on fin_code X40; it can only fire where X40 is observed,
|
||||
# which is exactly the recipe path.
|
||||
expect_true("SB195" %in% refs)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the second holdings bridge works too", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011
|
||||
# and Z78 in 2012, so both legs are exercised.
|
||||
r <- cog_balances("550000227544", c(2011, 2012),
|
||||
recipe = "cash_securities_z78_wide")
|
||||
codes <- attr(r, "provenance")$codes_summed$observed
|
||||
expect_true(all(c("X41", "Z78") %in% codes))
|
||||
expect_equal(sort(r$year), c(2011, 2012))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("an unknown recipe id is rejected", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84).
|
||||
# Without it the test is non-discriminating: deleting the validation call
|
||||
# leaves .recipe_components() returning 0 rows and comps$label[[1]]
|
||||
# throwing "subscript out of bounds", which a bare expect_error() accepts
|
||||
# while the user loses the curated "valid recipe ids are ..." message.
|
||||
expect_error(
|
||||
cog_balances("550000227544", 2019, recipe = "no_such_recipe"),
|
||||
class = "uscogdata_unknown_recipe"
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
# --- balance_caveats: GAAP disclosure + measured coverage windows ----------
|
||||
|
||||
test_that("balance_caveats is always present and flags the GAAP distinction", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", 2019)
|
||||
cav <- attr(r, "provenance")$balance_caveats
|
||||
expect_false(is.null(cav))
|
||||
expect_true(cav$not_gaap)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("coverage_window is computed from the corpus, not hardcoded", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
|
||||
cav <- attr(r, "provenance")$balance_caveats
|
||||
|
||||
# Read the "general" family's true year extent independently, via a
|
||||
# fresh DuckDB connection against the raw parquet files (never through
|
||||
# balance_long/.balance_caveats() itself, and never via arrow -- this
|
||||
# package reads parquet through DuckDB only, see CLAUDE.md). Replicates
|
||||
# the same predicates 26-balance_long.sql applies (category_type =
|
||||
# 'balance', NOT is_aggregate) so this is a faithful, independent
|
||||
# measurement rather than a re-statement of the view under test.
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet")
|
||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
||||
obs <- DBI::dbGetQuery(con2, sprintf(
|
||||
"SELECT MIN(l.year) AS y0, MAX(l.year) AS y1
|
||||
FROM read_parquet(%s, hive_partitioning = true) l
|
||||
JOIN read_parquet(%s) c USING (item_code)
|
||||
WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate",
|
||||
uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path)
|
||||
))
|
||||
|
||||
expect_identical(as.integer(cav$coverage_window$general),
|
||||
c(as.integer(obs$y0), as.integer(obs$y1)))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("coverage_window covers every corpus subtype, not just observed ones", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Deliberate contract (provenance-v1.json): the window block is corpus-
|
||||
# scoped so a caller can ask "is there a family I missed?", while
|
||||
# `truncated` is the observed-scoped field. A single-category query must
|
||||
# therefore still report every balance family in the mounted corpus.
|
||||
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
||||
expect_identical(unique(r$balance_subtype), "general")
|
||||
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
||||
all_subtypes <- DBI::dbGetQuery(con2, sprintf(
|
||||
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
|
||||
WHERE balance_subtype IS NOT NULL",
|
||||
uscogdata:::.sql_lit_chr(cats_path)
|
||||
))$balance_subtype
|
||||
|
||||
cav <- attr(r, "provenance")$balance_caveats
|
||||
expect_setequal(names(cav$coverage_window), all_subtypes)
|
||||
expect_true(length(all_subtypes) > 1L)
|
||||
# ...while `truncated` stays scoped to what this query actually observed.
|
||||
expect_true(all(cav$truncated %in% unique(r$balance_subtype)))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the corpus-constant coverage windows are memoised per session", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# The windows query has no govid/year predicate: its answer depends only
|
||||
# on which corpus is mounted, so re-running the full balance_long scan on
|
||||
# every call is pure waste (35% of verb runtime on the fixture). Same
|
||||
# memoise-and-invalidate pattern as .uscogdata_env$manifest.
|
||||
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
|
||||
suppressMessages(cog_balances("550000227544", 2019))
|
||||
memo <- uscogdata:::.uscogdata_env$balance_coverage_windows
|
||||
expect_false(is.null(memo))
|
||||
expect_true("general" %in% names(memo))
|
||||
|
||||
uscogdata:::cog_close()
|
||||
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("a request past a family's coverage window is flagged", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in
|
||||
# the LIVE corpus -- Census moved employee retirement reporting to the
|
||||
# Annual Survey of Public Pensions after that year. This bundled FIXTURE
|
||||
# doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present),
|
||||
# so the family's *observed* max here is 2012, not 2016. Either way the
|
||||
# requested span (2012, 2019) reaches past what the family covers in
|
||||
# THIS corpus, which is what makes .balance_caveats() flag it -- the
|
||||
# assertion below is about the fixture's measured window, not the FY2016
|
||||
# live-corpus cutoff.
|
||||
r <- cog_balances("550000227544", c(2012, 2019))
|
||||
cav <- attr(r, "provenance")$balance_caveats
|
||||
expect_true("employee_retirement" %in% cav$truncated)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the provenance schema documents balance_caveats", {
|
||||
sch <- jsonlite::fromJSON(
|
||||
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
|
||||
simplifyVector = FALSE
|
||||
)
|
||||
expect_true("balance_caveats" %in% names(sch$properties))
|
||||
})
|
||||
|
||||
test_that("cog_explain surfaces the balance caveats", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Asserted on the RENDERED text, not on prov$balance_caveats: the field
|
||||
# is already covered above, and the once-per-session cli_inform() means
|
||||
# cog_explain() is the only surface a caller who missed (or suppressed)
|
||||
# the first message can still audit.
|
||||
r <- suppressMessages(cog_balances("550000227544", c(2012, 2019)))
|
||||
# Both streams: cli routes most of its output through conditions that
|
||||
# land on stderr, so a stdout-only capture would be empty (the pattern
|
||||
# used throughout test-explain.R).
|
||||
out <- paste(c(capture.output(cog_explain(r)),
|
||||
capture.output(cog_explain(r), type = "message")),
|
||||
collapse = "\n")
|
||||
expect_match(out, "GAAP")
|
||||
expect_match(out, "employee_retirement")
|
||||
})
|
||||
})
|
||||
|
||||
test_that("cog_explain on a money-verb result has no balance caveat section", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- suppressMessages(cog_spending("550000227544", 2019))
|
||||
out <- paste(c(capture.output(cog_explain(r)),
|
||||
capture.output(cog_explain(r), type = "message")),
|
||||
collapse = "\n")
|
||||
# Guard against the capture itself being vacuous: the section must be
|
||||
# absent from output that demonstrably contains the rest of the report.
|
||||
expect_match(out, "Data vintage")
|
||||
expect_false(grepl("GAAP", out))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the caveat message fires once per session", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
expect_message(cog_balances("550000227544", 2019), "not.*GAAP")
|
||||
expect_no_message(cog_balances("550000227544", 2020))
|
||||
})
|
||||
})
|
||||
@@ -8,14 +8,32 @@ test_that("cog_categories returns all categories grouped by subtype", {
|
||||
expect_gt(nrow(r), 10L)
|
||||
# corpus preserves Census-native "expenditure" vocabulary; the API takes
|
||||
# "spending" as a friendlier alias.
|
||||
expect_setequal(unique(r$category_type), c("expenditure", "revenue"))
|
||||
#
|
||||
# `balance` joined as a third category_type with the cash-and-security
|
||||
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
|
||||
# money verb, so it surfaces every category_type the corpus carries -- the
|
||||
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
|
||||
# return a balance row.
|
||||
expect_setequal(unique(r$category_type),
|
||||
c("expenditure", "revenue", "balance"))
|
||||
})
|
||||
|
||||
test_that("cog_categories(type = 'spending') returns only expenditure rows", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_categories(type = "spending")
|
||||
expect_true(all(r$category_type == "expenditure"))
|
||||
expect_true(all(r$subtype %in% c("operations", "capital", "intergovernmental")))
|
||||
# "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
|
||||
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
|
||||
# carrying dollars now maps to a category.
|
||||
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
|
||||
# joined with the I/Q/Y flow batch -- the last two characters of Census's
|
||||
# expenditure taxonomy. `interest` is what makes the three-concept model
|
||||
# computable: primary = direct minus debt service.
|
||||
# Exclude pseudo-category which has NA for subtype
|
||||
r_crosswalk <- r[r$category != "All Categories", ]
|
||||
expect_true(all(r_crosswalk$subtype %in%
|
||||
c("operations", "capital", "intergovernmental", "assistance",
|
||||
"interest", "insurance_benefits")))
|
||||
})
|
||||
|
||||
test_that("cog_categories surfaces the intergovernmental spending subtype", {
|
||||
@@ -33,8 +51,16 @@ test_that("cog_categories(type = 'revenue') returns only revenue rows", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_categories(type = "revenue")
|
||||
expect_true(all(r$category_type == "revenue"))
|
||||
expect_true(all(r$subtype %in%
|
||||
c("own_source", "federal", "state", "local_aid")))
|
||||
# The four non-general subtypes are deliberately NOT own_source: Census's
|
||||
# General Revenue excludes insurance trust (Y01 alone is $1.31T corpus-wide,
|
||||
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
|
||||
# (A90) revenue by definition, which is what makes both of its published
|
||||
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
|
||||
# Exclude pseudo-category which has NA for subtype
|
||||
r_crosswalk <- r[r$category != "All Categories", ]
|
||||
expect_true(all(r_crosswalk$subtype %in%
|
||||
c("own_source", "federal", "state", "local_aid",
|
||||
"insurance_trust", "utility", "liquor_store")))
|
||||
})
|
||||
|
||||
test_that("cog_categories(pattern = ...) filters case-insensitively", {
|
||||
@@ -47,6 +73,8 @@ test_that("cog_categories(pattern = ...) filters case-insensitively", {
|
||||
test_that("cog_categories has one row per (category, subtype)", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_categories()
|
||||
# Exclude pseudo-category which is not a crosswalk entry
|
||||
r <- r[r$category != "All Categories", ]
|
||||
key <- paste(r$category, r$subtype, sep = "|")
|
||||
expect_equal(length(key), length(unique(key)))
|
||||
})
|
||||
@@ -54,6 +82,8 @@ test_that("cog_categories has one row per (category, subtype)", {
|
||||
test_that("cog_categories item_codes is non-empty comma-separated string", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_categories()
|
||||
# Exclude pseudo-category which has NA for n_codes and item_codes
|
||||
r <- r[r$category != "All Categories", ]
|
||||
expect_true(all(nzchar(r$item_codes)))
|
||||
expect_true(all(r$n_codes >= 1L))
|
||||
# n_codes should equal count of commas + 1
|
||||
@@ -71,3 +101,40 @@ test_that("cog_categories sorted by category_type, category, subtype", {
|
||||
test_that("cog_categories rejects invalid type", {
|
||||
expect_error(cog_categories(type = "both"), "type")
|
||||
})
|
||||
|
||||
test_that("cog_categories() surfaces balance subtypes", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
cc <- cog_categories()
|
||||
b <- cc[cc$category_type == "balance", ]
|
||||
expect_true(nrow(b) > 0L)
|
||||
|
||||
# Every balance row must carry its subtype. Before the COALESCE included
|
||||
# balance_subtype these were all NA, which silently made the balance
|
||||
# taxonomy undiscoverable -- cog-api derives its subtype vocabulary from
|
||||
# this function, so an NA here becomes an unusable API parameter.
|
||||
expect_false(any(is.na(b$subtype)))
|
||||
|
||||
# The exact set, read independently from the crosswalk rather than from
|
||||
# the function under test.
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
p <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
||||
want <- DBI::dbGetQuery(con2, sprintf(
|
||||
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
|
||||
WHERE category_type = 'balance' AND balance_subtype IS NOT NULL
|
||||
ORDER BY 1", uscogdata:::.sql_lit_chr(p)))$balance_subtype
|
||||
expect_true(length(want) > 1L)
|
||||
expect_identical(sort(unique(b$subtype)), sort(want))
|
||||
})
|
||||
})
|
||||
|
||||
test_that('cog_categories(type = "balance") filters to holdings', {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
b <- cog_categories(type = "balance")
|
||||
expect_true(nrow(b) > 0L)
|
||||
expect_identical(unique(b$category_type), "balance")
|
||||
expect_false(any(is.na(b$subtype)))
|
||||
})
|
||||
})
|
||||
|
||||
@@ -0,0 +1,193 @@
|
||||
# tests/testthat/test-complete.R
|
||||
#
|
||||
# uscogdata#18. The published corpus no longer stores the wide era's explicit
|
||||
# zeros (cog_pipeline#64, series break SB194), so absence means two different
|
||||
# things:
|
||||
#
|
||||
# <= FY2011 (dense_source) : cell absent => Census published $0
|
||||
# >= FY2012 (sparse_source): cell absent => not reported, unknown
|
||||
#
|
||||
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
|
||||
# row's `value_source` so the two are distinguishable. Expected row sets here
|
||||
# are built from the corpus parquet directly, never from the verb under test --
|
||||
# verifying what a filter does through that same filter proves nothing.
|
||||
|
||||
# The (subtype, category) cells that SHOULD exist for one government-year:
|
||||
# every code in force for that government's type, mapped through
|
||||
# summary_categories, matching the verb's crosswalk subtype scope (the
|
||||
# default concept, `primary`, is operations/capital/assistance -- see
|
||||
# uscogdata#11) and excluding aggregate-flagged codes (which
|
||||
# spending_long/revenue_long drop).
|
||||
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
|
||||
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
|
||||
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
|
||||
wt_raw_query(sprintf(
|
||||
"SELECT DISTINCT c.%s AS subtype, c.category
|
||||
FROM %s cs
|
||||
JOIN %s x ON x.govs_type = cs.type
|
||||
JOIN %s c ON c.item_code = cs.item_code
|
||||
WHERE x.canonical_govid = '%s'
|
||||
AND cs.year = %d
|
||||
AND NOT cs.is_aggregate
|
||||
AND c.category IS NOT NULL
|
||||
AND c.%s IN (%s)",
|
||||
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
|
||||
q("summary_categories.parquet"), govid, year,
|
||||
subtype_col, paste0("'", subtypes, "'", collapse = ",")
|
||||
))
|
||||
}
|
||||
|
||||
# The default expenditure concept's subtype scope, mirrored from
|
||||
# R/spending.R's .spend_subtypes_primary.
|
||||
primary_subtypes <- c("operations", "capital", "assistance")
|
||||
|
||||
test_that("complete = FALSE is the default and changes nothing", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
plain <- cog_spending("121011212191", 2011L)
|
||||
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
|
||||
expect_equal(nrow(plain), nrow(explicit))
|
||||
expect_false("value_source" %in% names(plain))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# FY2011 is dense_source: before sparsification this government carried a
|
||||
# row for every code in force, most of them $0. complete = TRUE must
|
||||
# reproduce that cell set exactly.
|
||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
expected <- raw_expected_cells("121011212191", 2011L,
|
||||
primary_subtypes, "spend_subtype")
|
||||
|
||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||
expect_setequal(key(r$spend_subtype, r$category),
|
||||
key(expected$subtype, expected$category))
|
||||
expect_gt(nrow(expected), 0L)
|
||||
|
||||
# Every filled cell in a dense-source year is a Census-published $0 --
|
||||
# never "unknown", which is what the modern era's absences mean.
|
||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
|
||||
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
plain <- cog_spending("121011212191", 2011L)
|
||||
full <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
|
||||
# Filling adds rows; it must never alter or drop one.
|
||||
expect_gt(nrow(full), nrow(plain))
|
||||
reported <- full[full$value_source == "reported", ]
|
||||
expect_equal(nrow(reported), nrow(plain))
|
||||
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
|
||||
# ... and the total is unchanged, because every added cell is $0.
|
||||
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("a sparse-source year's absences are unknown, not zero", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# FY2019 is sparse_source: an absent cell means the government did not
|
||||
# report, which is NOT a zero. Filling those with 0 would invent data --
|
||||
# the exact error the representation contract exists to prevent.
|
||||
r <- cog_spending("121011212191", 2019L, complete = TRUE)
|
||||
filled <- r[r$value_source != "reported", ]
|
||||
expect_gt(nrow(filled), 0L)
|
||||
expect_true(all(filled$value_source == "not_reported"))
|
||||
expect_true(all(is.na(filled$amt_nominal)))
|
||||
expect_false(any(r$value_source == "census_zero"))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("the fill is scoped to each government's own type", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Filling against the union of all types would invent cells for codes a
|
||||
# county can never report. Every filled category must be one that
|
||||
# code_set puts in force for type 1 (county) specifically.
|
||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
||||
county_cells <- raw_expected_cells("121011212191", 2011L,
|
||||
primary_subtypes, "spend_subtype")
|
||||
expect_true(all(r$category %in% county_cells$category))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE respects the category filter", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011L, category = "Police",
|
||||
complete = TRUE)
|
||||
expect_true(all(r$category == "Police"))
|
||||
expect_true("value_source" %in% names(r))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("cog_revenue() completes on its own flow", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
|
||||
expected <- raw_expected_cells("121011212191", 2011L,
|
||||
c("own_source", "federal", "state", "local_aid"),
|
||||
"revenue_subtype")
|
||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
||||
expect_setequal(key(r$revenue_subtype, r$category),
|
||||
key(expected$subtype, expected$category))
|
||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("provenance records the completion and its absence rule", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
|
||||
"provenance")
|
||||
expect_true(prov$completion$applied)
|
||||
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
|
||||
expect_gt(prov$completion$rows_filled, 0L)
|
||||
|
||||
off <- attr(cog_spending("121011212191", 2011L), "provenance")
|
||||
expect_false(off$completion$applied)
|
||||
expect_equal(off$completion$rows_filled, 0L)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE is refused where the fill would be guesswork", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# A recipe defines its own component codes and does not go through
|
||||
# summary_categories at all, so there is no grid to fill from.
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
# The intergovernmental leg keeps aggregate rows by design
|
||||
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, expenditure_concept = "total",
|
||||
complete = TRUE),
|
||||
class = "uscogdata_complete_unsupported"
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("complete = TRUE aborts on a corpus with no representation contract", {
|
||||
skip_if_no_corpus()
|
||||
# A corpus published before sparsification carries neither table, so there
|
||||
# is nothing to fill from and no rule saying what an absence means. That
|
||||
# must abort rather than guess.
|
||||
with_corpus_missing_representation({
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2011L, complete = TRUE),
|
||||
class = "uscogdata_representation_unavailable"
|
||||
)
|
||||
# ... while an ordinary query on the same corpus still works.
|
||||
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
|
||||
})
|
||||
})
|
||||
@@ -64,3 +64,109 @@ test_that(".resolve_url does not invent a slash for an empty setting", {
|
||||
withr::local_options(uscogdata.url = "")
|
||||
expect_equal(.resolve_url(), "")
|
||||
})
|
||||
|
||||
test_that("the default corpus URL is real, not a placeholder", {
|
||||
# setup.R points USCOGDATA_URL at the bundled fixture for the whole suite,
|
||||
# so both the env var and the option have to be cleared to see the default.
|
||||
withr::local_envvar(USCOGDATA_URL = NA)
|
||||
withr::local_options(uscogdata.url = NULL)
|
||||
url <- .resolve_url()
|
||||
expect_false(grepl("REPLACE_WITH", url, fixed = TRUE))
|
||||
expect_match(url, "^https://")
|
||||
expect_match(url, "/$")
|
||||
})
|
||||
|
||||
test_that("an explicitly-set sentinel URL still aborts", {
|
||||
# The guard must survive the default change: a user who half-edited a
|
||||
# copied config still gets the actionable error.
|
||||
withr::local_envvar(
|
||||
USCOGDATA_URL = "https://other.example/s/REPLACE_WITH_SHARE_TOKEN/x/"
|
||||
)
|
||||
expect_error(
|
||||
.check_url_configured(.resolve_url()),
|
||||
class = "uscogdata_url_not_configured"
|
||||
)
|
||||
})
|
||||
|
||||
test_that("DESCRIPTION carries release metadata", {
|
||||
skip_if_no_source_tree("DESCRIPTION")
|
||||
d <- read.dcf(source_tree_path("DESCRIPTION"))
|
||||
fields <- colnames(d)
|
||||
|
||||
expect_true(all(c("URL", "BugReports") %in% fields))
|
||||
expect_match(d[1, "Authors@R"], "Knowles", fixed = TRUE)
|
||||
expect_match(d[1, "Authors@R"], "0000-0003-0005-9478", fixed = TRUE)
|
||||
expect_match(d[1, "Authors@R"], "Civilytics Consulting LLC", fixed = TRUE)
|
||||
|
||||
# The gate in .validate_schema() accepts up to 7 and the published corpus
|
||||
# IS 7; DESCRIPTION must not claim otherwise.
|
||||
expect_equal(as.integer(d[1, "MaxCorpusSchema"]), 7L)
|
||||
|
||||
# Authors@R must actually parse -- a malformed person() call is only
|
||||
# caught at citation()/build time otherwise.
|
||||
people <- eval(parse(text = d[1, "Authors@R"]))
|
||||
expect_s3_class(people, "person")
|
||||
expect_true("cre" %in% unlist(lapply(people, function(p) p$role)))
|
||||
})
|
||||
|
||||
test_that("LICENSE and LICENSE.md name the same copyright holder", {
|
||||
skip_if_no_source_tree("LICENSE", "LICENSE.md")
|
||||
holder <- sub("^COPYRIGHT HOLDER:\\s*", "",
|
||||
grep("^COPYRIGHT HOLDER:", readLines(source_tree_path("LICENSE"),
|
||||
warn = FALSE), value = TRUE))
|
||||
full <- paste(readLines(source_tree_path("LICENSE.md"), warn = FALSE), collapse = "\n")
|
||||
|
||||
expect_equal(holder, "Civilytics Consulting LLC")
|
||||
expect_match(full, holder, fixed = TRUE)
|
||||
# usethis::use_mit_license() writes LICENSE.md but leaves an existing
|
||||
# LICENSE alone, which is how the two came to disagree in the first place.
|
||||
expect_match(full, "MIT License", fixed = TRUE)
|
||||
})
|
||||
|
||||
test_that("vignettes are not excluded from the build", {
|
||||
skip_if_no_source_tree(".Rbuildignore")
|
||||
ignore <- readLines(source_tree_path(".Rbuildignore"), warn = FALSE)
|
||||
expect_false(any(grepl("^\\^vignettes\\$$", ignore)))
|
||||
# The fixture is what lets R CMD check run offline with no credentials on
|
||||
# r-universe and GitHub Actions. It must never be excluded.
|
||||
expect_false(any(grepl("fixture_corpus", ignore, fixed = TRUE)))
|
||||
# doc/ and Meta/ ARE build artefacts of devtools::build_vignettes() and must
|
||||
# stay excluded -- R CMD build regenerates inst/doc/ from vignettes/ on its
|
||||
# own, and leaving them in earns a "non-standard file at top level" NOTE.
|
||||
expect_true(any(grepl("^\\^doc\\$$", ignore)))
|
||||
expect_true(any(grepl("^\\^Meta\\$$", ignore)))
|
||||
})
|
||||
|
||||
test_that("_pkgdown.yml indexes every exported topic", {
|
||||
skip_if_no_source_tree("_pkgdown.yml", "NAMESPACE")
|
||||
exports <- grep("^export\\(", readLines(source_tree_path("NAMESPACE"), warn = FALSE),
|
||||
value = TRUE)
|
||||
exports <- sub("^export\\((.*)\\)$", "\\1", exports)
|
||||
yml <- paste(readLines(source_tree_path("_pkgdown.yml"), warn = FALSE), collapse = "\n")
|
||||
missing <- exports[!vapply(exports,
|
||||
function(e) grepl(paste0("\\b", e, "\\b"), yml),
|
||||
logical(1))]
|
||||
# pkgdown errors on topics missing from the index, so an unlisted export
|
||||
# means the docs site does not build at all.
|
||||
expect_equal(missing, character(0))
|
||||
})
|
||||
|
||||
test_that("README is written for a stranger, not a repo insider", {
|
||||
skip_if_no_source_tree("README.md")
|
||||
r <- paste(readLines(source_tree_path("README.md"), warn = FALSE), collapse = "\n")
|
||||
|
||||
# No paths that only resolve inside a maintainer's checkout.
|
||||
expect_false(grepl("../cog_pipeline", r, fixed = TRUE))
|
||||
# A real, uncommented install line.
|
||||
expect_match(r, "install.packages", fixed = TRUE)
|
||||
expect_false(grepl("# pak::pkg_install", r, fixed = TRUE))
|
||||
# The errata most likely to produce a plausible-looking wrong answer.
|
||||
expect_match(r, "full US dollars", fixed = TRUE)
|
||||
# The release advice that conflicts with public CI is gone.
|
||||
expect_false(grepl("Rbuildignore", r, fixed = TRUE))
|
||||
# Both read paths documented.
|
||||
expect_match(r, "cog_mirror", fixed = TRUE)
|
||||
# cog_spending() has no default for `years`; a quickstart that omits it
|
||||
# errors on the reader's first call.
|
||||
expect_match(r, "years\\s*=", perl = TRUE)
|
||||
})
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
# tests/testthat/test-corpus-breaks.R
|
||||
#
|
||||
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
|
||||
# are caveats about the corpus itself rather than about one item code:
|
||||
#
|
||||
# SB085 1977 dollar precision across the 1976/1977 boundary
|
||||
# SB087 2002 imputation exclusion FY2002-2006
|
||||
# SB194 2012 dense -> sparse representation change
|
||||
# SB086 2017 government ID scheme change
|
||||
#
|
||||
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
|
||||
# and no row's item_code is ever the literal "ALL", so none of them could
|
||||
# ever reach a user. They now travel in their own provenance field,
|
||||
# `corpus_break_refs`, which keeps them distinguishable from the
|
||||
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
|
||||
# result, not one series).
|
||||
|
||||
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
|
||||
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
|
||||
# DoD 4 intended to reach users.
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_true("SB194" %in% prov$corpus_break_refs)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# 2019-2020 spans no catalogued corpus-wide break.
|
||||
r <- cog_spending("121011212191", 2019:2020, "Police")
|
||||
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
|
||||
})
|
||||
})
|
||||
|
||||
test_that("corpus_break_refs and series_break_refs stay disjoint", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_type(prov$series_break_refs, "character")
|
||||
expect_type(prov$corpus_break_refs, "character")
|
||||
# An ALL caveat must never masquerade as a break in a specific series.
|
||||
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
|
||||
expect_false("SB194" %in% prov$series_break_refs)
|
||||
})
|
||||
})
|
||||
|
||||
test_that(".build_corpus_break_refs matches on the break_year window alone", {
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
|
||||
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
|
||||
# series_breaks table is a full cross-vintage registry, so the matching
|
||||
# logic is testable there even though no long partition covers it.
|
||||
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||
con, years = 1975:1980, schema_version = 6L
|
||||
))
|
||||
# ... and does not fire for a range that misses it, unlike a filter keyed
|
||||
# on the era rather than the boundary.
|
||||
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
||||
con, years = 1978:1980, schema_version = 6L
|
||||
))
|
||||
|
||||
# Unlike code-specific refs, these do not depend on which codes a result
|
||||
# happens to contain -- that dependency is the whole defect.
|
||||
expect_setequal(
|
||||
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
|
||||
"SB087"
|
||||
)
|
||||
|
||||
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
|
||||
expect_equal(
|
||||
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
|
||||
character(0)
|
||||
)
|
||||
})
|
||||
|
||||
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
||||
out <- paste(c(
|
||||
capture.output(cog_explain(r)),
|
||||
capture.output(cog_explain(r), type = "message")
|
||||
), collapse = "\n")
|
||||
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
|
||||
expect_match(out, "SB194", fixed = TRUE)
|
||||
})
|
||||
})
|
||||
@@ -30,7 +30,6 @@ wt_coverage <- function(x) {
|
||||
}
|
||||
|
||||
test_that("multi-government aggregates disclose reporting coverage on every result", {
|
||||
testthat::skip("Blocked on uscogdata#13 (findings F-020, F-023)")
|
||||
|
||||
# -- F-020: geographic rollups -------------------------------------------
|
||||
# Wisconsin's city/village universe is 608 governments. On the bundled
|
||||
@@ -49,10 +48,22 @@ test_that("multi-government aggregates disclose reporting coverage on every resu
|
||||
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
|
||||
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
|
||||
|
||||
# Cross-check against the raw partitions, scoped to the SAME universe the
|
||||
# rollup was given -- the 608 govids above. Scoping instead on the long
|
||||
# table's own `type`/`fips_state` asks a different question and answers 595:
|
||||
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
|
||||
# identity, when they were townships) while the xwalk lists them as
|
||||
# govs_type = 2 (their present identity, as villages). Schema v6 made the
|
||||
# long table's geography present-harmonized and moved as-of-year to the
|
||||
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
|
||||
# in R/manifest.R. n_units_reporting counts against the requested universe,
|
||||
# so 597 is the number that answers "how many of the governments I asked
|
||||
# about reported".
|
||||
raw_2012 <- wt_raw_query(paste0(
|
||||
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
|
||||
"WHERE type = 2 AND fips_state = 55 AND year = 2012 ",
|
||||
"AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate"))
|
||||
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
|
||||
"AND canonical_govid IN (",
|
||||
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
|
||||
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
|
||||
|
||||
# -- F-023: peer cohorts --------------------------------------------------
|
||||
|
||||
@@ -13,9 +13,13 @@ test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
|
||||
}
|
||||
})
|
||||
|
||||
test_that("expenditure_concept defaults to direct and preserves today's numbers", {
|
||||
test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
|
||||
gov <- "010000226085" # Alabama state government
|
||||
base <- cog_spending(gov, years = 2019, category = "Police")
|
||||
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
|
||||
# Police maps only to operations/capital codes (E62/F62/G62), so the
|
||||
# direct concept's extra subtypes (interest, insurance_benefits) cannot
|
||||
# contribute and the two concepts must agree exactly here.
|
||||
expl <- cog_spending(gov, years = 2019, category = "Police",
|
||||
expenditure_concept = "direct")
|
||||
expect_equal(base$amt_nominal, expl$amt_nominal)
|
||||
@@ -59,7 +63,9 @@ test_that("the IG leg never includes the L-- family total", {
|
||||
codes <- DBI::dbGetQuery(con,
|
||||
"SELECT DISTINCT item_code FROM ig_long")$item_code
|
||||
expect_false(any(grepl("--$", codes)))
|
||||
expect_true(all(substr(codes, 1, 1) %in% c("M", "L")))
|
||||
# Q joined the IG family with the crosswalk-membership rewrite
|
||||
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
|
||||
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
|
||||
})
|
||||
|
||||
test_that("expenditure_concept rejects unknown values", {
|
||||
@@ -246,9 +252,12 @@ test_that("both cross-government verbs still accept the direct default", {
|
||||
})
|
||||
|
||||
test_that("provenance always records the expenditure concept", {
|
||||
d <- cog_spending("010000226085", years = 2019, category = "Police")
|
||||
p <- cog_spending("010000226085", years = 2019, category = "Police")
|
||||
d <- cog_spending("010000226085", years = 2019, category = "Police",
|
||||
expenditure_concept = "direct")
|
||||
t <- cog_spending("010000226085", years = 2019, category = "Police",
|
||||
expenditure_concept = "total")
|
||||
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
|
||||
expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
|
||||
expect_equal(attr(t, "provenance")$expenditure_concept, "total")
|
||||
# The note explains the non-obvious part: how legacy IG was assembled.
|
||||
@@ -298,8 +307,16 @@ test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to
|
||||
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
|
||||
# Direct/Total pairing. The flow-family gate in
|
||||
# .attach_ig_counterparts() must keep ig_recipe_id NULL here.
|
||||
#
|
||||
# Anchored on FL state government, not AL. Coverage is presence-based: a
|
||||
# recipe is only suggested when its component codes have rows for the
|
||||
# requested government-year. AL state's only FY2011 B47 cell was an
|
||||
# explicit zero, which the corpus no longer stores after sparsification
|
||||
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
|
||||
# FL state carries a real FY2011 B47 amount, so this exercises the guard
|
||||
# against a suggestion that genuinely fires.
|
||||
r <- suppressMessages(
|
||||
cog_spending("010000226085", years = c(2005, 2011), category = "IG Federal")
|
||||
cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
|
||||
)
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
expect_gt(length(sugg), 0L)
|
||||
|
||||
@@ -19,8 +19,6 @@
|
||||
# also check the FY2022 numbers above.
|
||||
|
||||
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
|
||||
testthat::skip("Blocked on uscogdata#11 (findings F-012, F-017, F-018)")
|
||||
|
||||
mad <- "552025209777" # MADISON CITY, WI
|
||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
||||
|
||||
@@ -61,15 +59,54 @@ test_that("expenditure concepts classify on spend_type, not item-code prefix", {
|
||||
|
||||
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
|
||||
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
|
||||
# payments. All four share the first letter `Y` and the spend_type
|
||||
# "Insurance Trust", so this pair of assertions is the concrete proof that
|
||||
# classification is no longer keyed on the first letter.
|
||||
# payments. All four share the first letter `Y`, so no first-letter allowlist
|
||||
# can route them. The proof that classification is crosswalk-keyed:
|
||||
# Y05 lands in `total` spending (insurance_benefits is inside `direct`),
|
||||
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
|
||||
# therefore can never appear in a spending result.
|
||||
#
|
||||
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
|
||||
# DEFAULT stays Census General Revenue and so excludes insurance-trust
|
||||
# revenue; Y01's revenue-side classification is asserted against the
|
||||
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
|
||||
# revenue concept argument is uscogdata#12.
|
||||
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
|
||||
spend_codes <- wt_codes_included(wi_total)
|
||||
rev_codes <- wt_codes_included(wi_revenue)
|
||||
|
||||
expect_true("Y05" %in% spend_codes)
|
||||
expect_false("Y05" %in% rev_codes)
|
||||
expect_true("Y01" %in% rev_codes)
|
||||
expect_false("Y01" %in% spend_codes)
|
||||
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
|
||||
|
||||
con <- uscogdata:::.ensure_session()
|
||||
y_class <- DBI::dbGetQuery(con,
|
||||
"SELECT item_code, category_type, spend_subtype, revenue_subtype
|
||||
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
|
||||
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
|
||||
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
|
||||
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
|
||||
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
|
||||
})
|
||||
|
||||
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
|
||||
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
|
||||
# balances) share first letters with flow codes, so this could never be
|
||||
# guaranteed under prefix classification; under crosswalk membership it
|
||||
# falls out structurally -- asserted here at the verb level, on a
|
||||
# government-year the fixture gives real balance rows (Wisconsin carries
|
||||
# Y07/Y08/Y21/Y61-type balances in FY2019).
|
||||
wi_state <- "550000227544"
|
||||
con <- uscogdata:::.ensure_session()
|
||||
balance <- DBI::dbGetQuery(con,
|
||||
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
|
||||
expect_gt(nrow(balance), 0L)
|
||||
|
||||
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
|
||||
rev <- cog_revenue(wi_state, 2019L)
|
||||
|
||||
expect_false(any(spend$category %in% balance$category))
|
||||
expect_false(any(rev$category %in% balance$category))
|
||||
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
|
||||
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
|
||||
})
|
||||
|
||||
@@ -83,7 +83,7 @@ test_that("cog_explain prints the expenditure concept (I1)", {
|
||||
capture.output(cog_explain(t)),
|
||||
capture.output(cog_explain(t), type = "message")
|
||||
), collapse = "\n")
|
||||
expect_true(grepl("Concept: direct", txt_d))
|
||||
expect_true(grepl("Concept: primary", txt_d))
|
||||
expect_true(grepl("Concept: total", txt_t))
|
||||
})
|
||||
|
||||
|
||||
@@ -0,0 +1,112 @@
|
||||
# tests/testthat/test-fixture-vintage.R
|
||||
#
|
||||
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
|
||||
# every test in this package -- plus the whole cog-api suite -- runs against
|
||||
# it. When the published corpus changes shape and the fixture does not, both
|
||||
# suites stay green against a corpus that no longer exists (uscogdata#18).
|
||||
#
|
||||
# These tests pin the structural facts that distinguish the current published
|
||||
# vintage from its predecessor, so a stale fixture fails loudly instead of
|
||||
# passing quietly. They assert shape, never dollar values: re-running
|
||||
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
|
||||
# keep them green.
|
||||
|
||||
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
|
||||
# not the package session: these assertions are about what the fixture
|
||||
# CONTAINS, and routing them through the reader's own views would let a
|
||||
# filter hide the very absence being checked.
|
||||
fixture_query <- function(sql, ...) {
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
path <- function(rel) {
|
||||
sprintf("read_parquet(%s)",
|
||||
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
|
||||
}
|
||||
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
|
||||
}
|
||||
|
||||
test_that("fixture ships every metadata table the publish tree does", {
|
||||
skip_if_no_corpus()
|
||||
# representation/code_set are what make a sparse corpus interpretable; a
|
||||
# fixture without them predates sparsification (cog_pipeline#64).
|
||||
expected <- c(
|
||||
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
|
||||
"census_collection_coverage.parquet", "code_set.parquet",
|
||||
"harmonization_map.parquet", "harmonization_recipes.parquet",
|
||||
"lineage_events.parquet", "representation.parquet",
|
||||
"series_breaks.parquet", "summary_categories.parquet"
|
||||
)
|
||||
on_disk <- basename(list.files(
|
||||
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
|
||||
))
|
||||
expect_true(all(expected %in% on_disk))
|
||||
|
||||
# The manifest must list them too -- consumers read the manifest, not ls().
|
||||
in_manifest <- with_fixture_corpus(
|
||||
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
|
||||
)
|
||||
expect_true(all(expected %in% in_manifest))
|
||||
})
|
||||
|
||||
test_that("fixture carries the dense/sparse representation contract", {
|
||||
skip_if_no_corpus()
|
||||
rep <- fixture_query(
|
||||
"SELECT year, representation, absence_means FROM %s
|
||||
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
|
||||
"data/representation.parquet"
|
||||
)
|
||||
expect_equal(nrow(rep), 4L)
|
||||
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
|
||||
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
|
||||
})
|
||||
|
||||
test_that("the fixture's wide era is sparse, not zero-padded", {
|
||||
skip_if_no_corpus()
|
||||
# FY2011 is a dense_source year: the corpus publishes only the cells Census
|
||||
# reported non-zero, and an absent cell means Census published $0. Before
|
||||
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
|
||||
# zeros. A single explicit zero here means the fixture predates the change.
|
||||
zeros_2011 <- fixture_query(
|
||||
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
|
||||
"data/long/year=2011/part-0.parquet"
|
||||
)$n
|
||||
expect_equal(zeros_2011, 0L)
|
||||
|
||||
# The modern era is a different regime: a reported zero there is real data
|
||||
# (the government filed $0), so zeros legitimately survive and must not be
|
||||
# asserted away.
|
||||
expect_gt(
|
||||
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
|
||||
0L
|
||||
)
|
||||
})
|
||||
|
||||
test_that("code_set covers every fixture year with the reader-spec columns", {
|
||||
skip_if_no_corpus()
|
||||
cs <- fixture_query(
|
||||
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
|
||||
"data/code_set.parquet"
|
||||
)
|
||||
expect_true(all(
|
||||
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
|
||||
%in% names(cs)
|
||||
))
|
||||
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
|
||||
})
|
||||
|
||||
test_that("every flow code carrying dollars has a category, J-prefix included", {
|
||||
skip_if_no_corpus()
|
||||
# The J (assistance/benefit) codes were uncategorised until the crosswalk
|
||||
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
|
||||
# duplication fix landed). Their absence is how a pre-crosswalk fixture
|
||||
# gives itself away.
|
||||
j <- fixture_query(
|
||||
"SELECT item_code, category, category_type, spend_subtype FROM %s
|
||||
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
|
||||
"data/summary_categories.parquet"
|
||||
)
|
||||
expect_true("J19" %in% j$item_code)
|
||||
expect_true(all(j$category_type == "expenditure"))
|
||||
expect_true(all(j$spend_subtype == "assistance"))
|
||||
expect_false(any(is.na(j$category)))
|
||||
})
|
||||
@@ -18,7 +18,6 @@
|
||||
# semantics, not a row the fix makes findable.
|
||||
|
||||
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
||||
testthat::skip("Blocked on uscogdata#16 (finding F-025)")
|
||||
|
||||
# -- correctness (1): a government must be findable by its own exact name ---
|
||||
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# Network-gated. Set USCOGDATA_LIVE_TEST=true to run.
|
||||
#
|
||||
# This file exists because the defect fixed for 0.3.0 -- no remote corpus was
|
||||
# readable at all, because DuckDB cannot expand a glob over generic HTTP --
|
||||
# survived precisely because every other test path used a LOCAL corpus (the
|
||||
# bundled fixture), and so did the API in production (a host mount). Nothing
|
||||
# ever exercised the package the way a new user does.
|
||||
skip_live <- function() {
|
||||
testthat::skip_if_not(
|
||||
identical(tolower(Sys.getenv("USCOGDATA_LIVE_TEST", "")), "true"),
|
||||
"live-corpus test: set USCOGDATA_LIVE_TEST=true to run"
|
||||
)
|
||||
}
|
||||
|
||||
# The suite's setup.R pins USCOGDATA_URL to the bundled fixture, so reaching
|
||||
# the default requires clearing both the env var and the option.
|
||||
with_default_corpus <- function(code) {
|
||||
withr::local_envvar(
|
||||
USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA,
|
||||
.local_envir = parent.frame()
|
||||
)
|
||||
withr::local_options(uscogdata.url = NULL, .local_envir = parent.frame())
|
||||
cog_close()
|
||||
withr::defer(cog_close(), envir = parent.frame())
|
||||
force(code)
|
||||
}
|
||||
|
||||
test_that("the package reads the public corpus with no configuration at all", {
|
||||
skip_live()
|
||||
with_default_corpus({
|
||||
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||||
expect_gt(nrow(g), 0)
|
||||
|
||||
s <- cog_spending(g$canonical_govid[1], years = 2022)
|
||||
expect_gt(nrow(s), 0)
|
||||
expect_true(all(c("amt_nominal", "year", "category") %in% names(s)))
|
||||
|
||||
# Amounts are full dollars, already x1000. A city's annual spending is
|
||||
# millions, not thousands -- this catches a regression that dropped or
|
||||
# doubled the conversion.
|
||||
expect_gt(sum(s$amt_nominal, na.rm = TRUE), 1e6)
|
||||
|
||||
p <- attr(s, "provenance")
|
||||
expect_true(isTRUE(p$transformations$units_conversion$applied))
|
||||
expect_equal(p$transformations$units_conversion$multiplier, 1000)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("a multi-decade query reads across many partitions", {
|
||||
skip_live()
|
||||
with_default_corpus({
|
||||
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||||
# `years` is required on cog_spending() -- there is no full-history
|
||||
# default at the reader level (the API's /profile route supplies one).
|
||||
s <- cog_spending(g$canonical_govid[1], years = 2000:2022)
|
||||
# Enumeration builds one read_parquet() path per requested partition. If
|
||||
# the list were truncated, or silently collapsed to a single file, the
|
||||
# returned span is what catches it.
|
||||
expect_gt(diff(range(s$year)), 10)
|
||||
expect_gt(length(unique(s$year)), 5)
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,64 @@
|
||||
test_that(".long_files_sql enumerates every partition the manifest lists", {
|
||||
manifest <- list(files = list(long_partitions = list(
|
||||
list(year = 2011L, path = "data/long/year=2011/part-0.parquet"),
|
||||
list(year = 2012L, path = "data/long/year=2012/part-0.parquet")
|
||||
)))
|
||||
expect_equal(
|
||||
uscogdata:::.long_files_sql("https://example.org/corpus/", manifest),
|
||||
paste0(
|
||||
"['https://example.org/corpus/data/long/year=2011/part-0.parquet',",
|
||||
"'https://example.org/corpus/data/long/year=2012/part-0.parquet']"
|
||||
)
|
||||
)
|
||||
})
|
||||
|
||||
test_that(".long_files_sql falls back to the glob when no partition list is present", {
|
||||
# test-views.R registers views with a hand-built manifest that has no
|
||||
# `files` element. That must keep working: the glob is valid for the
|
||||
# local paths such a manifest is used with.
|
||||
expect_equal(
|
||||
uscogdata:::.long_files_sql("/tmp/corpus/", list(schema_version = 4L)),
|
||||
"'/tmp/corpus/data/long/**/*.parquet'"
|
||||
)
|
||||
expect_equal(
|
||||
uscogdata:::.long_files_sql("/tmp/corpus/", list(files = list(long_partitions = list()))),
|
||||
"'/tmp/corpus/data/long/**/*.parquet'"
|
||||
)
|
||||
})
|
||||
|
||||
test_that("the enumerated list matches the bundled fixture's partition count", {
|
||||
skip_if_no_corpus()
|
||||
m <- jsonlite::fromJSON(
|
||||
file.path(fixture_corpus_path(), "manifest.json"), simplifyVector = FALSE
|
||||
)
|
||||
out <- uscogdata:::.long_files_sql(fixture_corpus_path(), m)
|
||||
expect_equal(
|
||||
lengths(regmatches(out, gregexpr("part-0\\.parquet", out))),
|
||||
length(m$files$long_partitions)
|
||||
)
|
||||
})
|
||||
|
||||
test_that("no view SQL survives rendering with an unsubstituted token", {
|
||||
# Introducing {long_files} broke four test sites that had hand-rolled the
|
||||
# {url} substitution -- each failed with a DuckDB parser error on the
|
||||
# surviving brace. This asserts the whole SQL directory renders clean, so
|
||||
# a future token cannot reintroduce that silently.
|
||||
sql_dir <- system.file("sql", package = "uscogdata")
|
||||
for (f in list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE)) {
|
||||
rendered <- uscogdata:::.render_view_sql(
|
||||
paste(readLines(f, warn = FALSE), collapse = "\n"), "/tmp/corpus/"
|
||||
)
|
||||
expect_false(grepl("\\{[a-z_]+\\}", rendered), label = basename(f))
|
||||
}
|
||||
})
|
||||
|
||||
test_that("registered `long` view reads through the enumerated list", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
con <- uscogdata:::.ensure_session()
|
||||
n <- DBI::dbGetQuery(con, "SELECT count(*) AS n FROM long")$n
|
||||
expect_gt(n, 0)
|
||||
yrs <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
|
||||
expect_true(all(c(2011, 2012, 2019, 2020) %in% yrs))
|
||||
})
|
||||
})
|
||||
@@ -4,12 +4,14 @@
|
||||
# protect users from silent failures when USCOGDATA_URL is misconfigured
|
||||
# or returns non-JSON content.
|
||||
|
||||
test_that("cog_open aborts with actionable error when URL is the placeholder default", {
|
||||
test_that("cog_open aborts with actionable error when URL contains the sentinel", {
|
||||
uscogdata:::cog_close()
|
||||
on.exit(uscogdata:::cog_close(), add = TRUE)
|
||||
|
||||
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
|
||||
withr::with_envvar(c(USCOGDATA_URL = placeholder), {
|
||||
# No longer the package default (that is the public HF corpus). This is a
|
||||
# user who copied a config template and did not finish editing it.
|
||||
sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
|
||||
withr::with_envvar(c(USCOGDATA_URL = sentinel_url), {
|
||||
expect_error(
|
||||
uscogdata:::cog_open(),
|
||||
class = "uscogdata_url_not_configured"
|
||||
@@ -35,8 +37,10 @@ test_that("placeholder guard error names both env var and option as remediation"
|
||||
uscogdata:::cog_close()
|
||||
on.exit(uscogdata:::cog_close(), add = TRUE)
|
||||
|
||||
placeholder <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
|
||||
withr::with_envvar(c(USCOGDATA_URL = placeholder), {
|
||||
# No longer the package default (that is the public HF corpus). This is a
|
||||
# user who copied a config template and did not finish editing it.
|
||||
sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
|
||||
withr::with_envvar(c(USCOGDATA_URL = sentinel_url), {
|
||||
msg <- tryCatch(uscogdata:::cog_open(), error = conditionMessage)
|
||||
expect_match(msg, "USCOGDATA_URL", fixed = TRUE)
|
||||
expect_match(msg, "uscogdata.url", fixed = TRUE)
|
||||
@@ -111,7 +115,7 @@ test_that("cog_manifest returns the active session's parsed manifest", {
|
||||
})
|
||||
})
|
||||
|
||||
test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others", {
|
||||
test_that(".validate_schema accepts schema_version 4 through 7, rejects others", {
|
||||
expect_silent(uscogdata:::.validate_schema(list(schema_version = 4L)))
|
||||
expect_silent(uscogdata:::.validate_schema(list(schema_version = 5L)))
|
||||
# v6 = FIPS geography harmonization (2026-07-22): _code -> _asof rename +
|
||||
@@ -119,12 +123,22 @@ test_that(".validate_schema accepts schema_version 4, 5 and 6, rejects others",
|
||||
# renamed columns and its geography comes from the xwalk, so v6 is accepted
|
||||
# without behavioural change -- see .validate_schema()'s note.
|
||||
expect_silent(uscogdata:::.validate_schema(list(schema_version = 6L)))
|
||||
# v7 = `data_year` APPENDED as column 29 (cog_pipeline #80, 2026-08-03), the
|
||||
# most recent fiscal year contributing to a collapsed key. Appended, never
|
||||
# inserted: canonical_govid stays at position 26, so nothing this package
|
||||
# reads shifts. Verified against the real v7 corpus before widening the
|
||||
# allow-list -- cog_spending()/cog_balances() return correctly for FY2024 AND
|
||||
# for FY2012, so the new column is inert here.
|
||||
expect_silent(uscogdata:::.validate_schema(list(schema_version = 7L)))
|
||||
expect_error(
|
||||
uscogdata:::.validate_schema(list(schema_version = 3L)),
|
||||
"schema_version"
|
||||
)
|
||||
# The upper bound still has to be ENFORCED, not just moved. Without this the
|
||||
# test would no longer prove that an unknown future schema is refused, and a
|
||||
# v8 corpus with a genuinely breaking change would sail through.
|
||||
expect_error(
|
||||
uscogdata:::.validate_schema(list(schema_version = 7L)),
|
||||
uscogdata:::.validate_schema(list(schema_version = 8L)),
|
||||
"schema_version"
|
||||
)
|
||||
})
|
||||
|
||||
@@ -13,10 +13,16 @@
|
||||
# is unaffected, so the fix is documentation: one sentence in @return.
|
||||
|
||||
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
|
||||
testthat::skip("Blocked on uscogdata#14 (finding F-021)")
|
||||
|
||||
rd <- paste(readLines(testthat::test_path("..", "..", "man", "cog_peer_compare.Rd"),
|
||||
warn = FALSE), collapse = " ")
|
||||
# man/ ships only in the source tree (the installed package carries a
|
||||
# compiled help database instead), so the prose assertions below cannot run
|
||||
# under R CMD check -- CI's earlier testthat::test_local() step enforces
|
||||
# them. The numeric pin further down needs only the corpus, but it lives in
|
||||
# the same test_that() as the sentence it protects, deliberately: they are
|
||||
# one claim, and splitting them would let the prose drift while a separate
|
||||
# test kept passing.
|
||||
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
|
||||
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
|
||||
|
||||
# The @return section must say the quantile is computed within each cell...
|
||||
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
|
||||
|
||||
@@ -214,3 +214,255 @@ test_that("no signposting under basis = 'raw'", {
|
||||
prov <- attr(r, "provenance")
|
||||
expect_length(prov$suggestions, 0L)
|
||||
})
|
||||
|
||||
# --- uscogdata#9: partial-coverage signposting ------------------------------
|
||||
|
||||
test_that("no recipe component is ever renamed by harmonization", {
|
||||
# The suppression trigger anti-joins the verb's long view on item_code.
|
||||
# That is only sound because harmonization never rewrites a recipe
|
||||
# component's code -- every component whose harmonized_code differs has
|
||||
# harmonized_code IS NULL (and is aggregate-flagged). If this ever fails,
|
||||
# .suppressed_components() would report reachable dollars as suppressed.
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
n <- DBI::dbGetQuery(con,
|
||||
"SELECT COUNT(*) AS renamed FROM long
|
||||
WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes)
|
||||
AND harmonized_code IS NOT NULL
|
||||
AND harmonized_code <> item_code")$renamed
|
||||
expect_equal(as.integer(n), 0L)
|
||||
})
|
||||
|
||||
test_that(".select_long_view maps annotated view bases to their long views", {
|
||||
expect_equal(
|
||||
uscogdata:::.select_long_view("spending_annotated", "harmonized"),
|
||||
"spending_long_harmonized")
|
||||
expect_equal(
|
||||
uscogdata:::.select_long_view("revenue_annotated", "harmonized"),
|
||||
"revenue_long_harmonized")
|
||||
expect_equal(
|
||||
uscogdata:::.select_long_view("spending_annotated", "raw"),
|
||||
"spending_long")
|
||||
})
|
||||
|
||||
test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", {
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
s <- uscogdata:::.suppressed_components(
|
||||
con,
|
||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
||||
govid = "061037123085", years = 2011L,
|
||||
long_view = "spending_long_harmonized",
|
||||
flow_prefixes = c("E", "F", "G"))
|
||||
|
||||
expect_s3_class(s, "tbl_df")
|
||||
expect_equal(nrow(s), 2L)
|
||||
s <- s[order(s$recipe_id), ]
|
||||
expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
|
||||
expect_equal(s$suppressed_amount, c(1803872000, 271589000))
|
||||
expect_equal(s$suppressed_codes, c("E67", "E68"))
|
||||
})
|
||||
|
||||
test_that(".suppressed_components finds nothing in a modern year", {
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
s <- uscogdata:::.suppressed_components(
|
||||
con,
|
||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
||||
govid = "061037123085", years = 2019L,
|
||||
long_view = "spending_long_harmonized",
|
||||
flow_prefixes = c("E", "F", "G"))
|
||||
expect_equal(nrow(s), 0L)
|
||||
})
|
||||
|
||||
test_that(".suppressed_components rejects a long_view outside the allowlist", {
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
expect_error(
|
||||
uscogdata:::.suppressed_components(
|
||||
con, candidates = "welfare_cash_e67_wide", govid = "061037123085",
|
||||
years = 2011L, long_view = "long; DROP TABLE x",
|
||||
flow_prefixes = c("E", "F", "G")),
|
||||
class = "uscogdata_internal_error")
|
||||
})
|
||||
|
||||
test_that(".suppressed_components never measures a component from the other flow family (I1)", {
|
||||
# uscogdata#9 review, finding I1: without the flow_prefixes filter, a
|
||||
# candidate recipe entirely outside the calling verb's own flow family is
|
||||
# ALWAYS absent from that verb's view (by construction), so it was always
|
||||
# reported as "suppressed" -- fabricating a dollar claim. E67/E68 are
|
||||
# Public Welfare EXPENDITURE codes; scoping the measurement to revenue's
|
||||
# own flow_prefixes must find nothing for them.
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
s <- uscogdata:::.suppressed_components(
|
||||
con,
|
||||
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
|
||||
govid = "061037123085", years = 2011L,
|
||||
long_view = "revenue_long_harmonized",
|
||||
flow_prefixes = c("T", "A", "U", "B", "C", "D"))
|
||||
expect_equal(nrow(s), 0L)
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", {
|
||||
# The bug: E74/E79 return rows for FY2011, so there is no row-absence gap,
|
||||
# so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were
|
||||
# dropped for being aggregate-published. LA County reports $3,185,943,000
|
||||
# and omits $2,075,461,000, a 39% understatement, silently.
|
||||
skip_if_no_corpus()
|
||||
r <- suppressMessages(
|
||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
|
||||
expect_length(sugg, 2L)
|
||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
||||
expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
|
||||
|
||||
e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]]
|
||||
expect_equal(e67$trigger, "suppressed_component")
|
||||
expect_equal(e67$suppressed_amount, 1803872000)
|
||||
expect_equal(e67$suppressed_years, 2011L)
|
||||
expect_equal(e67$suppressed_codes, "E67")
|
||||
expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'")
|
||||
|
||||
e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]]
|
||||
expect_equal(e68$trigger, "suppressed_component")
|
||||
expect_equal(e68$suppressed_amount, 271589000)
|
||||
expect_equal(e68$suppressed_codes, "E68")
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", {
|
||||
# Corrections is the case that already worked: zero rows in FY2011, so the
|
||||
# row-absence path fires. It must keep firing, keep trigger = "empty_year",
|
||||
# keep its IG counterpart -- and now also report what was suppressed.
|
||||
skip_if_no_corpus()
|
||||
r <- suppressMessages(
|
||||
cog_spending("061037123085", years = 2011L, category = "Corrections"))
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
|
||||
expect_length(sugg, 3L)
|
||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
||||
expect_setequal(ids, c("corrections_combined", "corrections_capital_combined",
|
||||
"corrections_other_capital_combined"))
|
||||
expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year"))
|
||||
|
||||
cc <- sugg[[which(ids == "corrections_combined")]]
|
||||
expect_equal(cc$suppressed_amount, 1371460000)
|
||||
expect_equal(cc$suppressed_codes, "E05")
|
||||
expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined")
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: the revenue verb inherits the same trigger", {
|
||||
# Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from
|
||||
# U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-`
|
||||
# rents and royalties -- the omission is LARGER than the reported figure.
|
||||
skip_if_no_corpus()
|
||||
r <- suppressMessages(
|
||||
cog_revenue("020000227749", years = 2011L,
|
||||
category = "Miscellaneous Revenue"))
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
|
||||
expect_length(sugg, 1L)
|
||||
expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide")
|
||||
expect_equal(sugg[[1]]$trigger, "suppressed_component")
|
||||
expect_equal(sugg[[1]]$suppressed_amount, 1899995000)
|
||||
expect_equal(sugg[[1]]$suppressed_codes, "U4-")
|
||||
# A revenue recipe must never be handed an M/L expenditure counterpart.
|
||||
expect_null(sugg[[1]]$ig_recipe_id)
|
||||
})
|
||||
|
||||
test_that("I1: cog_revenue never fabricates suppressed dollars for an expenditure-only recipe", {
|
||||
# uscogdata#9 review, finding I1: Corrections is an expenditure-only
|
||||
# category (E04/E05). cog_revenue() naturally returns zero rows for it, so
|
||||
# corrections_combined still fires as an empty_year suggestion (its own
|
||||
# generic join finds real E04/E05 data for this government) -- but before
|
||||
# the flow_prefixes fix, .suppressed_components() measured E04/E05 against
|
||||
# cog_revenue()'s OWN view (which can never contain an E-coded row by
|
||||
# construction) and reported the full $3,631,945,000 as "suppressed",
|
||||
# when cog_spending() for the same gov/years/category actually returns
|
||||
# $3,691,029,000 -- nothing was suppressed at all.
|
||||
skip_if_no_corpus()
|
||||
r <- suppressMessages(
|
||||
cog_revenue("061037123085", years = 2019:2020, category = "Corrections"))
|
||||
sugg <- attr(r, "provenance")$suggestions
|
||||
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
|
||||
expect_true("corrections_combined" %in% ids)
|
||||
|
||||
hit <- sugg[[which(ids == "corrections_combined")]]
|
||||
expect_equal(hit$suppressed_amount, 0)
|
||||
expect_equal(hit$suppressed_years, integer(0))
|
||||
expect_equal(hit$suppressed_codes, character(0))
|
||||
|
||||
# And cog_spending() for the identical gov/years/category is unaffected --
|
||||
# it actually finds the E04/E05 dollars the buggy measurement claimed were
|
||||
# excluded.
|
||||
sp <- suppressMessages(
|
||||
cog_spending("061037123085", years = 2019:2020, category = "Corrections"))
|
||||
expect_equal(sum(sp$amt_nominal), 3691029000)
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: no partial-coverage fire in a modern year", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare")
|
||||
expect_length(attr(r, "provenance")$suggestions, 0L)
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: leaf-and-classified wide-era families never fire", {
|
||||
# higher_ed_e18_wide and general_gov_e89_wide are the control group: their
|
||||
# components (E16/E18, E85/E89) are ordinary classified leaves even in the
|
||||
# wide era, so widening the trigger must leave them silent. This is the
|
||||
# measurement that refutes "it would fire on every category in every legacy
|
||||
# year" -- corpus-wide on the fixture, these two produce zero suppressed rows.
|
||||
skip_if_no_corpus()
|
||||
con <- uscogdata:::.ensure_session()
|
||||
n <- DBI::dbGetQuery(con,
|
||||
"SELECT COUNT(*) AS n
|
||||
FROM long l
|
||||
JOIN harmonization_recipes r
|
||||
ON l.item_code = r.component_code
|
||||
AND l.year BETWEEN r.year_min AND r.year_max
|
||||
WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide')
|
||||
AND l.amt <> 0
|
||||
AND NOT EXISTS (
|
||||
SELECT 1 FROM spending_long_harmonized v
|
||||
WHERE v.canonical_govid = l.canonical_govid
|
||||
AND v.year = l.year AND v.item_code = l.item_code)")$n
|
||||
expect_equal(as.integer(n), 0L)
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: the cli message reports the suppressed dollars", {
|
||||
skip_if_no_corpus()
|
||||
expect_message(
|
||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
||||
"1,803,872,000", fixed = TRUE)
|
||||
expect_message(
|
||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
||||
"FY2011", fixed = TRUE)
|
||||
expect_message(
|
||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
|
||||
"E67", fixed = TRUE)
|
||||
})
|
||||
|
||||
test_that("uscogdata#9: cog_explain() reports the suppressed dollars", {
|
||||
# cog_explain()'s whole "print" output -- including the Suggestions
|
||||
# section built from cli::cli_ul() -- is emitted on the message stream
|
||||
# (verified empirically 2026-08-04: capture.output(..., type = "output")
|
||||
# returns character(0) for this call; testthat::capture_messages() is what
|
||||
# actually carries it), so that is the stream this test captures.
|
||||
skip_if_no_corpus()
|
||||
r <- suppressMessages(
|
||||
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
|
||||
out <- paste(testthat::capture_messages(cog_explain(r)), collapse = "")
|
||||
expect_match(out, "271,589,000", fixed = TRUE)
|
||||
})
|
||||
|
||||
test_that("the provenance schema documents the suggestion trigger fields", {
|
||||
sch <- jsonlite::fromJSON(
|
||||
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
|
||||
simplifyVector = FALSE)
|
||||
props <- sch$properties$suggestions$items$properties
|
||||
expect_true(all(c("trigger", "suppressed_amount", "suppressed_years",
|
||||
"suppressed_codes") %in% names(props)))
|
||||
expect_setequal(unlist(props$trigger$enum),
|
||||
c("empty_year", "suppressed_component"))
|
||||
})
|
||||
|
||||
@@ -9,47 +9,131 @@
|
||||
# a published Census revenue concept exactly the way I89 sits inside Census's
|
||||
# Direct Expenditure concept (finding F-012).
|
||||
#
|
||||
# CAVEAT FOR WHOEVER PICKS THIS UP: the argument name below (`revenue_concept =
|
||||
# "total"`) is this test's *proposal*, not a settled decision. The owner's
|
||||
# 2026-07-28 resolution covers expenditure concepts only; no revenue-side
|
||||
# naming has been ruled on. If the eventual argument is named differently,
|
||||
# change the two calls here -- the asserted dollar invariants are what matter
|
||||
# and are independent of the naming.
|
||||
# RULED 2026-07-30. `revenue_concept = c("general", "total")` mirrors
|
||||
# `expenditure_concept`, and the two values are Census's two published revenue
|
||||
# concepts, related by the manual's own identity (section 4.3, which defines
|
||||
# the first by SUBTRACTING from the second):
|
||||
#
|
||||
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
||||
#
|
||||
# so `general` is the four general subtypes (own_source/federal/state/
|
||||
# local_aid) and `total` is every revenue subtype. Naming utility (A91-A94)
|
||||
# and liquor store (A90) separately is what makes BOTH computable -- before
|
||||
# cog_pipeline#79 they sat in own_source, so the default was really
|
||||
# "General + Utility + Liquor", a concept Census does not publish.
|
||||
#
|
||||
# Fixture reproducibility: Madison's own X-prefix revenue (FY1970-FY1986,
|
||||
# $15,098,000 nominal, $0 thereafter) is outside the bundled fixture's year
|
||||
# window (2011/2012/2019/2020), so the same invariant is asserted on Wisconsin
|
||||
# state government FY2012, where the fixture carries nonzero X01/X05/X08.
|
||||
# state government FY2012, where the fixture carries nonzero X01/X02/X05/X08.
|
||||
|
||||
test_that("cog_revenue() can return Census Total Revenue including Insurance Trust (prefix X)", {
|
||||
testthat::skip("Blocked on uscogdata#12 (finding F-014)")
|
||||
|
||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
||||
|
||||
# Revenue-shaped Employee Retirement codes, read from the RAW corpus rather
|
||||
# than through cog_revenue(), which is the filter under test:
|
||||
# X01 local employee contribution, X04/X05 contributions and transfers from
|
||||
# other governments, X08 earnings on investments.
|
||||
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X04", "X05", "X08"))
|
||||
expect_equal(x_revenue, 2038800) # 615,835 + 0 + 560,382 + 862,583 ($1,000s)
|
||||
# X01/X02 employee contributions, X05 contributions from other governments,
|
||||
# X08 total earnings on investments.
|
||||
#
|
||||
# X04 and X06 are deliberately NOT in this set, though an earlier draft of
|
||||
# this test included X04. Both are exhibit codes for INTRAgovernmental
|
||||
# transfers (the administering government paying into its own fund), which
|
||||
# X05's own definition excludes by name. Census agrees: its computed "Total
|
||||
# Emp Ret Rev" for this government-year is exactly the four codes below.
|
||||
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X02", "X05", "X08"))
|
||||
expect_equal(x_revenue, 2283883) # 615,835 + 245,083 + 560,382 + 862,583
|
||||
|
||||
# The Y-prefix insurance trust revenue (unemployment + workers comp), which
|
||||
# is the other half of the same Census concept.
|
||||
y_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("Y01", "Y11"))
|
||||
expect_equal(y_revenue, 1259785)
|
||||
|
||||
general <- cog_revenue(govid = wi_state, years = 2012L)
|
||||
expect_equal(attr(general, "provenance")$revenue_concept, "general")
|
||||
expect_equal(sum(general$amt_nominal), 31338293000)
|
||||
|
||||
total <- cog_revenue(govid = wi_state, years = 2012L, revenue_concept = "total")
|
||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal), x_revenue * 1000)
|
||||
expect_equal(sum(total$amt_nominal), 33377093000)
|
||||
expect_true(all(c("X01", "X05", "X08") %in% wt_codes_included(total)))
|
||||
expect_equal(attr(total, "provenance")$revenue_concept, "total")
|
||||
|
||||
# total - general is the whole insurance trust leg, X and Y together.
|
||||
# Asserted as a delta as well as a level so this stays correct however the
|
||||
# utility/liquor families land (both are $0 for WI state in FY2012).
|
||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
|
||||
(x_revenue + y_revenue) * 1000)
|
||||
expect_equal(sum(total$amt_nominal), 34881961000)
|
||||
expect_true(all(c("X01", "X02", "X05", "X08") %in% wt_codes_included(total)))
|
||||
|
||||
# Sibling codes under the SAME first letter must stay out: X11/X12 are
|
||||
# benefit payments (an expenditure) and X21/X30/X47 are cash and securities
|
||||
# holdings (a balance-sheet stock). This is the F-018 point restated on the
|
||||
# revenue side -- the split has to come from the crosswalk's spend_type, not
|
||||
# from the letter X.
|
||||
# revenue side -- the split comes from the crosswalk, not from the letter X.
|
||||
expect_false(any(c("X11", "X12", "X21", "X30", "X47") %in% wt_codes_included(total)))
|
||||
|
||||
# Every returned row still resolves to a category. summary_categories has
|
||||
# zero rows for prefix X today, so relaxing the prefix filter alone would
|
||||
# produce category = NA rows -- see census_of_governments_finance_pipeline#60.
|
||||
# Every returned row still resolves to a category (cog_pipeline#79 added the
|
||||
# X crosswalk rows; relaxing a prefix filter alone would have produced
|
||||
# category = NA rows).
|
||||
expect_false(any(is.na(total$category)))
|
||||
})
|
||||
|
||||
test_that("revenue_concept = 'general' is the default and is strict Census General Revenue", {
|
||||
wi_state <- "550000227544"
|
||||
default <- cog_revenue(govid = wi_state, years = 2012L)
|
||||
explicit <- cog_revenue(govid = wi_state, years = 2012L,
|
||||
revenue_concept = "general")
|
||||
expect_equal(sum(default$amt_nominal), sum(explicit$amt_nominal))
|
||||
|
||||
# General Revenue excludes utility, liquor store AND insurance trust
|
||||
# revenue. WI state carries $0 of utility/liquor in FY2012, so the level
|
||||
# assertion above cannot see those two -- assert the subtype scope directly.
|
||||
#
|
||||
# A subset, not setequal: `state` means "intergovernmental revenue FROM the
|
||||
# state government" (the C codes), which a STATE government does not receive
|
||||
# from itself, so it is legitimately absent here.
|
||||
expect_true(all(default$revenue_subtype %in%
|
||||
c("own_source", "federal", "state", "local_aid")))
|
||||
expect_false(any(c("utility", "liquor_store", "insurance_trust") %in%
|
||||
default$revenue_subtype))
|
||||
})
|
||||
|
||||
test_that("utility and liquor store revenue are inside `total` and outside `general`", {
|
||||
# A city, where utility revenue is material: this is the case the WI state
|
||||
# baseline structurally cannot exercise. Measured on the fixture, utility +
|
||||
# liquor is 15.9% of what cog_revenue() returned for type-2 governments
|
||||
# before the general/total split, so this is the largest behaviour change
|
||||
# the concept split introduces.
|
||||
con <- uscogdata:::.ensure_session()
|
||||
gov <- DBI::dbGetQuery(con,
|
||||
"SELECT canonical_govid, SUM(amt) amt FROM long
|
||||
WHERE year = 2012 AND type = 2 AND NOT is_aggregate
|
||||
AND item_code IN ('A91','A92','A93','A94')
|
||||
GROUP BY 1 ORDER BY amt DESC LIMIT 1")$canonical_govid
|
||||
|
||||
util_raw <- wt_raw_amt(gov, 2012L, codes = c("A90", "A91", "A92", "A93", "A94"))
|
||||
expect_gt(util_raw, 0)
|
||||
|
||||
general <- cog_revenue(govid = gov, years = 2012L)
|
||||
total <- cog_revenue(govid = gov, years = 2012L, revenue_concept = "total")
|
||||
|
||||
expect_false(any(c("utility", "liquor_store") %in% general$revenue_subtype))
|
||||
expect_true("utility" %in% total$revenue_subtype)
|
||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
|
||||
util_raw * 1000 +
|
||||
wt_raw_amt(gov, 2012L, codes = c("Y01", "Y11", "X01", "X02",
|
||||
"X05", "X08")) * 1000)
|
||||
})
|
||||
|
||||
test_that("revenue_concept rejects unknown values and never returns a balance row", {
|
||||
expect_error(
|
||||
cog_revenue("550000227544", years = 2012L, revenue_concept = "gross"),
|
||||
class = "uscogdata_invalid_revenue_concept"
|
||||
)
|
||||
|
||||
# uscogdata#25 restated for the widest revenue concept: stocks are not
|
||||
# flows, and `total` must not quietly admit the X/Y/W/Z balance families.
|
||||
con <- uscogdata:::.ensure_session()
|
||||
balance <- DBI::dbGetQuery(con,
|
||||
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
|
||||
total <- cog_revenue("550000227544", years = 2012L, revenue_concept = "total")
|
||||
expect_false(any(total$category %in% balance$category))
|
||||
expect_length(intersect(wt_codes_included(total), balance$item_code), 0L)
|
||||
})
|
||||
|
||||
@@ -0,0 +1,29 @@
|
||||
# Mirror of test-spending-pagination.R for cog_revenue(), which shares the
|
||||
# same .verb_spendrev()/.build_verb_sql() pushdown -- see that file for the
|
||||
# incident this fixes.
|
||||
|
||||
test_that("cog_revenue limit/offset page correctly and report total_rows", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_revenue("121011212191", years = 2019:2020, category = NULL)
|
||||
page <- cog_revenue("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = 5L, offset = 3L)
|
||||
expect_equal(nrow(page), 5L)
|
||||
expect_equal(page[c("year", "canonical_govid", "revenue_subtype", "category")],
|
||||
full[4:8, c("year", "canonical_govid", "revenue_subtype", "category")],
|
||||
ignore_attr = TRUE)
|
||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
||||
})
|
||||
|
||||
test_that("cog_revenue limit unset by default leaves total_rows absent", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_revenue("121011212191", 2020L, "Property Tax")
|
||||
expect_null(attr(r, "total_rows"))
|
||||
})
|
||||
|
||||
test_that("cog_revenue complete + limit conflict aborts the same way as cog_spending", {
|
||||
skip_if_no_corpus()
|
||||
expect_error(
|
||||
cog_revenue("121011212191", 2020L, "Property Tax", complete = TRUE, limit = 5L),
|
||||
class = "uscogdata_complete_pagination_conflict"
|
||||
)
|
||||
})
|
||||
@@ -80,8 +80,11 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
|
||||
|
||||
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
||||
skip_if_no_corpus()
|
||||
fl_state <- cog_gov_search("^FLORIDA$", type = "state")
|
||||
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||
# Unanchored: utility mode matches literally now, so "^...$" would be
|
||||
# searched for as characters rather than read as anchors (uscogdata#16).
|
||||
# Both still resolve to exactly one row once scoped by type/state.
|
||||
fl_state <- cog_gov_search("FLORIDA", type = "state")
|
||||
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||
r <- cog_geographic_rollup(
|
||||
govids = list(state = fl_state, county = broward),
|
||||
category = "Police", years = 2020L
|
||||
|
||||
@@ -0,0 +1,95 @@
|
||||
# cog-api's paginate() used to slice an already-fully-materialized result:
|
||||
# every page of a deep sweep re-ran the whole query and re-listified every
|
||||
# row, just to keep 1000 and discard the rest. For a 193,105-row fleet-wide
|
||||
# query walked 194 pages deep, that repeated the full cost 194 times and
|
||||
# wedged the production server for hours (2026-08-06 incident). limit/offset
|
||||
# here push the slice into the SQL itself, so a page costs O(limit), not
|
||||
# O(full result).
|
||||
|
||||
test_that("limit without offset returns the first page, matching the unpaginated head", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = 10L)
|
||||
expect_equal(nrow(page), 10L)
|
||||
expect_equal(page[c("year", "canonical_govid", "spend_subtype", "category")],
|
||||
full[1:10, c("year", "canonical_govid", "spend_subtype", "category")],
|
||||
ignore_attr = TRUE)
|
||||
})
|
||||
|
||||
test_that("offset skips ahead without gaps or overlap", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
||||
page2 <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = 10L, offset = 10L)
|
||||
expect_equal(nrow(page2), 10L)
|
||||
expect_equal(page2[c("year", "canonical_govid", "spend_subtype", "category")],
|
||||
full[11:20, c("year", "canonical_govid", "spend_subtype", "category")],
|
||||
ignore_attr = TRUE)
|
||||
})
|
||||
|
||||
test_that("walking every page reconstructs the unpaginated result exactly", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
||||
n <- nrow(full)
|
||||
limit <- 7L
|
||||
pages <- list()
|
||||
offset <- 0L
|
||||
repeat {
|
||||
p <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = limit, offset = offset)
|
||||
if (nrow(p) == 0L) break
|
||||
pages[[length(pages) + 1L]] <- p
|
||||
offset <- offset + limit
|
||||
if (offset > n + limit) stop("test runaway: paging did not terminate")
|
||||
}
|
||||
walked <- dplyr::bind_rows(pages)
|
||||
expect_equal(nrow(walked), n)
|
||||
key_cols <- c("year", "canonical_govid", "spend_subtype", "category", "amt_nominal")
|
||||
expect_equal(walked[key_cols], full[key_cols], ignore_attr = TRUE)
|
||||
})
|
||||
|
||||
test_that("total_rows attribute reports the full unpaginated count", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = 5L, offset = 0L)
|
||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
||||
})
|
||||
|
||||
test_that("offset past the end returns zero rows, not an error", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2019:2020, category = NULL)
|
||||
page <- cog_spending("121011212191", years = 2019:2020, category = NULL,
|
||||
limit = 10L, offset = nrow(full) + 100L)
|
||||
expect_equal(nrow(page), 0L)
|
||||
expect_equal(attr(page, "total_rows"), nrow(full))
|
||||
})
|
||||
|
||||
test_that("limit is unset by default -- unpaginated calls are unaffected", {
|
||||
skip_if_no_corpus()
|
||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
||||
expect_null(attr(r, "total_rows"))
|
||||
})
|
||||
|
||||
test_that("per_capita and adjust_to_year still apply correctly within a page", {
|
||||
skip_if_no_corpus()
|
||||
full <- cog_spending("121011212191", years = 2020L, category = NULL,
|
||||
per_capita = TRUE, adjust_to_year = 2022L)
|
||||
page <- cog_spending("121011212191", years = 2020L, category = NULL,
|
||||
per_capita = TRUE, adjust_to_year = 2022L,
|
||||
limit = 3L, offset = 2L)
|
||||
expect_equal(page[c("amt_nominal", "amt_real", "amt_per_capita_nominal",
|
||||
"amt_per_capita_real")],
|
||||
full[3:5, c("amt_nominal", "amt_real", "amt_per_capita_nominal",
|
||||
"amt_per_capita_real")],
|
||||
ignore_attr = TRUE)
|
||||
})
|
||||
|
||||
test_that("complete = TRUE with limit aborts -- pagination over a partial grid is undefined", {
|
||||
skip_if_no_corpus()
|
||||
expect_error(
|
||||
cog_spending("121011212191", 2020L, "Corrections", complete = TRUE, limit = 5L),
|
||||
class = "uscogdata_complete_pagination_conflict"
|
||||
)
|
||||
})
|
||||
@@ -98,7 +98,11 @@ test_that("cog_spending rejects invalid inputs", {
|
||||
|
||||
test_that("cog_spending accepts a cog_gov_search result directly", {
|
||||
skip_if_no_corpus()
|
||||
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||
# Unanchored: utility mode matches `name` as a literal substring now, so
|
||||
# "^...$" would be searched for as those characters rather than read as
|
||||
# anchors (uscogdata#16). Scoped by state and type, the bare name still
|
||||
# resolves to exactly one row.
|
||||
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
||||
expect_gt(nrow(picks), 0L)
|
||||
r <- cog_spending(picks, 2020L, "Corrections")
|
||||
expect_equal(unique(r$canonical_govid), "121011212191")
|
||||
@@ -244,24 +248,42 @@ test_that("basis defaults to 'harmonized' when not passed", {
|
||||
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
||||
# FL state government. The harmonization block is scoped by government,
|
||||
# year and flow prefix -- NOT by category -- so a Corrections query still
|
||||
# counts every E/F/G-prefixed row the harmonized basis drops for having
|
||||
# no harmonized_code. The three that apply here are E21/F21/G21
|
||||
# (Education NEC, SB184-186, "discontinued_na", wide-era window ending
|
||||
# FY2011); the other discontinued_na rulings live outside E/F/G.
|
||||
# See docs/phase_r_harmonization_review.md § 1.3/1.4 and cog_pipeline
|
||||
# data/harmonization_map.csv.
|
||||
r <- cog_spending("120000226351", 2011:2012, "Corrections")
|
||||
prov <- attr(r, "provenance")
|
||||
expect_equal(prov$basis, "harmonized")
|
||||
expect_true(prov$harmonization$applied)
|
||||
expect_true(prov$harmonization$na_rows_excluded >= 0L)
|
||||
expect_true(prov$harmonization$na_amount_excluded >= 0)
|
||||
# Data-verified for the v6 fixture (corpus 2026-07-22). The Task 18 map
|
||||
# extension added E/F/G-prefix discontinued_na rulings the earlier pin's
|
||||
# comment predated: E21/F21/G21 (Education NEC local, SB184-186,
|
||||
# "trivial; explicit-NA, full wide-era window"). Broward's 2011 legacy
|
||||
# partition zero-pads exactly those three codes, so this query now
|
||||
# excludes 3 NA-harmonized rows -- all with amt = 0, hence the excluded
|
||||
# AMOUNT stays exactly zero. (The other discontinued_na rulings -- S74,
|
||||
# Z61, X04, X06, the debt-detail family, L24 -- remain outside the
|
||||
# E/F/G/K prefixes.) See docs/phase_r_harmonization_review.md § 1.3/1.4
|
||||
# and cog_pipeline data/harmonization_map.csv E21/F21/G21 rows.
|
||||
expect_equal(prov$harmonization$na_rows_excluded, 3L)
|
||||
expect_equal(prov$harmonization$na_amount_excluded, 0)
|
||||
# $2,825,439 thousands of FY2011 E21 + F21 + G21, reported in full USD.
|
||||
# Pinning a non-zero amount is the point: the earlier Broward anchor's
|
||||
# three rows were all explicit zeros, so the AMOUNT accounting was
|
||||
# asserted only against 0 and could not have caught a bug.
|
||||
expect_equal(prov$harmonization$na_amount_excluded, 2825439 * 1000)
|
||||
})
|
||||
})
|
||||
|
||||
test_that("sparsification removed the wide era's zero-pads from the exclusion count", {
|
||||
skip_if_no_corpus()
|
||||
with_fixture_corpus({
|
||||
# Broward County FY2011 used to carry E21/F21/G21 rows of exactly $0 --
|
||||
# the wide era stored every government x every code, zeros included. The
|
||||
# published corpus no longer does (SB194, cog_pipeline#64), so there is
|
||||
# now nothing for the harmonized basis to exclude. Absence in a
|
||||
# dense_source year means Census published $0; it does not mean the
|
||||
# exclusion machinery stopped working, which the FL state anchor above
|
||||
# proves independently.
|
||||
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
||||
h <- attr(r, "provenance")$harmonization
|
||||
expect_true(h$applied)
|
||||
expect_equal(h$na_rows_excluded, 0L)
|
||||
expect_equal(h$na_amount_excluded, 0)
|
||||
})
|
||||
})
|
||||
|
||||
@@ -303,11 +325,13 @@ test_that("provenance$series_break_refs is a populated-when-applicable character
|
||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
||||
refs <- attr(r, "provenance")$series_break_refs
|
||||
expect_type(refs, "character")
|
||||
# No catalogued series_breaks_pq row falls inside this fixture's
|
||||
# 2011/2012/2019/2020 window for the codes this query touches (E04/G04)
|
||||
# -- data-verified; the mechanism itself is what's under test here, via
|
||||
# a query-shaped unit test in test-views.R since the fixture has no
|
||||
# positive case to pin against.
|
||||
# No catalogued code-specific series_breaks_pq row falls inside this
|
||||
# fixture's 2011/2012/2019/2020 window for the codes this query touches
|
||||
# (E04/G04) -- data-verified; the mechanism itself is what's under test
|
||||
# here, via a query-shaped unit test in test-views.R since the fixture
|
||||
# has no positive case to pin against. Corpus-wide ("ALL") entries never
|
||||
# appear in this field by construction -- they travel in
|
||||
# corpus_break_refs; see test-corpus-breaks.R.
|
||||
expect_equal(refs, character(0))
|
||||
})
|
||||
})
|
||||
|
||||
+108
-35
@@ -36,8 +36,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
||||
# {url} exactly as .register_views() does, and executes them -- plus
|
||||
# their 10-long.sql dependency -- against a synthetic hive-partitioned
|
||||
# parquet tree written to a temp dir. A regression in any predicate (e.g.
|
||||
# `NOT is_aggregate` dropped, the prefix list changed, the NULL guard
|
||||
# removed) would change which of the rows below survive.
|
||||
# `NOT is_aggregate` dropped, the crosswalk-membership subquery changed,
|
||||
# the NULL guard removed) would change which of the rows below survive.
|
||||
#
|
||||
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
|
||||
# PARQUET) rather than the arrow package: this package has no arrow
|
||||
@@ -61,34 +61,55 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
||||
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
|
||||
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
|
||||
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
||||
-- 'S74' is outside BOTH flow families (E/F/G/K spending and
|
||||
-- T/A/U/B/C/D revenue -- it mirrors the real corpus's own
|
||||
-- non-flow-type codes like S74/Z61), so it can only leak into
|
||||
-- EITHER view via the E/F/G/K or T/A/U/B/C/D prefix filter, never
|
||||
-- both at once -- a prefix drawn from the other view's own family
|
||||
-- (e.g. a real T-code for the spending row) would incorrectly
|
||||
-- leak into the other view's assertion below and not discriminate
|
||||
-- the predicate under test.
|
||||
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by the E/F/G/K prefix filter
|
||||
-- Revenue (T/A/U/B/C/D) rows, exercised against revenue_long_harmonized:
|
||||
-- 'S74' and 'Z61' are classified `balance` in the synthetic
|
||||
-- crosswalk below (mirroring the real corpus's own non-flow codes),
|
||||
-- so each is excluded from its view ONLY by the crosswalk-membership
|
||||
-- subquery -- the mechanism that replaced the prefix allowlists
|
||||
-- (uscogdata#11) and keeps balance stocks out of both flows
|
||||
-- (uscogdata#25).
|
||||
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by crosswalk membership (balance)
|
||||
-- Revenue rows, exercised against revenue_long_harmonized:
|
||||
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
|
||||
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
|
||||
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
|
||||
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
||||
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by the T/A/U/B/C/D prefix filter
|
||||
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by crosswalk membership (balance)
|
||||
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(part_path)))
|
||||
|
||||
# The flow views classify by membership in summary_categories, so the
|
||||
# synthetic corpus needs one too. Every flow code above is a member of its
|
||||
# own flow (so is_aggregate / NULL-harmonized exclusions stay the SOLE
|
||||
# excluder for those rows); S74/Z61 are members but classified balance, so
|
||||
# membership itself is what excludes them.
|
||||
DBI::dbExecute(write_con, sprintf("
|
||||
COPY (
|
||||
SELECT * FROM (VALUES
|
||||
('E36', 'Water Utilities', 'expenditure', 'operations', NULL),
|
||||
('E38', 'Water Utilities', 'expenditure', 'operations', NULL),
|
||||
('E05', 'Corrections', 'expenditure', 'operations', NULL),
|
||||
('E99', 'Other', 'expenditure', 'operations', NULL),
|
||||
('S74', 'Fund Balances', 'balance', NULL, NULL),
|
||||
('U11', 'Interest Earnings','revenue', NULL, 'own_source'),
|
||||
('U10', 'Interest Earnings','revenue', NULL, 'own_source'),
|
||||
('T29', 'Other Taxes', 'revenue', NULL, 'own_source'),
|
||||
('T88', 'Other Taxes', 'revenue', NULL, 'own_source'),
|
||||
('Z61', 'Fund Balances', 'balance', NULL, NULL)
|
||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
||||
|
||||
sql_dir <- system.file("sql", package = "uscogdata")
|
||||
.read_view_sql <- function(filename) {
|
||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
||||
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
|
||||
uscogdata:::.render_view_sql(txt, paste0(tmp, "/"))
|
||||
}
|
||||
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
|
||||
|
||||
@@ -97,8 +118,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
||||
GROUP BY item_code ORDER BY item_code"
|
||||
)
|
||||
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
|
||||
# harmonized_code), and spend-E (wrong prefix family) must all be gone,
|
||||
# and spend-A + spend-B must be folded together under E36.
|
||||
# harmonized_code), and spend-E (balance, not an expenditure member) must
|
||||
# all be gone, and spend-A + spend-B must be folded together under E36.
|
||||
expect_equal(nrow(spend), 1L)
|
||||
expect_equal(spend$item_code, "E36")
|
||||
expect_equal(spend$amt, 150)
|
||||
@@ -138,21 +159,37 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
|
||||
('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is
|
||||
('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis
|
||||
('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views
|
||||
('ig-D', 'L--', 55555, false, 'L--'), -- family total: excluded from BOTH views
|
||||
('ig-E', 'T29', 44444, false, 'T29') -- wrong prefix (revenue, not M/L): excluded from BOTH views
|
||||
('ig-D', 'L--', 55555, false, 'L--'), -- family total: deliberately NOT a crosswalk member, excluded from BOTH views
|
||||
('ig-E', 'T29', 44444, false, 'T29') -- revenue member, not intergovernmental: excluded from BOTH views
|
||||
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(part_path)))
|
||||
|
||||
# The IG views classify by summary_categories membership
|
||||
# (spend_subtype = 'intergovernmental'). L-- is deliberately absent --
|
||||
# exactly as it is from the real crosswalk -- which is what excludes it.
|
||||
DBI::dbExecute(write_con, sprintf("
|
||||
COPY (
|
||||
SELECT * FROM (VALUES
|
||||
('M04', 'Corrections', 'expenditure', 'intergovernmental', NULL),
|
||||
('M38', 'Health', 'expenditure', 'intergovernmental', NULL),
|
||||
('M36', 'Health', 'expenditure', 'intergovernmental', NULL),
|
||||
('M47', 'IG Other', 'expenditure', 'intergovernmental', NULL),
|
||||
('T29', 'Other Taxes', 'revenue', NULL, 'own_source')
|
||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
|
||||
) TO %s (FORMAT PARQUET)
|
||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
||||
|
||||
sql_dir <- system.file("sql", package = "uscogdata")
|
||||
.read_view_sql <- function(filename) {
|
||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
||||
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
|
||||
uscogdata:::.render_view_sql(txt, paste0(tmp, "/"))
|
||||
}
|
||||
|
||||
con <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("24-ig_long.sql"))
|
||||
DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql"))
|
||||
|
||||
@@ -160,8 +197,9 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
|
||||
"SELECT item_code, SUM(amt) AS amt FROM ig_long
|
||||
GROUP BY item_code ORDER BY item_code"
|
||||
)
|
||||
# L-- (family total) and T29 (wrong prefix) are gone; the aggregate row
|
||||
# M47 survives -- proof `NOT is_aggregate` is absent from ig_long.
|
||||
# L-- (family total, not a member) and T29 (revenue, not IG) are gone; the
|
||||
# aggregate row M47 survives -- proof `NOT is_aggregate` is absent from
|
||||
# ig_long.
|
||||
expect_equal(raw$item_code, c("M04", "M38", "M47"))
|
||||
expect_equal(raw$amt, c(100, 50, 99999))
|
||||
|
||||
@@ -177,11 +215,13 @@ test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmon
|
||||
})
|
||||
|
||||
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
||||
# No series_breaks_pq row falls inside the bundled fixture's 2011-2020
|
||||
# window (data-verified; see the "series_break_refs" test in
|
||||
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
|
||||
# 2011-2020 window (data-verified; see the "series_break_refs" test in
|
||||
# test-spending.R), so this proves the matching logic itself against the
|
||||
# live view + a synthetic year window that DOES hit a cataloged break
|
||||
# (SB075, fin_code E62, break_year 2005).
|
||||
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
|
||||
# separate path with its own coverage -- SB194 does sit at 2012, inside
|
||||
# the fixture window; see test-corpus-breaks.R.
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
@@ -295,7 +335,7 @@ test_that(".harmonization_view_files guard is necessary: registration against a
|
||||
sql_dir <- system.file("sql", package = "uscogdata")
|
||||
.read_view_sql <- function(filename) {
|
||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
||||
gsub("\\{url\\}", url, txt, fixed = FALSE)
|
||||
uscogdata:::.render_view_sql(txt, url)
|
||||
}
|
||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
||||
@@ -320,14 +360,32 @@ test_that(".harmonization_view_files guard is necessary: registration against a
|
||||
)
|
||||
})
|
||||
|
||||
test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
|
||||
test_that("spending_long carries exactly the non-IG expenditure crosswalk codes and excludes aggregates", {
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
prefixes <- DBI::dbGetQuery(con,
|
||||
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM spending_long"
|
||||
)$pfx
|
||||
expect_true(all(prefixes %in% c("E", "F", "G", "K")))
|
||||
|
||||
# Classification is crosswalk membership, not prefixes (uscogdata#11):
|
||||
# every row's code must classify as expenditure and never as
|
||||
# intergovernmental (which lives in ig_long).
|
||||
stray <- DBI::dbGetQuery(con,
|
||||
"SELECT DISTINCT s.item_code
|
||||
FROM spending_long s
|
||||
LEFT JOIN summary_categories c USING (item_code)
|
||||
WHERE c.category_type IS DISTINCT FROM 'expenditure'
|
||||
OR c.spend_subtype = 'intergovernmental'"
|
||||
)$item_code
|
||||
expect_length(stray, 0L)
|
||||
|
||||
# Balance codes are stocks, not flows -- they must never appear in a
|
||||
# spending result (uscogdata#25). Prefix filtering could not guarantee
|
||||
# this (W/X/Y/Z balance codes share letters with flow codes).
|
||||
balance_n <- DBI::dbGetQuery(con,
|
||||
"SELECT count(*) AS n FROM spending_long WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
|
||||
)"
|
||||
)$n
|
||||
expect_equal(balance_n, 0)
|
||||
|
||||
agg_count <- DBI::dbGetQuery(con,
|
||||
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
|
||||
@@ -335,14 +393,29 @@ test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
|
||||
expect_equal(agg_count, 0)
|
||||
})
|
||||
|
||||
test_that("revenue_long filters to T/A/U/B/C/D prefixes and excludes aggregates", {
|
||||
test_that("revenue_long carries exactly the revenue crosswalk codes and excludes aggregates", {
|
||||
skip_if_no_corpus()
|
||||
con <- cog_open()
|
||||
on.exit(cog_close())
|
||||
prefixes <- DBI::dbGetQuery(con,
|
||||
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM revenue_long"
|
||||
)$pfx
|
||||
expect_true(all(prefixes %in% c("T", "A", "U", "B", "C", "D")))
|
||||
|
||||
# The view carries EVERY revenue subtype; which of Census's two published
|
||||
# concepts a query returns is decided per `revenue_concept` in R
|
||||
# (uscogdata#12), exactly as `expenditure_concept` narrows spending_long.
|
||||
stray <- DBI::dbGetQuery(con,
|
||||
"SELECT DISTINCT s.item_code
|
||||
FROM revenue_long s
|
||||
LEFT JOIN summary_categories c USING (item_code)
|
||||
WHERE c.category_type IS DISTINCT FROM 'revenue'"
|
||||
)$item_code
|
||||
expect_length(stray, 0L)
|
||||
|
||||
# No balance stock ever appears in a revenue result (uscogdata#25).
|
||||
balance_n <- DBI::dbGetQuery(con,
|
||||
"SELECT count(*) AS n FROM revenue_long WHERE item_code IN (
|
||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
|
||||
)"
|
||||
)$n
|
||||
expect_equal(balance_n, 0)
|
||||
|
||||
agg_count <- DBI::dbGetQuery(con,
|
||||
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
|
||||
|
||||
@@ -13,6 +13,8 @@ knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
|
||||
|
||||
# Why per-year population matters
|
||||
|
||||
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
|
||||
|
||||
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
||||
|
||||
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
||||
|
||||
+101
-74
@@ -1,8 +1,8 @@
|
||||
---
|
||||
title: "Total spending: Direct, Total, and when each is right"
|
||||
title: "Total spending: Primary, Direct, Total, and when each is right"
|
||||
output: rmarkdown::html_vignette
|
||||
vignette: >
|
||||
%\VignetteIndexEntry{Total spending: Direct, Total, and when each is right}
|
||||
%\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right}
|
||||
%\VignetteEngine{knitr::rmarkdown}
|
||||
%\VignetteEncoding{UTF-8}
|
||||
---
|
||||
@@ -17,17 +17,35 @@ knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
|
||||
is about one government or several:
|
||||
|
||||
1. **"What did my county spend in total, a decade ago vs today?"** — one
|
||||
government, tracked over time. Either `direct` or `total` spending answers
|
||||
this correctly, as long as the same concept is used for both years.
|
||||
government, tracked over time. Any concept answers this correctly, as
|
||||
long as the same concept is used for both years.
|
||||
2. **"How do all the counties in my state compare, a decade ago vs today,
|
||||
against the neighboring state?"** — several governments, summed together.
|
||||
Here only `direct` gives the right answer; summing `total` across
|
||||
governments double-counts money that passes between them.
|
||||
Here only a non-intergovernmental concept (`primary` or `direct`) gives
|
||||
the right answer; summing `total` across governments double-counts money
|
||||
that passes between them.
|
||||
|
||||
`cog_spending()`'s `expenditure_concept` argument (`"direct"` or `"total"`)
|
||||
controls which of these a query answers. This vignette walks through both
|
||||
questions with code that actually runs against the package's bundled fixture
|
||||
corpus, then explains why the second question refuses `"total"` outright.
|
||||
`cog_spending()`'s `expenditure_concept` argument controls which of these a
|
||||
query answers, via three nested concepts defined as sets of the crosswalk's
|
||||
`spend_subtype` values (never item-code first letters — the letter `Y` alone
|
||||
spans revenue, expenditure, and balance codes):
|
||||
|
||||
- `"primary"` (the default) — the government's own service provision:
|
||||
`operations` + `capital` + `assistance`.
|
||||
- `"direct"` — Census's published Direct Expenditure: `primary` plus
|
||||
`interest` on debt and `insurance_benefits` (e.g. pension payments).
|
||||
- `"total"` — `direct` plus the `intergovernmental` leg.
|
||||
|
||||
This vignette walks through both questions with code that actually runs
|
||||
against the package's bundled fixture corpus, then explains why the second
|
||||
question refuses `"total"` outright.
|
||||
|
||||
Before any of the numbers below: every amount column here — `amt_nominal`,
|
||||
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
|
||||
dollars**. The raw Census files report **thousands of dollars** and the
|
||||
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
|
||||
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
|
||||
billion. Do not scale it again.
|
||||
|
||||
```{r}
|
||||
library(uscogdata)
|
||||
@@ -61,20 +79,23 @@ al_total <- cog_spending(
|
||||
al_total
|
||||
```
|
||||
|
||||
The `intergovernmental` rows are what `"total"` adds on top of `"direct"`
|
||||
(`capital` + `operations`): Alabama's own payments out to counties and
|
||||
cities for highway work. Because this query only ever concerns Alabama,
|
||||
including that piece is safe -- there's no other government's number it
|
||||
could be double-counted against.
|
||||
The `intergovernmental` rows are what `"total"` adds on top of the
|
||||
non-intergovernmental subtypes (here `capital` + `operations`): Alabama's
|
||||
own payments out to counties and cities for highway work. Because this
|
||||
query only ever concerns Alabama, including that piece is safe -- there's
|
||||
no other government's number it could be double-counted against.
|
||||
|
||||
`"direct"` (the default) answers the same trend question just as validly:
|
||||
`"primary"` (the default) answers the same trend question just as validly
|
||||
(for Highways, which maps only to operations/capital codes, `"primary"` and
|
||||
`"direct"` coincide -- there is no highway-specific interest or insurance
|
||||
benefit to add):
|
||||
|
||||
```{r}
|
||||
al_direct <- cog_spending(
|
||||
al_primary <- cog_spending(
|
||||
"010000226085", years = c(2012, 2020), category = "Highways"
|
||||
# expenditure_concept = "direct" is the default; shown here for contrast
|
||||
# expenditure_concept = "primary" is the default; shown here for contrast
|
||||
)
|
||||
al_direct
|
||||
al_primary
|
||||
```
|
||||
|
||||
Both are internally consistent series. What breaks the comparison is
|
||||
@@ -86,8 +107,8 @@ every year in the series.
|
||||
# Archetype 2: a cross-government rollup
|
||||
|
||||
`cog_geographic_rollup()` sums spending across state/county/city layers for
|
||||
a place. Its default -- and, as shown below, its *only* accepted value for
|
||||
`expenditure_concept` -- is `"direct"`:
|
||||
a place. Its default is `"primary"`, and (as shown below) it accepts only
|
||||
the non-intergovernmental concepts, `"primary"` and `"direct"`:
|
||||
|
||||
```{r}
|
||||
fl_rollup <- cog_geographic_rollup(
|
||||
@@ -138,15 +159,15 @@ shows up **twice** in the underlying corpus:
|
||||
the county is the government that actually lets the contract and pays the
|
||||
paving crew.
|
||||
|
||||
`direct` (item codes `E`/`F`/`G`) only ever counts the second of those --
|
||||
the government that actually did the spending. `total` (Direct plus the
|
||||
`M`/`L` intergovernmental legs) counts the first one *as well*, which is
|
||||
exactly right for describing Alabama's own budget: Alabama's `total`
|
||||
genuinely includes the $10M it committed to highways, whether it built the
|
||||
road itself or paid the county to. But sum `total` across Alabama **and**
|
||||
the county, and that $10M is counted twice -- once as Alabama's payment out,
|
||||
once as the county's spending in -- reporting $20M of highway work for $10M
|
||||
actually spent.
|
||||
`primary` and `direct` (the crosswalk's non-intergovernmental expenditure
|
||||
subtypes) only ever count the second of those -- the government that
|
||||
actually did the spending. `total` (Direct plus the intergovernmental leg)
|
||||
counts the first one *as well*, which is exactly right for describing
|
||||
Alabama's own budget: Alabama's `total` genuinely includes the $10M it
|
||||
committed to highways, whether it built the road itself or paid the county
|
||||
to. But sum `total` across Alabama **and** the county, and that $10M is
|
||||
counted twice -- once as Alabama's payment out, once as the county's
|
||||
spending in -- reporting $20M of highway work for $10M actually spent.
|
||||
|
||||
This is exactly the shape of query `cog_geographic_rollup()` exists to run
|
||||
(summing across layers of government), so it refuses `"total"` rather than
|
||||
@@ -161,48 +182,51 @@ share of a government's own Direct spending is:
|
||||
|
||||
| Government type | Intergovernmental / Direct |
|
||||
|---|---|
|
||||
| State | 16.7%-48.4% (varies by year; 24.0% pooled across all four) |
|
||||
| County | 3.4%-5.1% (varies by year) |
|
||||
| City | 2.6%-3.1% (varies by year) |
|
||||
| State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) |
|
||||
| County | 3.3%-4.8% (varies by year) |
|
||||
| City | 2.4%-2.9% (varies by year) |
|
||||
|
||||
So the Direct/Total choice matters overwhelmingly for **state** governments
|
||||
-- a state's Total genuinely differs from its Direct by a meaningful margin,
|
||||
while for a county or city the two are close. The state range is also far
|
||||
wider than a single flat figure would suggest: legacy wide-era years (2011:
|
||||
48.4%) carry proportionally more intergovernmental spending than the modern
|
||||
era (2019-2020: 16.7%-17.0%), so a state's Direct/Total gap can be nearly
|
||||
3x larger a decade earlier than it is today. That's also why the mistake
|
||||
this vignette warns about is easy to make unnoticed at the county/city level
|
||||
and costly at the state level: rolling up every government in a state using
|
||||
`total` instead of `direct` overstates the true figure -- measured at 7.6%
|
||||
for Alabama in FY2019, and 11.6% nationally.
|
||||
So the Direct/Total choice matters overwhelmingly for **state**
|
||||
governments -- a state's Total genuinely differs from its Direct by more
|
||||
than a third, while for a county or city the two are close. (The state
|
||||
share is much larger than pre-#11 measurements suggested, because the
|
||||
intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state
|
||||
payments to school systems -- for most states the single largest transfer
|
||||
they make.) That's also why the mistake this vignette warns about is easy
|
||||
to make unnoticed at the county/city level and costly at the state level:
|
||||
rolling up every government using `total` instead of `primary`/`direct`
|
||||
overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally.
|
||||
|
||||
# Why Total = Direct + M + L, not Direct + M
|
||||
# Why Total = Direct + M + L + Q, not Direct + M
|
||||
|
||||
It's tempting to assume `total` only needs to add `M`. But `M` and `L` are
|
||||
both money the queried government itself pays **out** -- they're not two
|
||||
different accounts of a receiving government's revenue. `M` is what it
|
||||
pays to other **local** governments (e.g. a county paying a city for a
|
||||
shared paving contract); `L` is what it pays **up** to its **state**
|
||||
government (e.g. a county's contribution to a state-administered program).
|
||||
A local government's Total genuinely includes both legs, because both are
|
||||
its own spending, just routed to a different kind of recipient. On the
|
||||
bundled fixture corpus (all 50 states, 2011/2012/2019/2020), `L` is 0 for
|
||||
state governments (a state has no "payments to the state government" leg of
|
||||
its own) but is 43%-51% the size of `M` for counties (varies by year) and
|
||||
144%-189% the size of `M` for cities (varies by year; 166% pooled across
|
||||
all four) -- so a `total` that omitted `L` would silently undercount Total
|
||||
specifically for local governments, and for cities `L` is often the
|
||||
*larger* of the two legs.
|
||||
`cog_spending(expenditure_concept = "total")` includes both legs (excluding
|
||||
the `L--` family-total rollup row, which would double-count its own
|
||||
components).
|
||||
It's tempting to assume `total` only needs to add `M`. But the
|
||||
intergovernmental leg has three families, all money the queried government
|
||||
itself pays **out** -- they're not different accounts of a receiving
|
||||
government's revenue. `M` is what it pays to other **local** governments
|
||||
(e.g. a county paying a city for a shared paving contract); `L` is what it
|
||||
pays **up** to its **state** government (e.g. a county's contribution to a
|
||||
state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments
|
||||
to **school systems** (K-12 and higher-ed aid -- for most states the
|
||||
single largest transfer they make, and the piece the pre-#11 prefix
|
||||
allowlist silently dropped, finding F-017). A government's Total genuinely
|
||||
includes every leg it pays, because each is its own spending, just routed
|
||||
to a different kind of recipient. On the bundled fixture corpus (all 50
|
||||
states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has
|
||||
no "payments to the state government" leg of its own) but is 43%-51% the
|
||||
size of `M` for counties (varies by year) and 144%-189% the size of `M`
|
||||
for cities (varies by year; 166% pooled across all four) -- so a `total`
|
||||
that omitted `L` would silently undercount Total specifically for local
|
||||
governments, and for cities `L` is often the *larger* of the two legs.
|
||||
`cog_spending(expenditure_concept = "total")` includes every leg
|
||||
(excluding the `L--` family-total rollup row, which would double-count its
|
||||
own components).
|
||||
|
||||
# Composition rules
|
||||
|
||||
- `expenditure_concept` (whose spending counts -- Direct vs Direct plus
|
||||
intergovernmental) is **orthogonal** to `basis` (which vintage of the
|
||||
item-code space a query resolves against -- `"harmonized"` vs `"raw"`).
|
||||
- `expenditure_concept` (whose spending counts -- Primary, Direct, or
|
||||
Direct plus intergovernmental) is **orthogonal** to `basis` (which
|
||||
vintage of the item-code space a query resolves against --
|
||||
`"harmonized"` vs `"raw"`).
|
||||
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
|
||||
valid, meaningful query, and so is every other pairing.
|
||||
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
|
||||
@@ -218,12 +242,15 @@ components).
|
||||
|
||||
# Summary
|
||||
|
||||
- Comparing one government to itself over time: `"direct"` or `"total"`
|
||||
both work -- pick one and hold it fixed across every year compared.
|
||||
- Comparing one government to itself over time: any concept works -- pick
|
||||
one and hold it fixed across every year compared.
|
||||
- Comparing or summing across governments -- counties within a state, a
|
||||
state against its neighbor, cities against counties: use `"direct"`.
|
||||
`cog_geographic_rollup()` and `cog_peer_compare()` enforce this by
|
||||
refusing `"total"`.
|
||||
- `"total"` = Direct (`E`/`F`/`G`) + intergovernmental (`M` to local
|
||||
governments + `L` to the state government, excluding the `L--`
|
||||
family-total row).
|
||||
state against its neighbor, cities against counties: use `"primary"`
|
||||
(the default) or `"direct"`. `cog_geographic_rollup()` and
|
||||
`cog_peer_compare()` enforce this by refusing `"total"`.
|
||||
- `"primary"` = operations + capital + assistance. `"direct"` = primary +
|
||||
interest on debt + insurance trust benefits (Census's published Direct
|
||||
Expenditure). `"total"` = direct + intergovernmental (`M` to local
|
||||
governments, `L` to the state government excluding the `L--`
|
||||
family-total row, and `Q11`/`Q12`/`Q18` state payments to school
|
||||
systems).
|
||||
|
||||
Reference in New Issue
Block a user