docs: rewrite README for a stranger
Reordered around a new user: what the data is, where it comes from, install, a quickstart that runs with no configuration, then the full-dollars warning and the concepts that decide whether a published number is right. Adds a 'Where the data comes from' section linking the API documentation site, the live API, the Hugging Face corpus and the Census source, so attribution and provenance are reachable from the top rather than implied. Drops the sibling-repo path, the commented-out install line, the Status block, and the release advice telling you to strip the fixture -- which would break the vignette and leave public CI unable to check without credentials. The quickstart passes years=; cog_spending() has no full-history default, so the obvious one-liner errors on a reader's first call.
This commit is contained in:
@@ -1,24 +1,116 @@
|
|||||||
# uscogdata
|
# uscogdata
|
||||||
|
|
||||||
Curated R reader for the Civilytics US Census of Governments finance corpus.
|
<!-- badges: start -->
|
||||||
|
[](https://civilytics.r-universe.dev/uscogdata)
|
||||||
|
[](LICENSE.md)
|
||||||
|
<!-- badges: end -->
|
||||||
|
|
||||||
Provides unit-level financial profiles, geographic rollups, and peer comparisons
|
A curated R reader for the Civilytics US Census of Governments finance corpus —
|
||||||
with auditable provenance and built-in cross-vintage correctness. Reads the
|
every dollar that US state, county, municipal and township governments reported
|
||||||
published corpus (Hive-partitioned parquet + manifest.json) directly from
|
raising and spending, from **FY1967 to FY2024**, in one queryable place.
|
||||||
Nextcloud via DuckDB httpfs — no local bulk downloads required.
|
|
||||||
|
|
||||||
## Status
|
The Census of Governments is the only nationwide source for local government
|
||||||
|
finance, and it is hard to use: item codes change meaning across vintages,
|
||||||
|
government identifiers were renumbered in 2017, and an absent value means
|
||||||
|
"published zero" in one era and "not reported" in the next. This package
|
||||||
|
handles each of those problems, and it tells you when it has — every result
|
||||||
|
carries provenance describing what was converted, what was aggregated, and
|
||||||
|
which known series breaks intersect your query.
|
||||||
|
|
||||||
Under active development (Phase 2 of the cog_pipeline project). See
|
**Scope:** government types 0–3 (state, county, municipality, township).
|
||||||
`../cog_pipeline/docs/reader-specification.md` for the reader contract this
|
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
|
||||||
package implements.
|
or FY1969. Special districts (type 4) and school districts (type 5) are
|
||||||
|
excluded pending validation.
|
||||||
|
|
||||||
## Installation
|
## Where the data comes from
|
||||||
|
|
||||||
|
The corpus is published and documented at the **[US Census of Governments
|
||||||
|
Finance API](https://pages.civilytics.org/cog-api/)**. Start there for how the
|
||||||
|
data was built, how the identifier and item-code reconciliation works, and what
|
||||||
|
the corpus does and does not cover.
|
||||||
|
|
||||||
|
- **[API documentation and walkthroughs](https://pages.civilytics.org/cog-api/)**
|
||||||
|
— reference, data dictionary, and worked examples such as the
|
||||||
|
[Southern states guide](https://pages.civilytics.org/cog-api/cog-api-south-guide.html)
|
||||||
|
- **[Live API](https://cog-api.civilytics.org/api/v1/)** — the same corpus over
|
||||||
|
HTTP, for Tableau, Python, or anything that isn't R
|
||||||
|
- **[Bulk corpus on Hugging Face](https://huggingface.co/datasets/civilytics/us-cog-finance)**
|
||||||
|
— CC-BY-4.0; the same parquet files this package reads
|
||||||
|
- **[Census Bureau source data](https://www.census.gov/programs-surveys/gov-finances.html)**
|
||||||
|
— the underlying public files
|
||||||
|
|
||||||
|
## Install
|
||||||
|
|
||||||
```r
|
```r
|
||||||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
install.packages("uscogdata",
|
||||||
|
repos = c("https://civilytics.r-universe.dev",
|
||||||
|
"https://cloud.r-project.org"))
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Or from source:
|
||||||
|
|
||||||
|
```r
|
||||||
|
pak::pkg_install("git::https://gitea.civilytics.org/Civilytics/uscogdata.git")
|
||||||
|
```
|
||||||
|
|
||||||
|
## Quickstart
|
||||||
|
|
||||||
|
No configuration, no credentials, no download. The package reads the published
|
||||||
|
corpus over HTTPS by default.
|
||||||
|
|
||||||
|
```r
|
||||||
|
library(uscogdata)
|
||||||
|
|
||||||
|
# Resolve a place name to a canonical government id
|
||||||
|
madison <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||||||
|
madison$canonical_govid
|
||||||
|
#> [1] "552025209777"
|
||||||
|
|
||||||
|
# Police spending, inflation-adjusted and per capita
|
||||||
|
spend <- cog_spending(
|
||||||
|
madison$canonical_govid,
|
||||||
|
years = 2012:2022,
|
||||||
|
category = "Police",
|
||||||
|
per_capita = TRUE,
|
||||||
|
adjust_to_year = 2023
|
||||||
|
)
|
||||||
|
|
||||||
|
# What did that result do to the numbers, and what should you know about them?
|
||||||
|
cog_explain(spend)
|
||||||
|
```
|
||||||
|
|
||||||
|
`years` is required — there is no implicit full-history default.
|
||||||
|
|
||||||
|
## Two ways to read the corpus
|
||||||
|
|
||||||
|
| | Remote (default) | Mirrored |
|
||||||
|
|---|---|---|
|
||||||
|
| Setup | none | `cog_mirror(dest)`, 190.6 MB once |
|
||||||
|
| Disk used | **0 MB** — HTTP range requests only | 190.6 MB |
|
||||||
|
| Per query | ~4 s (one government, one year)<br>~6 s (one government, 23 years) | local speed |
|
||||||
|
| Good for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility |
|
||||||
|
|
||||||
|
Nothing is written to disk in remote mode: DuckDB fetches the parquet footer,
|
||||||
|
works out which row groups it needs, and reads only those. Nothing is cached
|
||||||
|
between sessions either, so every query goes back to the network.
|
||||||
|
|
||||||
|
The default points at a public HuggingFace mirror of the corpus. If you would
|
||||||
|
rather not depend on a third party — for reproducibility, for an air-gapped
|
||||||
|
environment, or on principle — **the escape hatch is one function call**:
|
||||||
|
|
||||||
|
```r
|
||||||
|
cog_mirror("~/cog-corpus")
|
||||||
|
Sys.setenv(USCOGDATA_URL = "~/cog-corpus/")
|
||||||
|
```
|
||||||
|
|
||||||
|
After that, nothing in your analysis touches an external service.
|
||||||
|
|
||||||
|
### Configuration
|
||||||
|
|
||||||
|
- `USCOGDATA_URL` — corpus root: an HTTPS URL or a local path, **trailing slash required**
|
||||||
|
- `USCOGDATA_CACHE_DIR` — where the manifest is cached (default: user cache dir)
|
||||||
|
- `USCOGDATA_MANIFEST_TTL_SECS` — manifest re-fetch interval (default 3600)
|
||||||
|
|
||||||
## Amounts are in full US dollars
|
## Amounts are in full US dollars
|
||||||
|
|
||||||
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
||||||
@@ -29,111 +121,127 @@ own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
|||||||
you never have to. The conversion is recorded in every result:
|
you never have to. The conversion is recorded in every result:
|
||||||
|
|
||||||
```r
|
```r
|
||||||
r <- cog_spending("552025209777", 2020L)
|
attr(spend, "provenance")$transformations$units_conversion
|
||||||
attr(r, "provenance")$transformations$units_conversion
|
#> $applied TRUE
|
||||||
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
|
#> $source_unit "$1,000s (raw Census)"
|
||||||
|
#> $target_unit "$USD"
|
||||||
|
#> $multiplier 1000
|
||||||
```
|
```
|
||||||
|
|
||||||
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
||||||
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
|
`$1,000s` — which is true of the raw Census files and of the corpus's own `amt`
|
||||||
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
|
column — that rule does not apply to anything a `cog_*()` verb hands you.
|
||||||
twice overstates every figure by 1000x, and the result looks plausible rather
|
Applying it twice overstates every figure by 1000x, and the result looks
|
||||||
than obviously wrong.
|
plausible rather than obviously wrong.
|
||||||
|
|
||||||
## Configuration
|
## Concepts worth understanding before you publish a number
|
||||||
|
|
||||||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
### Primary vs Direct vs Total spending
|
||||||
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
|
||||||
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
|
||||||
|
|
||||||
## Primary vs Direct vs Total spending
|
|
||||||
|
|
||||||
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
|
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
|
||||||
controls whose spending a result counts. Concepts are defined as sets of the
|
controls *whose* spending a result counts. Concepts are defined as sets of the
|
||||||
crosswalk's `spend_subtype` values — never item-code first letters, which
|
crosswalk's `spend_subtype` values, never item-code first letters — the letter
|
||||||
cannot classify correctly (the letter `Y` alone spans revenue, expenditure,
|
`Y` alone spans revenue, expenditure and balance codes.
|
||||||
and balance codes):
|
|
||||||
|
|
||||||
- `"primary"` (the default) is the government's own service provision:
|
- **`"primary"`** (default) — the government's own service provision: current
|
||||||
current operations, capital outlay, and assistance payments.
|
operations, capital outlay, assistance payments.
|
||||||
- `"direct"` is Census's published Direct Expenditure: `primary` plus
|
- **`"direct"`** — Census's published Direct Expenditure: `primary` plus
|
||||||
interest on debt and insurance trust benefit payments (e.g. pensions).
|
interest on debt and insurance trust benefits (e.g. pensions).
|
||||||
- `"total"` additionally adds the intergovernmental leg — money handed to
|
- **`"total"`** — adds the intergovernmental leg, money handed to other
|
||||||
other governments to spend (`M`/`L` codes plus `Q11`/`Q12`/`Q18` state
|
governments to spend. Meaningful for one government's own budget over time,
|
||||||
payments to school systems) — which is meaningful for describing one
|
but it double-counts when summed across governments: a state's payment to a
|
||||||
government's own budget over time, but double-counts when summed across
|
county is the same dollar the county reports as its own direct spending.
|
||||||
governments (a state's payment to a county is the same dollar the county
|
|
||||||
reports as its own direct spending).
|
|
||||||
|
|
||||||
**Rule of thumb: any figure that spans more than one government uses
|
**Rule of thumb: any figure spanning more than one government uses `primary`
|
||||||
`primary` or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()`
|
or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce that
|
||||||
enforce this by refusing `expenditure_concept = "total"`. See
|
by refusing `"total"` outright. Worked examples in
|
||||||
`vignette("total-spending", package = "uscogdata")` for the full
|
`vignette("total-spending", package = "uscogdata")`.
|
||||||
explanation with worked examples.
|
|
||||||
|
|
||||||
## General vs Total revenue
|
### General vs Total revenue
|
||||||
|
|
||||||
`cog_revenue(..., revenue_concept = c("general", "total"))` selects between
|
`cog_revenue(..., revenue_concept = c("general", "total"))`:
|
||||||
Census's two published revenue concepts, again defined as crosswalk
|
|
||||||
`revenue_subtype` sets rather than item-code prefixes:
|
|
||||||
|
|
||||||
- `"general"` (the default) is Census **General Revenue**: own-source
|
- **`"general"`** (default) — Census General Revenue: own-source taxes,
|
||||||
(taxes, charges, miscellaneous) plus federal, state and local
|
charges and miscellaneous, plus federal, state and local aid.
|
||||||
intergovernmental aid.
|
- **`"total"`** — General plus utility revenue (`A91`–`A94`), liquor store
|
||||||
- `"total"` is Census **Total Revenue**: `general` plus utility revenue
|
revenue (`A90`), and insurance trust revenue.
|
||||||
(`A91`–`A94`), liquor store revenue (`A90`), and insurance trust revenue
|
|
||||||
(unemployment and workers' compensation `Y` codes plus the
|
|
||||||
employee-retirement `X` codes).
|
|
||||||
|
|
||||||
The manual defines the first by subtracting the other three from the second,
|
Census defines these by its own identity:
|
||||||
so the two are related by Census's own identity:
|
|
||||||
|
|
||||||
```
|
```
|
||||||
Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
||||||
```
|
```
|
||||||
|
|
||||||
Two things worth knowing before switching to `"total"`:
|
Two things to know before switching to `"total"`. **Utility revenue is large
|
||||||
|
for cities** — measured on the bundled fixture, utility plus liquor store is
|
||||||
|
15.9% of city revenue, against 1.2% for states and 1.7% for counties. And the
|
||||||
|
**employee-retirement (`X`) codes stop at FY2016**, when those systems moved to
|
||||||
|
the separate Annual Survey of Public Pensions, so a `"total"` series steps down
|
||||||
|
at the FY2016/FY2017 boundary for reasons of collection scope, not revenue
|
||||||
|
(series breaks `SB197`–`SB209`).
|
||||||
|
|
||||||
- **Utility revenue is large for cities.** Measured on the bundled fixture,
|
### Reporting coverage: the Census is only sometimes a census
|
||||||
utility plus liquor store revenue is 15.9% of city (type 2) revenue, versus
|
|
||||||
1.2% for states and 1.7% for counties. `general` excludes it by definition.
|
|
||||||
- **The employee-retirement (`X`) codes stop at FY2016**, when those systems
|
|
||||||
moved out of the annual finance file into the separate Annual Survey of
|
|
||||||
Public Pensions. A `"total"` series therefore steps down at the
|
|
||||||
FY2016/FY2017 seam for reasons of collection scope, not revenue (series
|
|
||||||
breaks `SB197`–`SB202`, in the corpus's `series_breaks` table).
|
|
||||||
|
|
||||||
## Developer notes
|
**The Census of Governments is a complete enumeration only in years ending in
|
||||||
|
2 and 7.** Every other year is a sample, and the sample varies enormously —
|
||||||
|
measured on the bundled fixture, Wisconsin's 608-city universe rolls up 597
|
||||||
|
governments in FY2012 and 112 in FY2019.
|
||||||
|
|
||||||
### Testing
|
A statewide total resting on a fifth of the universe looks exactly like one
|
||||||
|
resting on all of it, so every multi-government result now says which it is:
|
||||||
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
|
||||||
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
|
|
||||||
all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
|
|
||||||
at this fixture, so the full test suite runs offline with no network
|
|
||||||
dependency:
|
|
||||||
|
|
||||||
```r
|
```r
|
||||||
devtools::test() # uses bundled fixture, no credentials required
|
attr(rollup, "provenance")$coverage # per-year n_units_reporting, is_census_year
|
||||||
```
|
```
|
||||||
|
|
||||||
### Releasing against the live corpus
|
`cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` take a
|
||||||
|
`coverage` argument — `"all"` (default), `"census"` (census years only), or
|
||||||
|
`"consistent"` (only units reporting in every requested year, a balanced
|
||||||
|
panel).
|
||||||
|
|
||||||
Before cutting a release, run the test suite against the published corpus to
|
`n_units_reporting` is **category-conditional**, and it is not a response rate. A government that was surveyed and genuinely spends
|
||||||
catch any drift between the fixture and the real data:
|
nothing in the requested category is indistinguishable from one never surveyed.
|
||||||
|
|
||||||
|
### Absent cells mean two different things
|
||||||
|
|
||||||
|
Before FY2012, an absent cell means Census published `$0`. From FY2012 on, it
|
||||||
|
means not reported. `cog_spending(..., complete = TRUE)` fills the requested
|
||||||
|
grid and labels every row with which it is, via `value_source`:
|
||||||
|
|
||||||
|
| `value_source` | meaning | `amt_nominal` |
|
||||||
|
|---|---|---|
|
||||||
|
| `reported` | the corpus carries this cell | as published |
|
||||||
|
| `census_zero` | dense-source year (≤ FY2011), absent — Census published `$0` | `0` |
|
||||||
|
| `not_reported` | sparse-source year (≥ FY2012), absent — unknown | `NA` |
|
||||||
|
|
||||||
|
That `NA` is deliberate. Filling a modern absence with `0` would invent data.
|
||||||
|
|
||||||
|
### Series breaks surface on their own
|
||||||
|
|
||||||
|
Catalogued breaks that intersect your query appear in provenance whether or not
|
||||||
|
you went looking for them — `series_break_refs` for breaks in a specific item code, and
|
||||||
|
`corpus_break_refs` for caveats about the corpus as a whole (dollar precision
|
||||||
|
across the 1976/1977 boundary, the FY2017 identifier change, the FY2012
|
||||||
|
dense→sparse representation change). `cog_explain()` prints both.
|
||||||
|
|
||||||
|
## How to cite
|
||||||
|
|
||||||
```r
|
```r
|
||||||
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
|
citation("uscogdata")
|
||||||
devtools::test()
|
|
||||||
```
|
```
|
||||||
|
|
||||||
When the live-corpus run is clean, strip the fixture from the built package by
|
The corpus itself is published under CC-BY-4.0. Cite it as:
|
||||||
adding this line to `.Rbuildignore`:
|
|
||||||
|
|
||||||
```
|
> Civilytics Consulting. US Census of Governments finance corpus.
|
||||||
^inst/extdata/fixture_corpus$
|
> https://huggingface.co/datasets/civilytics/us-cog-finance
|
||||||
```
|
|
||||||
|
|
||||||
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when
|
## Contributing
|
||||||
the bundled fixture is absent, so no test code changes are needed for the
|
|
||||||
release run or after stripping the fixture.
|
Development happens on [Gitea](https://gitea.civilytics.org/Civilytics/uscogdata);
|
||||||
|
[GitHub](https://github.com/civilytics/uscogdata) is a mirror that accepts
|
||||||
|
issues and pull requests. See [CONTRIBUTING.md](CONTRIBUTING.md) for how a
|
||||||
|
patch gets from there to here.
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
MIT © Civilytics Consulting LLC. See [LICENSE.md](LICENSE.md).
|
||||||
|
|||||||
@@ -145,3 +145,23 @@ test_that("_pkgdown.yml indexes every exported topic", {
|
|||||||
# means the docs site does not build at all.
|
# means the docs site does not build at all.
|
||||||
expect_equal(missing, character(0))
|
expect_equal(missing, character(0))
|
||||||
})
|
})
|
||||||
|
|
||||||
|
test_that("README is written for a stranger, not a repo insider", {
|
||||||
|
skip_if_no_source_tree("README.md")
|
||||||
|
r <- paste(readLines(source_tree_path("README.md"), warn = FALSE), collapse = "\n")
|
||||||
|
|
||||||
|
# No paths that only resolve inside a maintainer's checkout.
|
||||||
|
expect_false(grepl("../cog_pipeline", r, fixed = TRUE))
|
||||||
|
# A real, uncommented install line.
|
||||||
|
expect_match(r, "install.packages", fixed = TRUE)
|
||||||
|
expect_false(grepl("# pak::pkg_install", r, fixed = TRUE))
|
||||||
|
# The errata most likely to produce a plausible-looking wrong answer.
|
||||||
|
expect_match(r, "full US dollars", fixed = TRUE)
|
||||||
|
# The release advice that conflicts with public CI is gone.
|
||||||
|
expect_false(grepl("Rbuildignore", r, fixed = TRUE))
|
||||||
|
# Both read paths documented.
|
||||||
|
expect_match(r, "cog_mirror", fixed = TRUE)
|
||||||
|
# cog_spending() has no default for `years`; a quickstart that omits it
|
||||||
|
# errors on the reader's first call.
|
||||||
|
expect_match(r, "years\\s*=", perl = TRUE)
|
||||||
|
})
|
||||||
|
|||||||
Reference in New Issue
Block a user