The three kodor/fix issues, taken over after a day with no branch, PR or comment on any of them. Batched because each is single-file with a committed acceptance test, and two share documentation surfaces. #16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight into regexp_matches() unescaped, while basket mode in the same file already routed it through .escape_regex() with the comment "so `name` is treated as a literal substring". Two failure modes, both HTTP 200 through the API: a government could not be found by its own complete name when that name contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a bare "." matched all 608 Wisconsin cities. Malformed pattern text reached the engine as an error, which cog-api surfaced as a 500 -- reachable by typing a real name one character at a time ("Athens-Clarke County (bal"). Utility mode now calls the escaper that already existed. Roxygen updated: utility mode is documented as a literal case-insensitive substring match, and the basket-mode "substring fallback" step no longer describes itself as a regex either. BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop working, because there is no regex left to anchor. Two existing tests used "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now search for those characters literally. Updated to the bare names, which still resolve to exactly one row each once scoped by state/type (verified, not assumed). There is no exact-match option in utility mode any more -- noted on the issue, since that is a real if small capability loss. #15 (F-004) -- the raw Census files report thousands of dollars; this package multiplies by 1000 and returns full US dollars. Correct, and already stated in ?cog_spending / ?cog_revenue @return, in provenance, and in cog-api's data-dictionary. Absent from every surface a reader meets FIRST. Added to README.md as its own section and to both vignettes' openings. The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule ("All raw `amt` values are in $1,000s") without scoping it to the raw column -- a reader applying that to amt_nominal overstates by 1000x and gets a plausible-looking number rather than an obvious error. Fixed there too; that directory has no git remote, so it rides in no PR and is left uncommitted for the owner. #14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately inside each (year, spend_subtype, category) cell, so a summary_p50 row is "the median peer's value in that one category", never "the value of the median peer's total" -- the median peer for Police and for Fire are usually different governments. Summing them across categories misstated a total-spending band by -32.7% to +251.0% across 24 years, with a sign flip at FY2012. The verb is right and its documented use (facet by role AND category) is unaffected, so the fix is @return prose plus a worked snippet showing the correct computation: sum each peer's own categories first, then take the quantile of those per-government totals. This is the R-side counterpart of cog-api#9, fixed on the API surface earlier today; the wording is deliberately consistent across the two. Note the phrase "not additive" has to stay on one roxygen source line -- the test greps the generated Rd, where a line wrap turns it into "not additive" and stops matching. Cost one red run to find. man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with no content change. The two Rd files kept carry only the edits above. Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips are #11, #12 and #13.
100 lines
3.7 KiB
Markdown
100 lines
3.7 KiB
Markdown
# uscogdata
|
|
|
|
Curated R reader for the Civilytics US Census of Governments finance corpus.
|
|
|
|
Provides unit-level financial profiles, geographic rollups, and peer comparisons
|
|
with auditable provenance and built-in cross-vintage correctness. Reads the
|
|
published corpus (Hive-partitioned parquet + manifest.json) directly from
|
|
Nextcloud via DuckDB httpfs — no local bulk downloads required.
|
|
|
|
## Status
|
|
|
|
Under active development (Phase 2 of the cog_pipeline project). See
|
|
`../cog_pipeline/docs/reader-specification.md` for the reader contract this
|
|
package implements.
|
|
|
|
## Installation
|
|
|
|
```r
|
|
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
|
```
|
|
|
|
## Amounts are in full US dollars
|
|
|
|
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
|
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
|
|
|
|
The raw Census source files report **thousands of dollars**, and the corpus's
|
|
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
|
you never have to. The conversion is recorded in every result:
|
|
|
|
```r
|
|
r <- cog_spending("552025209777", 2020L)
|
|
attr(r, "provenance")$transformations$units_conversion
|
|
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
|
|
```
|
|
|
|
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
|
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
|
|
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
|
|
twice overstates every figure by 1000x, and the result looks plausible rather
|
|
than obviously wrong.
|
|
|
|
## Configuration
|
|
|
|
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
|
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
|
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
|
|
|
## Direct vs Total spending
|
|
|
|
`cog_spending(..., expenditure_concept = c("direct", "total"))` controls
|
|
whose spending a result counts. `"direct"` (the default) is a government's
|
|
own current operations, capital outlay, and other direct spending. `"total"`
|
|
additionally adds in the intergovernmental legs — money it hands to other
|
|
governments to spend on its behalf — which is meaningful for describing one
|
|
government's own budget over time, but double-counts when summed across
|
|
governments (a state's payment to a county is the same dollar the county
|
|
reports as its own direct spending).
|
|
|
|
**Rule of thumb: any figure that spans more than one government uses
|
|
`direct`.** `cog_geographic_rollup()` and `cog_peer_compare()` enforce this
|
|
by refusing `expenditure_concept = "total"`. See
|
|
`vignette("total-spending", package = "uscogdata")` for the full
|
|
explanation with worked examples.
|
|
|
|
## Developer notes
|
|
|
|
### Testing
|
|
|
|
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
|
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
|
|
all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
|
|
at this fixture, so the full test suite runs offline with no network
|
|
dependency:
|
|
|
|
```r
|
|
devtools::test() # uses bundled fixture, no credentials required
|
|
```
|
|
|
|
### Releasing against the live corpus
|
|
|
|
Before cutting a release, run the test suite against the published corpus to
|
|
catch any drift between the fixture and the real data:
|
|
|
|
```r
|
|
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
|
|
devtools::test()
|
|
```
|
|
|
|
When the live-corpus run is clean, strip the fixture from the built package by
|
|
adding this line to `.Rbuildignore`:
|
|
|
|
```
|
|
^inst/extdata/fixture_corpus$
|
|
```
|
|
|
|
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when
|
|
the bundled fixture is absent, so no test code changes are needed for the
|
|
release run or after stripping the fixture.
|