Files
uscogdata/README.md
T
jared 1722dd81ba
R-CMD-check / check (push) Failing after 1m53s
docs: survey_weight is col 28 under schema v6 (was col 26 in v5)
Rebase onto the v6 main (bd53230) shifted survey_weight from col 26 to
col 28: v5→v6 inserted cog_legacy_state/cog_legacy_county at positions
10-11 (26→28 cols). Position confirmed against the regenerated v6
fixture and both corpus docs (reader-specification.md §3 'Long parquet
schema (28 columns)' row 28; data_dictionary.md '28-column schema v6'
row 28).
2026-07-23 09:48:36 -04:00

77 lines
2.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons
with auditable provenance and built-in cross-vintage correctness. Reads the
published corpus (Hive-partitioned parquet + manifest.json) directly from
Nextcloud via DuckDB httpfs — no local bulk downloads required.
## Status
Under active development (Phase 2 of the cog_pipeline project). See
`../cog_pipeline/docs/reader-specification.md` for the reader contract this
package implements.
## Installation
```r
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
```
## Configuration
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
## Raw-parquet caveat: `survey_weight` is not an aggregation weight
Users reading the corpus parquet directly (DuckDB, arrow) will see a
`survey_weight` column (schema v6, col 28). It is legacy Census IndFin
sample-design **metadata passed through verbatim** — the Census Bureau's own
source documentation says it "is for informational purposes only and should
not be used to derive any other statistics" (`_ReadMe_First_IndFin.txt`;
likewise `UserGuide.xls` Data User Note 8: "Do not use the weight field to
derive state or national totals"). The raw encoding is also inconsistent
across vintages (reciprocal scale most years, direct scale in 2003, a `1`
placeholder in 1967/70/71/73/2001, all-`0` in 2007–2012, `NA` for all
modern-source rows), so `sum(amt * survey_weight/10000)`-style expressions
produce silently wrong totals — including exact zeros for 2007–2012. Sum
`amt` unweighted; no uscogdata function reads this column. Full evidence:
`cog_pipeline/.superpowers/sdd/weight-semantics-findings.md`.
## Developer notes
### Testing
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50
states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` at this
fixture, so the full test suite runs offline with no network dependency:
```r
devtools::test() # uses bundled fixture, no credentials required
```
### Releasing against the live corpus
Before cutting a release, run the test suite against the published corpus to
catch any drift between the fixture and the real data:
```r
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()
```
When the live-corpus run is clean, strip the fixture from the built package by
adding this line to `.Rbuildignore`:
```
^inst/extdata/fixture_corpus$
```
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when
the bundled fixture is absent, so no test code changes are needed for the
release run or after stripping the fixture.