The v5 schema passes the legacy IndFin Weight column through verbatim as survey_weight. Census documents it as informational-only, and its encoding is inconsistent across vintages (reciprocal scale most years, direct in 2003, placeholder 1 in 1967-2001 gap years, all-0 in 2007-2012, NA modern), so weighting amt by it produces silently wrong totals. No uscogdata function reads the column; this warning is for direct DuckDB/arrow consumers. Evidence: cog_pipeline/.superpowers/sdd/weight-semantics-findings.md.
77 lines
2.9 KiB
Markdown
77 lines
2.9 KiB
Markdown
# uscogdata
|
||
|
||
Curated R reader for the Civilytics US Census of Governments finance corpus.
|
||
|
||
Provides unit-level financial profiles, geographic rollups, and peer comparisons
|
||
with auditable provenance and built-in cross-vintage correctness. Reads the
|
||
published corpus (Hive-partitioned parquet + manifest.json) directly from
|
||
Nextcloud via DuckDB httpfs — no local bulk downloads required.
|
||
|
||
## Status
|
||
|
||
Under active development (Phase 2 of the cog_pipeline project). See
|
||
`../cog_pipeline/docs/reader-specification.md` for the reader contract this
|
||
package implements.
|
||
|
||
## Installation
|
||
|
||
```r
|
||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
||
```
|
||
|
||
## Configuration
|
||
|
||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
||
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
||
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
||
|
||
## Raw-parquet caveat: `survey_weight` is not an aggregation weight
|
||
|
||
Users reading the corpus parquet directly (DuckDB, arrow) will see a
|
||
`survey_weight` column (schema v5, col 26). It is legacy Census IndFin
|
||
sample-design **metadata passed through verbatim** — the Census Bureau's own
|
||
source documentation says it "is for informational purposes only and should
|
||
not be used to derive any other statistics" (`_ReadMe_First_IndFin.txt`;
|
||
likewise `UserGuide.xls` Data User Note 8: "Do not use the weight field to
|
||
derive state or national totals"). The raw encoding is also inconsistent
|
||
across vintages (reciprocal scale most years, direct scale in 2003, a `1`
|
||
placeholder in 1967/70/71/73/2001, all-`0` in 2007–2012, `NA` for all
|
||
modern-source rows), so `sum(amt * survey_weight/10000)`-style expressions
|
||
produce silently wrong totals — including exact zeros for 2007–2012. Sum
|
||
`amt` unweighted; no uscogdata function reads this column. Full evidence:
|
||
`cog_pipeline/.superpowers/sdd/weight-semantics-findings.md`.
|
||
|
||
## Developer notes
|
||
|
||
### Testing
|
||
|
||
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
||
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50
|
||
states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` at this
|
||
fixture, so the full test suite runs offline with no network dependency:
|
||
|
||
```r
|
||
devtools::test() # uses bundled fixture, no credentials required
|
||
```
|
||
|
||
### Releasing against the live corpus
|
||
|
||
Before cutting a release, run the test suite against the published corpus to
|
||
catch any drift between the fixture and the real data:
|
||
|
||
```r
|
||
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
|
||
devtools::test()
|
||
```
|
||
|
||
When the live-corpus run is clean, strip the fixture from the built package by
|
||
adding this line to `.Rbuildignore`:
|
||
|
||
```
|
||
^inst/extdata/fixture_corpus$
|
||
```
|
||
|
||
The test suite is URL-agnostic — `setup.R` falls back to `USCOGDATA_URL` when
|
||
the bundled fixture is absent, so no test code changes are needed for the
|
||
release run or after stripping the fixture.
|