a640c9cc218838ab72033a95a53852cc9a037978
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice of the published corpus (OH+VT+WY fixture from cog_pipeline test profile plus all 50 states). Includes canonical_fips_xwalk.parquet, summary_categories.parquet, docs/, and a trimmed manifest.json. setup.R now points USCOGDATA_URL at the bundled fixture automatically, bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the existing .is_local_path() fast-path in manifest.R; no httpfs required. session is reset between test files via withr::defer(cog_close()). helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus() that checks the bundled fixture first, and with_fixture_corpus() for tests that need explicit session isolation. test-spending.R: adjust the inflate-column test to use 2019 (fixture year) instead of 2015 (absent from fixture). Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet data with real DuckDB queries and no network dependency.
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Languages
R
100%