Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under spend_subtype = intergovernmental (194 -> 260 rows). cog_categories() (R/categories.R) has no item-code prefix filter, so the new IG rows surface immediately as a third spend subtype; this broke test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is correct behavior, not a regression -- cog_categories() is a discovery verb documented to surface valid category values, and after Task 3 lands users will see spend_subtype = intergovernmental in cog_spending(expenditure_concept = total) results. Widened the subtype assertion and added positive coverage asserting the intergovernmental subtype and that it reuses existing functional categories (plus Other Education, pipeline #58). R/categories.R itself is unchanged -- its behavior was already right. Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new expectations).
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Developer notes
Testing
The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ —
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50
states. tests/testthat/setup.R automatically points USCOGDATA_URL at this
fixture, so the full test suite runs offline with no network dependency:
devtools::test() # uses bundled fixture, no credentials required
Releasing against the live corpus
Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()
When the live-corpus run is clean, strip the fixture from the built package by
adding this line to .Rbuildignore:
^inst/extdata/fixture_corpus$
The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when
the bundled fixture is absent, so no test code changes are needed for the
release run or after stripping the fixture.