Nine review items on the expenditure_concept = direct|total feature: - bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback: bool_and silently misreported $5,740,775,000 of aggregate-sourced IG dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense wide-era data puts a $0 leaf row in the same group as the real aggregate row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups across both tables) and correct for the IG leg. - Added a year-disjointness invariant test for the four legacy aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93), scoped to the aggregate flag rather than bare code presence (M89/L89 continue past 2011 as independent, non-aggregate leaves). - Extended the real-SQL-text/synthetic-parquet harness in test-views.R to pin ig_long/ig_long_harmonized's predicates directly (aggregate rows retained, NULL harmonized_code coalesced, L-- excluded), rather than relying on one fixture row's incidental shape. - Added a test proving the .harmonization_view_files schema-v5 guard is necessary (not just incidental) against a corpus whose `long` genuinely lacks a harmonized_code column, and rewrote the misleading "v5-only parquet files" comment to name both real reasons a file is gated. - Fixed an NA-fragile subtype filter, extended the expected-view-list test, guarded .verb_spendrev() against total on a non-spending view_base, added a roxygen caveat against summing total across levels of government, and replaced an uncheckable corpus-wide SQL comment figure with a fixture-verifiable one. Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing value changed).
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Developer notes
Testing
The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ —
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50
states. tests/testthat/setup.R automatically points USCOGDATA_URL at this
fixture, so the full test suite runs offline with no network dependency:
devtools::test() # uses bundled fixture, no credentials required
Releasing against the live corpus
Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()
When the live-corpus run is clean, strip the fixture from the built package by
adding this line to .Rbuildignore:
^inst/extdata/fixture_corpus$
The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when
the bundled fixture is absent, so no test code changes are needed for the
release run or after stripping the fixture.