c682e6547d7ca2ba4325f9f66fb45eff34615643
Three core query verbs over the spending_annotated / revenue_annotated DuckDB views. Each verb accepts vector govid, vector years, optional category filter, per_capita flag, and adjust_to_year for CPI-U real-dollar conversion (bundled index). Amounts are returned in full USD (SUM(amt) * 1000) so callers can freely rescale to millions/billions. The $1,000s -> $USD conversion is recorded in provenance$transformations$units_conversion. Every result carries an attr(., "provenance") list matching inst/schemas/provenance-v1.json. cog_explain() prints the structured form via cli or returns the raw list for MCP/JSON consumers. Also: .fetch_or_cache_manifest() now handles local fixture paths so tests can point USCOGDATA_FIXTURE_URL at the pipeline publish_cache/ without a working HTTP server. Tests: 80 pass / 0 fail. devtools::check() 0E/0W/2N (both notes pre-existing / environmental).
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Languages
R
100%