488d03d74b931f02eab65fdd168629b002d5be5d
Discovery verb over the summary_categories view, grouped one row per (category, subtype). Parallels cog_gov_search: analysts use it to find the valid `category` values to pass into cog_spending(), cog_revenue(), cog_geographic_rollup(). Columns: category, category_type, subtype, n_codes, item_codes (comma-separated, alphabetical). Optional filters: type = NULL | "spending" | "revenue" pattern = regex matched case-insensitively on category The user-facing "spending" alias is translated internally to the corpus-native "expenditure" so callers don't have to learn Census vocabulary, while the returned category_type column preserves the native value for auditability. Also: fix @noRd placement in session.R so devtools::document() stops warning. Tests: +16 new / 181 total pass. check 0E/0W/0N.
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Languages
R
100%