Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
uscogdata
Curated R reader for the Civilytics US Census of Governments finance corpus.
Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.
Status
Under active development (Phase 2 of the cog_pipeline project). See
../cog_pipeline/docs/reader-specification.md for the reader contract this
package implements.
Installation
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
Amounts are in full US dollars
Every amount column this package returns — amt_nominal, amt_real,
amt_per_capita_nominal, amt_per_capita_real — is in full US dollars.
The raw Census source files report thousands of dollars, and the corpus's
own amt column preserves that. The verbs multiply by 1000 on the way out, so
you never have to. The conversion is recorded in every result:
r <- cog_spending("552025209777", 2020L)
attr(r, "provenance")$transformations$units_conversion
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
Do not multiply again. If you have read elsewhere that COG amounts are in
$1,000s — true of the raw corpus, and of cog_explorer's conventions doc —
that rule does not apply to anything a cog_*() verb hands you. Applying it
twice overstates every figure by 1000x, and the result looks plausible rather
than obviously wrong.
Configuration
USCOGDATA_URL— corpus root URL (public Nextcloud share, trailing slash)USCOGDATA_CACHE_DIR— optional override for the manifest cache directoryUSCOGDATA_MANIFEST_TTL_SECS— optional manifest re-fetch TTL (default 3600)
Direct vs Total spending
cog_spending(..., expenditure_concept = c("direct", "total")) controls
whose spending a result counts. "direct" (the default) is a government's
own current operations, capital outlay, and other direct spending. "total"
additionally adds in the intergovernmental legs — money it hands to other
governments to spend on its behalf — which is meaningful for describing one
government's own budget over time, but double-counts when summed across
governments (a state's payment to a county is the same dollar the county
reports as its own direct spending).
Rule of thumb: any figure that spans more than one government uses
direct. cog_geographic_rollup() and cog_peer_compare() enforce this
by refusing expenditure_concept = "total". See
vignette("total-spending", package = "uscogdata") for the full
explanation with worked examples.
Developer notes
Testing
The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ —
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
all 50 states. tests/testthat/setup.R automatically points USCOGDATA_URL
at this fixture, so the full test suite runs offline with no network
dependency:
devtools::test() # uses bundled fixture, no credentials required
Releasing against the live corpus
Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:
Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()
When the live-corpus run is clean, strip the fixture from the built package by
adding this line to .Rbuildignore:
^inst/extdata/fixture_corpus$
The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when
the bundled fixture is absent, so no test code changes are needed for the
release run or after stripping the fixture.