jared 1d553a788f
R-CMD-check / check (push) Successful in 3m1s
R-CMD-check / check (pull_request) Successful in 3m1s
fix: surface ALL-scoped series breaks in provenance (#19)
.build_series_break_refs() matches `fin_code IN (<codes in the result>)`.
No row's item_code is ever the literal "ALL", so the four corpus-wide
entries could never match and reached no user:

  SB085  1977  dollar precision across the 1976/1977 boundary
  SB087  2002  imputation exclusion FY2002-2006
  SB194  2012  dense -> sparse representation change
  SB086  2017  government id scheme change

SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv
carries an ALL @ 2012 entry describing the representation change, SO
cog_explain() surfaces it". The entry shipped; the reader dropped it. A
query spanning FY2011 -> FY2012 crosses the boundary where an absent cell
stops meaning "Census published $0" and starts meaning "not reported", and
nothing said so.

Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs()
on the break_year window alone -- which codes a result happens to contain
is irrelevant to a caveat about the corpus. A separate field rather than
more entries in series_break_refs, because an ALL caveat qualifies the
whole result and folding the two together invites reading it as a caveat
about one series; .build_series_break_refs() now excludes 'ALL' explicitly
so the two stay disjoint by construction. cog_explain() prints them under
their own "Corpus-wide caveats" heading, and cog-api passes provenance
through verbatim, so the field reaches the API with no change there.

On the year rule: all four entries are BOUNDARY caveats -- their own
join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not
being comparable across FY2012, of pre- vs post-2017 ids -- so the same
`break_year BETWEEN min(years) AND max(years)` rule the code-specific path
uses is the right one, and matches the issue's DoD 1. The issue's DoD 3
also asks that a FY2011 query surface SB085; that cannot hold under DoD 1
and does not hold under any reading of SB085's text, whose boundary is
1976/1977. Tested with a range that actually spans it, and flagged on the
issue.

Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's
bundled fixture, which predates the break being catalogued.

Suite: 606 pass / 0 fail / 6 skip (was 594/0/6).
cog-api 357 / 0 / 8, unchanged.
2026-07-30 10:27:48 -04:00

uscogdata

Curated R reader for the Civilytics US Census of Governments finance corpus.

Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.

Status

Under active development (Phase 2 of the cog_pipeline project). See ../cog_pipeline/docs/reader-specification.md for the reader contract this package implements.

Installation

# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")

Configuration

  • USCOGDATA_URL — corpus root URL (public Nextcloud share, trailing slash)
  • USCOGDATA_CACHE_DIR — optional override for the manifest cache directory
  • USCOGDATA_MANIFEST_TTL_SECS — optional manifest re-fetch TTL (default 3600)

Direct vs Total spending

cog_spending(..., expenditure_concept = c("direct", "total")) controls whose spending a result counts. "direct" (the default) is a government's own current operations, capital outlay, and other direct spending. "total" additionally adds in the intergovernmental legs — money it hands to other governments to spend on its behalf — which is meaningful for describing one government's own budget over time, but double-counts when summed across governments (a state's payment to a county is the same dollar the county reports as its own direct spending).

Rule of thumb: any figure that spans more than one government uses direct. cog_geographic_rollup() and cog_peer_compare() enforce this by refusing expenditure_concept = "total". See vignette("total-spending", package = "uscogdata") for the full explanation with worked examples.

Developer notes

Testing

The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ — a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering all 50 states. tests/testthat/setup.R automatically points USCOGDATA_URL at this fixture, so the full test suite runs offline with no network dependency:

devtools::test()   # uses bundled fixture, no credentials required

Releasing against the live corpus

Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:

Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()

When the live-corpus run is clean, strip the fixture from the built package by adding this line to .Rbuildignore:

^inst/extdata/fixture_corpus$

The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when the bundled fixture is absent, so no test code changes are needed for the release run or after stripping the fixture.

S
Description
R reader for the Civilytics US Census of Governments finance corpus
Readme
18 MiB
Languages
R 100%