jared 267bc24fee feat: narrow harmonization signposting to per-code gap detection
.build_suggestions() previously flagged a recipe only when the WHOLE
category result had zero rows in a requested year, so a multi-code
category where one recipe component was genuinely gapped never fired
if any sibling code (same recipe or not) had data that year. Each
recipe's own in-category component is now checked individually -- a
component fires when it has no rows in a requested (in-scope) year the
recipe's own generic join otherwise covers, even when the overall
category result looks complete.

Decomposes .build_suggestions() into .category_recipe_components/
.recipe_meta/.component_presence/.recipe_coverage/.recipe_component_gapped
helpers, drops the now-unused `result` param, and rewrites the header
comment to describe the new, deliberately wider scope plus the
per-government `covered` guard that still filters recipes with no data
at all (ordinary reporting variance vs. a real format-boundary gap).

Tests pin the multi-code case the coarse check missed (Cleburne County
FY2012: G05 gapped, G04 covers, masked because E04/E05 have data) next
to the still-guarded no-recipe-coverage case (F04/F05 both absent), and
update the Broward 2019-2020 case to its new, correct expectation (fires
for corrections_combined/corrections_other_capital_combined, still
silent for corrections_capital_combined) plus a fresh true-full-coverage
negative case (Maricopa County).
2026-07-23 12:22:22 -04:00

uscogdata

Curated R reader for the Civilytics US Census of Governments finance corpus.

Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.

Status

Under active development (Phase 2 of the cog_pipeline project). See ../cog_pipeline/docs/reader-specification.md for the reader contract this package implements.

Installation

# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")

Configuration

  • USCOGDATA_URL — corpus root URL (public Nextcloud share, trailing slash)
  • USCOGDATA_CACHE_DIR — optional override for the manifest cache directory
  • USCOGDATA_MANIFEST_TTL_SECS — optional manifest re-fetch TTL (default 3600)

Developer notes

Testing

The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ — a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50 states. tests/testthat/setup.R automatically points USCOGDATA_URL at this fixture, so the full test suite runs offline with no network dependency:

devtools::test()   # uses bundled fixture, no credentials required

Releasing against the live corpus

Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:

Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()

When the live-corpus run is clean, strip the fixture from the built package by adding this line to .Rbuildignore:

^inst/extdata/fixture_corpus$

The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when the bundled fixture is absent, so no test code changes are needed for the release run or after stripping the fixture.

S
Description
R reader for the Civilytics US Census of Governments finance corpus
Readme
18 MiB
Languages
R 100%