jared 77f48047b1
R-CMD-check / check (pull_request) Successful in 3m5s
R-CMD-check / check (push) Successful in 3m1s
fix: exercise real harmonized-view SQL in tests; unambiguous recipe provenance
test-views.R's harmonized-view test previously ran a hand-rolled REPLACE
query with no WHERE clause, so a regression in any of
inst/sql/22-spending_long_harmonized.sql / 23-revenue_long_harmonized.sql's
three predicates (NOT is_aggregate, harmonized_code IS NOT NULL, the
E/F/G/K or T/A/U/B/C/D prefix filter) would go uncaught. Replaced it with a
test that reads the real SQL files off disk, substitutes {url} exactly as
.register_views() does, and executes them (plus their 10-long.sql
dependency) against a synthetic hive-partitioned parquet tree written via
DuckDB's own COPY ... TO (FORMAT PARQUET) (no arrow dependency, matching
this package's existing convention). Ten rows are crafted so each predicate
is independently falsifiable by a specific row; manually broke each
predicate in turn to confirm the test fails exactly as expected, then
restored the SQL files (see the task report for the RED-phase transcript).

Also fixes a provenance ambiguity: a recipe= query bypasses
spending_annotated(_harmonized)/revenue_annotated(_harmonized) entirely
(.run_recipe() joins `long` directly), so basis= has no effect on it, but
provenance was still reporting basis = "harmonized"/"raw" (whatever the
argument resolved to) with harmonization$applied = FALSE alongside it --
misleading, since it looks like harmonization was evaluated and found
nothing to exclude rather than "not applicable here." Recipe results now
report basis = "recipe" with an inert harmonization block carrying an
explicit note, regardless of what basis= was passed.
2026-07-18 23:55:48 -04:00

uscogdata

Curated R reader for the Civilytics US Census of Governments finance corpus.

Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.

Status

Under active development (Phase 2 of the cog_pipeline project). See ../cog_pipeline/docs/reader-specification.md for the reader contract this package implements.

Installation

# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")

Configuration

  • USCOGDATA_URL — corpus root URL (public Nextcloud share, trailing slash)
  • USCOGDATA_CACHE_DIR — optional override for the manifest cache directory
  • USCOGDATA_MANIFEST_TTL_SECS — optional manifest re-fetch TTL (default 3600)

Developer notes

Testing

The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ — a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50 states. tests/testthat/setup.R automatically points USCOGDATA_URL at this fixture, so the full test suite runs offline with no network dependency:

devtools::test()   # uses bundled fixture, no credentials required

Releasing against the live corpus

Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:

Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()

When the live-corpus run is clean, strip the fixture from the built package by adding this line to .Rbuildignore:

^inst/extdata/fixture_corpus$

The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when the bundled fixture is absent, so no test code changes are needed for the release run or after stripping the fixture.

S
Description
R reader for the Civilytics US Census of Governments finance corpus
Readme
18 MiB
Languages
R 100%