jared d238bc0a22 feat: report the coarse-vs-per-code subset relation in the signposting harness
The coarse and per-code signposting checks are partly DISJOINT, not nested:
coarse fires on queries per-code does not, so the coarse -> percode move
both adds and removes signposting. Every `*_delta_pp` the harness reports is
therefore a NET that can mask a coverage loss in either direction. The
staged-corpus headline (+1.875 pp, coarse 1/640 -> percode 13/640) sits on
top of Corrections losing coverage outright (0.05 -> 0.00, -5 pp).

The cause is structural, not sampling: coarse's coverage test is at recipe
grain and self-coverage-permissive, while per-code requires a DIFFERENT
component of the same recipe. When a whole category is empty in a year --
coarse's own trigger -- and the only covering evidence is the gapped
component's own wide-era aggregate row, per-code cannot fire by
construction. That case is already pinned as intended behaviour in
test-recipes.R; this change measures what it costs, it does not change it.

Measurement and disclosure only. R/suggestions.R is untouched -- which arm
ships is the human ruling at Checkpoint R3.

- header: replace the "noise trade" framing with an explicit statement that
  the checks are partly disjoint and every delta is a net
- .measure_subset_relation(): split the disagreement into violations
  (coarse fired, per-code silent -- coverage LOST) and additions, returning
  the offending rows, not just counts. No assertion: the violation set is
  genuinely non-empty and a stopifnot() would only break the harness that
  is supposed to surface it
- .measure_format_subset_report(): prominent HOLDS / *** VIOLATED ***
  section naming each offending (category, government, year)
- detail gains coarse_gap_years / coarse_recipes / percode_recipes;
  by_category gains n_coarse_only / n_percode_only so the two netted flows
  are visible per category
- new test-signposting-harness.R pins the reporting, including inversion
  guards and an end-to-end case (Broward FY2011 Corrections) where coarse
  fires and per-code does not

Tests: 518 PASS / 0 FAIL / 0 WARN / 0 SKIP (was 476). Mutation-checked:
inverting the violation direction fails 18 assertions, removing the
violation reporting fails 9.
2026-07-23 12:22:24 -04:00

uscogdata

Curated R reader for the Civilytics US Census of Governments finance corpus.

Provides unit-level financial profiles, geographic rollups, and peer comparisons with auditable provenance and built-in cross-vintage correctness. Reads the published corpus (Hive-partitioned parquet + manifest.json) directly from Nextcloud via DuckDB httpfs — no local bulk downloads required.

Status

Under active development (Phase 2 of the cog_pipeline project). See ../cog_pipeline/docs/reader-specification.md for the reader contract this package implements.

Installation

# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")

Configuration

  • USCOGDATA_URL — corpus root URL (public Nextcloud share, trailing slash)
  • USCOGDATA_CACHE_DIR — optional override for the manifest cache directory
  • USCOGDATA_MANIFEST_TTL_SECS — optional manifest re-fetch TTL (default 3600)

Raw-parquet caveat: survey_weight is not an aggregation weight

Users reading the corpus parquet directly (DuckDB, arrow) will see a survey_weight column (schema v5, col 26). It is legacy Census IndFin sample-design metadata passed through verbatim — the Census Bureau's own source documentation says it "is for informational purposes only and should not be used to derive any other statistics" (_ReadMe_First_IndFin.txt; likewise UserGuide.xls Data User Note 8: "Do not use the weight field to derive state or national totals"). The raw encoding is also inconsistent across vintages (reciprocal scale most years, direct scale in 2003, a 1 placeholder in 1967/70/71/73/2001, all-0 in 2007–2012, NA for all modern-source rows), so sum(amt * survey_weight/10000)-style expressions produce silently wrong totals — including exact zeros for 2007–2012. Sum amt unweighted; no uscogdata function reads this column. Full evidence: cog_pipeline/.superpowers/sdd/weight-semantics-findings.md.

Developer notes

Testing

The package ships a bundled fixture corpus at inst/extdata/fixture_corpus/ — a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50 states. tests/testthat/setup.R automatically points USCOGDATA_URL at this fixture, so the full test suite runs offline with no network dependency:

devtools::test()   # uses bundled fixture, no credentials required

Releasing against the live corpus

Before cutting a release, run the test suite against the published corpus to catch any drift between the fixture and the real data:

Sys.setenv(USCOGDATA_URL = "<published-corpus-url-with-trailing-slash>")
devtools::test()

When the live-corpus run is clean, strip the fixture from the built package by adding this line to .Rbuildignore:

^inst/extdata/fixture_corpus$

The test suite is URL-agnostic — setup.R falls back to USCOGDATA_URL when the bundled fixture is absent, so no test code changes are needed for the release run or after stripping the fixture.

S
Description
R reader for the Civilytics US Census of Governments finance corpus
Readme
18 MiB
Languages
R 100%