Compare commits

..
Author SHA1 Message Date
jared 4bedf857e9 docs: implementation plan for the uscogdata 0.1.0 public release
R-CMD-check / check (push) Successful in 4m23s
Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest
enumeration, working default URL, and the live-corpus test whose absence
let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite
README, NEWS and CONTRIBUTING.

Distribution mechanics stay out of scope -- r-universe publishes check
results on registration, so it comes after final verification is green.
2026-08-08 16:54:18 -04:00
jared a9de5ba5ab docs: design spec for the uscogdata 0.1.0 public release
Covers the P0 finding that the package cannot read the corpus remotely at
all -- no working default URL, and Hive globs are unsupported over generic
HTTP by DuckDB 1.5.5. Fix is manifest-driven file enumeration (measured:
46,148,034 rows over plain https, identical to the hf:// glob) plus a
working public default.

Also: seven release-readiness fixes, a README restructured for a stranger,
NEWS rewritten as an initial release rather than a pre-release churn log,
and the Gitea-canonical/GitHub-mirror/r-universe distribution mechanics.
2026-08-08 16:46:39 -04:00
jared d0b4bae3cc docs: plan for partial-coverage signposting (#9) 2026-08-04 21:31:15 -04:00
9 changed files with 1383 additions and 560 deletions
-33
View File
@@ -1,38 +1,5 @@
# uscogdata 0.1.0 (development)
## Signposting now catches partially-suppressed categories
* A coverage suggestion used to fire only when a category returned **no rows
at all** in a requested year. That missed the more dangerous case: a
category that still returns rows while silently dropping component codes
the wide era publishes only as aggregates (#9). `cog_spending(category =
"Public Welfare")` for FY2011 returned a plausible figure that omitted
`E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true
$5,261,404,000, a 39% understatement, with `provenance$suggestions` empty.
* Suggestions now also fire on **partial** coverage, and every suggestion
carries `trigger` (`"empty_year"` or `"suppressed_component"`),
`suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller
can see how much is missing and decide whether to re-run with the recipe.
* `cog_revenue()` gets the same fix through the shared verb path. Alaska's
FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping
$1,899,995,000 of aggregate-published `U4-` rents and royalties.
* The trigger stays recipe-driven, so it only fires where a harmonization
recipe actually exists to name the fix. `higher_ed_e18_wide` and
`general_gov_e89_wide` stay silent in every year measured on the bundled
fixture, because their components are ordinary classified leaves even
pre-2012.
* The `suppressed_component` trigger (and any `suppressed_amount`/
`suppressed_codes` an `empty_year` fire also carries) is scoped to the
calling verb's own flow family: `cog_spending()` only ever measures E/F/G
component dollars, `cog_revenue()` only T/A/U/B/C/D. A component from the
OTHER flow family reports `suppressed_amount = 0` rather than a fabricated
claim. The `empty_year` trigger itself is not flow-scoped -- a category
belonging to the other flow (e.g. `cog_spending(category = "IG Local")`)
still returns zero rows and can still fire, in any year including modern
ones, naming the recipe whose own generic join finds real data for this
government. That is a mis-scoped query, not a corpus-format gap, so its
`suppressed_amount` is correctly 0.
## New: `cog_balances()` for cash-and-security holdings
* New `cog_balances()` exposes the 14 cash-and-security holding codes
+1 -8
View File
@@ -125,15 +125,8 @@ cog_explain <- function(result, format = c("print", "list")) {
if (length(prov$suggestions) > 0L) {
cli::cli_h2("Suggestions")
sugg_lines <- vapply(prov$suggestions, function(s) {
line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label,
s$available_years[1], s$available_years[2], s$hint)
if (isTRUE(s$suppressed_amount > 0)) {
line <- paste0(line, sprintf(" [$%s excluded from %s: %s]",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
line
}, character(1))
cli::cli_ul(sugg_lines)
}
+1 -14
View File
@@ -395,8 +395,7 @@ cog_spending <- function(govid, years, category = NULL,
}
suggestions <- .build_suggestions(con, govid, years, category,
direct_leg_result,
resolved$basis, flow_prefixes,
.select_long_view(view_base, resolved$basis))
resolved$basis, flow_prefixes)
}
# C1(b): when expenditure_concept = "total", flag any row where the IG
@@ -513,18 +512,6 @@ cog_spending <- function(govid, years, category = NULL,
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
}
#' The `*_long`/`*_long_harmonized` view behind an annotated view base --
#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()`
#' anti-joins the LONG view rather than the annotated one: they have identical
#' row membership (the annotated views are the long views plus LEFT JOINs, see
#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the
#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule
#' the suppression test is asking about.
#' @noRd
.select_long_view <- function(view_base, basis) {
.select_view(sub("_annotated$", "_long", view_base), basis)
}
#' @noRd
.select_ig_view <- function(basis) {
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
+16 -82
View File
@@ -1,16 +1,8 @@
# R/suggestions.R
# Recipe-component-driven signposting. When a basis = "harmonized" query for
# a category comes back incomplete in some requested year -- and a
# harmonization recipe would actually fill it for this government -- surface
# that recipe as a suggestion. "Incomplete" has two forms, and a recipe
# qualifies on either:
# 1. empty_year -- the result has no rows at all in that year.
# 2. suppressed_component -- the result HAS rows, but a component code
# carries dollars the verb's own long view structurally excludes
# (aggregate-published, or absent from summary_categories). This is
# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping
# aggregate-only E67/E68, so form 1 never fired and the caller got a
# number a third too low with no signpost at all.
# Recipe-component-driven signposting: when a basis = "harmonized" query for
# a category comes back with a coverage gap in some requested years (the
# result has no rows at all in that year) that a harmonization recipe would
# actually fill for this government, surface that recipe as a suggestion.
#
# This is deliberately keyed off the recipe catalog's component codes, not
# off harmonization_map rows: no live map row carries a non-blank
@@ -56,15 +48,11 @@
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
#' lookup scoped to the calling verb's own flow family.
#' @param long_view Name of the verb's own long view (from
#' `.select_long_view()`), passed through to `.suppressed_components()` to
#' measure the second qualifying path (uscogdata#9).
#' @return List of `list(recipe_id, label, available_years, hint,
#' ig_recipe_id, trigger, suppressed_amount, suppressed_years,
#' suppressed_codes)`, possibly empty.
#' ig_recipe_id)`, possibly empty.
#' @noRd
.build_suggestions <- function(con, govid, years, category, result, basis,
flow_prefixes, long_view) {
flow_prefixes) {
if (!identical(basis, "harmonized") || is.null(category)) return(list())
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
@@ -98,28 +86,7 @@
unique(as.integer(result$year))
}
gap_years <- setdiff(as.integer(years), result_years)
# Path 2 (uscogdata#9): component dollars this government holds that the
# verb's own view structurally excludes. Measured across ALL requested
# years, not just gap years -- the whole point is that a year with rows can
# still be missing dollars. Scoped to the calling verb's own flow_prefixes
# (I1) -- see `.suppressed_components()`'s own roxygen for why.
#
# This runs unconditionally whenever there are candidates -- an earlier
# revision of this fix wave tried a free, in-memory pre-check
# (`.needs_suppression_query()`) to skip the round trip on an already-
# covered path, but a scoped re-review measured it against the fixture and
# found it didn't pay for itself (it skipped ~3% of healthy calls, ~0% of
# the multi-govid batch shape it was meant to help, at a net cost increase
# once its own always-run metadata query was counted) while adding an
# untested exactness invariant -- that `result$codes_included` and this
# anti-join share the harmonized `item_code` space -- whose silent
# violation would kill signposting, the exact failure class uscogdata#9
# exists to prevent. Owner's call: keep this simple; a batch-aware
# optimization, if one is worth building, is a separate issue.
supp <- .suppressed_components(con, candidates, govid, years, long_view, flow_prefixes)
if (length(gap_years) == 0L && nrow(supp) == 0L) return(list())
if (length(gap_years) == 0L) return(list())
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
"SELECT recipe_id, any_value(label) AS label,
@@ -130,12 +97,11 @@
.sql_lit_chr(candidates)
)))
# Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join
# covers for this government, restricted to the gap years.
covered <- if (length(gap_years) == 0L) {
data.frame(recipe_id = character(0), year = integer(0))
} else {
DBI::dbGetQuery(con, sprintf(
# Which (recipe_id, year) pairs the recipe's own generic join actually
# covers for this government, restricted to the gap years -- the same
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
# no is_aggregate filter), just checking existence instead of summing.
covered <- DBI::dbGetQuery(con, sprintf(
"SELECT DISTINCT r.recipe_id, l.year
FROM long l
JOIN harmonization_recipes r
@@ -150,35 +116,16 @@
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(gap_years, collapse = ",")
))
}
suggestions <- list()
for (rid in candidates) {
empty_hit <- rid %in% covered$recipe_id
s_rows <- supp[supp$recipe_id == rid, , drop = FALSE]
supp_hit <- nrow(s_rows) > 0L
if (!empty_hit && !supp_hit) next
if (!rid %in% covered$recipe_id) next
m <- meta[meta$recipe_id == rid, ]
suggestions[[length(suggestions) + 1L]] <- list(
recipe_id = rid,
label = m$label[[1]],
available_years = c(as.integer(m$year_min), as.integer(m$year_max)),
hint = sprintf("re-run with recipe = '%s'", rid),
# An empty year is the stronger claim -- the category returned nothing
# at all -- so it wins when both paths qualify. The suppressed_* fields
# are still populated, so an empty_year fire also reports its dollars.
trigger = if (empty_hit) "empty_year" else "suppressed_component",
suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0,
suppressed_years = if (supp_hit) {
sort(unique(as.integer(s_rows$year)))
} else {
integer(0)
},
suppressed_codes = if (supp_hit) {
sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE))))
} else {
character(0)
}
hint = sprintf("re-run with recipe = '%s'", rid)
)
}
.attach_ig_counterparts(con, suggestions, flow_prefixes)
@@ -285,25 +232,12 @@
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
#' continuation line is appended naming the intergovernmental counterpart
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
#' same bullet under cli, not a new bullet). Same treatment for
#' `suppressed_amount` (uscogdata#9): only present when dollars were
#' actually measured as excluded (an `empty_year` fire can carry them too --
#' see `.build_suggestions()` -- so this keys off the amount, not `trigger`).
#' same bullet under cli, not a new bullet).
#' @noRd
.inform_suggestions <- function(suggestions) {
bullets <- vapply(suggestions, function(s) {
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
s$available_years[1], s$available_years[2], s$hint)
# Only present when dollars were actually measured as excluded. An
# empty_year fire can carry them too -- the year had no rows AND the
# component was suppressed -- which is strictly more informative.
if (isTRUE(s$suppressed_amount > 0)) {
bullet <- paste0(bullet, sprintf(
"\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk",
formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","),
paste0("FY", s$suppressed_years, collapse = ", "),
paste(s$suppressed_codes, collapse = ", ")))
}
if (!is.null(s$ig_recipe_id)) {
bullet <- paste0(bullet, sprintf(
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
@@ -311,7 +245,7 @@
bullet
}, character(1))
cli::cli_inform(c(
i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:",
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
stats::setNames(bullets, rep("*", length(bullets)))
))
}
-115
View File
@@ -1,115 +0,0 @@
# R/suppression.R
# Split out of R/suggestions.R (2026-08-05) to keep files under the project's
# 400-line limit. Owns the second qualifying path for coverage signposting
# (uscogdata#9): measuring, per government, the component dollars the
# calling verb's own long view structurally excludes (aggregate-published,
# or absent from summary_categories). See R/suggestions.R for the
# orchestrator (`.build_suggestions()`) that calls this and the full
# uscogdata#9 background.
#' Measure, per (recipe, year), the component dollars this government holds
#' that the calling verb's own long view structurally excludes.
#'
#' This is the second qualifying path for a suggestion (uscogdata#9). The
#' first -- row absence -- only fires when a category returns NOTHING in a
#' requested year, which is how Corrections behaves in the wide era. Public
#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows,
#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967-
#' 2011, and absent from `summary_categories` entirely) are dropped. The
#' caller gets a plausible number a third too low, silently.
#'
#' "Structurally excluded" is decided by anti-joining the verb's REAL long
#' view rather than restating its WHERE clause, so this stays correct if
#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That
#' anti-join is keyed on `item_code`, which is sound only because
#' harmonization never renames a recipe component -- asserted by the "no
#' recipe component is ever renamed by harmonization" test in
#' tests/testthat/test-recipes.R.
#'
#' Note what this deliberately does NOT count as suppressed: a component
#' excluded from the RESULT for scoping reasons -- because it belongs to a
#' different `category`, or because `expenditure_concept` narrowed the
#' subtypes -- is still present in the view, so it never fires. Suggesting a
#' recipe is a coverage fix, not a category redefinition.
#'
#' `flow_prefixes` (uscogdata#9 review, finding I1) restricts the measured
#' components to the CALLING VERB's own flow family (`c("E","F","G")` for
#' spending, `c("T","A","U","B","C","D")` for revenue). Without this, a
#' candidate recipe belonging to the OTHER flow family is always absent from
#' this verb's view (by construction -- `cog_revenue()`'s view never carries
#' an E-coded row) and so was always reported as "suppressed", fabricating a
#' dollar claim across flow families (`cog_revenue(category = "Corrections")`
#' claimed $3.63B excluded that `cog_spending()` reports and fully accounts
#' for). Filtering on `LEFT(r.component_code, 1)` also drops M/L-prefixed
#' components from measurement under `cog_spending()` (`flow_prefixes` never
#' includes "M"/"L") -- harmless today, because a recipe's own M/L components
#' (e.g. `corrections_ig_local_combined`'s M04/M05) are present in the view
#' in every year they exist and so never fired as suppressed anyway, but
#' worth recording since this filter is now the thing relied on to prevent
#' it.
#'
#' @param con Active DuckDB connection.
#' @param candidates Character vector of recipe ids to measure.
#' @param govid Character vector of canonical_govid values.
#' @param years Integer vector of requested years.
#' @param long_view Name of the verb's long view, from `.select_long_view()`.
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
#' `.build_suggestions()`). Only recipe components whose first character is
#' in this set are measured.
#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US
#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when
#' nothing is suppressed.
#' @noRd
.suppressed_components <- function(con, candidates, govid, years, long_view,
flow_prefixes) {
empty <- tibble::tibble(
recipe_id = character(0), year = numeric(0),
suppressed_amount = numeric(0), suppressed_codes = character(0)
)
if (length(candidates) == 0L) return(empty)
# long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can
# never be quoted safely. It is always internally derived from a fixed
# view_base, so an off-allowlist value is a programming error, not input.
if (!long_view %in% c("spending_long", "spending_long_harmonized",
"revenue_long", "revenue_long_harmonized")) {
cli::cli_abort(
"Internal error: unexpected `long_view` {.val {long_view}}.",
class = "uscogdata_internal_error"
)
}
sql <- sprintf(
"SELECT r.recipe_id,
l.year,
SUM(l.amt) * 1000.0 AS suppressed_amount,
string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code)
AS suppressed_codes
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
AND (r.gov_type_scope = 'all'
OR (r.gov_type_scope = 'state' AND l.type = 0)
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
WHERE r.recipe_id IN (%1$s)
AND l.canonical_govid IN (%2$s)
AND l.year IN (%3$s)
AND l.amt <> 0
AND LEFT(r.component_code, 1) IN (%5$s)
AND NOT EXISTS (
SELECT 1 FROM %4$s v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year
AND v.item_code = l.item_code
AND v.year IN (%3$s) -- restated: enables partition pruning (I3a)
AND v.canonical_govid IN (%2$s) -- restated: pushes the govid filter (I3a)
)
GROUP BY 1, 2
ORDER BY 1, 2",
.sql_lit_chr(candidates), .sql_lit_chr(govid),
paste(as.integer(years), collapse = ","), long_view,
.sql_lit_chr(flow_prefixes)
)
tibble::as_tibble(DBI::dbGetQuery(con, sql))
}
+1 -42
View File
@@ -33,48 +33,7 @@
},
"harmonization": { "type": "object" },
"recipe": { "type": ["object", "null"] },
"suggestions": {
"type": "array",
"description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).",
"items": {
"type": "object",
"required": ["recipe_id", "label", "available_years", "hint", "ig_recipe_id",
"trigger", "suppressed_amount", "suppressed_years", "suppressed_codes"],
"properties": {
"recipe_id": { "type": "string" },
"label": { "type": "string" },
"available_years": {
"type": "array",
"items": { "type": "integer" },
"description": "[year_min, year_max] of the recipe's component coverage."
},
"hint": { "type": "string" },
"ig_recipe_id": {
"type": ["string", "null"],
"description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes."
},
"trigger": {
"type": "string",
"enum": ["empty_year", "suppressed_component"],
"description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars this government reports in the requested years that the verb's underlying long view structurally excludes -- aggregate-published, carrying no harmonized code, or absent from summary_categories. This is NOT the same thing as 'excluded from the result': a component present in the view under a different category (a scoping choice, e.g. a different `category` or a narrower `expenditure_concept`) contributes 0 and never fires. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way, using the same underlying-view measurement, and can be 0 even on an 'empty_year' fire."
},
"suppressed_amount": {
"type": "number",
"description": "Full US dollars this government reports, in the recipe's component codes, in the requested years, that the verb's underlying long view structurally excludes (aggregate-published, carrying no harmonized code, or absent from summary_categories) -- summed across those years. This is NOT the same quantity as 'what the result excludes': a component present in the view under a different category or a narrower `expenditure_concept` is scoped out on purpose, counts as 0 here, and is not suppression. 0 does not always mean full coverage -- see 'trigger' and 'empty_year'. May be negative where Census publishes a negative `amt` for the excluded rows."
},
"suppressed_years": {
"type": "array",
"items": { "type": "integer" },
"description": "The requested years contributing to suppressed_amount."
},
"suppressed_codes": {
"type": "array",
"items": { "type": "string" },
"description": "The excluded component item codes, sorted."
}
}
}
},
"suggestions": { "type": "array" },
"scope": { "type": "object" },
"codes_summed": { "type": "object" },
"aggregate_fallback": { "type": ["object", "null"] },
File diff suppressed because it is too large Load Diff
+286
View File
@@ -0,0 +1,286 @@
# `uscogdata` 0.1.0 — public release
**Date:** 2026-08-08 · **Status:** design, awaiting approval
**Scope:** release-readiness, README, NEWS. Distribution mechanics recorded here as
decided, sequenced after the package is clean.
`uscogdata` is feature-complete and the corpus it reads has been public on
HuggingFace since 2026-08-07 (294 downloads as of this writing). The API built on
it is live. What does not exist is a public *package*: the repo is private, there
is no install path, and — measured, not assumed — **a stranger who installed it
today could not read the corpus at all.**
This spec covers making that untrue.
## Decisions locked
| Decision | Choice |
|---|---|
| Canonical source | `gitea.civilytics.org/Civilytics/uscogdata`, flipped public |
| Public mirror | `github.com/civilytics/uscogdata` — issues, PRs, multi-OS check, CDN |
| Mirror mechanism | Gitea Actions non-force `git push` (not a push mirror) |
| Binaries | `civilytics.r-universe.dev`, registry pinned to a release tag |
| Author of record | Jared E. Knowles `<jared@civilytics.com>`, ORCID `0000-0003-0005-9478` |
| Copyright | Civilytics Consulting LLC (`cph`, `fnd`) |
| License | MIT (package) · CC-BY-4.0 (corpus) |
| Corrections intake | Deferred — see *Out of scope* |
| Other packages | Parked until this one walks the path end to end |
## P0 — the corpus is unreachable
Two independent faults, either of which alone is fatal.
**No corpus URL exists.** `R/config.R` defaults to the literal
`REPLACE_WITH_SHARE_TOKEN` sentinel, and no file in the repo supplies a working
one. A new user calling any verb gets `uscogdata_url_not_configured` with no path
to resolution.
**Remote reads are broken regardless.** Every partitioned view globs:
```sql
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true)
```
DuckDB 1.5.5 refuses globs over generic HTTP. Its suggested
`allow_asterisks_in_http_paths` escape hatch does not help — it forwards the
literal `**/*` as a filename and 404s, because plain HTTP exposes no directory
listing to expand against.
The package therefore works only against a **local path**. That is how the API
runs it (`CORPUS_HOST_PATH` is a host mount on maxwell) and how the tests run
(bundled fixture), which is why the fault went unnoticed. The README's headline
claim — *"Reads the published corpus directly from Nextcloud via DuckDB httpfs —
no local bulk downloads required"* — is currently false.
### Fix: enumerate from the manifest, do not glob
`manifest.json` already lists every partition under `files.long_partitions[]`
with `path`, `year`, `sha256`, `row_count` and `size_bytes` — 56 of them.
Substituting an explicit file list for the glob was measured against the
published corpus on 2026-08-08:
| Path | Result |
|---|---|
| `https://…/data/long/**/*.parquet` (default) | error — globs unsupported over HTTP |
| same, `allow_asterisks_in_http_paths = true` | error — literal `**/*` 404s |
| `hf://datasets/civilytics/us-cog-finance/…` glob | 46,148,034 rows |
| **explicit list over plain https** | **46,148,034 rows** |
`hive_partitioning = true` still recovers `year` from the paths under
enumeration, so no downstream view or verb changes.
Enumeration is preferred over `hf://` deliberately. It is **host-agnostic** —
Nextcloud, HuggingFace, or any static server take the same code path — where
`hf://` would tie the default to one vendor's protocol and still need
special-casing, since manifest fetching goes through `httr2`, which cannot speak
`hf://`. Enumeration also *removes* a dependency (globbing) rather than adding
one, and the manifest's per-file `sha256` becomes available for integrity
checking later.
Views are registered from `inst/sql/` with `{url}` substitution in
`R/views.R:.register_views()`. The list must be built once per session from the
already-fetched manifest and substituted the same way, so the change is confined
to view registration and does not touch verb code.
### Fix: ship a working default
`R/config.R`'s default becomes the public HuggingFace `resolve/main/` URL:
CC-BY-4.0, no token to publish, CDN-backed, and it keeps maxwell's uplink out of
the path — the same reasoning behind the GitHub mirror and r-universe.
This means `library(uscogdata)` followed by a verb works with **zero
configuration**, which is what makes the package demonstrable in a README and
later in a post. `USCOGDATA_URL` and `options(uscogdata.url=)` continue to
override, so the Nextcloud copy and local mirrors are unaffected.
The `uscogdata_url_not_configured` error class stays — it still fires for an
explicitly-set empty or placeholder URL — but ceases to be the default
experience.
### Consequence: `cog_mirror()` is promoted
Measured cost of the remote default, from efron on a good connection:
| | |
|---|---|
| Whole corpus | **190.6 MB**, 56 partitions, 46,148,034 rows, FY1967–FY2024 |
| One government, one year | 1.5 s |
| One government, all 56 years | 2.8 s |
| Disk written | **0.00 MB** — range requests only; `external_file_cache` is in-memory |
Nothing persists locally beyond the shared `httpfs` extension in `~/.duckdb` (a
few MB, once per machine, across all DuckDB use). Costs are RAM and per-query
bandwidth, since nothing caches between sessions.
Those timings are raw scans. Real verbs additionally join crosswalks, resolve
categories and assemble provenance, so end-to-end verb latency will be higher and
**must be re-measured once the fix lands** — it cannot be measured today.
The corpus being only 190.6 MB makes `cog_mirror()` a first-class option rather
than a developer footnote. The README presents **both paths**:
- **Remote (default, zero setup)** — trying it out, teaching, one-off questions.
- **Mirrored (`cog_mirror()`, 190 MB once)** — repeated or heavy analysis,
offline work, reproducibility, or preferring not to depend on HuggingFace.
The second is also the honest answer to the vendor-dependency question raised by
defaulting to HuggingFace: **the escape hatch is one function call and 190 MB**,
after which no analysis touches an external service. The README says so
explicitly. That is the difference between a convenience default and lock-in.
## Release-readiness fixes
| # | Issue | Fix |
|---|---|---|
| 1 | `MaxCorpusSchema: 5` in DESCRIPTION; `.validate_schema()` accepts `4,5,6,7`; published corpus is **7** | `MaxCorpusSchema: 7` |
| 2 | `^vignettes$` in `.Rbuildignore` — both vignettes absent from the installed package, while README tells users to run `vignette("total-spending")` | Remove `^vignettes$`, `^doc$`, `^Meta$`. Both vignettes build offline (`total-spending` reads the bundled fixture; `population-denominators` is `eval = FALSE`) |
| 3 | `_pkgdown.yml` reference index covers 6 of 14 exports — pkgdown errors on missing topics | Add `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`; set `url:` |
| 4 | No `URL:` / `BugReports:` in DESCRIPTION | Add both, pointing at the GitHub mirror |
| 5 | No `LICENSE.md`; `LICENSE` holder reads `Civilytics` | `usethis::use_mit_license("Civilytics Consulting LLC")` |
| 6 | README instructs stripping the fixture at release | Delete that section — see below |
| 7 | `Authors@R` is an org with no human | Jared E. Knowles `aut`/`cre` + ORCID; Civilytics Consulting LLC `cph`/`fnd` |
**On #6.** The advice to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`
is CRAN-sized thinking (5 MB limit) and this package is not going to CRAN.
Stripping the 15 MB fixture would break `total-spending.Rmd`, which reads from
it, and would leave r-universe and GitHub Actions unable to run the 28 test files
without a corpus credential. **The fixture is what lets `R CMD check` pass
anywhere with zero secrets** — precisely what public CI needs. It ships.
## README
The current README addresses someone standing inside the repo tree: status reads
"Under active development (Phase 2 of the cog_pipeline project)", it points at
`../cog_pipeline/docs/reader-specification.md`, the install line is commented
out, and developer, testing and release sections sit above anything a user needs.
Restructured around a stranger, in this order:
1. **What this is** — one paragraph, and what the corpus covers (types 0–3,
FY1967–FY2024, 46M rows, 190.6 MB).
2. **Install** — r-universe first (binaries), git second.
3. **Quickstart that actually runs** — resolve a government, get its history,
print provenance. No configuration step.
4. **Two ways to read the corpus** — remote default vs `cog_mirror()`, with the
measured numbers and the independence note.
5. **Amounts are in full US dollars** — kept near the top. This is the errata
most likely to produce a wrong answer that looks plausible.
6. **Concepts** — primary/direct/total spending, general/total revenue,
coverage. Condensed, linking to the vignettes for the full treatment.
7. **How to cite** — `citation("uscogdata")`, corpus CC-BY-4.0 attribution.
8. **Contributing** — canonical-on-Gitea PR flow.
Developer notes, testing instructions and release procedure move to
`CONTRIBUTING.md`. Every path reference to a sibling repo is removed or replaced
with a URL that resolves for someone who has only this repo.
## NEWS.md
The current NEWS is a pre-release churn log: changes described relative to states
no user has seen ("Breaking: corpus schema_version 4", "the package now
requires…"), newest-first across the package's entire pre-release development
(2026-04-23 to 2026-08-04, 140 commits). To a newcomer evaluating whether to
depend on the package, it reads as instability.
**0.1.0 is rewritten as an initial release**: what the package does, what the
corpus covers, and the caveats that are genuinely load-bearing. The pre-release
history is not preserved in NEWS — it is in git, where it belongs.
The substantive content is migrated, not deleted. These are hard-won and belong
in documentation rather than buried in a changelog:
| Content | Destination |
|---|---|
| Coverage disclosure on multi-government aggregates (census vs sample years) | README concepts + `cog_geographic_rollup()` docs |
| `complete = TRUE` three-way absence semantics (`reported` / `census_zero` / `not_reported`) | `cog_spending()` / `cog_revenue()` docs |
| Series-break and corpus-break surfacing | README + `cog_explain()` docs |
| $1,000s → full dollars conversion | README, already prominent |
| Per-year F-33 population denominators | `population-denominators` vignette, already there |
This also makes NEWS reusable as raw material for the release announcement,
which is the stated downstream purpose.
## Distribution mechanics
Recorded as decided; executed after the package is clean and checks are green.
**Sequence matters.** r-universe publishes check results the moment a package is
registered. Registering before the fixes above land means a red badge on day one,
which is a worse first impression than a week's delay.
1. `gitleaks` over full history. A coarse grep found nothing across 140 commits
and the default corpus URL is still the placeholder sentinel, but a proper
scan is the gate on an irreversible action.
2. Flip the Gitea repo public. Disable Gitea issues on it, so there is exactly
one inbox.
3. Create `github.com/civilytics/uscogdata`. Add `.github/workflows/` for the
Windows/macOS/Linux `R CMD check` matrix — the platforms the Gitea runner
cannot provide, and which this package has never been tested on despite
depending on duckdb and httr2. Gitea reads `.gitea/workflows`, GitHub reads
`.github/workflows`; both live in one tree without colliding.
4. Gitea Actions workflow pushing to GitHub **without `--force`**, so divergence
fails loudly in CI rather than silently overwriting.
5. Add `jared@civilytics.com` as a verified secondary email on the GitHub
account — r-universe links maintainer identity by matching DESCRIPTION's email
against registered GitHub emails, and the association only takes effect on the
next build.
6. Tag `v0.1.0`. Create `github.com/civilytics/civilytics.r-universe.dev` with a
`packages.json` pinned to the tag, pointing at the GitHub mirror rather than
Gitea so clone traffic stays off maxwell. Install the r-universe app.
### PR flow
Never press Merge on GitHub. A merge there is overwritten by the next sync, the
PR still displays "Merged", and nothing says otherwise.
```sh
git remote add github https://github.com/civilytics/uscogdata.git
git config --add remote.github.fetch '+refs/pull/*/head:refs/remotes/github/pr/*'
git fetch github
git switch -c pr-42 github/pr/42 # test
git switch main && git merge --no-ff pr-42
git push origin main # Gitea -> mirror -> GitHub
```
GitHub auto-closes a PR as merged once its head commit becomes an ancestor of the
base branch, so `--no-ff` — which preserves the contributor's SHAs — makes the PR
close itself when the mirror pushes. **For external PRs, merge; do not squash or
rebase.** Squashing rewrites the SHAs, the auto-close never fires, and closing by
hand reads to a first-time contributor as rejection.
`CONTRIBUTING.md` states this, and a GitHub Action comments it on incoming PRs.
No CLA; no DCO.
## Verification
The release is not done until all of these pass:
1. `R CMD check --as-cran` clean on Linux, and on Windows and macOS via the
GitHub matrix. This package has never been checked on the latter two.
2. Full test suite (28 files) green against the **bundled fixture**, offline,
with no credentials — the property public CI depends on.
3. Full test suite green against the **live corpus**, which additionally
exercises the enumeration fix that the fixture's local path cannot.
4. `pkgdown::build_site()` completes.
5. Both vignettes present in the built tarball and
`vignette("total-spending", package = "uscogdata")` resolves from an
installed copy.
6. **Cold-start check on a machine that has never seen this package:** install
from r-universe, `library(uscogdata)`, run the README quickstart verbatim with
no environment variables set. This is the only test that catches the P0 class
of fault, and its absence is why the fault survived.
7. End-to-end verb latency re-measured against the live corpus and the README's
numbers updated if they moved.
## Out of scope
- **Corrections intake.** Deferred by decision. Consequence: the release cannot
invite data-error reports or make the "traceable and correctable" claim that
most distinguishes this corpus from Census's own files. `BugReports:` points at
package issues only. A verified correction should eventually terminate as a
`lineage_event` or `series_break` row so it propagates through provenance to
every consumer — that design is unstarted.
- **Announcement posts.** Deferred. The API announcement is gated on corrections
landing and merits a Civic Pulse edition.
- **The rest of the R package backlog.** Parked until this one completes the path.
- **`cog_pipeline` publication.** Stays private.
-252
View File
@@ -214,255 +214,3 @@ test_that("no signposting under basis = 'raw'", {
prov <- attr(r, "provenance")
expect_length(prov$suggestions, 0L)
})
# --- uscogdata#9: partial-coverage signposting ------------------------------
test_that("no recipe component is ever renamed by harmonization", {
# The suppression trigger anti-joins the verb's long view on item_code.
# That is only sound because harmonization never rewrites a recipe
# component's code -- every component whose harmonized_code differs has
# harmonized_code IS NULL (and is aggregate-flagged). If this ever fails,
# .suppressed_components() would report reachable dollars as suppressed.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS renamed FROM long
WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes)
AND harmonized_code IS NOT NULL
AND harmonized_code <> item_code")$renamed
expect_equal(as.integer(n), 0L)
})
test_that(".select_long_view maps annotated view bases to their long views", {
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "harmonized"),
"spending_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("revenue_annotated", "harmonized"),
"revenue_long_harmonized")
expect_equal(
uscogdata:::.select_long_view("spending_annotated", "raw"),
"spending_long")
})
test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2011L,
long_view = "spending_long_harmonized",
flow_prefixes = c("E", "F", "G"))
expect_s3_class(s, "tbl_df")
expect_equal(nrow(s), 2L)
s <- s[order(s$recipe_id), ]
expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
expect_equal(s$suppressed_amount, c(1803872000, 271589000))
expect_equal(s$suppressed_codes, c("E67", "E68"))
})
test_that(".suppressed_components finds nothing in a modern year", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2019L,
long_view = "spending_long_harmonized",
flow_prefixes = c("E", "F", "G"))
expect_equal(nrow(s), 0L)
})
test_that(".suppressed_components rejects a long_view outside the allowlist", {
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
expect_error(
uscogdata:::.suppressed_components(
con, candidates = "welfare_cash_e67_wide", govid = "061037123085",
years = 2011L, long_view = "long; DROP TABLE x",
flow_prefixes = c("E", "F", "G")),
class = "uscogdata_internal_error")
})
test_that(".suppressed_components never measures a component from the other flow family (I1)", {
# uscogdata#9 review, finding I1: without the flow_prefixes filter, a
# candidate recipe entirely outside the calling verb's own flow family is
# ALWAYS absent from that verb's view (by construction), so it was always
# reported as "suppressed" -- fabricating a dollar claim. E67/E68 are
# Public Welfare EXPENDITURE codes; scoping the measurement to revenue's
# own flow_prefixes must find nothing for them.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
s <- uscogdata:::.suppressed_components(
con,
candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"),
govid = "061037123085", years = 2011L,
long_view = "revenue_long_harmonized",
flow_prefixes = c("T", "A", "U", "B", "C", "D"))
expect_equal(nrow(s), 0L)
})
test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", {
# The bug: E74/E79 return rows for FY2011, so there is no row-absence gap,
# so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were
# dropped for being aggregate-published. LA County reports $3,185,943,000
# and omits $2,075,461,000, a 39% understatement, silently.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 2L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide"))
e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]]
expect_equal(e67$trigger, "suppressed_component")
expect_equal(e67$suppressed_amount, 1803872000)
expect_equal(e67$suppressed_years, 2011L)
expect_equal(e67$suppressed_codes, "E67")
expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'")
e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]]
expect_equal(e68$trigger, "suppressed_component")
expect_equal(e68$suppressed_amount, 271589000)
expect_equal(e68$suppressed_codes, "E68")
})
test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", {
# Corrections is the case that already worked: zero rows in FY2011, so the
# row-absence path fires. It must keep firing, keep trigger = "empty_year",
# keep its IG counterpart -- and now also report what was suppressed.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Corrections"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 3L)
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_setequal(ids, c("corrections_combined", "corrections_capital_combined",
"corrections_other_capital_combined"))
expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year"))
cc <- sugg[[which(ids == "corrections_combined")]]
expect_equal(cc$suppressed_amount, 1371460000)
expect_equal(cc$suppressed_codes, "E05")
expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined")
})
test_that("uscogdata#9: the revenue verb inherits the same trigger", {
# Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from
# U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-`
# rents and royalties -- the omission is LARGER than the reported figure.
skip_if_no_corpus()
r <- suppressMessages(
cog_revenue("020000227749", years = 2011L,
category = "Miscellaneous Revenue"))
sugg <- attr(r, "provenance")$suggestions
expect_length(sugg, 1L)
expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide")
expect_equal(sugg[[1]]$trigger, "suppressed_component")
expect_equal(sugg[[1]]$suppressed_amount, 1899995000)
expect_equal(sugg[[1]]$suppressed_codes, "U4-")
# A revenue recipe must never be handed an M/L expenditure counterpart.
expect_null(sugg[[1]]$ig_recipe_id)
})
test_that("I1: cog_revenue never fabricates suppressed dollars for an expenditure-only recipe", {
# uscogdata#9 review, finding I1: Corrections is an expenditure-only
# category (E04/E05). cog_revenue() naturally returns zero rows for it, so
# corrections_combined still fires as an empty_year suggestion (its own
# generic join finds real E04/E05 data for this government) -- but before
# the flow_prefixes fix, .suppressed_components() measured E04/E05 against
# cog_revenue()'s OWN view (which can never contain an E-coded row by
# construction) and reported the full $3,631,945,000 as "suppressed",
# when cog_spending() for the same gov/years/category actually returns
# $3,691,029,000 -- nothing was suppressed at all.
skip_if_no_corpus()
r <- suppressMessages(
cog_revenue("061037123085", years = 2019:2020, category = "Corrections"))
sugg <- attr(r, "provenance")$suggestions
ids <- vapply(sugg, function(s) s$recipe_id, character(1))
expect_true("corrections_combined" %in% ids)
hit <- sugg[[which(ids == "corrections_combined")]]
expect_equal(hit$suppressed_amount, 0)
expect_equal(hit$suppressed_years, integer(0))
expect_equal(hit$suppressed_codes, character(0))
# And cog_spending() for the identical gov/years/category is unaffected --
# it actually finds the E04/E05 dollars the buggy measurement claimed were
# excluded.
sp <- suppressMessages(
cog_spending("061037123085", years = 2019:2020, category = "Corrections"))
expect_equal(sum(sp$amt_nominal), 3691029000)
})
test_that("uscogdata#9: no partial-coverage fire in a modern year", {
skip_if_no_corpus()
r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare")
expect_length(attr(r, "provenance")$suggestions, 0L)
})
test_that("uscogdata#9: leaf-and-classified wide-era families never fire", {
# higher_ed_e18_wide and general_gov_e89_wide are the control group: their
# components (E16/E18, E85/E89) are ordinary classified leaves even in the
# wide era, so widening the trigger must leave them silent. This is the
# measurement that refutes "it would fire on every category in every legacy
# year" -- corpus-wide on the fixture, these two produce zero suppressed rows.
skip_if_no_corpus()
con <- uscogdata:::.ensure_session()
n <- DBI::dbGetQuery(con,
"SELECT COUNT(*) AS n
FROM long l
JOIN harmonization_recipes r
ON l.item_code = r.component_code
AND l.year BETWEEN r.year_min AND r.year_max
WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide')
AND l.amt <> 0
AND NOT EXISTS (
SELECT 1 FROM spending_long_harmonized v
WHERE v.canonical_govid = l.canonical_govid
AND v.year = l.year AND v.item_code = l.item_code)")$n
expect_equal(as.integer(n), 0L)
})
test_that("uscogdata#9: the cli message reports the suppressed dollars", {
skip_if_no_corpus()
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"1,803,872,000", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"FY2011", fixed = TRUE)
expect_message(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"),
"E67", fixed = TRUE)
})
test_that("uscogdata#9: cog_explain() reports the suppressed dollars", {
# cog_explain()'s whole "print" output -- including the Suggestions
# section built from cli::cli_ul() -- is emitted on the message stream
# (verified empirically 2026-08-04: capture.output(..., type = "output")
# returns character(0) for this call; testthat::capture_messages() is what
# actually carries it), so that is the stream this test captures.
skip_if_no_corpus()
r <- suppressMessages(
cog_spending("061037123085", years = 2011L, category = "Public Welfare"))
out <- paste(testthat::capture_messages(cog_explain(r)), collapse = "")
expect_match(out, "271,589,000", fixed = TRUE)
})
test_that("the provenance schema documents the suggestion trigger fields", {
sch <- jsonlite::fromJSON(
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
simplifyVector = FALSE)
props <- sch$properties$suggestions$items$properties
expect_true(all(c("trigger", "suppressed_amount", "suppressed_years",
"suppressed_codes") %in% names(props)))
expect_setequal(unlist(props$trigger$enum),
c("empty_year", "suppressed_component"))
})