feat: uscogdata 0.3.0 — first public release #41
@@ -1,3 +1,63 @@
|
||||
# uscogdata 0.3.0
|
||||
|
||||
First public release.
|
||||
|
||||
`uscogdata` provides curated R verbs over the Civilytics US Census of
|
||||
Governments finance corpus: unit-level financial profiles, geographic rollups
|
||||
and peer comparisons, with auditable provenance on every result.
|
||||
|
||||
## What it covers
|
||||
|
||||
Government types 0-3 (state, county, municipality, township), FY1967-FY2024 --
|
||||
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
|
||||
or FY1969. Special districts (type 4) and school districts (type 5) are out of
|
||||
scope pending validation.
|
||||
|
||||
## The verbs
|
||||
|
||||
`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and holdings;
|
||||
`cog_gov_search()` to resolve place names (including basket mode for many at
|
||||
once); `cog_find_peers()` and `cog_peer_compare()` for cohorts;
|
||||
`cog_geographic_rollup()` for aggregates; `cog_categories()`, `cog_recipes()`,
|
||||
`cog_manifest()` and `cog_explain()` for metadata and provenance; and
|
||||
`cog_mirror()` for a local copy of the corpus.
|
||||
|
||||
## Reading the corpus now works out of the box
|
||||
|
||||
* The package reads the published corpus over HTTPS **with no configuration**.
|
||||
Previously the default was a placeholder sentinel and no document in the
|
||||
package supplied a working URL, so a new user had no path to a session.
|
||||
* Remote reads work at all. The partitioned view used a glob, and DuckDB
|
||||
cannot expand a glob over generic HTTP -- there is no directory listing to
|
||||
expand against. Partition paths are now enumerated from the corpus manifest,
|
||||
which is host-agnostic: an HTTPS mirror, a Nextcloud share and a local
|
||||
`cog_mirror()` copy all take the same path.
|
||||
* Nothing is written to disk in remote mode; DuckDB fetches only the row
|
||||
groups a query needs.
|
||||
|
||||
## Four things to know before your first query
|
||||
|
||||
* **Amounts are in full US dollars.** The raw Census files report thousands;
|
||||
the verbs multiply by 1000 on the way out. Do not multiply again.
|
||||
* **Multi-government aggregates disclose their coverage.** The Census is a
|
||||
complete enumeration only in years ending in 2 and 7; every other year is a
|
||||
sample. Every such result carries `provenance$coverage` with per-year
|
||||
`n_units_reporting`.
|
||||
* **Absence means two different things.** Before FY2012 an absent cell means
|
||||
Census published $0; from FY2012 it means not reported. `complete = TRUE`
|
||||
labels which.
|
||||
* **Series breaks reach you unasked.** Catalogued breaks intersecting your
|
||||
query appear in provenance and in `cog_explain()`.
|
||||
|
||||
## Known limits
|
||||
|
||||
* Special districts (type 4) and school districts (type 5) are out of scope.
|
||||
* Per-capita rollups exclude governments with no F-33 population, which is by
|
||||
design but does silently narrow a rollup.
|
||||
* `n_units_reporting` is category-conditional and is not a response rate.
|
||||
* Employee-retirement (`X`) codes stop at FY2016, when those systems moved to
|
||||
the Annual Survey of Public Pensions.
|
||||
|
||||
# uscogdata 0.2.0
|
||||
|
||||
## New features
|
||||
@@ -39,262 +99,3 @@
|
||||
is not a response rate: a government that was surveyed and genuinely spends
|
||||
nothing in the requested category is indistinguishable from one never
|
||||
surveyed (uscogdata#36).
|
||||
|
||||
# uscogdata 0.1.0 (development)
|
||||
|
||||
## Signposting now catches partially-suppressed categories
|
||||
|
||||
* A coverage suggestion used to fire only when a category returned **no rows
|
||||
at all** in a requested year. That missed the more dangerous case: a
|
||||
category that still returns rows while silently dropping component codes
|
||||
the wide era publishes only as aggregates (#9). `cog_spending(category =
|
||||
"Public Welfare")` for FY2011 returned a plausible figure that omitted
|
||||
`E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true
|
||||
$5,261,404,000, a 39% understatement, with `provenance$suggestions` empty.
|
||||
* Suggestions now also fire on **partial** coverage, and every suggestion
|
||||
carries `trigger` (`"empty_year"` or `"suppressed_component"`),
|
||||
`suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller
|
||||
can see how much is missing and decide whether to re-run with the recipe.
|
||||
* `cog_revenue()` gets the same fix through the shared verb path. Alaska's
|
||||
FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping
|
||||
$1,899,995,000 of aggregate-published `U4-` rents and royalties.
|
||||
* The trigger stays recipe-driven, so it only fires where a harmonization
|
||||
recipe actually exists to name the fix. `higher_ed_e18_wide` and
|
||||
`general_gov_e89_wide` stay silent in every year measured on the bundled
|
||||
fixture, because their components are ordinary classified leaves even
|
||||
pre-2012.
|
||||
* The `suppressed_component` trigger (and any `suppressed_amount`/
|
||||
`suppressed_codes` an `empty_year` fire also carries) is scoped to the
|
||||
calling verb's own flow family: `cog_spending()` only ever measures E/F/G
|
||||
component dollars, `cog_revenue()` only T/A/U/B/C/D. A component from the
|
||||
OTHER flow family reports `suppressed_amount = 0` rather than a fabricated
|
||||
claim. The `empty_year` trigger itself is not flow-scoped -- a category
|
||||
belonging to the other flow (e.g. `cog_spending(category = "IG Local")`)
|
||||
still returns zero rows and can still fire, in any year including modern
|
||||
ones, naming the recipe whose own generic join finds real data for this
|
||||
government. That is a mis-scoped query, not a corpus-format gap, so its
|
||||
`suppressed_amount` is correctly 0.
|
||||
|
||||
## New: `cog_balances()` for cash-and-security holdings
|
||||
|
||||
* New `cog_balances()` exposes the 14 cash-and-security holding codes
|
||||
(`category_type = "balance"`): fund balances, retirement system holdings and
|
||||
insurance trust balances (#25). Holdings are a stock, not a flow, so the verb
|
||||
has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and
|
||||
no `subtype` argument either -- for holdings, `category` is a strict
|
||||
coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly
|
||||
the `general` family (`W01`/`W31`/`W61`).
|
||||
* `cog_balances()` results carry `provenance$balance_caveats`, recording that
|
||||
Census holdings are gross rather than GAAP fund balance, and the measured
|
||||
coverage window of each subtype family.
|
||||
|
||||
## Multi-government aggregates now disclose their reporting coverage
|
||||
|
||||
* The Census of Governments is a **complete census only in years ending in 2
|
||||
and 7**; every other year is a sample, and the sample varies enormously. On
|
||||
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
|
||||
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
|
||||
return value said nothing about, so a statewide total resting on a fifth of
|
||||
the universe looked exactly like one resting on all of it.
|
||||
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
|
||||
`coverage`:
|
||||
|
||||
| value | effect |
|
||||
|---|---|
|
||||
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
|
||||
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
|
||||
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
|
||||
|
||||
* **Regardless of mode**, every result now carries `provenance$coverage` with
|
||||
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
|
||||
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
|
||||
section. So the default mode can no longer mislead silently.
|
||||
* `is_census_year` is a statement about the **survey calendar**, never a claim
|
||||
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
|
||||
cities report. `n_units_reporting` is the number that tells the truth.
|
||||
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
|
||||
it is the subject of the comparison, not a member of the cohort — and the
|
||||
`summary_*` quantiles are computed after the filter, so they describe the
|
||||
cohort actually returned. `n_units_reporting` counts peers only, against the
|
||||
cohort size.
|
||||
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
|
||||
is `NULL`: `"census"` snaps to the most recent census year with an observed
|
||||
population, so a cohort is not built from a sample year in which most of the
|
||||
candidate universe is absent.
|
||||
|
||||
## `complete = TRUE`: absent cells, labelled with why they are absent
|
||||
|
||||
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
|
||||
(today's behaviour). With `complete = TRUE` the requested grid is filled
|
||||
from the corpus's `code_set` table and every row carries a new
|
||||
`value_source` column:
|
||||
|
||||
| `value_source` | meaning | `amt_nominal` |
|
||||
|---|---|---|
|
||||
| `reported` | the corpus carries this cell | as published |
|
||||
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
|
||||
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
|
||||
|
||||
The `NA` is deliberate and is the whole point: filling a modern absence
|
||||
with `0` would invent data, which is precisely the error the corpus's
|
||||
representation contract exists to prevent.
|
||||
* This restores information the reader lost when the corpus was sparsified
|
||||
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
|
||||
had begun returning nothing at all — and improves on what came before it,
|
||||
since the pre-sparsification corpus could not distinguish a published zero
|
||||
from an unreported cell either.
|
||||
* The grid is scoped to each government's **own type**, so a county is never
|
||||
filled with cells only a state can report.
|
||||
* Needs a corpus published from 2026-07-29 onward (when `representation` and
|
||||
`code_set` began shipping); aborts with class
|
||||
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
|
||||
listing those tables rather than on `schema_version`, which was never
|
||||
bumped for the change. Not available with `recipe` or
|
||||
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
|
||||
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
|
||||
`absence_means` rule; `cog_explain()` prints a "Completion" section.
|
||||
|
||||
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
|
||||
|
||||
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
|
||||
corpus as a whole rather than about one item code. `series_break_refs` is
|
||||
built by matching `fin_code` against the item codes in the result, and no
|
||||
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
|
||||
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
|
||||
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
|
||||
representation change at FY2012) and `SB086` (the government id scheme
|
||||
change at FY2017).
|
||||
* Provenance gains `corpus_break_refs`, selected on the break-year window
|
||||
alone and disjoint from `series_break_refs` by construction, so a consumer
|
||||
can tell a whole-result caveat from a break in one series. `cog_explain()`
|
||||
prints them under their own "Corpus-wide caveats" heading. cog-api passes
|
||||
provenance through verbatim, so the field appears there without an API
|
||||
change.
|
||||
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
|
||||
crosses the boundary where an absent cell stops meaning "Census published
|
||||
`$0`" and starts meaning "not reported", and until now nothing said so.
|
||||
|
||||
## Bundled fixture regenerated against the sparsified corpus
|
||||
|
||||
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
|
||||
2026-07-29 (`pipeline_commit 83f9715`, schema v6). The wide era no longer
|
||||
stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of
|
||||
which none are `$0`. **Absence now means two different things** — in a
|
||||
`dense_source` year (≤ FY2011) an absent cell means Census published `$0`;
|
||||
in a `sparse_source` year (≥ FY2012) it means not reported. The corpus
|
||||
carries that rule in two new tables the fixture now ships,
|
||||
`representation.parquet` and `code_set.parquet`, alongside
|
||||
`census_collection_coverage.parquet` and `lineage_events.parquet`
|
||||
(all ten publish-tree metadata tables, up from six). Catalogued upstream
|
||||
as series break `SB194`.
|
||||
* `cog_categories()` gains an `assistance` spending subtype: the J-prefix
|
||||
aid/benefit codes (`J19`, `J67`, `J68`, `J85`) are categorised now that
|
||||
the upstream crosswalk covers every flow code carrying dollars.
|
||||
* Two consequences worth knowing about, both visible in provenance rather
|
||||
than in returned dollars. The harmonization block's `na_rows_excluded`
|
||||
counts only rows that exist, so wide-era codes that were zero-padded no
|
||||
longer appear there. Coverage-gap `suggestions` are presence-based for the
|
||||
same reason, so a recipe whose component codes were all `$0` for a given
|
||||
government-year is no longer suggested for it.
|
||||
* `tests/testthat/test-fixture-vintage.R` pins these structural facts, so a
|
||||
fixture left behind by a future publish fails loudly instead of letting the
|
||||
suite pass against a corpus that no longer exists.
|
||||
|
||||
## Breaking: corpus schema_version 4 (Phase P canonical ids)
|
||||
|
||||
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
|
||||
`MaxCorpusSchema` in `DESCRIPTION` are both `4`); older corpora built
|
||||
against schema 3 are rejected by `cog_open()` with a clear version-mismatch
|
||||
error. `canonical_govid` is now uniformly 12 characters across every
|
||||
vintage the corpus covers (previously a mix of 9-char legacy ids and
|
||||
12-char FIPS ids depending on source year) — **every hardcoded
|
||||
`canonical_govid` literal from a pre-Phase-P corpus is now invalid** and
|
||||
must be re-resolved via `cog_gov_search()` or the new `canonical_alias`
|
||||
lookup table. `canonical_fips_xwalk` gains four columns
|
||||
(`legacy_govs_id`, `census_geoid`, `id_source`; `confidence` is renamed to
|
||||
`pop_confidence`) and a companion `canonical_alias` table ships in the
|
||||
corpus for mapping legacy/alternate ids onto the current canonical
|
||||
namespace. The bundled fixture corpus (`inst/extdata/fixture_corpus/`) has
|
||||
been regenerated against the Phase P publish tree, now ships the full
|
||||
`canonical_fips_xwalk` and `canonical_alias` master tables alongside the
|
||||
2019-2020 long partitions, and is reproducible via
|
||||
`data-raw/regenerate_fixture_corpus.R`.
|
||||
|
||||
## Clearer errors when `USCOGDATA_URL` is unconfigured or returns non-JSON
|
||||
|
||||
* `cog_open()` now aborts with the `uscogdata_url_not_configured` error
|
||||
class when the resolved corpus URL still contains the placeholder
|
||||
`REPLACE_WITH_SHARE_TOKEN` sentinel (or is empty). The message lists both
|
||||
remediation paths (`Sys.setenv(USCOGDATA_URL = ...)` and
|
||||
`options(uscogdata.url = ...)`) and points at the bundled fixture for
|
||||
offline testing. Previously the package proceeded to fetch the placeholder
|
||||
URL, cached the resulting HTML welcome page, and failed downstream with a
|
||||
cryptic `jsonlite` lexical-error.
|
||||
* `.fetch_or_cache_manifest()` now parses the HTTP response body before
|
||||
persisting it. Non-JSON responses (login pages, 404 HTML) raise
|
||||
`uscogdata_invalid_manifest` with the URL, Content-Type, and underlying
|
||||
parse error — and never write to the on-disk cache.
|
||||
* Manifest cache writes are now atomic (write to `manifest.json.tmp.<pid>`
|
||||
in `cache_dir`, then `file.rename` over the target), so an interrupted
|
||||
fetch cannot replace a previously-good cache.
|
||||
* Existing caches with non-JSON content (poisoned by the prior code path)
|
||||
are silently refetched instead of returning a parse error to the caller.
|
||||
* Local `USCOGDATA_URL` paths whose `manifest.json` is not valid JSON now
|
||||
surface the same `uscogdata_invalid_manifest` class with file context.
|
||||
|
||||
## Per-capita denominators now use per-year Census F-33 population
|
||||
|
||||
* `cog_spending()` and `cog_revenue()` previously divided all years' amounts
|
||||
by a single ACS 2018-2022 estimate (`canonical_fips_xwalk.population_acs`),
|
||||
producing biased per-capita values for time-series analysis. They now
|
||||
divide by the F-33 `population` recorded on each gov-year via the new
|
||||
`gov_population_yearly` view. Result tibbles gain a `pop_source` column
|
||||
with values `"census_f33"` or `"unavailable"`. `notes` is updated to
|
||||
concatenate multiple notes with `"; "`.
|
||||
|
||||
## Peer cohorts can be set to a chosen year
|
||||
|
||||
* `cog_find_peers()` adds a `year` argument (default: most recent year for
|
||||
which the target has an observed population in `gov_population_yearly`).
|
||||
The returned column previously named `population_acs` is now `population`
|
||||
and reflects the cohort year's vintage. The cohort year is attached to the
|
||||
returned tibble as `attr(x, "cohort_year")`.
|
||||
* `cog_peer_compare()` now stamps a `cohort_year` column on its result (read
|
||||
from the peers tibble's attribute) and records `cohort_year` plus
|
||||
`cohort_govids` in provenance. When the caller supplies a bare character
|
||||
vector instead of a `cog_find_peers()` result, `cohort_year` is `NA`.
|
||||
|
||||
## Rollups exclude govs missing population
|
||||
|
||||
* `cog_geographic_rollup(per_capita = TRUE)` drops rows whose government has
|
||||
`pop_source == "unavailable"` and records the dropped govids in
|
||||
`provenance$rollup$excluded_govids`. This excludes special districts
|
||||
(type 4) and school districts (type 5) from per-capita rollups by design.
|
||||
|
||||
## New: vignette and provenance metadata
|
||||
|
||||
* New vignette `population-denominators` covers the four population sources,
|
||||
the type-4/5 coverage gap, the popyear quirk, and how to build moving-window
|
||||
peer cohorts manually.
|
||||
* Provenance gains `transformations$per_capita$popyear_range` and
|
||||
`pop_source_counts`. `cog_explain()` renders both.
|
||||
|
||||
## New features
|
||||
|
||||
* `cog_gov_search()` gains a **basket mode**: passing vector `name`
|
||||
/ `state` / `type` arguments resolves multiple place names in one
|
||||
call and returns a tibble of canonical rows in input order, ready
|
||||
to pipe into `cog_spending()` / `cog_revenue()`. Per-row resolution
|
||||
follows an exact-then-substring matching algorithm with deterministic
|
||||
disambiguation; ambiguous and missing entries are surfaced via a
|
||||
sidecar audit tibble plus a single console summary message.
|
||||
* New exports `cog_basket_resolution()` and `cog_basket_unresolved()`
|
||||
expose the basket sidecar for iterative query refinement.
|
||||
|
||||
## Breaking changes
|
||||
|
||||
* The first formal of `cog_gov_search()` was renamed from `pattern`
|
||||
to `name`. All existing call sites in `cog_explorer/` and the
|
||||
package itself use positional first-arg, so this rename is
|
||||
non-breaking in practice. Callers that pass `pattern = ...` by name
|
||||
must update to `name = ...`.
|
||||
|
||||
+24
@@ -11,6 +11,30 @@
|
||||
#' returns `result` invisibly for chaining. `"list"` returns the raw
|
||||
#' provenance list (identical to `attr(result, "provenance")`).
|
||||
#' @return Either `result` (invisibly) or the provenance list.
|
||||
#' @section Two kinds of series break:
|
||||
#' Catalogued breaks reach you without being asked for, in two disjoint
|
||||
#' fields, because a caveat about one series and a caveat about the whole
|
||||
#' corpus are different claims:
|
||||
#'
|
||||
#' * **`series_break_refs`** — breaks matched against the item codes actually
|
||||
#' present in this result. A break in one code you queried.
|
||||
#' * **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
|
||||
#' which are statements about the corpus rather than about any one code:
|
||||
#' dollar precision across the 1976/1977 boundary (`SB085`), imputation
|
||||
#' exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
|
||||
#' representation change (`SB194`), and the FY2017 government-identifier
|
||||
#' change (`SB086`). These are selected on the break-year window alone.
|
||||
#'
|
||||
#' `SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
|
||||
#' crosses the boundary where an absent cell stops meaning "Census published
|
||||
#' $0" and starts meaning "not reported".
|
||||
#' @section Other provenance blocks:
|
||||
#' `transformations$units_conversion` records the `$1,000s`-to-dollars
|
||||
#' multiply that every amount column has already had applied.
|
||||
#' `transformations$per_capita` records the population denominator and its
|
||||
#' year range. `coverage` and `coverage_mode` appear on multi-government
|
||||
#' results (see [cog_geographic_rollup()]). `completion` appears when
|
||||
#' `complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
|
||||
#' @export
|
||||
cog_explain <- function(result, format = c("print", "list")) {
|
||||
format <- match.arg(format)
|
||||
|
||||
@@ -21,3 +21,33 @@ Prints the structured provenance attached to a tibble returned by any
|
||||
`cog_*` verb, or returns it as a list for downstream use (MCP tools,
|
||||
dashboards, JSON export).
|
||||
}
|
||||
\section{Two kinds of series break}{
|
||||
|
||||
Catalogued breaks reach you without being asked for, in two disjoint
|
||||
fields, because a caveat about one series and a caveat about the whole
|
||||
corpus are different claims:
|
||||
|
||||
* **`series_break_refs`** — breaks matched against the item codes actually
|
||||
present in this result. A break in one code you queried.
|
||||
* **`corpus_break_refs`** — breaks catalogued with `fin_code = "ALL"`,
|
||||
which are statements about the corpus rather than about any one code:
|
||||
dollar precision across the 1976/1977 boundary (`SB085`), imputation
|
||||
exclusion from FY2002 (`SB087`), the FY2012 dense-to-sparse
|
||||
representation change (`SB194`), and the FY2017 government-identifier
|
||||
change (`SB086`). These are selected on the break-year window alone.
|
||||
|
||||
`SB194` is the one most likely to matter: a query spanning FY2011 to FY2012
|
||||
crosses the boundary where an absent cell stops meaning "Census published
|
||||
$0" and starts meaning "not reported".
|
||||
}
|
||||
|
||||
\section{Other provenance blocks}{
|
||||
|
||||
`transformations$units_conversion` records the `$1,000s`-to-dollars
|
||||
multiply that every amount column has already had applied.
|
||||
`transformations$per_capita` records the population denominator and its
|
||||
year range. `coverage` and `coverage_mode` appear on multi-government
|
||||
results (see [cog_geographic_rollup()]). `completion` appears when
|
||||
`complete = TRUE`. `balance_caveats` appears on [cog_balances()] results.
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user