Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
515ab3b019
|
||
|
|
d238bc0a22
|
||
|
|
e53aeb9643
|
||
|
|
9244e08085
|
||
|
|
1b2294e3a0
|
||
|
|
91b64b9b8b
|
||
|
|
267bc24fee
|
@@ -16,4 +16,3 @@
|
|||||||
^Meta$
|
^Meta$
|
||||||
^\.gitea$
|
^\.gitea$
|
||||||
^CLAUDE\.md$
|
^CLAUDE\.md$
|
||||||
^\.superpowers$
|
|
||||||
|
|||||||
@@ -9,6 +9,3 @@ docs/
|
|||||||
/Meta/
|
/Meta/
|
||||||
.DS_Store
|
.DS_Store
|
||||||
/.quarto/
|
/.quarto/
|
||||||
|
|
||||||
# SDD working artifacts (ledger, briefs, review packages) — plans/ stays tracked
|
|
||||||
.superpowers/sdd/
|
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -28,23 +28,8 @@ USCOGDATA_URL (local path or https://)
|
|||||||
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
|
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
|
||||||
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
|
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
|
||||||
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
|
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
|
||||||
- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
|
- `inst/sql/` — 7 SQL view definitions: `long`, `spending_long`, `revenue_long`, `canonical_fips_xwalk`, `summary_categories`, `spending_annotated`, `revenue_annotated`
|
||||||
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
|
|
||||||
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
|
|
||||||
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
|
|
||||||
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
|
|
||||||
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
|
|
||||||
`revenue_annotated`/`ig_annotated`), and metadata views
|
|
||||||
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
|
|
||||||
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
|
|
||||||
`representation`, `code_set`)
|
|
||||||
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
|
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
|
||||||
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
|
|
||||||
**stock** (a balance at a point in time) rather than a **flow** (activity
|
|
||||||
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
|
|
||||||
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
|
|
||||||
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
|
|
||||||
disclosure + measured per-subtype coverage windows).
|
|
||||||
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
|
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
|
||||||
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
|
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
|
||||||
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
|
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
|
||||||
@@ -60,31 +45,28 @@ USCOGDATA_URL (local path or https://)
|
|||||||
Any value without `://` is treated as a local path by `.is_local_path()` and reads
|
Any value without `://` is treated as a local path by `.is_local_path()` and reads
|
||||||
`manifest.json` directly from disk (no HTTP, no TTL cache).
|
`manifest.json` directly from disk (no HTTP, no TTL cache).
|
||||||
|
|
||||||
## Current State (2026-08-03)
|
## Current State (2026-04-27)
|
||||||
|
|
||||||
**Version:** 0.1.0 (pre-release)
|
**Version:** 0.1.0 (pre-release)
|
||||||
**Branch:** `feat/cog-balances-25`, commit `fde62eb`
|
**Branch:** `main`, commit `d65e9fe`
|
||||||
**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
|
**Tests:** 181 PASS / 0 FAIL / 0 SKIP
|
||||||
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
|
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
|
||||||
|
|
||||||
### Completed (Tasks 2.1–2.7)
|
### Completed (Tasks 2.1–2.7)
|
||||||
|
|
||||||
All **14** exports implemented and tested (measured from `NAMESPACE`):
|
All 8 exported verbs implemented and tested:
|
||||||
`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
|
`cog_spending`, `cog_revenue`, `cog_explain`, `cog_geographic_rollup`,
|
||||||
`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
|
`cog_find_peers`, `cog_peer_compare`, `cog_gov_search`, `cog_mirror`,
|
||||||
`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
|
plus `cog_categories`.
|
||||||
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
|
|
||||||
|
|
||||||
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
|
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (3.6 MB, years
|
||||||
2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
|
2019+2020, all 50 states). Tests run fully offline — no credentials needed.
|
||||||
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
|
|
||||||
|
|
||||||
### Remaining to v0.1 release
|
### Remaining to v0.1 release
|
||||||
|
|
||||||
1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
|
1. **Task 2.8 — Docs:** roxygen `@param`/`@return`/`@examples` on all exports;
|
||||||
`README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
|
full `README.md`; `_pkgdown.yml`; `devtools::document()` + `pkgdown::build_site()`.
|
||||||
`total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
|
Vignettes can be stubbed for v0.1.
|
||||||
`pkgdown::build_site()` has never been run (no `docs/`).
|
|
||||||
|
|
||||||
2. **Phase 3 — cog_explorer bridge:** create
|
2. **Phase 3 — cog_explorer bridge:** create
|
||||||
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
|
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
|
||||||
@@ -117,10 +99,6 @@ devtools::test()
|
|||||||
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
|
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
|
||||||
- Return value is always a `tbl_df` with a `provenance` attribute
|
- Return value is always a `tbl_df` with a `provenance` attribute
|
||||||
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
|
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
|
||||||
- SQL has two layers. **View definitions** live in `inst/sql/` and are
|
- SQL lives in `inst/sql/` — never inline SQL strings in R files
|
||||||
registered by `.register_views()`, which globs the directory in sorted order
|
|
||||||
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
|
|
||||||
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
|
|
||||||
a numbered `.sql` file; build a query in R.
|
|
||||||
- No arrow dependency — DuckDB reads parquet natively
|
- No arrow dependency — DuckDB reads parquet natively
|
||||||
- `withr` is a Suggests-only dep; only used in tests
|
- `withr` is a Suggests-only dep; only used in tests
|
||||||
|
|||||||
@@ -1,6 +1,5 @@
|
|||||||
# Generated by roxygen2: do not edit by hand
|
# Generated by roxygen2: do not edit by hand
|
||||||
|
|
||||||
export(cog_balances)
|
|
||||||
export(cog_basket_resolution)
|
export(cog_basket_resolution)
|
||||||
export(cog_basket_unresolved)
|
export(cog_basket_unresolved)
|
||||||
export(cog_categories)
|
export(cog_categories)
|
||||||
|
|||||||
@@ -1,130 +1,5 @@
|
|||||||
# uscogdata 0.1.0 (development)
|
# uscogdata 0.1.0 (development)
|
||||||
|
|
||||||
## New: `cog_balances()` for cash-and-security holdings
|
|
||||||
|
|
||||||
* New `cog_balances()` exposes the 14 cash-and-security holding codes
|
|
||||||
(`category_type = "balance"`): fund balances, retirement system holdings and
|
|
||||||
insurance trust balances (#25). Holdings are a stock, not a flow, so the verb
|
|
||||||
has no `expenditure_concept` / `revenue_concept` / `complete` arguments, and
|
|
||||||
no `subtype` argument either -- for holdings, `category` is a strict
|
|
||||||
coarsening of `balance_subtype`, so `category = "Fund Balances"` is exactly
|
|
||||||
the `general` family (`W01`/`W31`/`W61`).
|
|
||||||
* `cog_balances()` results carry `provenance$balance_caveats`, recording that
|
|
||||||
Census holdings are gross rather than GAAP fund balance, and the measured
|
|
||||||
coverage window of each subtype family.
|
|
||||||
|
|
||||||
## Multi-government aggregates now disclose their reporting coverage
|
|
||||||
|
|
||||||
* The Census of Governments is a **complete census only in years ending in 2
|
|
||||||
and 7**; every other year is a sample, and the sample varies enormously. On
|
|
||||||
the bundled fixture, Wisconsin's 608-city universe rolls up **597**
|
|
||||||
governments in FY2012 and **112** in FY2019 — an 18%-to-98% swing the
|
|
||||||
return value said nothing about, so a statewide total resting on a fifth of
|
|
||||||
the universe looked exactly like one resting on all of it.
|
|
||||||
* `cog_geographic_rollup()`, `cog_peer_compare()` and `cog_find_peers()` gain
|
|
||||||
`coverage`:
|
|
||||||
|
|
||||||
| value | effect |
|
|
||||||
|---|---|
|
|
||||||
| `"all"` (default) | every unit that reported that year — unchanged behaviour |
|
|
||||||
| `"census"` | census years only; aborts if the range holds none rather than returning nothing |
|
|
||||||
| `"consistent"` | only units reporting in *every* requested year — a balanced panel |
|
|
||||||
|
|
||||||
* **Regardless of mode**, every result now carries `provenance$coverage` with
|
|
||||||
per-year `n_units_reporting`, `n_units_expected` and `is_census_year`, plus
|
|
||||||
`provenance$coverage_mode`. `cog_explain()` prints a "Reporting coverage"
|
|
||||||
section. So the default mode can no longer mislead silently.
|
|
||||||
* `is_census_year` is a statement about the **survey calendar**, never a claim
|
|
||||||
of completeness: FY1967 is a census year in which only 97 of Wisconsin's 608
|
|
||||||
cities report. `n_units_reporting` is the number that tells the truth.
|
|
||||||
* On `cog_peer_compare()` the target is exempt from `"consistent"` balancing —
|
|
||||||
it is the subject of the comparison, not a member of the cohort — and the
|
|
||||||
`summary_*` quantiles are computed after the filter, so they describe the
|
|
||||||
cohort actually returned. `n_units_reporting` counts peers only, against the
|
|
||||||
cohort size.
|
|
||||||
* On `cog_find_peers()`, `coverage` governs the cohort **vintage** when `year`
|
|
||||||
is `NULL`: `"census"` snaps to the most recent census year with an observed
|
|
||||||
population, so a cohort is not built from a sample year in which most of the
|
|
||||||
candidate universe is absent.
|
|
||||||
|
|
||||||
## `complete = TRUE`: absent cells, labelled with why they are absent
|
|
||||||
|
|
||||||
* `cog_spending()` and `cog_revenue()` gain `complete`, defaulting to `FALSE`
|
|
||||||
(today's behaviour). With `complete = TRUE` the requested grid is filled
|
|
||||||
from the corpus's `code_set` table and every row carries a new
|
|
||||||
`value_source` column:
|
|
||||||
|
|
||||||
| `value_source` | meaning | `amt_nominal` |
|
|
||||||
|---|---|---|
|
|
||||||
| `reported` | the corpus carries this cell | as published |
|
|
||||||
| `census_zero` | dense-source year (≤ FY2011), cell absent — Census published `$0` | `0` |
|
|
||||||
| `not_reported` | sparse-source year (≥ FY2012), cell absent — unknown | `NA` |
|
|
||||||
|
|
||||||
The `NA` is deliberate and is the whole point: filling a modern absence
|
|
||||||
with `0` would invent data, which is precisely the error the corpus's
|
|
||||||
representation contract exists to prevent.
|
|
||||||
* This restores information the reader lost when the corpus was sparsified
|
|
||||||
(`SB194`, cog_pipeline#64) — a wide-era query whose cells were all `$0`
|
|
||||||
had begun returning nothing at all — and improves on what came before it,
|
|
||||||
since the pre-sparsification corpus could not distinguish a published zero
|
|
||||||
from an unreported cell either.
|
|
||||||
* The grid is scoped to each government's **own type**, so a county is never
|
|
||||||
filled with cells only a state can report.
|
|
||||||
* Needs a corpus published from 2026-07-29 onward (when `representation` and
|
|
||||||
`code_set` began shipping); aborts with class
|
|
||||||
`uscogdata_representation_unavailable` otherwise. Gated on the manifest
|
|
||||||
listing those tables rather than on `schema_version`, which was never
|
|
||||||
bumped for the change. Not available with `recipe` or
|
|
||||||
`expenditure_concept = "total"` — neither draws its cells from `code_set`.
|
|
||||||
* `provenance$completion` reports `applied`, `rows_filled`, and the per-year
|
|
||||||
`absence_means` rule; `cog_explain()` prints a "Completion" section.
|
|
||||||
|
|
||||||
## Corpus-wide series breaks now reach users (`corpus_break_refs`)
|
|
||||||
|
|
||||||
* Four catalogued series breaks carry `fin_code = "ALL"` — caveats about the
|
|
||||||
corpus as a whole rather than about one item code. `series_break_refs` is
|
|
||||||
built by matching `fin_code` against the item codes in the result, and no
|
|
||||||
row's `item_code` is ever the literal `"ALL"`, so **none of them could ever
|
|
||||||
be surfaced**: `SB085` (dollar precision across the 1976/1977 boundary),
|
|
||||||
`SB087` (imputation exclusion from FY2002), `SB194` (the dense → sparse
|
|
||||||
representation change at FY2012) and `SB086` (the government id scheme
|
|
||||||
change at FY2017).
|
|
||||||
* Provenance gains `corpus_break_refs`, selected on the break-year window
|
|
||||||
alone and disjoint from `series_break_refs` by construction, so a consumer
|
|
||||||
can tell a whole-result caveat from a break in one series. `cog_explain()`
|
|
||||||
prints them under their own "Corpus-wide caveats" heading. cog-api passes
|
|
||||||
provenance through verbatim, so the field appears there without an API
|
|
||||||
change.
|
|
||||||
* `SB194` is the one that made this urgent: a query spanning FY2011 → FY2012
|
|
||||||
crosses the boundary where an absent cell stops meaning "Census published
|
|
||||||
`$0`" and starts meaning "not reported", and until now nothing said so.
|
|
||||||
|
|
||||||
## Bundled fixture regenerated against the sparsified corpus
|
|
||||||
|
|
||||||
* `inst/extdata/fixture_corpus/` now tracks the corpus published on
|
|
||||||
2026-07-29 (`pipeline_commit 83f9715`, schema v6). The wide era no longer
|
|
||||||
stores explicit zeros: FY2011 fell from 2,864,212 rows to 496,004, of
|
|
||||||
which none are `$0`. **Absence now means two different things** — in a
|
|
||||||
`dense_source` year (≤ FY2011) an absent cell means Census published `$0`;
|
|
||||||
in a `sparse_source` year (≥ FY2012) it means not reported. The corpus
|
|
||||||
carries that rule in two new tables the fixture now ships,
|
|
||||||
`representation.parquet` and `code_set.parquet`, alongside
|
|
||||||
`census_collection_coverage.parquet` and `lineage_events.parquet`
|
|
||||||
(all ten publish-tree metadata tables, up from six). Catalogued upstream
|
|
||||||
as series break `SB194`.
|
|
||||||
* `cog_categories()` gains an `assistance` spending subtype: the J-prefix
|
|
||||||
aid/benefit codes (`J19`, `J67`, `J68`, `J85`) are categorised now that
|
|
||||||
the upstream crosswalk covers every flow code carrying dollars.
|
|
||||||
* Two consequences worth knowing about, both visible in provenance rather
|
|
||||||
than in returned dollars. The harmonization block's `na_rows_excluded`
|
|
||||||
counts only rows that exist, so wide-era codes that were zero-padded no
|
|
||||||
longer appear there. Coverage-gap `suggestions` are presence-based for the
|
|
||||||
same reason, so a recipe whose component codes were all `$0` for a given
|
|
||||||
government-year is no longer suggested for it.
|
|
||||||
* `tests/testthat/test-fixture-vintage.R` pins these structural facts, so a
|
|
||||||
fixture left behind by a future publish fails loudly instead of letting the
|
|
||||||
suite pass against a corpus that no longer exists.
|
|
||||||
|
|
||||||
## Breaking: corpus schema_version 4 (Phase P canonical ids)
|
## Breaking: corpus schema_version 4 (Phase P canonical ids)
|
||||||
|
|
||||||
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
|
* The package now requires corpus `schema_version = 4` (`MinCorpusSchema` /
|
||||||
|
|||||||
@@ -1,122 +0,0 @@
|
|||||||
# R/balance_caveats.R
|
|
||||||
#
|
|
||||||
# The four caveats from cog_pipeline/docs/data_dictionary.md § Cash and
|
|
||||||
# security holdings. Each one silently invalidates an obvious analysis, so
|
|
||||||
# they travel in provenance (machine-readable, for cog-api#26) rather than
|
|
||||||
# living only in prose.
|
|
||||||
#
|
|
||||||
# Two of the four are already carried by the code-driven series-break
|
|
||||||
# builders and are deliberately NOT duplicated here:
|
|
||||||
# * SB195/SB196 -- X40/X41 book -> market at FY2002 -- fire via
|
|
||||||
# series_break_refs on the recipe path, the only path that observes those
|
|
||||||
# codes.
|
|
||||||
# What remains is the GAAP distinction (a constant) and the coverage windows
|
|
||||||
# (measured, never hardcoded, so they stay correct as the corpus grows).
|
|
||||||
|
|
||||||
#' Per-subtype observed year extents, plus which requested families are
|
|
||||||
#' truncated relative to the requested span.
|
|
||||||
#' @noRd
|
|
||||||
.balance_caveats <- function(con, codes_observed, years) {
|
|
||||||
cw <- .balance_coverage_windows(con)
|
|
||||||
|
|
||||||
observed_subtypes <- if (length(codes_observed) == 0L) {
|
|
||||||
character(0)
|
|
||||||
} else {
|
|
||||||
DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT DISTINCT balance_subtype FROM summary_categories
|
|
||||||
WHERE item_code IN (%s) AND balance_subtype IS NOT NULL",
|
|
||||||
.sql_lit_chr(codes_observed)
|
|
||||||
))$balance_subtype
|
|
||||||
}
|
|
||||||
|
|
||||||
# A family is "truncated" when the caller asked for years outside the span
|
|
||||||
# that family actually covers -- the FY2016 employee-retirement termination
|
|
||||||
# and the FY2021 end of the W family are both this shape.
|
|
||||||
truncated <- character(0)
|
|
||||||
if (length(years) > 0L) {
|
|
||||||
for (s in observed_subtypes) {
|
|
||||||
w <- cw[[s]]
|
|
||||||
if (is.null(w)) next
|
|
||||||
if (max(years) > w[2] || min(years) < w[1]) truncated <- c(truncated, s)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
list(
|
|
||||||
not_gaap = TRUE,
|
|
||||||
not_gaap_note = paste0(
|
|
||||||
"Census holdings are gross -- no liabilities are netted -- and are NOT ",
|
|
||||||
"GAAP fund balance. A reserve ratio built from them overstates what is ",
|
|
||||||
"actually available."
|
|
||||||
),
|
|
||||||
coverage_window = cw,
|
|
||||||
truncated = sort(unique(truncated))
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Per-subtype [min year, max year] extents for EVERY balance subtype in the
|
|
||||||
#' mounted corpus, memoised for the session.
|
|
||||||
#'
|
|
||||||
#' The query carries no govid and no year predicate -- its answer is a property
|
|
||||||
#' of the mounted corpus alone and cannot change between calls -- but it scans
|
|
||||||
#' the whole of `balance_long`, which measured 35% of `cog_balances()` runtime
|
|
||||||
#' on the bundled fixture and would be a per-request throughput ceiling once
|
|
||||||
#' cog-api#26 serves this verb over HTTP. Memoised in `.uscogdata_env` and
|
|
||||||
#' invalidated by `cog_close()`, the same pattern as `.uscogdata_env$manifest`.
|
|
||||||
#'
|
|
||||||
#' Scope is deliberately corpus-wide rather than query-scoped: a caller asking
|
|
||||||
#' "is there a family I missed?" needs every window. The observed-scoped field
|
|
||||||
#' is `truncated`. Documented as such in inst/schemas/provenance-v1.json.
|
|
||||||
#' @noRd
|
|
||||||
.balance_coverage_windows <- function(con) {
|
|
||||||
cached <- .uscogdata_env$balance_coverage_windows
|
|
||||||
if (!is.null(cached)) return(cached)
|
|
||||||
|
|
||||||
windows <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT c.balance_subtype AS subtype,
|
|
||||||
MIN(l.year) AS year_min,
|
|
||||||
MAX(l.year) AS year_max
|
|
||||||
FROM balance_long l
|
|
||||||
JOIN summary_categories c USING (item_code)
|
|
||||||
WHERE c.balance_subtype IS NOT NULL
|
|
||||||
GROUP BY 1
|
|
||||||
ORDER BY 1"
|
|
||||||
)
|
|
||||||
|
|
||||||
cw <- stats::setNames(
|
|
||||||
lapply(seq_len(nrow(windows)),
|
|
||||||
function(i) as.integer(c(windows$year_min[i], windows$year_max[i]))),
|
|
||||||
windows$subtype
|
|
||||||
)
|
|
||||||
.uscogdata_env$balance_coverage_windows <- cw
|
|
||||||
cw
|
|
||||||
}
|
|
||||||
|
|
||||||
#' TRUE the first time `key` is seen this session, FALSE thereafter.
|
|
||||||
#' Reset by cog_close().
|
|
||||||
#' @noRd
|
|
||||||
.balance_caveat_once <- function(key) {
|
|
||||||
seen <- .uscogdata_env$balance_caveats_shown
|
|
||||||
if (is.null(seen)) seen <- character(0)
|
|
||||||
if (key %in% seen) return(FALSE)
|
|
||||||
.uscogdata_env$balance_caveats_shown <- c(seen, key)
|
|
||||||
TRUE
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Emit at most one message per caveat class per session.
|
|
||||||
#' @noRd
|
|
||||||
.emit_balance_caveats <- function(caveats) {
|
|
||||||
if (.balance_caveat_once("not_gaap")) {
|
|
||||||
cli::cli_inform(c(
|
|
||||||
"!" = "Census holdings are gross and are {.strong not} GAAP fund balance.",
|
|
||||||
"i" = "No liabilities are netted; a reserve ratio built from them overstates available funds."
|
|
||||||
))
|
|
||||||
}
|
|
||||||
if (length(caveats$truncated) > 0L &&
|
|
||||||
.balance_caveat_once("coverage_window")) {
|
|
||||||
cli::cli_inform(c(
|
|
||||||
"!" = "Requested years extend beyond what {.val {caveats$truncated}} actually covers.",
|
|
||||||
"i" = "See {.code provenance$balance_caveats$coverage_window}."
|
|
||||||
))
|
|
||||||
}
|
|
||||||
invisible(NULL)
|
|
||||||
}
|
|
||||||
-159
@@ -1,159 +0,0 @@
|
|||||||
# R/balances.R
|
|
||||||
#
|
|
||||||
# Cash and security holdings. A third verb rather than an argument on a money
|
|
||||||
# verb because holdings are a STOCK -- a balance at a point in time -- while
|
|
||||||
# cog_spending()/cog_revenue() return FLOWS over a fiscal year. The money
|
|
||||||
# verbs' whole argument vocabulary (expenditure_concept, revenue_concept,
|
|
||||||
# complete=) describes flows and is meaningless here, so this deliberately
|
|
||||||
# does NOT route through .verb_spendrev().
|
|
||||||
|
|
||||||
#' Cash and security holdings for one or more governments
|
|
||||||
#'
|
|
||||||
#' Returns Census cash-and-security holdings (`category_type = "balance"`):
|
|
||||||
#' fund balances, retirement system holdings and insurance trust balances.
|
|
||||||
#'
|
|
||||||
#' @section Holdings are not GAAP fund balance:
|
|
||||||
#' Census holdings are **gross** -- no liabilities are netted -- so a reserve
|
|
||||||
#' ratio built from them overstates what is actually available. They are not
|
|
||||||
#' comparable to a GAAP fund balance from an ACFR.
|
|
||||||
#'
|
|
||||||
#' @param govid Canonical govid(s): a character vector, or a data frame with a
|
|
||||||
#' `canonical_govid` column (e.g. from [cog_gov_search()]).
|
|
||||||
#' @param years Integer vector of fiscal years.
|
|
||||||
#' @param category Optional character vector of categories to keep. One of
|
|
||||||
#' `"Fund Balances"`, `"Insurance Trust Balances"`,
|
|
||||||
#' `"Retirement System Holdings"`. There is deliberately no `subtype`
|
|
||||||
#' argument: for holdings, `category` is a strict coarsening of
|
|
||||||
#' `balance_subtype` (unlike the money verbs, where the two axes cross), so
|
|
||||||
#' every combination would be either redundant or empty.
|
|
||||||
#' `category = "Fund Balances"` is exactly the `general` family
|
|
||||||
#' (`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
|
||||||
#' one `dplyr::filter()` away.
|
|
||||||
#' @param per_capita Divide holdings by population. Note this is a **stock per
|
|
||||||
#' resident** (reserves per person), which is *not* comparable to
|
|
||||||
#' [cog_spending()]'s per-capita figures -- those are a flow per person.
|
|
||||||
#' @param adjust_to_year Deflate to this year's dollars (CPI-U).
|
|
||||||
#' @param basis Accepted for uniformity with the money verbs, but currently a
|
|
||||||
#' **no-op**: `harmonization_map` carries no balance-code rows, so harmonized
|
|
||||||
#' and raw space are identical for holdings. Reported in
|
|
||||||
#' `provenance$basis_note`.
|
|
||||||
#' @param recipe Optional harmonization recipe id (see [cog_recipes()]).
|
|
||||||
#' `"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
|
|
||||||
#' wide era to the modern one.
|
|
||||||
#'
|
|
||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
|
||||||
#' `balance_subtype`, `category`, `amt_nominal`, `codes_included`,
|
|
||||||
#' `aggregate_fallback`, plus optional `amt_per_capita_nominal` and
|
|
||||||
#' `pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
|
|
||||||
#' `adjust_to_year` is set), and optional `amt_per_capita_real` (only when
|
|
||||||
#' **both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
|
|
||||||
#' nominal per-capita column to deflate otherwise). Amounts are full US
|
|
||||||
#' dollars.
|
|
||||||
#'
|
|
||||||
#' Carries a `provenance` attribute matching
|
|
||||||
#' `inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
|
|
||||||
#' `not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
|
|
||||||
#' every balance subtype in the mounted corpus, not only the observed ones)
|
|
||||||
#' and `truncated` (the observed subtypes whose coverage falls short of the
|
|
||||||
#' requested years). `expenditure_concept`/`revenue_concept` are `NA` --
|
|
||||||
#' holdings are a stock, not a flow, so neither concept vocabulary applies.
|
|
||||||
#' @export
|
|
||||||
cog_balances <- function(govid, years, category = NULL,
|
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
|
||||||
basis = c("harmonized", "raw"), recipe = NULL) {
|
|
||||||
call <- match.call()
|
|
||||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
|
||||||
# Coerce FIRST, validate second: .validate_verb_inputs() asserts
|
|
||||||
# is.character(govid), and a data-frame govid (cog_gov_search() output) has
|
|
||||||
# not been unwrapped yet at this point.
|
|
||||||
govid <- .coerce_govid_input(govid)
|
|
||||||
# The money verbs' validator, reused rather than re-implemented (R/spending.R).
|
|
||||||
# It covers the exact superset cog_balances() needs -- including the
|
|
||||||
# recipe/category mutual-exclusivity guard -- so a second local copy would
|
|
||||||
# only be a place for the two to drift apart. This is the same kind of
|
|
||||||
# helper reuse as .build_verb_sql()/.attach_per_capita() below; it does NOT
|
|
||||||
# route the verb through .verb_spendrev(), which stays deliberately unused
|
|
||||||
# here because its flow vocabulary is meaningless for a stock.
|
|
||||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
|
||||||
recipe)
|
|
||||||
years <- as.integer(years)
|
|
||||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
|
||||||
|
|
||||||
con <- .ensure_session()
|
|
||||||
.require_balance_support(con)
|
|
||||||
scope <- .check_govids_in_scope(govid)
|
|
||||||
|
|
||||||
basis_note <- paste0(
|
|
||||||
"`basis` has no effect on holdings: harmonization_map carries no ",
|
|
||||||
"balance-code rows, so harmonized and raw space are identical here."
|
|
||||||
)
|
|
||||||
|
|
||||||
manifest <- .uscogdata_env$manifest
|
|
||||||
recipe_block <- NULL
|
|
||||||
category_for_prov <- category
|
|
||||||
|
|
||||||
if (!is.null(recipe)) {
|
|
||||||
.require_schema_v5(con, manifest, "recipe =")
|
|
||||||
.validate_recipe_id(con, recipe)
|
|
||||||
comps <- .recipe_components(con, recipe)
|
|
||||||
recipe_label <- comps$label[[1]]
|
|
||||||
result <- .run_recipe(con, recipe, govid, years)
|
|
||||||
sql <- attr(result, "sql_query")
|
|
||||||
result <- .shape_recipe_result(result, "balance_subtype", recipe_label)
|
|
||||||
recipe_block <- list(
|
|
||||||
recipe_id = recipe, label = recipe_label,
|
|
||||||
components = .df_to_row_list(comps)
|
|
||||||
)
|
|
||||||
category_for_prov <- recipe_label
|
|
||||||
} else {
|
|
||||||
sql <- .build_verb_sql("balance_annotated", "balance_subtype",
|
|
||||||
govid, years, category,
|
|
||||||
ig_view = NULL, subtype_scope = NULL)
|
|
||||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
|
||||||
}
|
|
||||||
|
|
||||||
# Order matters (matches .verb_spendrev()): per-capita first, so
|
|
||||||
# .attach_real_dollars() deflates the nominal per-capita column into
|
|
||||||
# amt_per_capita_real rather than needing amt_per_capita_nominal recomputed.
|
|
||||||
if (isTRUE(per_capita)) result <- .attach_per_capita(result, con, govid)
|
|
||||||
if (!is.null(adjust_to_year)) {
|
|
||||||
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
|
||||||
}
|
|
||||||
|
|
||||||
prov <- .build_provenance(
|
|
||||||
verb = "cog_balances", call = call, govid = govid, years = years,
|
|
||||||
category = category_for_prov, per_capita = per_capita,
|
|
||||||
adjust_to_year = adjust_to_year, result = result, sql = sql,
|
|
||||||
subtype_col = "balance_subtype",
|
|
||||||
basis = basis, basis_note = basis_note,
|
|
||||||
# Neither concept vocabulary applies to a stock.
|
|
||||||
expenditure_concept = NA_character_,
|
|
||||||
revenue_concept = NA_character_,
|
|
||||||
recipe = recipe_block
|
|
||||||
)
|
|
||||||
prov$scope$govids_found <- scope$found
|
|
||||||
prov$scope$govids_missing <- scope$missing
|
|
||||||
|
|
||||||
prov$balance_caveats <- .balance_caveats(
|
|
||||||
con, prov$codes_summed$observed, years
|
|
||||||
)
|
|
||||||
.emit_balance_caveats(prov$balance_caveats)
|
|
||||||
|
|
||||||
attr(result, "provenance") <- prov
|
|
||||||
result
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Abort unless the mounted corpus classifies balance codes.
|
|
||||||
#'
|
|
||||||
#' `balance_subtype` arrived with cog_pipeline #76/#77 without a
|
|
||||||
#' schema_version bump, so the check is on the column, not the version.
|
|
||||||
#' @noRd
|
|
||||||
.require_balance_support <- function(con) {
|
|
||||||
if (.corpus_has_balance_subtype(con)) return(invisible(TRUE))
|
|
||||||
cli::cli_abort(
|
|
||||||
c("This corpus does not classify cash and security holdings.",
|
|
||||||
i = "`summary_categories` has no {.field balance_subtype} column.",
|
|
||||||
i = "Republish from cog_pipeline at #76/#77 or later."),
|
|
||||||
class = "uscogdata_no_balance_support"
|
|
||||||
)
|
|
||||||
}
|
|
||||||
@@ -44,18 +44,11 @@
|
|||||||
|
|
||||||
#' Count + sum item-level rows that basis="harmonized" excludes because they
|
#' Count + sum item-level rows that basis="harmonized" excludes because they
|
||||||
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
|
#' carry no harmonized_code (discontinued / not-yet-ruled codes) within the
|
||||||
#' calling verb's crosswalk scope (`subtype_col` values in `subtype_scope` --
|
#' requested flow type (spending or revenue), govids, and years. Only
|
||||||
#' the same subtype-membership classification the verb SQL uses, never
|
#' meaningful when the resolved basis is "harmonized"; returns an
|
||||||
#' item-code prefixes), govids, and years. Only meaningful when the resolved
|
#' applied = FALSE stub otherwise (raw basis never excludes rows this way).
|
||||||
#' basis is "harmonized"; returns an applied = FALSE stub otherwise (raw
|
|
||||||
#' basis never excludes rows this way).
|
|
||||||
#'
|
|
||||||
#' The intergovernmental leg is deliberately outside this count even for
|
|
||||||
#' expenditure_concept = "total": ig_long_harmonized COALESCEs rather than
|
|
||||||
#' drops NULL-harmonized rows, so harmonization never excludes an IG row.
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.build_harmonization_block <- function(con, govid, years, resolved,
|
.build_harmonization_block <- function(con, govid, years, resolved, flow_prefixes) {
|
||||||
subtype_col, subtype_scope) {
|
|
||||||
if (!identical(resolved$basis, "harmonized")) {
|
if (!identical(resolved$basis, "harmonized")) {
|
||||||
return(list(
|
return(list(
|
||||||
applied = FALSE,
|
applied = FALSE,
|
||||||
@@ -70,11 +63,9 @@
|
|||||||
FROM long
|
FROM long
|
||||||
WHERE canonical_govid IN (%s) AND year IN (%s)
|
WHERE canonical_govid IN (%s) AND year IN (%s)
|
||||||
AND NOT is_aggregate AND harmonized_code IS NULL
|
AND NOT is_aggregate AND harmonized_code IS NULL
|
||||||
AND item_code IN (
|
AND LEFT(item_code, 1) IN (%s)",
|
||||||
SELECT item_code FROM summary_categories WHERE %s IN (%s)
|
|
||||||
)",
|
|
||||||
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
|
.sql_lit_chr(govid), paste(as.integer(years), collapse = ","),
|
||||||
subtype_col, .sql_lit_chr(subtype_scope)
|
.sql_lit_chr(flow_prefixes)
|
||||||
)
|
)
|
||||||
na <- DBI::dbGetQuery(con, sql)
|
na <- DBI::dbGetQuery(con, sql)
|
||||||
|
|
||||||
|
|||||||
-149
@@ -1,149 +0,0 @@
|
|||||||
# R/complete.R
|
|
||||||
#
|
|
||||||
# `complete = TRUE` on the money verbs. Fills the requested grid so that a
|
|
||||||
# cell the corpus does not carry still appears, labelled with WHY it is
|
|
||||||
# missing.
|
|
||||||
#
|
|
||||||
# The corpus stopped storing the wide era's explicit zeros
|
|
||||||
# (cog_pipeline#64, series break SB194), which made absence ambiguous:
|
|
||||||
#
|
|
||||||
# <= FY2011 dense_source absent => Census published $0 (census_zero)
|
|
||||||
# >= FY2012 sparse_source absent => not reported, unknown (not_reported)
|
|
||||||
#
|
|
||||||
# Before sparsification a wide-era query whose cells were all $0 came back as
|
|
||||||
# explicit $0 rows; afterwards it came back empty, with nothing to say which
|
|
||||||
# of the two meanings applied. This restores that -- and improves on it,
|
|
||||||
# because the pre-sparsification corpus could not distinguish the two either.
|
|
||||||
#
|
|
||||||
# `census_zero` fills carry `amt_nominal = 0`; `not_reported` fills carry NA.
|
|
||||||
# That difference is the entire point: writing 0 into a modern absence would
|
|
||||||
# invent data, which is the error the representation contract exists to stop.
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.abort_complete_unsupported <- function(reason, alternative) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.code complete = TRUE} is not supported for this query.",
|
|
||||||
x = reason,
|
|
||||||
i = alternative
|
|
||||||
), class = "uscogdata_complete_unsupported")
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.require_representation <- function(con, manifest) {
|
|
||||||
needed <- c("representation.parquet", "code_set.parquet")
|
|
||||||
missing <- needed[!vapply(needed, function(f) .corpus_has_table(manifest, f),
|
|
||||||
logical(1))]
|
|
||||||
if (length(missing) == 0L) return(invisible(TRUE))
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"This corpus does not publish the representation contract.",
|
|
||||||
x = "Missing: {.file {missing}}.",
|
|
||||||
i = "{.code complete = TRUE} needs those tables to know whether an absent cell means Census published $0 or means the government did not report.",
|
|
||||||
i = "They ship with corpora published from 2026-07-29 onward; re-point {.envvar USCOGDATA_URL} at a current corpus, or omit {.code complete}."
|
|
||||||
), class = "uscogdata_representation_unavailable")
|
|
||||||
}
|
|
||||||
|
|
||||||
#' The cells a government-year COULD carry: every code in force for that
|
|
||||||
#' government's own type, mapped through `summary_categories`, restricted to
|
|
||||||
#' the calling verb's crosswalk subtype scope (the same subtype-membership
|
|
||||||
#' classification the verb SQL itself uses -- e.g. the `primary` concept's
|
|
||||||
#' operations/capital/assistance) and (when given) its category filter.
|
|
||||||
#'
|
|
||||||
#' Scoped by `govs_type` deliberately. Filling against the union of all types
|
|
||||||
#' would invent cells that the government can never report -- a county row for
|
|
||||||
#' "state IG transfer to school districts" -- and those inventions would then
|
|
||||||
#' be indistinguishable from real census zeros.
|
|
||||||
#'
|
|
||||||
#' `NOT cs.is_aggregate` mirrors `spending_long` / `revenue_long`, which drop
|
|
||||||
#' aggregate rows. Without it the grid would offer cells the verb structurally
|
|
||||||
#' never returns, so every one of them would fill as a phantom $0.
|
|
||||||
#' @noRd
|
|
||||||
.completion_grid_sql <- function(subtype_col, govid, years, category,
|
|
||||||
subtype_scope) {
|
|
||||||
category_pred <- if (is.null(category)) {
|
|
||||||
""
|
|
||||||
} else {
|
|
||||||
sprintf("AND c.category IN (%s)", .sql_lit_chr(category))
|
|
||||||
}
|
|
||||||
sprintf(
|
|
||||||
"SELECT DISTINCT
|
|
||||||
cs.year,
|
|
||||||
x.canonical_govid,
|
|
||||||
x.gov_name,
|
|
||||||
c.%1$s AS subtype_value,
|
|
||||||
c.category,
|
|
||||||
r.absence_means
|
|
||||||
FROM code_set cs
|
|
||||||
JOIN canonical_fips_xwalk x ON x.govs_type = cs.type
|
|
||||||
JOIN summary_categories c ON c.item_code = cs.item_code
|
|
||||||
JOIN representation r ON r.year = cs.year
|
|
||||||
WHERE x.canonical_govid IN (%2$s)
|
|
||||||
AND cs.year IN (%3$s)
|
|
||||||
AND NOT cs.is_aggregate
|
|
||||||
AND c.category IS NOT NULL
|
|
||||||
AND c.%1$s IN (%4$s)
|
|
||||||
%5$s",
|
|
||||||
subtype_col, .sql_lit_chr(govid),
|
|
||||||
paste(as.integer(years), collapse = ","),
|
|
||||||
.sql_lit_chr(subtype_scope), category_pred
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Fill `result` out to the full grid, stamping `value_source` on every row.
|
|
||||||
#'
|
|
||||||
#' Returns the completed tibble with a `.completion` attribute carrying the
|
|
||||||
#' provenance block. Reported rows are passed through untouched -- filling
|
|
||||||
#' must never alter or drop what the corpus actually published.
|
|
||||||
#' @noRd
|
|
||||||
.complete_result <- function(result, con, subtype_col, govid, years, category,
|
|
||||||
subtype_scope) {
|
|
||||||
grid <- tibble::as_tibble(DBI::dbGetQuery(
|
|
||||||
con, .completion_grid_sql(subtype_col, govid, years, category, subtype_scope)
|
|
||||||
))
|
|
||||||
|
|
||||||
result$value_source <- rep("reported", nrow(result))
|
|
||||||
if (nrow(grid) == 0L) {
|
|
||||||
attr(result, ".completion") <- list(
|
|
||||||
applied = TRUE, rows_filled = 0L, absence_means = list()
|
|
||||||
)
|
|
||||||
return(result)
|
|
||||||
}
|
|
||||||
|
|
||||||
names(grid)[names(grid) == "subtype_value"] <- subtype_col
|
|
||||||
key <- function(d) {
|
|
||||||
paste(d$year, d$canonical_govid, d[[subtype_col]], d$category, sep = "\r")
|
|
||||||
}
|
|
||||||
missing <- grid[!key(grid) %in% key(result), , drop = FALSE]
|
|
||||||
|
|
||||||
if (nrow(missing) > 0L) {
|
|
||||||
filled <- tibble::tibble(
|
|
||||||
year = as.integer(missing$year),
|
|
||||||
canonical_govid = as.character(missing$canonical_govid),
|
|
||||||
gov_name = as.character(missing$gov_name),
|
|
||||||
category = as.character(missing$category),
|
|
||||||
# census_zero is a value Census published; not_reported is unknown and
|
|
||||||
# must stay NA. Collapsing the two to 0 is the defect, not the fill.
|
|
||||||
amt_nominal = ifelse(missing$absence_means == "census_zero",
|
|
||||||
0, NA_real_),
|
|
||||||
codes_included = NA_character_,
|
|
||||||
aggregate_fallback = NA,
|
|
||||||
value_source = as.character(missing$absence_means)
|
|
||||||
)
|
|
||||||
filled[[subtype_col]] <- as.character(missing[[subtype_col]])
|
|
||||||
if ("notes" %in% names(result)) filled$notes <- NA_character_
|
|
||||||
|
|
||||||
result <- dplyr::bind_rows(result, filled)
|
|
||||||
result <- result[order(result$year, result$canonical_govid,
|
|
||||||
result[[subtype_col]], result$category), ,
|
|
||||||
drop = FALSE]
|
|
||||||
}
|
|
||||||
|
|
||||||
rules <- unique(grid[, c("year", "absence_means")])
|
|
||||||
attr(result, ".completion") <- list(
|
|
||||||
applied = TRUE,
|
|
||||||
rows_filled = nrow(missing),
|
|
||||||
absence_means = stats::setNames(
|
|
||||||
as.list(as.character(rules$absence_means)), as.character(rules$year)
|
|
||||||
)
|
|
||||||
)
|
|
||||||
result
|
|
||||||
}
|
|
||||||
+1
-24
@@ -21,30 +21,7 @@
|
|||||||
.uscogdata_defaults[[key]]
|
.uscogdata_defaults[[key]]
|
||||||
}
|
}
|
||||||
|
|
||||||
#' Resolve the corpus URL, guaranteeing the trailing slash the package assumes.
|
.resolve_url <- function() .cfg("url")
|
||||||
#'
|
|
||||||
#' Every consumer builds locations by CONCATENATION -- `paste0(url,
|
|
||||||
#' "manifest.json")` in manifest.R, `paste0(url, e$path)` in mirror.R, and the
|
|
||||||
#' parquet glob in views.R -- and mirror.R:104 documents the invariant outright
|
|
||||||
#' ('url ends in "/"'). Nothing enforced it, so a URL entered without the slash
|
|
||||||
#' failed silently and misleadingly:
|
|
||||||
#'
|
|
||||||
#' HTTPS -> ".../downloadmanifest.json"; the host answers with an HTML 404
|
|
||||||
#' page, which lands in the JSON parser as the lexical error
|
|
||||||
#' reported in issue #3 -- pointing the user at "login page / wrong
|
|
||||||
#' share" when the real cause was one missing character.
|
|
||||||
#' local -> ".../corpusdata/long/**/*.parquet" and a DuckDB "No files found".
|
|
||||||
#'
|
|
||||||
#' Normalizing here fixes every consumer at once, rather than each call site
|
|
||||||
#' re-deriving the same invariant. An empty setting is passed through
|
|
||||||
#' untouched so manifest.R's "not configured" guard still fires instead of the
|
|
||||||
#' value degrading into a bare "/" filesystem root.
|
|
||||||
#' @noRd
|
|
||||||
.resolve_url <- function() {
|
|
||||||
url <- .cfg("url")
|
|
||||||
if (is.null(url) || !nzchar(url) || grepl("/$", url)) return(url)
|
|
||||||
paste0(url, "/")
|
|
||||||
}
|
|
||||||
|
|
||||||
.resolve_cache_dir <- function() {
|
.resolve_cache_dir <- function() {
|
||||||
v <- .cfg("cache_dir")
|
v <- .cfg("cache_dir")
|
||||||
|
|||||||
-107
@@ -1,107 +0,0 @@
|
|||||||
# R/coverage.R
|
|
||||||
#
|
|
||||||
# Reporting-coverage disclosure for the multi-government verbs (uscogdata#13,
|
|
||||||
# findings F-020 and F-023).
|
|
||||||
#
|
|
||||||
# The Census of Governments is a COMPLETE CENSUS only in years ending in 2 and
|
|
||||||
# 7. Every other year is a sample, and the sample varies enormously: on the
|
|
||||||
# bundled fixture, Wisconsin's 608-city universe reports 597 governments in
|
|
||||||
# FY2012 and 112 in FY2019. Summing "whatever reported" across those years is
|
|
||||||
# what the verbs have always done -- correctly -- but the return value said
|
|
||||||
# nothing about it, so a statewide total resting on 18% of the universe looked
|
|
||||||
# exactly like one resting on 98%.
|
|
||||||
#
|
|
||||||
# Owner's settled design: a `coverage` argument selecting WHICH units to
|
|
||||||
# include, plus always-on metadata saying how many there were either way. The
|
|
||||||
# principle behind it: using these verbs correctly must not require the caller
|
|
||||||
# to know the survey calendar.
|
|
||||||
|
|
||||||
# Years ending in 2 or 7 are full censuses of every government; all others are
|
|
||||||
# samples.
|
|
||||||
.CENSUS_YEAR_ENDINGS <- c(2L, 7L)
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.is_census_year <- function(years) {
|
|
||||||
as.integer(years) %% 10L %in% .CENSUS_YEAR_ENDINGS
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.validate_coverage <- function(coverage) {
|
|
||||||
tryCatch(
|
|
||||||
match.arg(coverage, c("all", "census", "consistent")),
|
|
||||||
error = function(e) {
|
|
||||||
cli::cli_abort(
|
|
||||||
"`coverage` must be one of {.val all}, {.val census} or {.val consistent}.",
|
|
||||||
class = "uscogdata_invalid_coverage", parent = e
|
|
||||||
)
|
|
||||||
}
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Restrict `years` to census years for `coverage = "census"`.
|
|
||||||
#'
|
|
||||||
#' Aborts rather than returning an empty result when the requested range holds
|
|
||||||
#' no census year: silently handing back zero rows for a query the caller
|
|
||||||
#' believes they made is the failure mode this whole issue is about.
|
|
||||||
#' @noRd
|
|
||||||
.apply_census_years <- function(years, coverage, verb) {
|
|
||||||
if (!identical(coverage, "census")) return(as.integer(years))
|
|
||||||
keep <- as.integer(years)[.is_census_year(years)]
|
|
||||||
if (length(keep) == 0L) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.code coverage = \"census\"} leaves no years to query.",
|
|
||||||
x = "None of the requested years end in 2 or 7: {.val {sort(unique(as.integer(years)))}}.",
|
|
||||||
i = "Census of Governments years ending in 2 or 7 are complete censuses; all others are samples.",
|
|
||||||
i = "Use {.code coverage = \"all\"} (the default) to keep every requested year, or request a census year."
|
|
||||||
), class = "uscogdata_no_census_years")
|
|
||||||
}
|
|
||||||
sort(keep)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Keep only units that report in EVERY requested year (a balanced panel).
|
|
||||||
#'
|
|
||||||
#' `id_col` is the government identifier; `keep_ids` are rows exempt from the
|
|
||||||
#' filter (the peer-comparison target, which is the subject of the comparison
|
|
||||||
#' rather than a member of the cohort being balanced).
|
|
||||||
#' @noRd
|
|
||||||
.filter_consistent <- function(result, years, id_col = "canonical_govid",
|
|
||||||
keep_ids = character(0)) {
|
|
||||||
years <- unique(as.integer(years))
|
|
||||||
if (nrow(result) == 0L || length(years) <= 1L) return(result)
|
|
||||||
ids <- setdiff(unique(result[[id_col]]), c(NA, keep_ids))
|
|
||||||
present <- vapply(ids, function(g) {
|
|
||||||
all(years %in% unique(as.integer(result$year[result[[id_col]] == g])))
|
|
||||||
}, logical(1))
|
|
||||||
consistent <- c(ids[present], keep_ids)
|
|
||||||
result[result[[id_col]] %in% consistent | is.na(result[[id_col]]), ,
|
|
||||||
drop = FALSE]
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Per-year coverage metadata, always attached regardless of mode.
|
|
||||||
#'
|
|
||||||
#' Built from the REQUESTED years rather than the years present in the result,
|
|
||||||
#' so a year in which nothing reported still appears -- with
|
|
||||||
#' `n_units_reporting = 0`, which is precisely the disclosure a silently
|
|
||||||
#' missing year fails to make.
|
|
||||||
#'
|
|
||||||
#' `n_units_reporting` describes the result the caller actually received, so
|
|
||||||
#' under `coverage = "consistent"` it reports the balanced count. `is_census_year`
|
|
||||||
#' is a statement about the SURVEY CALENDAR, never a claim of completeness:
|
|
||||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities report.
|
|
||||||
#' `n_units_reporting` is the number that tells the truth.
|
|
||||||
#' @noRd
|
|
||||||
.coverage_table <- function(result, years, n_expected,
|
|
||||||
id_col = "canonical_govid", rows = NULL) {
|
|
||||||
years <- sort(unique(as.integer(years)))
|
|
||||||
src <- if (is.null(rows)) result else rows
|
|
||||||
reporting <- vapply(years, function(y) {
|
|
||||||
ids <- src[[id_col]][as.integer(src$year) == y]
|
|
||||||
length(unique(ids[!is.na(ids)]))
|
|
||||||
}, integer(1))
|
|
||||||
tibble::tibble(
|
|
||||||
year = years,
|
|
||||||
n_units_reporting = as.integer(reporting),
|
|
||||||
n_units_expected = rep(as.integer(n_expected), length(years)),
|
|
||||||
is_census_year = .is_census_year(years)
|
|
||||||
)
|
|
||||||
}
|
|
||||||
-96
@@ -60,34 +60,6 @@ cog_explain <- function(result, format = c("print", "list")) {
|
|||||||
cli::cli_text("Basis: {prov$basis}{note}")
|
cli::cli_text("Basis: {prov$basis}{note}")
|
||||||
}
|
}
|
||||||
|
|
||||||
# Each verb reports its OWN concept. Both fields are always present (each
|
|
||||||
# defaults to its concept's default), so printing `expenditure_concept`
|
|
||||||
# unconditionally would tell a cog_revenue() caller "Concept: primary",
|
|
||||||
# which names a spending concept their result has nothing to do with.
|
|
||||||
if (identical(prov$verb, "cog_revenue")) {
|
|
||||||
if (!is.null(prov$revenue_concept)) {
|
|
||||||
cli::cli_text("Concept: {prov$revenue_concept} revenue")
|
|
||||||
}
|
|
||||||
} else if (identical(prov$verb, "cog_balances")) {
|
|
||||||
# Both concept fields are deliberately NA here (a stock has no flow
|
|
||||||
# concept). Printing the raw NA reads as a missing value rather than an
|
|
||||||
# intentional one, so say what it means instead.
|
|
||||||
cli::cli_text("Concept: not applicable (holdings are a stock, not a flow)")
|
|
||||||
} else if (!is.null(prov$expenditure_concept)) {
|
|
||||||
concept_note <- if (!is.null(prov$expenditure_concept_note) &&
|
|
||||||
!is.na(prov$expenditure_concept_note)) {
|
|
||||||
sprintf(" (%s)", prov$expenditure_concept_note)
|
|
||||||
} else {
|
|
||||||
""
|
|
||||||
}
|
|
||||||
cli::cli_text("Concept: {prov$expenditure_concept}{concept_note}")
|
|
||||||
if (isTRUE(prov$expenditure_concept_direct_suppressed)) {
|
|
||||||
cli::cli_alert_warning(
|
|
||||||
"Direct leg unavailable for at least one requested (year, category) -- affected rows report intergovernmental dollars alone, not Direct + IG. See each row's notes."
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
cli::cli_h2("Codes observed")
|
cli::cli_h2("Codes observed")
|
||||||
codes <- prov$codes_summed$observed
|
codes <- prov$codes_summed$observed
|
||||||
if (length(codes) == 0L) {
|
if (length(codes) == 0L) {
|
||||||
@@ -131,79 +103,11 @@ cog_explain <- function(result, format = c("print", "list")) {
|
|||||||
cli::cli_ul(sugg_lines)
|
cli::cli_ul(sugg_lines)
|
||||||
}
|
}
|
||||||
|
|
||||||
if (!is.null(prov$coverage) && nrow(prov$coverage) > 0L) {
|
|
||||||
cli::cli_h2("Reporting coverage")
|
|
||||||
cli::cli_text("Mode: {prov$coverage_mode %||% 'all'}")
|
|
||||||
cov <- prov$coverage
|
|
||||||
cli::cli_ul(sprintf(
|
|
||||||
"%d: %d of %d units reporting (%.0f%%) -- %s year",
|
|
||||||
cov$year, cov$n_units_reporting, cov$n_units_expected,
|
|
||||||
100 * cov$n_units_reporting / pmax(cov$n_units_expected, 1L),
|
|
||||||
ifelse(cov$is_census_year, "census", "sample")
|
|
||||||
))
|
|
||||||
if (any(!cov$is_census_year)) {
|
|
||||||
cli::cli_text(
|
|
||||||
"Note: the Census of Governments is a complete census only in years ending in 2 or 7; every other year is a sample."
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
if (isTRUE(prov$completion$applied)) {
|
|
||||||
cli::cli_h2("Completion")
|
|
||||||
cli::cli_text(
|
|
||||||
"Filled {prov$completion$rows_filled} absent cell(s) from the corpus code set."
|
|
||||||
)
|
|
||||||
rules <- prov$completion$absence_means
|
|
||||||
if (length(rules) > 0L) {
|
|
||||||
cli::cli_ul(vapply(names(rules), function(y) {
|
|
||||||
sprintf("%s: an absent cell means %s", y,
|
|
||||||
if (identical(rules[[y]], "census_zero")) {
|
|
||||||
"Census published $0 (filled as 0)"
|
|
||||||
} else {
|
|
||||||
"the government did not report (filled as NA, not 0)"
|
|
||||||
})
|
|
||||||
}, character(1)))
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
if (length(prov$series_break_refs) > 0L) {
|
if (length(prov$series_break_refs) > 0L) {
|
||||||
cli::cli_h2("Series breaks")
|
cli::cli_h2("Series breaks")
|
||||||
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
cli::cli_ul(.series_break_story_lines(prov$series_break_refs))
|
||||||
}
|
}
|
||||||
|
|
||||||
# Kept in a section of its own: these qualify the whole result, so folding
|
|
||||||
# them in with the per-code breaks above would invite reading them as a
|
|
||||||
# caveat about one series.
|
|
||||||
if (length(prov$corpus_break_refs) > 0L) {
|
|
||||||
cli::cli_h2("Corpus-wide caveats")
|
|
||||||
cli::cli_ul(.series_break_story_lines(prov$corpus_break_refs))
|
|
||||||
}
|
|
||||||
|
|
||||||
# Balance results only (NULL on money-verb provenance, so they are
|
|
||||||
# unaffected). This is the ONLY on-demand surface for the GAAP disclosure:
|
|
||||||
# .emit_balance_caveats() fires at most once per session, and is routinely
|
|
||||||
# consumed by a suppressMessages() call or by a knitted chunk nobody reads,
|
|
||||||
# so a caller who deliberately audits a result with cog_explain() must still
|
|
||||||
# be told.
|
|
||||||
bc <- prov$balance_caveats
|
|
||||||
if (!is.null(bc)) {
|
|
||||||
cli::cli_h2("Holdings caveats")
|
|
||||||
if (!is.null(bc$not_gaap_note)) cli::cli_alert_warning(bc$not_gaap_note)
|
|
||||||
if (length(bc$truncated) > 0L) {
|
|
||||||
cli::cli_text(
|
|
||||||
"Requested years extend beyond what these families actually cover:"
|
|
||||||
)
|
|
||||||
cli::cli_ul(vapply(bc$truncated, function(s) {
|
|
||||||
w <- bc$coverage_window[[s]]
|
|
||||||
if (length(w) == 2L) {
|
|
||||||
sprintf("%s: covered %s-%s in this corpus", s, w[1], w[2])
|
|
||||||
} else {
|
|
||||||
s
|
|
||||||
}
|
|
||||||
}, character(1)))
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
cli::cli_h2("Transformations")
|
cli::cli_h2("Transformations")
|
||||||
uc <- prov$transformations$units_conversion
|
uc <- prov$transformations$units_conversion
|
||||||
if (isTRUE(uc$applied)) {
|
if (isTRUE(uc$applied)) {
|
||||||
|
|||||||
@@ -19,13 +19,6 @@
|
|||||||
#' target's population at `year` to produce absolute bounds. If `FALSE`,
|
#' target's population at `year` to produce absolute bounds. If `FALSE`,
|
||||||
#' `pop_range` is interpreted as absolute population counts.
|
#' `pop_range` is interpreted as absolute population counts.
|
||||||
#' @param max_peers Integer cap on the number of peers returned.
|
#' @param max_peers Integer cap on the number of peers returned.
|
||||||
#' @param coverage Survey-cycle handling; see [cog_peer_compare()]. Here it
|
|
||||||
#' governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
|
||||||
#' most recent census year with an observed population, so a cohort is not
|
|
||||||
#' built from a sample year in which most of the candidate universe is
|
|
||||||
#' absent. `"consistent"` needs a year range, which cohort selection does not
|
|
||||||
#' have, so it selects like `"all"` and is carried on the result as
|
|
||||||
#' `attr(x, "coverage")` for [cog_peer_compare()].
|
|
||||||
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||||
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
|
#' `population`, `pop_ratio`, `rank`. The cohort year is attached as
|
||||||
#' `attr(x, "cohort_year")`.
|
#' `attr(x, "cohort_year")`.
|
||||||
@@ -36,9 +29,7 @@ cog_find_peers <- function(target_govid,
|
|||||||
same_state = FALSE,
|
same_state = FALSE,
|
||||||
pop_range = c(0.7, 1.3),
|
pop_range = c(0.7, 1.3),
|
||||||
is_ratio = TRUE,
|
is_ratio = TRUE,
|
||||||
max_peers = 10L,
|
max_peers = 10L) {
|
||||||
coverage = c("all", "census", "consistent")) {
|
|
||||||
coverage <- .validate_coverage(coverage)
|
|
||||||
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
||||||
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
||||||
}
|
}
|
||||||
@@ -68,7 +59,7 @@ cog_find_peers <- function(target_govid,
|
|||||||
))
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
cohort_year <- .resolve_cohort_year(con, target_govid, year, coverage)
|
cohort_year <- .resolve_cohort_year(con, target_govid, year)
|
||||||
|
|
||||||
pop_sql <- sprintf(
|
pop_sql <- sprintf(
|
||||||
"SELECT population FROM gov_population_yearly
|
"SELECT population FROM gov_population_yearly
|
||||||
@@ -116,34 +107,12 @@ cog_find_peers <- function(target_govid,
|
|||||||
attr(peers, "cohort_year") <- as.integer(cohort_year)
|
attr(peers, "cohort_year") <- as.integer(cohort_year)
|
||||||
attr(peers, "pop_range") <- as.numeric(pop_range)
|
attr(peers, "pop_range") <- as.numeric(pop_range)
|
||||||
attr(peers, "is_ratio") <- isTRUE(is_ratio)
|
attr(peers, "is_ratio") <- isTRUE(is_ratio)
|
||||||
attr(peers, "coverage") <- coverage
|
|
||||||
attr(peers, "is_census_year") <- .is_census_year(cohort_year)
|
|
||||||
peers
|
peers
|
||||||
}
|
}
|
||||||
|
|
||||||
# `coverage` picks the cohort vintage when the caller did not name one.
|
|
||||||
# "census" snaps to the most recent CENSUS year with an observed population,
|
|
||||||
# so a cohort is not silently built from a sample year in which most of the
|
|
||||||
# candidate universe is absent. "consistent" is a comparison-time concept --
|
|
||||||
# it needs a year RANGE, which cohort selection does not have -- so it selects
|
|
||||||
# like "all" here and is carried on the result for cog_peer_compare().
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.resolve_cohort_year <- function(con, target_govid, year,
|
.resolve_cohort_year <- function(con, target_govid, year) {
|
||||||
coverage = "all") {
|
|
||||||
if (!is.null(year)) return(as.integer(year))
|
if (!is.null(year)) return(as.integer(year))
|
||||||
if (identical(coverage, "census")) {
|
|
||||||
sql <- sprintf(
|
|
||||||
"SELECT MAX(year) AS y FROM gov_population_yearly
|
|
||||||
WHERE canonical_govid = %s AND year %% 10 IN (2, 7)",
|
|
||||||
.sql_lit_chr(target_govid)
|
|
||||||
)
|
|
||||||
y <- DBI::dbGetQuery(con, sql)$y
|
|
||||||
if (length(y) > 0L && !is.na(y)) return(as.integer(y))
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.code coverage = \"census\"} found no census year with an observed population for {target_govid}.",
|
|
||||||
i = "Pass an explicit {.arg year}, or use {.code coverage = \"all\"}."
|
|
||||||
), class = "uscogdata_no_census_years")
|
|
||||||
}
|
|
||||||
sql <- sprintf(
|
sql <- sprintf(
|
||||||
"SELECT MAX(year) AS y FROM gov_population_yearly
|
"SELECT MAX(year) AS y FROM gov_population_yearly
|
||||||
WHERE canonical_govid = %s",
|
WHERE canonical_govid = %s",
|
||||||
@@ -164,9 +133,7 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
#' [cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||||
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
#' appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||||
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
#' `summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||||
#' call. Those summary rows are quantiles **within each category**, not
|
#' call.
|
||||||
#' quantiles of each peer's total — see the `@return` section before summing
|
|
||||||
#' them.
|
|
||||||
#'
|
#'
|
||||||
#' @param target_govid Character scalar.
|
#' @param target_govid Character scalar.
|
||||||
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
#' @param peers A tibble from [cog_find_peers()] or a character vector of
|
||||||
@@ -176,36 +143,6 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' @param per_capita Default `TRUE` — peer compare usually normalizes by
|
#' @param per_capita Default `TRUE` — peer compare usually normalizes by
|
||||||
#' population.
|
#' population.
|
||||||
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
|
#' @param adjust_to_year Integer base year for CPI-U conversion or `NULL`.
|
||||||
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
|
|
||||||
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
|
||||||
#' refused here because combining Total across peer sets counts
|
|
||||||
#' intergovernmental transfers twice; `"primary"` and `"direct"` combine
|
|
||||||
#' safely.
|
|
||||||
#' @param coverage How to handle the Census of Governments survey cycle,
|
|
||||||
#' which is a **complete census only in years ending in 2 and 7** -- every
|
|
||||||
#' other year is a sample, and the sample varies enormously (on the bundled
|
|
||||||
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
|
||||||
#' and 112 in FY2019).
|
|
||||||
#'
|
|
||||||
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
|
||||||
#' behaviour, so existing code keeps working.
|
|
||||||
#' * `"census"` -- census years only. Aborts if the requested range holds
|
|
||||||
#' none, rather than silently returning nothing.
|
|
||||||
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
|
||||||
#' a balanced panel.
|
|
||||||
#'
|
|
||||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
|
||||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
|
||||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
|
||||||
#' statement about the **survey calendar**, never a claim of completeness:
|
|
||||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
|
||||||
#' report. `n_units_reporting` is the number that tells the truth.
|
|
||||||
#'
|
|
||||||
#' The comparison target is exempt from `"consistent"` balancing -- it is the
|
|
||||||
#' subject of the comparison, not a member of the cohort -- and the
|
|
||||||
#' `summary_*` quantiles are computed AFTER the filter, so they describe the
|
|
||||||
#' cohort actually returned. `n_units_reporting` counts peers only, against
|
|
||||||
#' the cohort size: "3 of your 15 peers reported in FY2019".
|
|
||||||
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
|
#' @return Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||||
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
#' column taking values `"target"`, `"peer"`, `"summary_p25"`,
|
||||||
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
#' `"summary_p50"`, or `"summary_p75"`, `target_rank` (target's rank
|
||||||
@@ -214,43 +151,10 @@ cog_find_peers <- function(target_govid,
|
|||||||
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
#' `attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||||
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
#' vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||||
#' `cohort_year`, and `cohort_govids`.
|
#' `cohort_year`, and `cohort_govids`.
|
||||||
#'
|
|
||||||
#' **The `summary_*` rows are per-category quantiles: they are not additive.**
|
|
||||||
#' Each one is computed **within each `(year, spend_subtype,
|
|
||||||
#' category)` cell** across the peer set, so a `summary_p50` row is *the
|
|
||||||
#' median peer's value in that one category*, not *the value of the median
|
|
||||||
#' peer's total*. The median peer for Police and the median peer for Fire
|
|
||||||
#' are usually different governments, so summing `summary_*` rows across
|
|
||||||
#' categories does not give any peer's total and misstates the band it
|
|
||||||
#' appears to describe — measured at −32.7% to +251.0% across 24 years on
|
|
||||||
#' one cohort, with a sign flip at FY2012.
|
|
||||||
#'
|
|
||||||
#' Facet by `role` **and** `category` (the documented use, and what the
|
|
||||||
#' rows are built for). For a genuine "median peer's total spending" line,
|
|
||||||
#' sum each peer's own categories first and take the quantile of those
|
|
||||||
#' per-government totals:
|
|
||||||
#'
|
|
||||||
#' ```r
|
|
||||||
#' library(dplyr)
|
|
||||||
#' cmp |>
|
|
||||||
#' filter(role %in% c("target", "peer")) |>
|
|
||||||
#' group_by(year, role, canonical_govid) |>
|
|
||||||
#' summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
|
||||||
#' filter(role == "peer") |>
|
|
||||||
#' group_by(year) |>
|
|
||||||
#' summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
|
||||||
#' ```
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_peer_compare <- function(target_govid, peers, category, years,
|
cog_peer_compare <- function(target_govid, peers, category, years,
|
||||||
per_capita = TRUE, adjust_to_year = NULL,
|
per_capita = TRUE, adjust_to_year = NULL) {
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
coverage = c("all", "census", "consistent")) {
|
|
||||||
call <- match.call()
|
call <- match.call()
|
||||||
expenditure_concept <- match.arg(expenditure_concept)
|
|
||||||
coverage <- .validate_coverage(coverage)
|
|
||||||
if (identical(expenditure_concept, "total")) {
|
|
||||||
.abort_concept_not_aggregatable("cog_peer_compare")
|
|
||||||
}
|
|
||||||
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
if (!is.character(target_govid) || length(target_govid) != 1L) {
|
||||||
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
cli::cli_abort("`target_govid` must be a length-1 character string.")
|
||||||
}
|
}
|
||||||
@@ -270,21 +174,9 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
|||||||
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
|
peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)]
|
||||||
all_govids <- unique(c(target_govid, peer_govids))
|
all_govids <- unique(c(target_govid, peer_govids))
|
||||||
|
|
||||||
years <- .apply_census_years(years, coverage, "cog_peer_compare")
|
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||||
|
|
||||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
|
|
||||||
expenditure_concept = expenditure_concept)
|
|
||||||
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
|
r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer")
|
||||||
|
|
||||||
# The target is exempt from balancing: it is the subject of the comparison,
|
|
||||||
# not a member of the cohort being balanced, and dropping it would leave a
|
|
||||||
# peer comparison with nothing to compare. Filtering happens BEFORE the
|
|
||||||
# quantiles below, so a "consistent" cohort's summary rows describe that
|
|
||||||
# cohort rather than the unbalanced one.
|
|
||||||
if (identical(coverage, "consistent")) {
|
|
||||||
r <- .filter_consistent(r, years, keep_ids = target_govid)
|
|
||||||
}
|
|
||||||
|
|
||||||
value_col <- .peer_value_col(per_capita, adjust_to_year)
|
value_col <- .peer_value_col(per_capita, adjust_to_year)
|
||||||
|
|
||||||
summary_rows <- .peer_summary_rows(r, value_col)
|
summary_rows <- .peer_summary_rows(r, value_col)
|
||||||
@@ -305,14 +197,6 @@ cog_peer_compare <- function(target_govid, peers, category, years,
|
|||||||
canonical_govid = target_govid,
|
canonical_govid = target_govid,
|
||||||
gov_name = unique(r$gov_name[r$role == "target"])
|
gov_name = unique(r$gov_name[r$role == "target"])
|
||||||
)
|
)
|
||||||
# Counted over PEER rows only, against the cohort size: "3 of your 15 peers
|
|
||||||
# reported in FY2019". Including the target would inflate every count by one
|
|
||||||
# and make a cohort that has entirely stopped reporting look non-empty.
|
|
||||||
prov$coverage_mode <- coverage
|
|
||||||
prov$coverage <- .coverage_table(
|
|
||||||
out, years, length(peer_govids),
|
|
||||||
rows = r[r$role == "peer", , drop = FALSE]
|
|
||||||
)
|
|
||||||
attr(out, "provenance") <- prov
|
attr(out, "provenance") <- prov
|
||||||
out
|
out
|
||||||
}
|
}
|
||||||
|
|||||||
+2
-28
@@ -6,13 +6,8 @@
|
|||||||
per_capita, adjust_to_year, result, sql,
|
per_capita, adjust_to_year, result, sql,
|
||||||
subtype_col, basis = NA_character_,
|
subtype_col, basis = NA_character_,
|
||||||
basis_note = NA_character_,
|
basis_note = NA_character_,
|
||||||
expenditure_concept = "primary",
|
|
||||||
expenditure_concept_note = NA_character_,
|
|
||||||
expenditure_concept_direct_suppressed = FALSE,
|
|
||||||
revenue_concept = "general",
|
|
||||||
harmonization = NULL, recipe = NULL,
|
harmonization = NULL, recipe = NULL,
|
||||||
suggestions = list(),
|
suggestions = list()) {
|
||||||
completion = NULL) {
|
|
||||||
manifest <- .uscogdata_env$manifest
|
manifest <- .uscogdata_env$manifest
|
||||||
|
|
||||||
codes <- result[["codes_included"]]
|
codes <- result[["codes_included"]]
|
||||||
@@ -39,20 +34,11 @@
|
|||||||
|
|
||||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||||
con <- .uscogdata_env$con
|
con <- .uscogdata_env$con
|
||||||
have_con <- !is.null(con) && DBI::dbIsValid(con)
|
break_refs <- if (!is.null(con) && DBI::dbIsValid(con)) {
|
||||||
break_refs <- if (have_con) {
|
|
||||||
.build_series_break_refs(con, codes_observed, years, schema_version)
|
.build_series_break_refs(con, codes_observed, years, schema_version)
|
||||||
} else {
|
} else {
|
||||||
character(0)
|
character(0)
|
||||||
}
|
}
|
||||||
# Corpus-wide caveats travel separately: they qualify the whole result
|
|
||||||
# rather than one series, and they do not depend on codes_observed (see
|
|
||||||
# .build_corpus_break_refs()).
|
|
||||||
corpus_refs <- if (have_con) {
|
|
||||||
.build_corpus_break_refs(con, years, schema_version)
|
|
||||||
} else {
|
|
||||||
character(0)
|
|
||||||
}
|
|
||||||
|
|
||||||
list(
|
list(
|
||||||
verb = verb,
|
verb = verb,
|
||||||
@@ -65,10 +51,6 @@
|
|||||||
category = category,
|
category = category,
|
||||||
basis = basis,
|
basis = basis,
|
||||||
basis_note = basis_note,
|
basis_note = basis_note,
|
||||||
expenditure_concept = expenditure_concept,
|
|
||||||
expenditure_concept_note = expenditure_concept_note,
|
|
||||||
expenditure_concept_direct_suppressed = isTRUE(expenditure_concept_direct_suppressed),
|
|
||||||
revenue_concept = revenue_concept,
|
|
||||||
harmonization = harmonization %||% list(
|
harmonization = harmonization %||% list(
|
||||||
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
|
applied = FALSE, na_rows_excluded = 0L, na_amount_excluded = 0,
|
||||||
note = NA_character_
|
note = NA_character_
|
||||||
@@ -128,14 +110,6 @@
|
|||||||
)
|
)
|
||||||
),
|
),
|
||||||
series_break_refs = break_refs,
|
series_break_refs = break_refs,
|
||||||
corpus_break_refs = corpus_refs,
|
|
||||||
# What `complete = TRUE` filled, and the rule it filled by. Always
|
|
||||||
# present so a consumer can read `completion$applied` without testing
|
|
||||||
# for the key -- an absent block and applied = FALSE would otherwise be
|
|
||||||
# indistinguishable from an older reader version.
|
|
||||||
completion = completion %||% list(
|
|
||||||
applied = FALSE, rows_filled = 0L, absence_means = list()
|
|
||||||
),
|
|
||||||
manifest = list(
|
manifest = list(
|
||||||
schema_version = as.integer(manifest$schema_version),
|
schema_version = as.integer(manifest$schema_version),
|
||||||
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
pipeline_commit = manifest$pipeline_commit %||% NA_character_,
|
||||||
|
|||||||
+3
-37
@@ -8,46 +8,14 @@
|
|||||||
#' multiplies by 1000 and records the conversion in `provenance`).
|
#' multiplies by 1000 and records the conversion in `provenance`).
|
||||||
#'
|
#'
|
||||||
#' @inheritParams cog_spending
|
#' @inheritParams cog_spending
|
||||||
#' @param revenue_concept Which of Census's two published revenue concepts to
|
|
||||||
#' return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
|
||||||
#' values -- never as item-code first letters, which cannot classify
|
|
||||||
#' correctly (prefix `Y` spans revenue, expenditure and balance codes, and
|
|
||||||
#' prefix `X` does the same):
|
|
||||||
#'
|
|
||||||
#' * `"general"` (default) -- Census General Revenue: `own_source` +
|
|
||||||
#' `federal` + `state` + `local_aid`. The manual defines this concept by
|
|
||||||
#' subtraction (section 4.3: *"General revenue comprises all revenue
|
|
||||||
#' except that classified as liquor store, utility, or insurance trust
|
|
||||||
#' revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
|
|
||||||
#' insurance trust revenue are all excluded.
|
|
||||||
#' * `"total"` -- Census Total Revenue: every revenue subtype, i.e.
|
|
||||||
#' `general` plus utility, liquor store, and insurance trust revenue
|
|
||||||
#' (`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
|
|
||||||
#' `X01`/`X02`/`X05`/`X08`).
|
|
||||||
#'
|
|
||||||
#' The two are related by Census's own identity, `Total Revenue = General +
|
|
||||||
#' Utility + Liquor Store + Insurance Trust`.
|
|
||||||
#'
|
|
||||||
#' Note that the employee-retirement (`X`) codes stop at FY2016, when those
|
|
||||||
#' systems moved out of the annual finance file into the separate Annual
|
|
||||||
#' Survey of Public Pensions, so a `"total"` series steps down at the
|
|
||||||
#' FY2016/FY2017 seam for reasons that are about collection scope rather
|
|
||||||
#' than revenue (series breaks `SB197`-`SB202`).
|
|
||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
#' `revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||||
#' and `value_source` when `complete = TRUE`.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_revenue <- function(govid, years, category = NULL,
|
cog_revenue <- function(govid, years, category = NULL,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||||
revenue_concept = c("general", "total"),
|
|
||||||
complete = FALSE) {
|
|
||||||
# flow_prefixes no longer classifies rows (crosswalk revenue_subtype
|
|
||||||
# membership does -- General Revenue, i.e. everything except
|
|
||||||
# insurance_trust) -- it only scopes the recipe-suggestion machinery to
|
|
||||||
# this verb's recipe families (see R/suggestions.R).
|
|
||||||
.verb_spendrev(
|
.verb_spendrev(
|
||||||
verb = "cog_revenue",
|
verb = "cog_revenue",
|
||||||
view_base = "revenue_annotated",
|
view_base = "revenue_annotated",
|
||||||
@@ -60,8 +28,6 @@ cog_revenue <- function(govid, years, category = NULL,
|
|||||||
per_capita = per_capita,
|
per_capita = per_capita,
|
||||||
adjust_to_year = adjust_to_year,
|
adjust_to_year = adjust_to_year,
|
||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe,
|
recipe = recipe
|
||||||
revenue_concept = revenue_concept,
|
|
||||||
complete = complete
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|||||||
+2
-49
@@ -25,31 +25,6 @@
|
|||||||
#' population from `gov_population_yearly`. Govs with missing population
|
#' population from `gov_population_yearly`. Govs with missing population
|
||||||
#' are excluded from the result.
|
#' are excluded from the result.
|
||||||
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
|
#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`.
|
||||||
#' @param expenditure_concept `"primary"` (default), `"direct"`, or
|
|
||||||
#' `"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
|
||||||
#' refused here because combining Total across multiple layers of
|
|
||||||
#' government double-counts intergovernmental transfers (a state's payment
|
|
||||||
#' to a school district is the same dollar the district reports as its own
|
|
||||||
#' Direct spending); `"primary"` and `"direct"` combine safely.
|
|
||||||
#' @param coverage How to handle the Census of Governments survey cycle,
|
|
||||||
#' which is a **complete census only in years ending in 2 and 7** -- every
|
|
||||||
#' other year is a sample, and the sample varies enormously (on the bundled
|
|
||||||
#' fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
|
||||||
#' and 112 in FY2019).
|
|
||||||
#'
|
|
||||||
#' * `"all"` (default) -- every unit that reported that year. Unchanged
|
|
||||||
#' behaviour, so existing code keeps working.
|
|
||||||
#' * `"census"` -- census years only. Aborts if the requested range holds
|
|
||||||
#' none, rather than silently returning nothing.
|
|
||||||
#' * `"consistent"` -- only units reporting in *every* requested year, giving
|
|
||||||
#' a balanced panel.
|
|
||||||
#'
|
|
||||||
#' Regardless of mode, `provenance$coverage` always carries per-year
|
|
||||||
#' `n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
|
||||||
#' `provenance$coverage_mode` records the mode. `is_census_year` is a
|
|
||||||
#' statement about the **survey calendar**, never a claim of completeness:
|
|
||||||
#' FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
|
||||||
#' report. `n_units_reporting` is the number that tells the truth.
|
|
||||||
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` /
|
||||||
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`,
|
||||||
@@ -58,15 +33,8 @@
|
|||||||
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
#' and `rollup$included_govids` / `rollup$excluded_govids`.
|
||||||
#' @export
|
#' @export
|
||||||
cog_geographic_rollup <- function(govids, category, years,
|
cog_geographic_rollup <- function(govids, category, years,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL) {
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
coverage = c("all", "census", "consistent")) {
|
|
||||||
call <- match.call()
|
call <- match.call()
|
||||||
expenditure_concept <- match.arg(expenditure_concept)
|
|
||||||
coverage <- .validate_coverage(coverage)
|
|
||||||
if (identical(expenditure_concept, "total")) {
|
|
||||||
.abort_concept_not_aggregatable("cog_geographic_rollup")
|
|
||||||
}
|
|
||||||
.validate_rollup_layers(govids)
|
.validate_rollup_layers(govids)
|
||||||
|
|
||||||
govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]")
|
govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]")
|
||||||
@@ -80,21 +48,11 @@ cog_geographic_rollup <- function(govids, category, years,
|
|||||||
layer = rep(layer_names, lengths(govids))
|
layer = rep(layer_names, lengths(govids))
|
||||||
)
|
)
|
||||||
|
|
||||||
# coverage = "census" drops non-census years BEFORE the query rather than
|
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year)
|
||||||
# after: a sample year's rows are not wanted at all, and fetching them only
|
|
||||||
# to discard them would also let them into the coverage table.
|
|
||||||
years <- .apply_census_years(years, coverage, "cog_geographic_rollup")
|
|
||||||
|
|
||||||
r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year,
|
|
||||||
expenditure_concept = expenditure_concept)
|
|
||||||
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
|
r <- dplyr::left_join(r, layer_map, by = "canonical_govid",
|
||||||
relationship = "many-to-many")
|
relationship = "many-to-many")
|
||||||
r$scope_note <- .rollup_scope_note(r$layer)
|
r$scope_note <- .rollup_scope_note(r$layer)
|
||||||
|
|
||||||
if (identical(coverage, "consistent")) {
|
|
||||||
r <- .filter_consistent(r, years)
|
|
||||||
}
|
|
||||||
|
|
||||||
excluded <- character(0)
|
excluded <- character(0)
|
||||||
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
|
if (isTRUE(per_capita) && "pop_source" %in% names(r)) {
|
||||||
drop <- r$pop_source == "unavailable"
|
drop <- r$pop_source == "unavailable"
|
||||||
@@ -113,11 +71,6 @@ cog_geographic_rollup <- function(govids, category, years,
|
|||||||
included_govids = included,
|
included_govids = included,
|
||||||
excluded_govids = excluded
|
excluded_govids = excluded
|
||||||
)
|
)
|
||||||
# n_units_expected is the universe the CALLER named -- the govids passed in
|
|
||||||
# -- not the national universe. That is what makes the ratio meaningful:
|
|
||||||
# "597 of the 608 Wisconsin cities you asked about reported in FY2012".
|
|
||||||
prov$coverage_mode <- coverage
|
|
||||||
prov$coverage <- .coverage_table(r, years, length(unique(all_govids)))
|
|
||||||
attr(r, "provenance") <- prov
|
attr(r, "provenance") <- prov
|
||||||
|
|
||||||
r
|
r
|
||||||
|
|||||||
+7
-19
@@ -6,11 +6,8 @@
|
|||||||
#' the cross-vintage canonical-government registry. Operates in two modes:
|
#' the cross-vintage canonical-government registry. Operates in two modes:
|
||||||
#'
|
#'
|
||||||
#' * **Utility mode** (single `name`, the original behavior): returns all
|
#' * **Utility mode** (single `name`, the original behavior): returns all
|
||||||
#' rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
#' rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||||
#' substring**, sorted by `population_acs` descending. Useful for
|
#' `population_acs` descending. Useful for exploratory lookups.
|
||||||
#' exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
|
||||||
#' government is findable by its own complete name even when that name
|
|
||||||
#' contains parentheses or a period.
|
|
||||||
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
#' * **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||||
#' single canonical govid and returns a tibble in input order, suitable
|
#' single canonical govid and returns a tibble in input order, suitable
|
||||||
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
#' for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||||
@@ -22,8 +19,7 @@
|
|||||||
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
#' 1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||||
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
#' 2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||||
#' Single hit -> resolved. Multiple -> step 4.
|
#' Single hit -> resolved. Multiple -> step 4.
|
||||||
#' 3. **Substring fallback:** case-insensitive literal substring against
|
#' 3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||||
#' `gov_name` (metacharacters escaped).
|
|
||||||
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
#' Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||||
#' `status = "no_match"`. Multiple hits -> step 4.
|
#' `status = "no_match"`. Multiple hits -> step 4.
|
||||||
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
#' 4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||||
@@ -52,7 +48,7 @@
|
|||||||
#' [cog_spending()], [cog_revenue()].
|
#' [cog_spending()], [cog_revenue()].
|
||||||
#' @examples
|
#' @examples
|
||||||
#' \dontrun{
|
#' \dontrun{
|
||||||
#' # Utility mode — exploratory substring lookup
|
#' # Utility mode — exploratory regex lookup
|
||||||
#' cog_gov_search("broward", state = "FL")
|
#' cog_gov_search("broward", state = "FL")
|
||||||
#'
|
#'
|
||||||
#' # Basket mode — resolve a known cohort
|
#' # Basket mode — resolve a known cohort
|
||||||
@@ -102,16 +98,9 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
|||||||
if (!is.character(name) || length(name) != 1L) {
|
if (!is.character(name) || length(name) != 1L) {
|
||||||
cli::cli_abort("`name` must be a length-1 character string.")
|
cli::cli_abort("`name` must be a length-1 character string.")
|
||||||
}
|
}
|
||||||
# Escaped, so `name` is a literal case-insensitive substring -- the same
|
|
||||||
# treatment basket mode has always given it. Interpolating it raw made a
|
|
||||||
# government unfindable by its own name whenever that name contains a
|
|
||||||
# metacharacter (FREDONIA (BRISCOE) CITY), turned a bare "." into a
|
|
||||||
# match-everything wildcard, and let malformed pattern text reach the
|
|
||||||
# engine as an error -- which cog-api surfaced as a 500, reachable by
|
|
||||||
# typing a real name one character at a time (uscogdata#16, F-025).
|
|
||||||
preds <- c(preds,
|
preds <- c(preds,
|
||||||
sprintf("regexp_matches(gov_name, %s, 'i')",
|
sprintf("regexp_matches(gov_name, %s, 'i')",
|
||||||
.sql_lit_chr(.escape_regex(name))))
|
.sql_lit_chr(name)))
|
||||||
}
|
}
|
||||||
if (!is.null(state)) {
|
if (!is.null(state)) {
|
||||||
st_fips <- .coerce_state_to_fips(state)
|
st_fips <- .coerce_state_to_fips(state)
|
||||||
@@ -147,9 +136,8 @@ cog_gov_search <- function(name = NULL, state = NULL, type = NULL) {
|
|||||||
#' @noRd
|
#' @noRd
|
||||||
.escape_regex <- function(x) {
|
.escape_regex <- function(x) {
|
||||||
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
# Backslash-escape POSIX regex metacharacters so `name` is treated as a
|
||||||
# literal substring in the DuckDB regexp_matches call. Used by BOTH modes:
|
# literal substring in the DuckDB regexp_matches call (substring fallback
|
||||||
# utility mode used to interpolate raw, which was a defect rather than a
|
# only; utility-mode intentionally preserves regex behavior).
|
||||||
# feature -- see the call site and uscogdata#16.
|
|
||||||
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
gsub("([\\^$.|?*+(){}\\[\\]])", "\\\\\\1", x, perl = TRUE)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+1
-33
@@ -14,41 +14,9 @@
|
|||||||
sql <- sprintf(
|
sql <- sprintf(
|
||||||
"SELECT DISTINCT break_id
|
"SELECT DISTINCT break_id
|
||||||
FROM series_breaks_pq
|
FROM series_breaks_pq
|
||||||
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
|
WHERE fin_code IN (%s) AND break_year BETWEEN %d AND %d
|
||||||
AND break_year BETWEEN %d AND %d
|
|
||||||
ORDER BY break_id",
|
ORDER BY break_id",
|
||||||
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
||||||
)
|
)
|
||||||
DBI::dbGetQuery(con, sql)$break_id
|
DBI::dbGetQuery(con, sql)$break_id
|
||||||
}
|
}
|
||||||
|
|
||||||
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
|
|
||||||
#' `"ALL"` rather than an item code. They qualify the whole result, so they
|
|
||||||
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
|
|
||||||
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
|
|
||||||
#' before uscogdata#19. Selection is on the break_year window alone: which
|
|
||||||
#' codes a result happens to contain is irrelevant to a caveat about the
|
|
||||||
#' corpus.
|
|
||||||
#'
|
|
||||||
#' All four catalogued entries are *boundary* caveats (dollar precision
|
|
||||||
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
|
|
||||||
#' representation change at 2012, the id scheme change at 2017), so the same
|
|
||||||
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
|
|
||||||
#' path uses is the right one -- a request that never crosses the boundary
|
|
||||||
#' is not affected by it.
|
|
||||||
#'
|
|
||||||
#' Returned separately from `series_break_refs` so a consumer can tell a
|
|
||||||
#' whole-result caveat from a break in one series; the two are disjoint by
|
|
||||||
#' construction.
|
|
||||||
#' @noRd
|
|
||||||
.build_corpus_break_refs <- function(con, years, schema_version) {
|
|
||||||
if (schema_version < 5L || length(years) == 0L) return(character(0))
|
|
||||||
sql <- sprintf(
|
|
||||||
"SELECT DISTINCT break_id
|
|
||||||
FROM series_breaks_pq
|
|
||||||
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
|
|
||||||
ORDER BY break_id",
|
|
||||||
min(as.integer(years)), max(as.integer(years))
|
|
||||||
)
|
|
||||||
DBI::dbGetQuery(con, sql)$break_id
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -95,7 +95,4 @@ cog_close <- function() {
|
|||||||
}
|
}
|
||||||
.uscogdata_env$con <- NULL
|
.uscogdata_env$con <- NULL
|
||||||
.uscogdata_env$manifest <- NULL
|
.uscogdata_env$manifest <- NULL
|
||||||
.uscogdata_env$balance_caveats_shown <- NULL
|
|
||||||
# Memoised corpus-constant; a different corpus may be mounted next.
|
|
||||||
.uscogdata_env$balance_coverage_windows <- NULL
|
|
||||||
}
|
}
|
||||||
|
|||||||
+17
-515
@@ -1,57 +1,5 @@
|
|||||||
# R/spending.R
|
# R/spending.R
|
||||||
|
|
||||||
# The three expenditure concepts (uscogdata#11), as sets of the crosswalk's
|
|
||||||
# `spend_subtype` values. Classification is crosswalk membership, never
|
|
||||||
# item-code first letters: prefix Y alone spans revenue (Y01/Y02),
|
|
||||||
# expenditure (Y05/Y06) and balance codes, so no first-letter allowlist can
|
|
||||||
# route it (finding F-018).
|
|
||||||
#
|
|
||||||
# primary = operations + capital + assistance (the default)
|
|
||||||
# direct = primary + interest + insurance_benefits (Census Direct Expenditure)
|
|
||||||
# total = direct + intergovernmental (via the ig_* views)
|
|
||||||
#
|
|
||||||
# Census manual section 5.2.2.1: Direct Expenditure is ALL expenditure other
|
|
||||||
# than intergovernmental -- including payments to retirees, i.e. insurance
|
|
||||||
# trust benefits. Verified against Census's own published FY2020 state
|
|
||||||
# aggregates (20statetypepu.txt): `total` reproduces the published
|
|
||||||
# expenditure sum to the dollar; omitting insurance benefits understates
|
|
||||||
# California's Direct by 10.9%.
|
|
||||||
.spend_subtypes_primary <- c("operations", "capital", "assistance")
|
|
||||||
.spend_subtypes_direct <- c(.spend_subtypes_primary, "interest", "insurance_benefits")
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.expenditure_concept_subtypes <- function(concept) {
|
|
||||||
switch(concept,
|
|
||||||
primary = .spend_subtypes_primary,
|
|
||||||
# "total" = the direct subtypes here PLUS the intergovernmental leg,
|
|
||||||
# which travels through the ig_* views rather than this scope (see
|
|
||||||
# .build_verb_sql()).
|
|
||||||
direct = ,
|
|
||||||
total = .spend_subtypes_direct
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
# The two revenue concepts (uscogdata#12), again as crosswalk subtype sets.
|
|
||||||
# Census's manual section 4.3 defines the first by SUBTRACTING from the second
|
|
||||||
# -- "General revenue comprises all revenue except that classified as liquor
|
|
||||||
# store, utility, or insurance trust revenue" -- giving the identity
|
|
||||||
#
|
|
||||||
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
|
||||||
#
|
|
||||||
# Verified against Census's own computed concept fields (IndFin FY2012,
|
|
||||||
# Wisconsin state): 31,410,686 + 0 + 0 + 4,469,906 = 35,880,592, exact.
|
|
||||||
.revenue_subtypes_general <- c("own_source", "federal", "state", "local_aid")
|
|
||||||
.revenue_subtypes_total <- c(.revenue_subtypes_general, "utility",
|
|
||||||
"liquor_store", "insurance_trust")
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.revenue_concept_subtypes <- function(concept) {
|
|
||||||
switch(concept,
|
|
||||||
general = .revenue_subtypes_general,
|
|
||||||
total = .revenue_subtypes_total
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Summarized spending by category
|
#' Summarized spending by category
|
||||||
#'
|
#'
|
||||||
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
#' One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||||
@@ -94,91 +42,20 @@
|
|||||||
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
|
#' `basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||||
#' FALSE`, pointing at the `recipe` block instead) rather than a
|
#' FALSE`, pointing at the `recipe` block instead) rather than a
|
||||||
#' possibly-misleading `"harmonized"`/`"raw"` value.
|
#' possibly-misleading `"harmonized"`/`"raw"` value.
|
||||||
#' @param expenditure_concept Which spending concept to return. Concepts are
|
|
||||||
#' defined as sets of the crosswalk's `spend_subtype` values -- never as
|
|
||||||
#' item-code first letters, which cannot classify correctly (prefix `Y`
|
|
||||||
#' alone spans revenue, expenditure, and balance codes):
|
|
||||||
#'
|
|
||||||
#' * `"primary"` (default) -- the government's own service provision:
|
|
||||||
#' `operations` + `capital` + `assistance` subtypes.
|
|
||||||
#' * `"direct"` -- Census's published Direct Expenditure: `primary` plus
|
|
||||||
#' `interest` (interest on debt) and `insurance_benefits` (insurance
|
|
||||||
#' trust benefit payments, e.g. pensions -- Census manual section
|
|
||||||
#' 5.2.2.1 includes payments to retirees in Direct).
|
|
||||||
#' * `"total"` -- `direct` plus the intergovernmental leg: payments to
|
|
||||||
#' local governments (`M` codes), to the state government (`L` codes,
|
|
||||||
#' excluding the `L--` family-total rollup), and state payments to
|
|
||||||
#' school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
|
|
||||||
#' `spend_subtype == "intergovernmental"`. Requires the active corpus's
|
|
||||||
#' `summary_categories` to carry M/L rows (added by cog_pipeline PR
|
|
||||||
#' #59); aborts with class `uscogdata_ig_categories_unsupported` on an
|
|
||||||
#' older corpus rather than silently under-reporting. Mutually
|
|
||||||
#' exclusive with `recipe` (a recipe already defines its own component
|
|
||||||
#' codes).
|
|
||||||
#'
|
|
||||||
#' **Do not sum `"total"` results across levels of government** (e.g.
|
|
||||||
#' state + county + city): a state's `M12` payment to a school district is
|
|
||||||
#' the same dollar the district reports as its own direct `E12`, so
|
|
||||||
#' summing both double-counts it. This matters in particular with
|
|
||||||
#' [cog_geographic_rollup()], which sums across exactly that kind of
|
|
||||||
#' multi-layer government set.
|
|
||||||
#'
|
|
||||||
#' In the legacy wide era (<= FY2011), some functions are published ONLY
|
|
||||||
#' as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
|
||||||
#' split), which the Direct leg excludes by construction but the IG leg
|
|
||||||
#' deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
|
||||||
#' query, any (year, category) where this leaves intergovernmental rows
|
|
||||||
#' with NO Direct counterpart is flagged: the affected rows' `notes`
|
|
||||||
#' name the harmonization recipe that recovers the missing Direct
|
|
||||||
#' component (when one exists), and
|
|
||||||
#' `provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
|
||||||
#' figure in those rows is the intergovernmental leg alone, not Direct +
|
|
||||||
#' IG.
|
|
||||||
#' @param complete If `TRUE`, fill the requested grid so that a cell the
|
|
||||||
#' corpus does not carry still appears, labelled with **why** it is
|
|
||||||
#' missing, and add a `value_source` column to every row:
|
|
||||||
#'
|
|
||||||
#' * `"reported"` — the corpus carries this cell.
|
|
||||||
#' * `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
|
||||||
#' Census published `$0`. `amt_nominal` is `0`.
|
|
||||||
#' * `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
|
||||||
#' government did not report, and the value is unknown. `amt_nominal` is
|
|
||||||
#' `NA`, **not** `0` — writing a zero there would invent data.
|
|
||||||
#'
|
|
||||||
#' The grid comes from the corpus's `code_set` table, scoped to each
|
|
||||||
#' government's own type, so a county is never filled with cells only a
|
|
||||||
#' state can report. Reported rows are passed through untouched.
|
|
||||||
#'
|
|
||||||
#' Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
|
||||||
#' not appear). Needs a corpus published from 2026-07-29 onward, which is
|
|
||||||
#' when `representation`/`code_set` began shipping; aborts with class
|
|
||||||
#' `uscogdata_representation_unavailable` otherwise. Not available with
|
|
||||||
#' `recipe` or with `expenditure_concept = "total"` (class
|
|
||||||
#' `uscogdata_complete_unsupported`) — neither draws its cells from
|
|
||||||
#' `code_set`.
|
|
||||||
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
#' @return Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
#' optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
#' optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||||
#' and `value_source` when `complete = TRUE`.
|
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||||
#' Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
|
||||||
#' whose `completion` block reports `applied`, `rows_filled`, and the
|
|
||||||
#' per-year `absence_means` rule that was applied.
|
|
||||||
#' @export
|
#' @export
|
||||||
cog_spending <- function(govid, years, category = NULL,
|
cog_spending <- function(govid, years, category = NULL,
|
||||||
per_capita = FALSE, adjust_to_year = NULL,
|
per_capita = FALSE, adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
complete = FALSE) {
|
|
||||||
# flow_prefixes no longer classifies rows (crosswalk subtype membership
|
|
||||||
# does, per expenditure_concept) -- it only scopes the recipe-suggestion
|
|
||||||
# machinery to this verb's recipe families (see R/suggestions.R; the
|
|
||||||
# catalog only has E/F/G-component direct-expenditure recipes).
|
|
||||||
.verb_spendrev(
|
.verb_spendrev(
|
||||||
verb = "cog_spending",
|
verb = "cog_spending",
|
||||||
view_base = "spending_annotated",
|
view_base = "spending_annotated",
|
||||||
subtype_col = "spend_subtype",
|
subtype_col = "spend_subtype",
|
||||||
flow_prefixes = c("E", "F", "G"),
|
flow_prefixes = c("E", "F", "G", "K"),
|
||||||
call = match.call(),
|
call = match.call(),
|
||||||
govid = govid,
|
govid = govid,
|
||||||
years = years,
|
years = years,
|
||||||
@@ -186,127 +63,28 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
per_capita = per_capita,
|
per_capita = per_capita,
|
||||||
adjust_to_year = adjust_to_year,
|
adjust_to_year = adjust_to_year,
|
||||||
basis = basis,
|
basis = basis,
|
||||||
recipe = recipe,
|
recipe = recipe
|
||||||
expenditure_concept = expenditure_concept,
|
|
||||||
complete = complete
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.abort_concept_not_aggregatable <- function(verb) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"{.code expenditure_concept = \"total\"} cannot be used in {.fn {verb}}.",
|
|
||||||
"*" = "Use {.code expenditure_concept = \"primary\"} (the default) or \\
|
|
||||||
{.code \"direct\"} for any comparison or sum that spans more than \\
|
|
||||||
one government.",
|
|
||||||
"i" = "Why: Census \"Total\" is a government's own Direct spending PLUS the \\
|
|
||||||
money it hands to other governments. The receiving government reports \\
|
|
||||||
that same dollar again as its own Direct when it actually spends it, \\
|
|
||||||
so combining Total across governments double-counts intergovernmental \\
|
|
||||||
transfers.",
|
|
||||||
"i" = "For one government's own Total, use \\
|
|
||||||
{.code cog_spending(expenditure_concept = \"total\")}."
|
|
||||||
), class = "uscogdata_concept_not_aggregatable")
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.verb_spendrev <- function(verb, view_base, subtype_col, flow_prefixes, call,
|
.verb_spendrev <- function(verb, view_base, subtype_col, flow_prefixes, call,
|
||||||
govid, years, category,
|
govid, years, category,
|
||||||
per_capita, adjust_to_year,
|
per_capita, adjust_to_year,
|
||||||
basis = c("harmonized", "raw"), recipe = NULL,
|
basis = c("harmonized", "raw"), recipe = NULL) {
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
revenue_concept = c("general", "total"),
|
|
||||||
complete = FALSE) {
|
|
||||||
basis_explicit <- length(basis) == 1L
|
basis_explicit <- length(basis) == 1L
|
||||||
basis <- match.arg(basis, c("harmonized", "raw"))
|
basis <- match.arg(basis, c("harmonized", "raw"))
|
||||||
# match.arg() itself throws a base `simpleError`, not an rlang-classed
|
|
||||||
# condition; wrap it so an invalid expenditure_concept aborts consistently
|
|
||||||
# with the rest of this package's validation (cli::cli_abort -> rlang_error).
|
|
||||||
expenditure_concept <- tryCatch(
|
|
||||||
match.arg(expenditure_concept, c("primary", "direct", "total")),
|
|
||||||
error = function(e) {
|
|
||||||
cli::cli_abort(
|
|
||||||
"`expenditure_concept` must be one of {.val primary}, {.val direct}, or {.val total}.",
|
|
||||||
class = "uscogdata_invalid_expenditure_concept",
|
|
||||||
parent = e
|
|
||||||
)
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
revenue_concept <- tryCatch(
|
|
||||||
match.arg(revenue_concept, c("general", "total")),
|
|
||||||
error = function(e) {
|
|
||||||
cli::cli_abort(
|
|
||||||
"`revenue_concept` must be one of {.val general} or {.val total}.",
|
|
||||||
class = "uscogdata_invalid_revenue_concept",
|
|
||||||
parent = e
|
|
||||||
)
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
# The concept's subtype scope. Every code path below -- the verb SQL, the
|
|
||||||
# harmonization exclusion count, and the complete = TRUE grid -- is scoped
|
|
||||||
# by crosswalk subtype membership, never by item-code prefix. The
|
|
||||||
# expenditure "total" concept's extra intergovernmental leg is the one
|
|
||||||
# exception: it travels through the ig_* views rather than this scope,
|
|
||||||
# because its legacy rows are aggregate-flagged.
|
|
||||||
subtype_scope <- if (identical(subtype_col, "spend_subtype")) {
|
|
||||||
.expenditure_concept_subtypes(expenditure_concept)
|
|
||||||
} else {
|
|
||||||
.revenue_concept_subtypes(revenue_concept)
|
|
||||||
}
|
|
||||||
|
|
||||||
govid <- .coerce_govid_input(govid, arg = "govid")
|
govid <- .coerce_govid_input(govid, arg = "govid")
|
||||||
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
.validate_verb_inputs(govid, years, category, per_capita, adjust_to_year,
|
||||||
recipe)
|
recipe)
|
||||||
|
|
||||||
if (!is.null(recipe) && identical(expenditure_concept, "total")) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
"`recipe` and `expenditure_concept = \"total\"` are mutually exclusive.",
|
|
||||||
i = "A recipe defines its own component codes; pass one or the other.",
|
|
||||||
i = "For a recipe's intergovernmental counterpart, use the matching IG recipe (e.g. `corrections_ig_local_combined`)."
|
|
||||||
), class = "uscogdata_recipe_concept_conflict")
|
|
||||||
}
|
|
||||||
|
|
||||||
# .verb_spendrev() is shared with cog_revenue(), which never exposes
|
|
||||||
# expenditure_concept and always resolves it to the default -- so nothing
|
|
||||||
# on the public API can reach this today. But it's a cheap guard against a
|
|
||||||
# future call (direct or via a modified cog_revenue()) that would UNION
|
|
||||||
# the IG leg's expenditure M/L/Q rows into a revenue result, which has no
|
|
||||||
# matching IG view and no sensible meaning.
|
|
||||||
if (identical(expenditure_concept, "total") &&
|
|
||||||
!identical(view_base, "spending_annotated")) {
|
|
||||||
cli::cli_abort(
|
|
||||||
paste0(
|
|
||||||
"`expenditure_concept = \"total\"` is only supported for spending ",
|
|
||||||
"(view_base = \"spending_annotated\"); got view_base = ",
|
|
||||||
"{.val {view_base}}."
|
|
||||||
),
|
|
||||||
class = "uscogdata_expenditure_concept_unsupported"
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
complete <- isTRUE(complete)
|
|
||||||
if (complete && !is.null(recipe)) {
|
|
||||||
.abort_complete_unsupported(
|
|
||||||
"A recipe defines its own component codes and never goes through `summary_categories`, so there is no grid to fill from.",
|
|
||||||
"Query the recipe without `complete`, or use a category query with `complete = TRUE`."
|
|
||||||
)
|
|
||||||
}
|
|
||||||
if (complete && identical(expenditure_concept, "total")) {
|
|
||||||
.abort_complete_unsupported(
|
|
||||||
"The intergovernmental leg deliberately keeps aggregate-flagged rows (see `inst/sql/24-ig_long.sql`), so its cells are not the ones `code_set` describes.",
|
|
||||||
"Use `expenditure_concept = \"direct\"` with `complete = TRUE`, or drop `complete`."
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
years <- as.integer(years)
|
years <- as.integer(years)
|
||||||
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
if (!is.null(adjust_to_year)) adjust_to_year <- as.integer(adjust_to_year)
|
||||||
|
|
||||||
con <- .ensure_session()
|
con <- .ensure_session()
|
||||||
manifest <- .uscogdata_env$manifest
|
manifest <- .uscogdata_env$manifest
|
||||||
scope <- .check_govids_in_scope(govid)
|
scope <- .check_govids_in_scope(govid)
|
||||||
if (complete) .require_representation(con, manifest)
|
|
||||||
|
|
||||||
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
resolved <- .resolve_basis(basis, basis_explicit, manifest)
|
||||||
|
|
||||||
@@ -327,34 +105,17 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
category_for_prov <- recipe_label
|
category_for_prov <- recipe_label
|
||||||
} else {
|
} else {
|
||||||
view <- .select_view(view_base, resolved$basis)
|
view <- .select_view(view_base, resolved$basis)
|
||||||
ig_view <- if (identical(expenditure_concept, "total")) {
|
sql <- .build_verb_sql(view, subtype_col, govid, years, category)
|
||||||
.require_ig_categories(con)
|
|
||||||
.select_ig_view(resolved$basis)
|
|
||||||
} else {
|
|
||||||
NULL
|
|
||||||
}
|
|
||||||
sql <- .build_verb_sql(view, subtype_col, govid, years, category, ig_view,
|
|
||||||
subtype_scope)
|
|
||||||
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||||
}
|
}
|
||||||
|
|
||||||
# Fill BEFORE per_capita / inflation so the added cells get the same
|
|
||||||
# treatment as reported ones: a census_zero stays $0 per capita and in real
|
|
||||||
# dollars, and a not_reported stays NA through both rather than becoming a
|
|
||||||
# spurious 0.
|
|
||||||
completion <- list(applied = FALSE, rows_filled = 0L, absence_means = list())
|
|
||||||
if (complete) {
|
|
||||||
result <- .complete_result(result, con, subtype_col, govid, years,
|
|
||||||
category, subtype_scope)
|
|
||||||
completion <- attr(result, ".completion")
|
|
||||||
attr(result, ".completion") <- NULL
|
|
||||||
}
|
|
||||||
|
|
||||||
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
if (per_capita) result <- .attach_per_capita(result, con, govid)
|
||||||
if (!is.null(adjust_to_year)) {
|
if (!is.null(adjust_to_year)) {
|
||||||
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
result <- .attach_real_dollars(result, adjust_to_year, per_capita)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
result$notes <- .notes_column(result)
|
||||||
|
|
||||||
# A recipe result doesn't go through spending_annotated(_harmonized) /
|
# A recipe result doesn't go through spending_annotated(_harmonized) /
|
||||||
# revenue_annotated(_harmonized) at all -- .run_recipe()'s generic join
|
# revenue_annotated(_harmonized) at all -- .run_recipe()'s generic join
|
||||||
# reads `long` directly -- so `basis` and the `harmonization` exclusion
|
# reads `long` directly -- so `basis` and the `harmonization` exclusion
|
||||||
@@ -378,66 +139,9 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
basis_for_prov <- resolved$basis
|
basis_for_prov <- resolved$basis
|
||||||
basis_note_for_prov <- resolved$note
|
basis_note_for_prov <- resolved$note
|
||||||
harmonization <- .build_harmonization_block(
|
harmonization <- .build_harmonization_block(
|
||||||
con, govid, years, resolved, subtype_col, subtype_scope
|
con, govid, years, resolved, flow_prefixes
|
||||||
)
|
)
|
||||||
# C1(a): gap detection must run against the Direct leg alone. `result`
|
suggestions <- .build_suggestions(con, govid, years, category, resolved$basis)
|
||||||
# can also carry UNION'd intergovernmental rows (expenditure_concept =
|
|
||||||
# "total"), and the wide era (<= FY2011) routinely has legacy IG dollars
|
|
||||||
# surviving (ig_long deliberately keeps aggregate rows) for a
|
|
||||||
# (year, category) whose legacy Direct dollars were suppressed (spending_
|
|
||||||
# long/spending_long_harmonized both filter NOT is_aggregate). Passing
|
|
||||||
# the UNION'd result here would let a surviving IG row count as coverage
|
|
||||||
# and silently cancel the recipe-hint suggestion that should fire.
|
|
||||||
direct_leg_result <- if (identical(expenditure_concept, "total")) {
|
|
||||||
result[!(result[[subtype_col]] %in% "intergovernmental"), , drop = FALSE]
|
|
||||||
} else {
|
|
||||||
result
|
|
||||||
}
|
|
||||||
suggestions <- .build_suggestions(con, govid, years, category,
|
|
||||||
direct_leg_result,
|
|
||||||
resolved$basis, flow_prefixes)
|
|
||||||
}
|
|
||||||
|
|
||||||
# C1(b): when expenditure_concept = "total", flag any row where the IG
|
|
||||||
# leg has dollars but the Direct leg has none for that same (year,
|
|
||||||
# canonical_govid, category) AND a harmonization recipe actually recovers
|
|
||||||
# the missing Direct dollars for that exact triple -- see
|
|
||||||
# .detect_direct_suppressed() for why bare Direct-row absence alone is NOT
|
|
||||||
# sufficient (the dominant real cause is a government that simply has no
|
|
||||||
# direct spending in that category, which is correct, ordinary data). When
|
|
||||||
# a covering recipe is found, both the row-level notes and the provenance
|
|
||||||
# say so rather than pass silently as a plausible Total.
|
|
||||||
direct_suppressed_info <- if (identical(expenditure_concept, "total")) {
|
|
||||||
.detect_direct_suppressed(con, result, subtype_col)
|
|
||||||
} else {
|
|
||||||
list(flag = rep(FALSE, nrow(result)), notes = rep(NA_character_, nrow(result)))
|
|
||||||
}
|
|
||||||
direct_suppressed <- direct_suppressed_info$flag
|
|
||||||
direct_suppressed_flag <- isTRUE(any(direct_suppressed))
|
|
||||||
|
|
||||||
result$notes <- .notes_column(result, direct_suppressed_info$notes)
|
|
||||||
|
|
||||||
# Determine expenditure_concept_note: only non-empty for "total", explains
|
|
||||||
# how the IG leg was assembled from legacy-era aggregates. When the Direct
|
|
||||||
# leg is suppressed for at least one requested (year, category), append an
|
|
||||||
# explicit warning rather than let the base note's "Total = Direct + IG"
|
|
||||||
# framing stand unqualified for rows where that arithmetic didn't happen.
|
|
||||||
expenditure_concept_note_for_prov <- if (identical(expenditure_concept, "total")) {
|
|
||||||
base_note <- "Total = Direct + intergovernmental (M to local govts + L to state govts). Legacy-era IG is assembled from aggregate-flagged rows, which are year-disjoint from their modern leaf components; the L-- family total is excluded."
|
|
||||||
if (direct_suppressed_flag) {
|
|
||||||
paste0(
|
|
||||||
base_note,
|
|
||||||
" NOTE: for at least one requested (year, category) the Direct leg ",
|
|
||||||
"has NO rows in this corpus (a legacy aggregate-only family) -- the ",
|
|
||||||
"affected result rows report the intergovernmental leg alone, not ",
|
|
||||||
"Direct + IG. See `expenditure_concept_direct_suppressed` and each ",
|
|
||||||
"affected row's `notes`."
|
|
||||||
)
|
|
||||||
} else {
|
|
||||||
base_note
|
|
||||||
}
|
|
||||||
} else {
|
|
||||||
NA_character_
|
|
||||||
}
|
}
|
||||||
|
|
||||||
prov <- .build_provenance(
|
prov <- .build_provenance(
|
||||||
@@ -453,14 +157,9 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
subtype_col = subtype_col,
|
subtype_col = subtype_col,
|
||||||
basis = basis_for_prov,
|
basis = basis_for_prov,
|
||||||
basis_note = basis_note_for_prov,
|
basis_note = basis_note_for_prov,
|
||||||
expenditure_concept = expenditure_concept,
|
|
||||||
expenditure_concept_note = expenditure_concept_note_for_prov,
|
|
||||||
expenditure_concept_direct_suppressed = direct_suppressed_flag,
|
|
||||||
revenue_concept = revenue_concept,
|
|
||||||
harmonization = harmonization,
|
harmonization = harmonization,
|
||||||
recipe = recipe_block,
|
recipe = recipe_block,
|
||||||
suggestions = suggestions,
|
suggestions = suggestions
|
||||||
completion = completion
|
|
||||||
)
|
)
|
||||||
prov$scope$govids_found <- scope$found
|
prov$scope$govids_found <- scope$found
|
||||||
prov$scope$govids_missing <- scope$missing
|
prov$scope$govids_missing <- scope$missing
|
||||||
@@ -512,43 +211,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
|
if (identical(basis, "harmonized")) paste0(view_base, "_harmonized") else view_base
|
||||||
}
|
}
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.select_ig_view <- function(basis) {
|
|
||||||
if (identical(basis, "harmonized")) "ig_annotated_harmonized" else "ig_annotated"
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Abort unless the active corpus's `summary_categories` actually carries
|
|
||||||
#' intergovernmental (M/L) rows.
|
|
||||||
#'
|
|
||||||
#' The 66 M/L category rows arrived via cog_pipeline PR #59 with NO
|
|
||||||
#' `schema_version` bump (`DESCRIPTION` still declares `MinCorpusSchema: 4`),
|
|
||||||
#' so `schema_version` alone cannot gate `expenditure_concept = "total"` --
|
|
||||||
#' a pre-#59 corpus can validly report schema_version 4, 5, or 6 and still
|
|
||||||
#' have zero M/L rows in `summary_categories`. Against such a corpus,
|
|
||||||
#' `ig_annotated`'s LEFT JOIN to `summary_categories` silently produces NA
|
|
||||||
#' `category`/`spend_subtype` for every IG row: with a `category` filter
|
|
||||||
#' this returns 0 rows (reads as "no intergovernmental spending" rather than
|
|
||||||
#' "can't tell"), and with `category = NULL` every IG dollar collapses into
|
|
||||||
#' one NA-subtype group that is invisible to the `spend_subtype ==
|
|
||||||
#' "intergovernmental"` filter this package's own tests, roxygen, and
|
|
||||||
#' vignette all rely on. Checking the data directly (rather than
|
|
||||||
#' schema_version) is the only reliable gate.
|
|
||||||
#' @noRd
|
|
||||||
.require_ig_categories <- function(con, what = "expenditure_concept = \"total\"") {
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')"
|
|
||||||
)$n
|
|
||||||
if (identical(as.integer(n), 0L)) {
|
|
||||||
cli::cli_abort(c(
|
|
||||||
sprintf("%s requires a corpus with intergovernmental category rows.", what),
|
|
||||||
x = "The active corpus's `summary_categories` has no M/L (intergovernmental) rows.",
|
|
||||||
i = "This corpus predates the intergovernmental category rows added by cog_pipeline PR #59.",
|
|
||||||
i = "Point USCOGDATA_URL at a newer corpus that includes the M/L summary_categories rows."
|
|
||||||
), class = "uscogdata_ig_categories_unsupported")
|
|
||||||
}
|
|
||||||
invisible(TRUE)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.sql_lit_chr <- function(x) {
|
.sql_lit_chr <- function(x) {
|
||||||
safe <- gsub("'", "''", x, fixed = TRUE)
|
safe <- gsub("'", "''", x, fixed = TRUE)
|
||||||
@@ -556,8 +218,7 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
}
|
}
|
||||||
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.build_verb_sql <- function(view, subtype_col, govid, years, category,
|
.build_verb_sql <- function(view, subtype_col, govid, years, category) {
|
||||||
ig_view = NULL, subtype_scope = NULL) {
|
|
||||||
govid_lit <- .sql_lit_chr(govid)
|
govid_lit <- .sql_lit_chr(govid)
|
||||||
years_lit <- paste(as.integer(years), collapse = ",")
|
years_lit <- paste(as.integer(years), collapse = ",")
|
||||||
category_pred <- if (is.null(category)) {
|
category_pred <- if (is.null(category)) {
|
||||||
@@ -566,39 +227,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
sprintf("AND category IN (%s)", .sql_lit_chr(category))
|
sprintf("AND category IN (%s)", .sql_lit_chr(category))
|
||||||
}
|
}
|
||||||
|
|
||||||
# The concept's subtype allowlist (see .expenditure_concept_subtypes()).
|
|
||||||
# The base views carry every subtype of their flow (spending_annotated has
|
|
||||||
# all five non-IG expenditure subtypes); the concept narrows here. For
|
|
||||||
# "total", the IG leg's rows are 'intergovernmental', so that value joins
|
|
||||||
# the allowlist exactly when ig_view is present.
|
|
||||||
subtype_pred <- if (is.null(subtype_scope)) {
|
|
||||||
""
|
|
||||||
} else {
|
|
||||||
scope <- if (is.null(ig_view)) subtype_scope else c(subtype_scope, "intergovernmental")
|
|
||||||
sprintf("AND %s IN (%s)", subtype_col, .sql_lit_chr(scope))
|
|
||||||
}
|
|
||||||
|
|
||||||
# expenditure_concept = "total" adds the intergovernmental leg. UNION ALL,
|
|
||||||
# never UNION: the two legs are disjoint by crosswalk subtype (the direct
|
|
||||||
# view excludes 'intergovernmental'; the IG view is only that), so
|
|
||||||
# de-duplication would be pure cost, and a silent row-drop if two
|
|
||||||
# governments ever reported identical values.
|
|
||||||
source_expr <- if (is.null(ig_view)) {
|
|
||||||
view
|
|
||||||
} else {
|
|
||||||
sprintf("(SELECT * FROM %s UNION ALL SELECT * FROM %s)", view, ig_view)
|
|
||||||
}
|
|
||||||
|
|
||||||
# bool_or(), not bool_and(): a no-op for the Direct/revenue legs (those
|
|
||||||
# views filter NOT is_aggregate, so no row in any group is ever aggregate),
|
|
||||||
# but load-bearing for the IG leg, which deliberately keeps aggregate rows
|
|
||||||
# (see inst/sql/24-ig_long.sql). The wide era is dense -- every government
|
|
||||||
# has a row for every code in a family, most of them $0 -- so a $0 leaf
|
|
||||||
# commonly lands in the same (year, gov, subtype, category) group as the
|
|
||||||
# real aggregate row. bool_and() would then read FALSE for that group even
|
|
||||||
# though its dollars came entirely from an aggregate row, silently
|
|
||||||
# suppressing the "Aggregate fallback applied" note on exactly the rows
|
|
||||||
# this feature exists to surface.
|
|
||||||
sprintf(
|
sprintf(
|
||||||
"SELECT
|
"SELECT
|
||||||
year,
|
year,
|
||||||
@@ -608,15 +236,14 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
category,
|
category,
|
||||||
SUM(amt) * 1000.0 AS amt_nominal,
|
SUM(amt) * 1000.0 AS amt_nominal,
|
||||||
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
|
string_agg(DISTINCT item_code, ',' ORDER BY item_code) AS codes_included,
|
||||||
bool_or(is_aggregate) AS aggregate_fallback
|
bool_and(is_aggregate) AS aggregate_fallback
|
||||||
FROM %2$s
|
FROM %2$s
|
||||||
WHERE canonical_govid IN (%3$s)
|
WHERE canonical_govid IN (%3$s)
|
||||||
AND year IN (%4$s)
|
AND year IN (%4$s)
|
||||||
%5$s
|
%5$s
|
||||||
%6$s
|
|
||||||
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
|
GROUP BY year, canonical_govid, gov_name, xwalk_gov_name, %1$s, category
|
||||||
ORDER BY year, canonical_govid, %1$s, category",
|
ORDER BY year, canonical_govid, %1$s, category",
|
||||||
subtype_col, source_expr, govid_lit, years_lit, category_pred, subtype_pred
|
subtype_col, view, govid_lit, years_lit, category_pred
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -669,131 +296,11 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
result
|
result
|
||||||
}
|
}
|
||||||
|
|
||||||
#' Detect rows where expenditure_concept = "total" is reporting the
|
|
||||||
#' intergovernmental leg with NO Direct counterpart in the same (year,
|
|
||||||
#' canonical_govid, category) group AND a harmonization recipe actually
|
|
||||||
#' recovers the missing Direct dollars for that exact (year, canonical_govid,
|
|
||||||
#' category) triple.
|
|
||||||
#'
|
|
||||||
#' Bare Direct-row absence is deliberately NOT sufficient on its own: the
|
|
||||||
#' dominant real cause of "no Direct sibling row" is a government that simply
|
|
||||||
#' has no direct spending in that category (e.g. a state that funds K-12
|
|
||||||
#' entirely through school districts), which is correct, ordinary data, not
|
|
||||||
#' suppression. Genuine suppression -- a legacy aggregate-only family whose
|
|
||||||
#' Direct-leg basis query excludes it by construction (spending_long/
|
|
||||||
#' spending_long_harmonized both filter NOT is_aggregate) -- always has a
|
|
||||||
#' covering harmonization recipe, because that is exactly what the recipe
|
|
||||||
#' catalog exists to recover (see R/suggestions.R and `cog_recipes()`). So
|
|
||||||
#' checking "does a recipe actually cover this triple" cleanly separates the
|
|
||||||
#' two cases instead of conflating them.
|
|
||||||
#'
|
|
||||||
#' Returns `list(flag, notes)`, both the same length as `result`: `flag` is
|
|
||||||
#' `TRUE` only for the `spend_subtype == "intergovernmental"` row(s) in a
|
|
||||||
#' suppressed group, and `notes` names the recovering recipe(s) for those
|
|
||||||
#' rows (`NA` everywhere else).
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.detect_direct_suppressed <- function(con, result, subtype_col) {
|
.notes_column <- function(result) {
|
||||||
n <- nrow(result)
|
|
||||||
empty_notes <- rep(NA_character_, n)
|
|
||||||
if (n == 0L) return(list(flag = logical(0), notes = character(0)))
|
|
||||||
is_ig <- result[[subtype_col]] %in% "intergovernmental"
|
|
||||||
if (!any(is_ig)) return(list(flag = rep(FALSE, n), notes = empty_notes))
|
|
||||||
|
|
||||||
key <- paste(result$year, result$canonical_govid, result$category, sep = "\r")
|
|
||||||
has_direct <- key %in% unique(key[!is_ig])
|
|
||||||
candidate <- is_ig & !has_direct
|
|
||||||
|
|
||||||
flag <- rep(FALSE, n)
|
|
||||||
notes <- empty_notes
|
|
||||||
if (!any(candidate)) return(list(flag = flag, notes = notes))
|
|
||||||
|
|
||||||
idx <- which(candidate)
|
|
||||||
rows <- unique(result[idx, c("year", "canonical_govid", "category")])
|
|
||||||
covering <- .covering_recipes(con, rows)
|
|
||||||
cov_key <- paste(covering$year, covering$canonical_govid, covering$category,
|
|
||||||
sep = "\r")
|
|
||||||
|
|
||||||
for (i in idx) {
|
|
||||||
k <- paste(result$year[i], result$canonical_govid[i], result$category[i],
|
|
||||||
sep = "\r")
|
|
||||||
m <- match(k, cov_key)
|
|
||||||
if (is.na(m)) next
|
|
||||||
ids <- covering$recipe_ids[[m]]
|
|
||||||
if (length(ids) == 0L) next
|
|
||||||
flag[i] <- TRUE
|
|
||||||
notes[i] <- sprintf(
|
|
||||||
"Direct component is unavailable through this basis for this year; recover it via recipe = '%s' (see cog_recipes()).",
|
|
||||||
paste(sort(unique(ids)), collapse = "', '")
|
|
||||||
)
|
|
||||||
}
|
|
||||||
list(flag = flag, notes = notes)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' For each (year, canonical_govid, category) triple potentially affected by
|
|
||||||
#' a suppressed Direct leg, find the harmonization recipe(s) that (a) cover
|
|
||||||
#' this `category` (share a component item_code via `summary_categories`,
|
|
||||||
#' excluding any recipe that is itself entirely intergovernmental M/L -- the
|
|
||||||
#' same exclusion `.build_suggestions()` applies, see I2) and (b) actually
|
|
||||||
#' produce a `long` row for this exact (canonical_govid, year) via the same
|
|
||||||
#' generic join `.run_recipe()` uses (component year_min/year_max +
|
|
||||||
#' gov_type_scope, no is_aggregate filter -- a recipe's whole point is to
|
|
||||||
#' recover data that's aggregate-only). Adds a list-column `recipe_ids`
|
|
||||||
#' (possibly length-0) to `rows`.
|
|
||||||
#' @noRd
|
|
||||||
.covering_recipes <- function(con, rows) {
|
|
||||||
rows$recipe_ids <- vector("list", nrow(rows))
|
|
||||||
cats <- unique(rows$category[!is.na(rows$category)])
|
|
||||||
if (length(cats) == 0L) return(rows)
|
|
||||||
|
|
||||||
cand <- DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT DISTINCT sc.category, r.recipe_id
|
|
||||||
FROM harmonization_recipes r
|
|
||||||
JOIN summary_categories sc ON sc.item_code = r.component_code
|
|
||||||
WHERE sc.category IN (%s)
|
|
||||||
AND r.recipe_id NOT IN (
|
|
||||||
SELECT DISTINCT recipe_id FROM harmonization_recipes
|
|
||||||
WHERE LEFT(component_code, 1) IN ('M', 'L')
|
|
||||||
)",
|
|
||||||
.sql_lit_chr(cats)
|
|
||||||
))
|
|
||||||
if (nrow(cand) == 0L) return(rows)
|
|
||||||
|
|
||||||
recipe_ids_all <- unique(cand$recipe_id)
|
|
||||||
govids <- unique(rows$canonical_govid)
|
|
||||||
years <- unique(rows$year)
|
|
||||||
covered <- DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT DISTINCT r.recipe_id, l.canonical_govid, l.year
|
|
||||||
FROM long l
|
|
||||||
JOIN harmonization_recipes r
|
|
||||||
ON l.item_code = r.component_code
|
|
||||||
AND l.year BETWEEN r.year_min AND r.year_max
|
|
||||||
AND (r.gov_type_scope = 'all'
|
|
||||||
OR (r.gov_type_scope = 'state' AND l.type = 0)
|
|
||||||
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
|
|
||||||
WHERE r.recipe_id IN (%s)
|
|
||||||
AND l.canonical_govid IN (%s)
|
|
||||||
AND l.year IN (%s)",
|
|
||||||
.sql_lit_chr(recipe_ids_all), .sql_lit_chr(govids), paste(years, collapse = ",")
|
|
||||||
))
|
|
||||||
|
|
||||||
for (i in seq_len(nrow(rows))) {
|
|
||||||
cat_i <- rows$category[i]
|
|
||||||
if (is.na(cat_i)) next
|
|
||||||
cat_recipe_ids <- cand$recipe_id[cand$category == cat_i]
|
|
||||||
if (length(cat_recipe_ids) == 0L) next
|
|
||||||
sub <- covered[covered$canonical_govid == rows$canonical_govid[i] &
|
|
||||||
covered$year == rows$year[i] &
|
|
||||||
covered$recipe_id %in% cat_recipe_ids, ]
|
|
||||||
rows$recipe_ids[[i]] <- sort(unique(sub$recipe_id))
|
|
||||||
}
|
|
||||||
rows
|
|
||||||
}
|
|
||||||
|
|
||||||
#' @noRd
|
|
||||||
.notes_column <- function(result, direct_suppressed_notes = NULL) {
|
|
||||||
n <- nrow(result)
|
n <- nrow(result)
|
||||||
if (n == 0L) return(character(0))
|
if (n == 0L) return(character(0))
|
||||||
parts <- vector("list", 3L)
|
parts <- vector("list", 2L)
|
||||||
agg <- result[["aggregate_fallback"]]
|
agg <- result[["aggregate_fallback"]]
|
||||||
parts[[1]] <- if (!is.null(agg)) {
|
parts[[1]] <- if (!is.null(agg)) {
|
||||||
ifelse(agg %in% TRUE,
|
ifelse(agg %in% TRUE,
|
||||||
@@ -810,11 +317,6 @@ cog_spending <- function(govid, years, category = NULL,
|
|||||||
} else {
|
} else {
|
||||||
rep(NA_character_, n)
|
rep(NA_character_, n)
|
||||||
}
|
}
|
||||||
parts[[3]] <- if (!is.null(direct_suppressed_notes)) {
|
|
||||||
direct_suppressed_notes
|
|
||||||
} else {
|
|
||||||
rep(NA_character_, n)
|
|
||||||
}
|
|
||||||
out <- character(n)
|
out <- character(n)
|
||||||
for (i in seq_len(n)) {
|
for (i in seq_len(n)) {
|
||||||
pieces <- vapply(parts, `[[`, character(1), i)
|
pieces <- vapply(parts, `[[`, character(1), i)
|
||||||
|
|||||||
+153
-189
@@ -1,8 +1,9 @@
|
|||||||
# R/suggestions.R
|
# R/suggestions.R
|
||||||
# Recipe-component-driven signposting: when a basis = "harmonized" query for
|
# Recipe-component-driven signposting: when a basis = "harmonized" query for
|
||||||
# a category comes back with a coverage gap in some requested years (the
|
# a category asks for a code that is itself a harmonization recipe
|
||||||
# result has no rows at all in that year) that a harmonization recipe would
|
# component, and that specific code has no rows in some requested years
|
||||||
# actually fill for this government, surface that recipe as a suggestion.
|
# while the recipe's own generic join would still fill those years for this
|
||||||
|
# government, surface that recipe as a suggestion.
|
||||||
#
|
#
|
||||||
# This is deliberately keyed off the recipe catalog's component codes, not
|
# This is deliberately keyed off the recipe catalog's component codes, not
|
||||||
# off harmonization_map rows: no live map row carries a non-blank
|
# off harmonization_map rows: no live map row carries a non-blank
|
||||||
@@ -12,83 +13,58 @@
|
|||||||
# suggestion off of, just a leaf-code absence a recipe happens to fill).
|
# suggestion off of, just a leaf-code absence a recipe happens to fill).
|
||||||
# See docs/phase_r_harmonization_review.md § 0.3.
|
# See docs/phase_r_harmonization_review.md § 0.3.
|
||||||
#
|
#
|
||||||
# Scope is deliberately narrow: signposting only runs when the caller
|
# Scope is deliberately narrow in one respect and, as of Phase R3 Task 19c,
|
||||||
|
# deliberately WIDE in another: signposting only runs when the caller
|
||||||
# supplied a `category` (an un-scoped, all-categories query has no single
|
# supplied a `category` (an un-scoped, all-categories query has no single
|
||||||
# coverage question to answer) and only flags a recipe when the ACTUAL
|
# coverage question to answer), but within that category it now checks
|
||||||
# result has zero rows in a requested year AND the candidate recipe's own
|
# EACH recipe component that is itself a category member individually,
|
||||||
# generic join (same join .run_recipe() uses, including its wide-era
|
# rather than asking whether the whole category *result* has zero rows
|
||||||
# aggregate rows) produces at least one row for this government in that
|
# that year. A recipe fires when one of its own components has zero rows
|
||||||
# year. Checking presence per-government (not corpus-wide) avoids false
|
# for this government in a requested year, AND SOME OTHER component of that
|
||||||
# positives from ordinary reporting variance -- most governments don't use
|
# SAME recipe -- excluding the gapped one itself -- has a row (same join
|
||||||
# every sibling code in a multi-code category every year, and that is not
|
# .run_recipe() uses, aggregate rows included) for that year. This is the
|
||||||
# a format-boundary gap worth signposting.
|
# literal review-doc § 0.3 criterion: "...has no rows ... but other
|
||||||
#
|
# components do." A component's OWN aggregate-only row does not satisfy
|
||||||
# C1(a): for expenditure_concept = "total" callers, `result` here must
|
# its own gap (self-coverage is not "other components"); only a genuinely
|
||||||
# already be the Direct-leg subset (the caller filters out
|
# different sibling component can. This fires even if OTHER, unrelated
|
||||||
# spend_subtype == "intergovernmental" rows before calling in). A gap year
|
# codes in the same category have full data that year and the overall
|
||||||
# is "the requested year has no Direct rows", never "no rows at all" --
|
# result looks complete. That is a deliberate narrowing of the R2-era
|
||||||
# an IG row surviving on a legacy aggregate that Direct excludes must not
|
# false-positive guard: most governments don't use every sibling code in a
|
||||||
# read as coverage and cancel the very suggestion that would recover it.
|
# multi-code category every year, and per-code detection WILL flag some of
|
||||||
|
# that as a "gap" even though it's really just a government not having
|
||||||
|
# that particular sub-type of spending, not a format-boundary artifact.
|
||||||
|
# The remaining guard against ordinary reporting variance is the
|
||||||
|
# per-government, per-OTHER-component `covered` check below (a component
|
||||||
|
# is only flagged when a DIFFERENT component of the SAME recipe -- not
|
||||||
|
# some unrelated code, and not the gapped component's own aggregate row --
|
||||||
|
# actually has something to offer in that year); it no longer tries to
|
||||||
|
# avoid noise from sibling *codes*, only from a recipe with genuinely
|
||||||
|
# nothing else to contribute. The acceptable noise level this trade
|
||||||
|
# produces is a product decision, measured (not tuned here) by
|
||||||
|
# data-raw/measure_signposting_rate.R and ruled on at Checkpoint R3.
|
||||||
|
|
||||||
#' Build the `prov$suggestions` list for a (non-recipe) basis = "harmonized"
|
#' Recipe components that are classified under the requested category --
|
||||||
#' verb call: recipes whose generic join would fill a real gap in `result`.
|
#' the codes a category-scoped query actually "requests". A recipe can
|
||||||
#'
|
#' have components outside the category (e.g. general_gov_e89_wide's E85
|
||||||
#' @param con Active DuckDB connection.
|
#' leg has no category assignment); those never trigger on their own, they
|
||||||
#' @param govid Character vector of canonical_govid values (the verb's raw
|
#' just were never part of what this query asked for.
|
||||||
#' `govid`).
|
|
||||||
#' @param years Integer vector of requested years.
|
|
||||||
#' @param category `category` argument as passed to the verb (character
|
|
||||||
#' vector or `NULL`; suggestions are only computed when non-NULL).
|
|
||||||
#' @param result The verb's already-computed result tibble (post basis
|
|
||||||
#' query, pre per_capita/adjust_to_year), pre-filtered to the Direct leg
|
|
||||||
#' only when the caller's `expenditure_concept = "total"` (see C1(a)).
|
|
||||||
#' @param basis The *resolved* basis (`"harmonized"` or `"raw"`).
|
|
||||||
#' @param flow_prefixes The calling verb's own flow-type prefixes (e.g.
|
|
||||||
#' `c("E", "F", "G")` for `cog_spending()`, `c("T", "A", "U", "B", "C",
|
|
||||||
#' "D")` for `cog_revenue()` -- see `.verb_spendrev()`). Passed through to
|
|
||||||
#' `.attach_ig_counterparts()` to keep the intergovernmental-counterpart
|
|
||||||
#' lookup scoped to the calling verb's own flow family.
|
|
||||||
#' @return List of `list(recipe_id, label, available_years, hint,
|
|
||||||
#' ig_recipe_id)`, possibly empty.
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.build_suggestions <- function(con, govid, years, category, result, basis,
|
.category_recipe_components <- function(con, category) {
|
||||||
flow_prefixes) {
|
DBI::dbGetQuery(con, sprintf(
|
||||||
if (!identical(basis, "harmonized") || is.null(category)) return(list())
|
"SELECT DISTINCT r.recipe_id, r.component_code, r.year_min, r.year_max,
|
||||||
|
r.gov_type_scope
|
||||||
# Exclude any recipe that is ITSELF an intergovernmental (M/L) recipe --
|
FROM harmonization_recipes r
|
||||||
# i.e. every one of its own component codes is M/L-prefixed. Without this,
|
JOIN summary_categories sc
|
||||||
# a category whose summary_categories rows span both a Direct family
|
ON sc.item_code = r.component_code AND sc.category IN (%s)",
|
||||||
# (e.g. E04/E05, "Corrections") and its M/L counterpart (M04/M05, same
|
|
||||||
# category since Task 1) makes the M/L recipe itself (e.g.
|
|
||||||
# `corrections_ig_local_combined`) a raw top-level candidate for a plain
|
|
||||||
# (Direct) cog_spending() call -- following that hint would silently
|
|
||||||
# return intergovernmental dollars under `expenditure_concept = "direct"`
|
|
||||||
# provenance. This is a stronger, unconditional exclusion than the
|
|
||||||
# flow-prefix gate below/in `.attach_ig_counterparts()`: an M/L recipe
|
|
||||||
# should never be suggested as a coverage-gap filler for EITHER verb, not
|
|
||||||
# just kept from being named as the *counterpart* of another suggestion.
|
|
||||||
candidates <- DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT DISTINCT recipe_id FROM harmonization_recipes
|
|
||||||
WHERE component_code IN (
|
|
||||||
SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)
|
|
||||||
)
|
|
||||||
AND recipe_id NOT IN (
|
|
||||||
SELECT DISTINCT recipe_id FROM harmonization_recipes
|
|
||||||
WHERE LEFT(component_code, 1) IN ('M', 'L')
|
|
||||||
)",
|
|
||||||
.sql_lit_chr(category)
|
.sql_lit_chr(category)
|
||||||
))$recipe_id
|
))
|
||||||
if (length(candidates) == 0L) return(list())
|
}
|
||||||
|
|
||||||
result_years <- if (is.null(result) || nrow(result) == 0L) {
|
#' Label + overall year coverage for a set of recipe ids (the suggestion's
|
||||||
integer(0)
|
#' `label`/`available_years`).
|
||||||
} else {
|
#' @noRd
|
||||||
unique(as.integer(result$year))
|
.recipe_meta <- function(con, candidates) {
|
||||||
}
|
tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
|
||||||
gap_years <- setdiff(as.integer(years), result_years)
|
|
||||||
if (length(gap_years) == 0L) return(list())
|
|
||||||
|
|
||||||
meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf(
|
|
||||||
"SELECT recipe_id, any_value(label) AS label,
|
"SELECT recipe_id, any_value(label) AS label,
|
||||||
MIN(year_min) AS year_min, MAX(year_max) AS year_max
|
MIN(year_min) AS year_min, MAX(year_max) AS year_max
|
||||||
FROM harmonization_recipes
|
FROM harmonization_recipes
|
||||||
@@ -96,13 +72,48 @@
|
|||||||
GROUP BY recipe_id",
|
GROUP BY recipe_id",
|
||||||
.sql_lit_chr(candidates)
|
.sql_lit_chr(candidates)
|
||||||
)))
|
)))
|
||||||
|
}
|
||||||
|
|
||||||
# Which (recipe_id, year) pairs the recipe's own generic join actually
|
#' Which (recipe_id, component_code, year) triples have at least one
|
||||||
# covers for this government, restricted to the gap years -- the same
|
#' NOT-aggregate row for these governments -- i.e. that specific requested
|
||||||
# join .run_recipe() uses (component year_min/year_max + gov_type_scope,
|
#' code itself has data, scoped exactly like .run_recipe()'s join
|
||||||
# no is_aggregate filter), just checking existence instead of summing.
|
#' (component year_min/year_max + gov_type_scope). NOT-aggregate mirrors
|
||||||
covered <- DBI::dbGetQuery(con, sprintf(
|
#' what basis = "harmonized" itself excludes: an aggregate-only year is a
|
||||||
"SELECT DISTINCT r.recipe_id, l.year
|
#' gap for that code exactly as it would be in a plain category query.
|
||||||
|
#' @noRd
|
||||||
|
.component_presence <- function(con, candidates, govid, years_lit) {
|
||||||
|
DBI::dbGetQuery(con, sprintf(
|
||||||
|
"SELECT DISTINCT r.recipe_id, r.component_code, l.year
|
||||||
|
FROM long l
|
||||||
|
JOIN harmonization_recipes r
|
||||||
|
ON l.item_code = r.component_code
|
||||||
|
AND l.year BETWEEN r.year_min AND r.year_max
|
||||||
|
AND (r.gov_type_scope = 'all'
|
||||||
|
OR (r.gov_type_scope = 'state' AND l.type = 0)
|
||||||
|
OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3))
|
||||||
|
WHERE NOT l.is_aggregate
|
||||||
|
AND r.recipe_id IN (%s)
|
||||||
|
AND l.canonical_govid IN (%s)
|
||||||
|
AND l.year IN (%s)",
|
||||||
|
.sql_lit_chr(candidates), .sql_lit_chr(govid), years_lit
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Which (recipe_id, component_code, year) triples have at least one row
|
||||||
|
#' (aggregate rows included) for these governments -- the same scoping
|
||||||
|
#' .run_recipe()'s join uses (component year_min/year_max + gov_type_scope),
|
||||||
|
#' just checking existence instead of summing. Kept at per-component grain
|
||||||
|
#' (not unioned across the whole recipe, unlike the R2/R3-pre-fix version of
|
||||||
|
#' this function) so a gap check can require the covering evidence to come
|
||||||
|
#' from a DIFFERENT component -- review-doc § 0.3's "other components", not
|
||||||
|
#' the gapped component's own aggregate row. This is the per-government
|
||||||
|
#' guard against ordinary reporting variance: a recipe with genuinely
|
||||||
|
#' nothing to offer from any OTHER component (aggregate or leaf) never
|
||||||
|
#' fires.
|
||||||
|
#' @noRd
|
||||||
|
.recipe_coverage <- function(con, candidates, govid, years_lit) {
|
||||||
|
DBI::dbGetQuery(con, sprintf(
|
||||||
|
"SELECT DISTINCT r.recipe_id, r.component_code, l.year
|
||||||
FROM long l
|
FROM long l
|
||||||
JOIN harmonization_recipes r
|
JOIN harmonization_recipes r
|
||||||
ON l.item_code = r.component_code
|
ON l.item_code = r.component_code
|
||||||
@@ -113,13 +124,68 @@
|
|||||||
WHERE r.recipe_id IN (%s)
|
WHERE r.recipe_id IN (%s)
|
||||||
AND l.canonical_govid IN (%s)
|
AND l.canonical_govid IN (%s)
|
||||||
AND l.year IN (%s)",
|
AND l.year IN (%s)",
|
||||||
.sql_lit_chr(candidates), .sql_lit_chr(govid),
|
.sql_lit_chr(candidates), .sql_lit_chr(govid), years_lit
|
||||||
paste(gap_years, collapse = ",")
|
|
||||||
))
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
#' TRUE if recipe `rid` has at least one requested component with an
|
||||||
|
#' in-scope requested year that has no data (`present`), in a year some
|
||||||
|
#' OTHER component of the same recipe is otherwise fillable (`covered`,
|
||||||
|
#' excluding the component under test) -- the per-code gap the R2
|
||||||
|
#' whole-result check couldn't see, covered by another component the way
|
||||||
|
#' review-doc § 0.3 specifies (not by the gapped component's own aggregate
|
||||||
|
#' row -- that is self-coverage, not "other components", and must not
|
||||||
|
#' count).
|
||||||
|
#' @noRd
|
||||||
|
.recipe_component_gapped <- function(rid, requested, present, covered, years) {
|
||||||
|
comps <- requested[requested$recipe_id == rid, , drop = FALSE]
|
||||||
|
for (i in seq_len(nrow(comps))) {
|
||||||
|
this_code <- comps$component_code[i]
|
||||||
|
in_scope <- years[years >= comps$year_min[i] & years <= comps$year_max[i]]
|
||||||
|
if (length(in_scope) == 0L) next
|
||||||
|
has_data <- present$year[
|
||||||
|
present$recipe_id == rid & present$component_code == this_code
|
||||||
|
]
|
||||||
|
gap_years <- setdiff(in_scope, has_data)
|
||||||
|
if (length(gap_years) == 0L) next
|
||||||
|
other_covered_years <- covered$year[
|
||||||
|
covered$recipe_id == rid & covered$component_code != this_code
|
||||||
|
]
|
||||||
|
if (any(gap_years %in% other_covered_years)) return(TRUE)
|
||||||
|
}
|
||||||
|
FALSE
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Build the `prov$suggestions` list for a (non-recipe) basis = "harmonized"
|
||||||
|
#' verb call: recipes whose generic join would fill a real per-code gap for
|
||||||
|
#' the requested category.
|
||||||
|
#'
|
||||||
|
#' @param con Active DuckDB connection.
|
||||||
|
#' @param govid Character vector of canonical_govid values (the verb's raw
|
||||||
|
#' `govid`).
|
||||||
|
#' @param years Integer vector of requested years.
|
||||||
|
#' @param category `category` argument as passed to the verb (character
|
||||||
|
#' vector or `NULL`; suggestions are only computed when non-NULL).
|
||||||
|
#' @param basis The *resolved* basis (`"harmonized"` or `"raw"`).
|
||||||
|
#' @return List of `list(recipe_id, label, available_years, hint)`, possibly
|
||||||
|
#' empty.
|
||||||
|
#' @noRd
|
||||||
|
.build_suggestions <- function(con, govid, years, category, basis) {
|
||||||
|
if (!identical(basis, "harmonized") || is.null(category)) return(list())
|
||||||
|
|
||||||
|
requested <- .category_recipe_components(con, category)
|
||||||
|
if (nrow(requested) == 0L) return(list())
|
||||||
|
candidates <- unique(requested$recipe_id)
|
||||||
|
years_int <- as.integer(years)
|
||||||
|
years_lit <- paste(years_int, collapse = ",")
|
||||||
|
|
||||||
|
meta <- .recipe_meta(con, candidates)
|
||||||
|
present <- .component_presence(con, candidates, govid, years_lit)
|
||||||
|
covered <- .recipe_coverage(con, candidates, govid, years_lit)
|
||||||
|
|
||||||
suggestions <- list()
|
suggestions <- list()
|
||||||
for (rid in candidates) {
|
for (rid in candidates) {
|
||||||
if (!rid %in% covered$recipe_id) next
|
if (!.recipe_component_gapped(rid, requested, present, covered, years_int)) next
|
||||||
m <- meta[meta$recipe_id == rid, ]
|
m <- meta[meta$recipe_id == rid, ]
|
||||||
suggestions[[length(suggestions) + 1L]] <- list(
|
suggestions[[length(suggestions) + 1L]] <- list(
|
||||||
recipe_id = rid,
|
recipe_id = rid,
|
||||||
@@ -128,121 +194,19 @@
|
|||||||
hint = sprintf("re-run with recipe = '%s'", rid)
|
hint = sprintf("re-run with recipe = '%s'", rid)
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
.attach_ig_counterparts(con, suggestions, flow_prefixes)
|
suggestions
|
||||||
}
|
|
||||||
|
|
||||||
#' Attach `ig_recipe_id` to each suggestion: the intergovernmental-expenditure
|
|
||||||
#' recipe (an M-to-local or L-to-state recipe) whose component codes cover
|
|
||||||
#' exactly the same set of function suffixes as the firing recipe's own
|
|
||||||
#' components, e.g. `corrections_combined`'s {E04, E05} -> suffixes {"04",
|
|
||||||
#' "05"} matches `corrections_ig_local_combined`'s {M04, M05} -> the same
|
|
||||||
#' {"04", "05"}. `NULL` when no such recipe exists, which also covers the
|
|
||||||
#' case where the firing recipe already IS the IG recipe (self-matches are
|
|
||||||
#' excluded, so an IG recipe never names itself as its own counterpart).
|
|
||||||
#'
|
|
||||||
#' Matching is deliberately an exact set match, not "any suffix in common":
|
|
||||||
#' the two-digit suffix only means the same "function" across recipes that
|
|
||||||
#' share the underlying Census functional-classification scheme (E/F/G/L/M
|
|
||||||
#' all use "04"/"05" for corrections). M/L "combined other" codes (47/89/
|
|
||||||
#' 91-94) reuse digits for an unrelated catch-all construct, so e.g.
|
|
||||||
#' `general_gov_e89_wide`'s {E85, E89} -> {"85", "89"} must NOT match
|
|
||||||
#' `ige_local_m89_wide`'s {"89", "91", "92", "93"} on the shared "89" alone.
|
|
||||||
#' Checked by hand against the full harmonization_recipes catalog: only the
|
|
||||||
#' corrections family (E/F/G/M, suffixes 04/05) has an exact-set match in
|
|
||||||
#' this corpus.
|
|
||||||
#'
|
|
||||||
#' Exact-set suffix matching is NOT enough on its own, though: the same
|
|
||||||
#' reused-digit problem exists ACROSS the revenue-side IG families too.
|
|
||||||
#' `ig_local_d47_wide` (D47/D94, suffixes {"47","94"}) is an exact-set match
|
|
||||||
#' for `ige_local_m47_wide` (M47/M94, same suffixes) even though one is
|
|
||||||
#' intergovernmental REVENUE received from local governments and the other is
|
|
||||||
#' intergovernmental EXPENDITURE paid to local governments -- unrelated flows
|
|
||||||
#' that happen to reuse "47"/"94" for their own "transit/utilities" and
|
|
||||||
#' "other/combined" catch-alls. `ig_federal_b47_wide`, `ig_state_c47_wide`,
|
|
||||||
#' and their `*_89` siblings all collide the same way. None of this is
|
|
||||||
#' reachable via `cog_revenue()` in the bundled fixture today (its B/C/D
|
|
||||||
#' recipes never happen to have a covered gap year for any fixture govid),
|
|
||||||
#' but it IS reachable via a mis-scoped `cog_spending()` call on a
|
|
||||||
#' revenue-only category, e.g. `cog_spending(gov, category = "IG Federal")`
|
|
||||||
#' fires `ig_federal_b47_wide`/`ig_federal_b89_wide` for real in the fixture
|
|
||||||
#' -- so this is a live, not merely theoretical, gap.
|
|
||||||
#'
|
|
||||||
#' Two flow-family checks close this, both required (see
|
|
||||||
#' `tests/testthat/test-expenditure-concept.R`, "revenue-flavored ... never
|
|
||||||
#' receives an M/L counterpart" tests, for the pairwise verification):
|
|
||||||
#' 1. `own_prefix %in% flow_prefixes`: the firing recipe's own component
|
|
||||||
#' codes must belong to the calling verb's own flow family (the same
|
|
||||||
#' `flow_prefixes` `.build_harmonization_block()` uses, see
|
|
||||||
#' `R/basis.R`). This blocks a recipe surfaced through a mis-scoped
|
|
||||||
#' category from ever reaching the M/L search, e.g. `cog_spending()`'s
|
|
||||||
#' flow_prefixes are `c("E","F","G")`, which `ig_federal_b47_wide`'s own
|
|
||||||
#' `"B"` is not part of.
|
|
||||||
#' 2. `own_prefix %in% c("E","F","G")`: M/L only ever pairs with the
|
|
||||||
#' DIRECT-expenditure family, never with revenue (`cog_revenue()`'s
|
|
||||||
#' flow_prefixes already fold B/C/D in as ordinary revenue -- there is
|
|
||||||
#' no separate "Total" bolt-on for revenue the way `expenditure_concept`
|
|
||||||
#' adds one for spending) and never with ANOTHER M/L recipe (without
|
|
||||||
#' this check, `ige_local_m47_wide` would wrongly match sibling
|
|
||||||
#' `ige_state_l47_wide` on their shared {"47","94"} suffix set).
|
|
||||||
#' Condition 1 alone does not catch this: under `cog_revenue()`,
|
|
||||||
#' `ig_federal_b47_wide`'s own `"B"` IS inside revenue's own
|
|
||||||
#' `flow_prefixes`, so only this second, family-specific check blocks
|
|
||||||
#' the search.
|
|
||||||
#' @noRd
|
|
||||||
.attach_ig_counterparts <- function(con, suggestions, flow_prefixes) {
|
|
||||||
if (length(suggestions) == 0L) return(suggestions)
|
|
||||||
|
|
||||||
comp <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT recipe_id, component_code FROM harmonization_recipes")
|
|
||||||
comp$prefix <- substr(comp$component_code, 1L, 1L)
|
|
||||||
comp$suffix <- substr(comp$component_code, 2L, nchar(comp$component_code))
|
|
||||||
suffix_sets <- lapply(split(comp$suffix, comp$recipe_id), function(x) sort(unique(x)))
|
|
||||||
prefix_sets <- lapply(split(comp$prefix, comp$recipe_id), function(x) sort(unique(x)))
|
|
||||||
|
|
||||||
ig_recipe_ids <- unique(comp$recipe_id[comp$prefix %in% c("M", "L")])
|
|
||||||
|
|
||||||
find_counterpart <- function(rid) {
|
|
||||||
own_prefix <- prefix_sets[[rid]]
|
|
||||||
own_suffix <- suffix_sets[[rid]]
|
|
||||||
if (is.null(own_prefix) || is.null(own_suffix)) return(NULL)
|
|
||||||
if (!all(own_prefix %in% flow_prefixes)) return(NULL)
|
|
||||||
if (!all(own_prefix %in% c("E", "F", "G"))) return(NULL)
|
|
||||||
for (cand in ig_recipe_ids) {
|
|
||||||
if (identical(cand, rid)) next
|
|
||||||
if (setequal(suffix_sets[[cand]], own_suffix)) return(cand)
|
|
||||||
}
|
|
||||||
NULL
|
|
||||||
}
|
|
||||||
|
|
||||||
lapply(suggestions, function(s) {
|
|
||||||
# `s$ig_recipe_id <- NULL` would DELETE the element rather than set it
|
|
||||||
# (standard R list-assignment gotcha), leaving no-match entries missing
|
|
||||||
# the key entirely instead of carrying it as NULL. Single-bracket
|
|
||||||
# assignment with a wrapped list preserves a NULL-valued element so the
|
|
||||||
# field is always present, per the brief's "NULL when there is none".
|
|
||||||
s["ig_recipe_id"] <- list(find_counterpart(s$recipe_id))
|
|
||||||
s
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
#' Emit the single cli::cli_inform() message summarizing all suggestions
|
#' Emit the single cli::cli_inform() message summarizing all suggestions
|
||||||
#' for a verb call (the brief's "one message", not one per suggestion).
|
#' for a verb call (the brief's "one message", not one per suggestion).
|
||||||
#' Bullet text is pre-formatted plain text (no cli/glue `{}` markup) since
|
#' Bullet text is pre-formatted plain text (no cli/glue `{}` markup) since
|
||||||
#' recipe ids/labels are untrusted-ish data values, not literal call-site
|
#' recipe ids/labels are untrusted-ish data values, not literal call-site
|
||||||
#' expressions. When a suggestion has an `ig_recipe_id`, one indented
|
#' expressions.
|
||||||
#' continuation line is appended naming the intergovernmental counterpart
|
|
||||||
#' recipe (embedded `\n` renders as a hanging-indent continuation of the
|
|
||||||
#' same bullet under cli, not a new bullet).
|
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.inform_suggestions <- function(suggestions) {
|
.inform_suggestions <- function(suggestions) {
|
||||||
bullets <- vapply(suggestions, function(s) {
|
bullets <- vapply(suggestions, function(s) {
|
||||||
bullet <- sprintf("%s (%d-%d): %s", s$recipe_id,
|
sprintf("%s (%d-%d): %s", s$recipe_id,
|
||||||
s$available_years[1], s$available_years[2], s$hint)
|
s$available_years[1], s$available_years[2], s$hint)
|
||||||
if (!is.null(s$ig_recipe_id)) {
|
|
||||||
bullet <- paste0(bullet, sprintf(
|
|
||||||
"\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id))
|
|
||||||
}
|
|
||||||
bullet
|
|
||||||
}, character(1))
|
}, character(1))
|
||||||
cli::cli_inform(c(
|
cli::cli_inform(c(
|
||||||
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
|
i = "Coverage gap detected for the requested years; a harmonization recipe may fill it:",
|
||||||
|
|||||||
@@ -1,83 +1,25 @@
|
|||||||
# R/views.R
|
# R/views.R
|
||||||
|
|
||||||
# SQL files that cannot be registered unconditionally against a v4 corpus,
|
# SQL files whose view definitions read schema-v5-only parquet tables
|
||||||
# for one of two distinct reasons -- both fail at CREATE VIEW time (DuckDB
|
|
||||||
# resolves a view's source schema eagerly, even though it defers execution),
|
|
||||||
# so a v4 corpus can't tolerate either unconditionally:
|
|
||||||
#
|
|
||||||
# (a) Missing FILE. 33-/34-/35- read_parquet() a v5-only parquet table
|
|
||||||
# (harmonization_map.parquet, harmonization_recipes.parquet,
|
# (harmonization_map.parquet, harmonization_recipes.parquet,
|
||||||
# series_breaks.parquet) that doesn't exist at all on a v4 corpus --
|
# series_breaks.parquet) or select from views built on top of them. DuckDB's
|
||||||
# "IO Error: No files found".
|
# read_parquet() resolves the file at CREATE VIEW time (even for a view, it
|
||||||
#
|
# still needs the source schema) and errors immediately -- "IO Error: No
|
||||||
# (b) Missing COLUMN. 22-/23-/25- reference `long.harmonized_code`, a
|
# files found" -- if the path doesn't exist, so these cannot be registered
|
||||||
# column that does not exist on a v4 corpus's `long` table (harmonized
|
# unconditionally against a v4 corpus the way the rest of inst/sql/ is.
|
||||||
# space was introduced in schema v5) -- "Binder Error: Referenced
|
# Registration is therefore gated on manifest$schema_version >= 5; verb-level
|
||||||
# column harmonized_code not found". 42-/43-/45- are on this list only
|
# *usage* of the resulting views is separately gated by .resolve_basis() /
|
||||||
# because they SELECT s.* FROM the (a)/(b) views above, so they'd fail
|
# .require_schema_v5().
|
||||||
# to resolve their own source view if it weren't already skipped.
|
|
||||||
#
|
|
||||||
# Registration is therefore gated on manifest$schema_version >= 5 for all of
|
|
||||||
# them; verb-level *usage* of the resulting views is separately gated by
|
|
||||||
# .resolve_basis() / .require_schema_v5().
|
|
||||||
.harmonization_view_files <- c(
|
.harmonization_view_files <- c(
|
||||||
"22-spending_long_harmonized.sql",
|
"22-spending_long_harmonized.sql",
|
||||||
"23-revenue_long_harmonized.sql",
|
"23-revenue_long_harmonized.sql",
|
||||||
"25-ig_long_harmonized.sql",
|
|
||||||
"33-harmonization_map.sql",
|
"33-harmonization_map.sql",
|
||||||
"34-harmonization_recipes.sql",
|
"34-harmonization_recipes.sql",
|
||||||
"35-series_breaks_pq.sql",
|
"35-series_breaks_pq.sql",
|
||||||
"42-spending_annotated_harmonized.sql",
|
"42-spending_annotated_harmonized.sql",
|
||||||
"43-revenue_annotated_harmonized.sql",
|
"43-revenue_annotated_harmonized.sql"
|
||||||
"45-ig_annotated_harmonized.sql"
|
|
||||||
)
|
)
|
||||||
|
|
||||||
# The representation contract (cog_pipeline#64): two parquet tables that say
|
|
||||||
# what an ABSENT cell means in a given year. Gated on manifest PRESENCE, not
|
|
||||||
# on schema_version, because the sparsification that introduced them did not
|
|
||||||
# bump the version -- the pre-sparsification corpus this package shipped
|
|
||||||
# against until 2026-07-30 was already schema v6 and carried neither table.
|
|
||||||
# Keying off the version number would therefore register a view over a file
|
|
||||||
# that does not exist and fail at CREATE VIEW time on exactly the corpora this
|
|
||||||
# check exists to tolerate.
|
|
||||||
.representation_view_files <- c(
|
|
||||||
"36-representation.sql" = "representation.parquet",
|
|
||||||
"37-code_set.sql" = "code_set.parquet"
|
|
||||||
)
|
|
||||||
|
|
||||||
# Cash and security holdings (uscogdata#25). 46- selects
|
|
||||||
# `c.balance_subtype`, a column that arrived with cog_pipeline #76/#77 and
|
|
||||||
# WITHOUT a schema_version bump -- so neither existing gate applies:
|
|
||||||
# .harmonization_view_files keys on schema_version, .representation_view_files
|
|
||||||
# on the presence of a FILE. Here the discriminator is a COLUMN on a table
|
|
||||||
# that exists either way. CREATE VIEW resolves its source schema eagerly, so
|
|
||||||
# on an older corpus 46- would fail at registration with "Binder Error:
|
|
||||||
# Referenced column balance_subtype not found" rather than at query time.
|
|
||||||
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
|
|
||||||
|
|
||||||
#' Does the mounted corpus's `summary_categories` carry `balance_subtype`?
|
|
||||||
#' Probed against the live connection rather than the manifest, because the
|
|
||||||
#' manifest describes files, not columns.
|
|
||||||
#' @noRd
|
|
||||||
.corpus_has_balance_subtype <- function(con) {
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n FROM information_schema.columns
|
|
||||||
WHERE table_name = 'summary_categories'
|
|
||||||
AND column_name = 'balance_subtype'"
|
|
||||||
)$n
|
|
||||||
isTRUE(as.integer(n) > 0L)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Does the mounted corpus publish `file` (e.g. "code_set.parquet")?
|
|
||||||
#' Reads the manifest's metadata list rather than stat-ing the URL, so it
|
|
||||||
#' works identically for a local fixture and a remote share.
|
|
||||||
#' @noRd
|
|
||||||
.corpus_has_table <- function(manifest, file) {
|
|
||||||
paths <- vapply(manifest$files$metadata %||% list(),
|
|
||||||
function(f) as.character(f$path %||% ""), character(1))
|
|
||||||
file %in% basename(paths)
|
|
||||||
}
|
|
||||||
|
|
||||||
#' Register DuckDB views from inst/sql/ SQL files
|
#' Register DuckDB views from inst/sql/ SQL files
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.register_views <- function(con, url, manifest) {
|
.register_views <- function(con, url, manifest) {
|
||||||
@@ -85,11 +27,7 @@
|
|||||||
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
files <- sort(list.files(sql_dir, pattern = "\\.sql$", full.names = TRUE))
|
||||||
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
schema_version <- suppressWarnings(as.integer(manifest$schema_version %||% 0L))
|
||||||
for (f in files) {
|
for (f in files) {
|
||||||
base <- basename(f)
|
if (basename(f) %in% .harmonization_view_files && schema_version < 5L) next
|
||||||
if (base %in% .harmonization_view_files && schema_version < 5L) next
|
|
||||||
if (base %in% names(.representation_view_files) &&
|
|
||||||
!.corpus_has_table(manifest, .representation_view_files[[base]])) next
|
|
||||||
if (base %in% .balance_view_files && !.corpus_has_balance_subtype(con)) next
|
|
||||||
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
sql <- paste(readLines(f, warn = FALSE), collapse = "\n")
|
||||||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||||||
DBI::dbExecute(con, sql)
|
DBI::dbExecute(con, sql)
|
||||||
|
|||||||
@@ -19,99 +19,36 @@ package implements.
|
|||||||
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
# pak::pkg_install("gitea.civilytics.org/Civilytics/uscogdata")
|
||||||
```
|
```
|
||||||
|
|
||||||
## Amounts are in full US dollars
|
|
||||||
|
|
||||||
Every amount column this package returns — `amt_nominal`, `amt_real`,
|
|
||||||
`amt_per_capita_nominal`, `amt_per_capita_real` — is in **full US dollars**.
|
|
||||||
|
|
||||||
The raw Census source files report **thousands of dollars**, and the corpus's
|
|
||||||
own `amt` column preserves that. The verbs multiply by 1000 on the way out, so
|
|
||||||
you never have to. The conversion is recorded in every result:
|
|
||||||
|
|
||||||
```r
|
|
||||||
r <- cog_spending("552025209777", 2020L)
|
|
||||||
attr(r, "provenance")$transformations$units_conversion
|
|
||||||
#> $applied TRUE $source_unit "$1,000s (raw Census)" $target_unit "$USD" $multiplier 1000
|
|
||||||
```
|
|
||||||
|
|
||||||
**Do not multiply again.** If you have read elsewhere that COG amounts are in
|
|
||||||
`$1,000s` — true of the raw corpus, and of `cog_explorer`'s conventions doc —
|
|
||||||
that rule does not apply to anything a `cog_*()` verb hands you. Applying it
|
|
||||||
twice overstates every figure by 1000x, and the result looks plausible rather
|
|
||||||
than obviously wrong.
|
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
- `USCOGDATA_URL` — corpus root URL (public Nextcloud share, trailing slash)
|
||||||
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
- `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory
|
||||||
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
- `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600)
|
||||||
|
|
||||||
## Primary vs Direct vs Total spending
|
## Raw-parquet caveat: `survey_weight` is not an aggregation weight
|
||||||
|
|
||||||
`cog_spending(..., expenditure_concept = c("primary", "direct", "total"))`
|
Users reading the corpus parquet directly (DuckDB, arrow) will see a
|
||||||
controls whose spending a result counts. Concepts are defined as sets of the
|
`survey_weight` column (schema v6, col 28). It is legacy Census IndFin
|
||||||
crosswalk's `spend_subtype` values — never item-code first letters, which
|
sample-design **metadata passed through verbatim** — the Census Bureau's own
|
||||||
cannot classify correctly (the letter `Y` alone spans revenue, expenditure,
|
source documentation says it "is for informational purposes only and should
|
||||||
and balance codes):
|
not be used to derive any other statistics" (`_ReadMe_First_IndFin.txt`;
|
||||||
|
likewise `UserGuide.xls` Data User Note 8: "Do not use the weight field to
|
||||||
- `"primary"` (the default) is the government's own service provision:
|
derive state or national totals"). The raw encoding is also inconsistent
|
||||||
current operations, capital outlay, and assistance payments.
|
across vintages (reciprocal scale most years, direct scale in 2003, a `1`
|
||||||
- `"direct"` is Census's published Direct Expenditure: `primary` plus
|
placeholder in 1967/70/71/73/2001, all-`0` in 2007–2012, `NA` for all
|
||||||
interest on debt and insurance trust benefit payments (e.g. pensions).
|
modern-source rows), so `sum(amt * survey_weight/10000)`-style expressions
|
||||||
- `"total"` additionally adds the intergovernmental leg — money handed to
|
produce silently wrong totals — including exact zeros for 2007–2012. Sum
|
||||||
other governments to spend (`M`/`L` codes plus `Q11`/`Q12`/`Q18` state
|
`amt` unweighted; no uscogdata function reads this column. Full evidence:
|
||||||
payments to school systems) — which is meaningful for describing one
|
`cog_pipeline/.superpowers/sdd/weight-semantics-findings.md`.
|
||||||
government's own budget over time, but double-counts when summed across
|
|
||||||
governments (a state's payment to a county is the same dollar the county
|
|
||||||
reports as its own direct spending).
|
|
||||||
|
|
||||||
**Rule of thumb: any figure that spans more than one government uses
|
|
||||||
`primary` or `direct`.** `cog_geographic_rollup()` and `cog_peer_compare()`
|
|
||||||
enforce this by refusing `expenditure_concept = "total"`. See
|
|
||||||
`vignette("total-spending", package = "uscogdata")` for the full
|
|
||||||
explanation with worked examples.
|
|
||||||
|
|
||||||
## General vs Total revenue
|
|
||||||
|
|
||||||
`cog_revenue(..., revenue_concept = c("general", "total"))` selects between
|
|
||||||
Census's two published revenue concepts, again defined as crosswalk
|
|
||||||
`revenue_subtype` sets rather than item-code prefixes:
|
|
||||||
|
|
||||||
- `"general"` (the default) is Census **General Revenue**: own-source
|
|
||||||
(taxes, charges, miscellaneous) plus federal, state and local
|
|
||||||
intergovernmental aid.
|
|
||||||
- `"total"` is Census **Total Revenue**: `general` plus utility revenue
|
|
||||||
(`A91`–`A94`), liquor store revenue (`A90`), and insurance trust revenue
|
|
||||||
(unemployment and workers' compensation `Y` codes plus the
|
|
||||||
employee-retirement `X` codes).
|
|
||||||
|
|
||||||
The manual defines the first by subtracting the other three from the second,
|
|
||||||
so the two are related by Census's own identity:
|
|
||||||
|
|
||||||
```
|
|
||||||
Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
|
||||||
```
|
|
||||||
|
|
||||||
Two things worth knowing before switching to `"total"`:
|
|
||||||
|
|
||||||
- **Utility revenue is large for cities.** Measured on the bundled fixture,
|
|
||||||
utility plus liquor store revenue is 15.9% of city (type 2) revenue, versus
|
|
||||||
1.2% for states and 1.7% for counties. `general` excludes it by definition.
|
|
||||||
- **The employee-retirement (`X`) codes stop at FY2016**, when those systems
|
|
||||||
moved out of the annual finance file into the separate Annual Survey of
|
|
||||||
Public Pensions. A `"total"` series therefore steps down at the
|
|
||||||
FY2016/FY2017 seam for reasons of collection scope, not revenue (series
|
|
||||||
breaks `SB197`–`SB202`, in the corpus's `series_breaks` table).
|
|
||||||
|
|
||||||
## Developer notes
|
## Developer notes
|
||||||
|
|
||||||
### Testing
|
### Testing
|
||||||
|
|
||||||
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
The package ships a bundled fixture corpus at `inst/extdata/fixture_corpus/` —
|
||||||
a 15 MB four-year slice (2011, 2012, 2019, 2020) of the full corpus covering
|
a 3.6 MB two-year slice (2019 + 2020) of the full corpus covering all 50
|
||||||
all 50 states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL`
|
states. `tests/testthat/setup.R` automatically points `USCOGDATA_URL` at this
|
||||||
at this fixture, so the full test suite runs offline with no network
|
fixture, so the full test suite runs offline with no network dependency:
|
||||||
dependency:
|
|
||||||
|
|
||||||
```r
|
```r
|
||||||
devtools::test() # uses bundled fixture, no credentials required
|
devtools::test() # uses bundled fixture, no credentials required
|
||||||
|
|||||||
@@ -3,12 +3,6 @@ template:
|
|||||||
bootstrap: 5
|
bootstrap: 5
|
||||||
|
|
||||||
reference:
|
reference:
|
||||||
- title: Financial data
|
|
||||||
desc: Spending, revenue and balance-sheet holdings for one or more governments.
|
|
||||||
contents:
|
|
||||||
- cog_spending
|
|
||||||
- cog_revenue
|
|
||||||
- cog_balances
|
|
||||||
- title: Search & basket
|
- title: Search & basket
|
||||||
desc: Resolve place names into canonical govids.
|
desc: Resolve place names into canonical govids.
|
||||||
contents:
|
contents:
|
||||||
|
|||||||
@@ -0,0 +1,563 @@
|
|||||||
|
# data-raw/measure_signposting_rate.R
|
||||||
|
#
|
||||||
|
# Phase R3 Task 19c: measures the harmonization-signposting suggestion rate
|
||||||
|
# under THREE `.build_suggestions()` implementations, over a realistic query
|
||||||
|
# battery:
|
||||||
|
# every summary_categories category
|
||||||
|
# x a 3-year pre/post-2012 span (the wide-aggregate -> modern-leaf
|
||||||
|
# format-boundary window; falls back to the widest span the corpus
|
||||||
|
# actually supports if it can't fill a full 3+3 design -- see
|
||||||
|
# .measure_year_span())
|
||||||
|
# x up to N_GOV sampled governments (seeded, deterministic)
|
||||||
|
#
|
||||||
|
# The three arms, oldest to newest:
|
||||||
|
# - "coarse" (git ref b0df1ec, the merged R2 tip): a year counts as
|
||||||
|
# gapped only when the WHOLE category result has zero rows that year.
|
||||||
|
# - "selfcov" (git ref da72bf3, Task 19c's first per-code pass, since
|
||||||
|
# amended after review): per-code, but a component's gap could be
|
||||||
|
# satisfied by ANY component of the recipe INCLUDING ITSELF -- so a
|
||||||
|
# code whose only representation in a year was its own wide-era
|
||||||
|
# aggregate row satisfied its own coverage check. Flagged in review as
|
||||||
|
# not matching review-doc S: 0.3's literal criterion ("... has no rows
|
||||||
|
# ... but OTHER components do") and fixed in the next commit.
|
||||||
|
# - "percode" (live code): per-code, requiring a genuinely DIFFERENT
|
||||||
|
# sibling component to supply the covering evidence -- the shipped,
|
||||||
|
# corrected implementation.
|
||||||
|
#
|
||||||
|
# READ THIS BEFORE QUOTING ANY DELTA FROM THIS SCRIPT
|
||||||
|
# -----------------------------------------------------
|
||||||
|
# The coarse and per-code checks are PARTLY DISJOINT, not nested. Per-code
|
||||||
|
# is NOT a strict widening of coarse: there are queries coarse fires on that
|
||||||
|
# per-code does not, so moving coarse -> percode both ADDS and REMOVES
|
||||||
|
# signposting. Every `*_delta_pp` figure this script reports -- overall and
|
||||||
|
# per category -- is therefore a NET of those two flows and can mask a
|
||||||
|
# coverage loss in either direction. A headline "+X pp" can sit on top of
|
||||||
|
# categories that lost coverage outright (a NEGATIVE corrected_delta_pp),
|
||||||
|
# and a category-level zero can be an add and a loss cancelling. Read
|
||||||
|
# `$subset_relation` (printed under "Subset relation" below) alongside any
|
||||||
|
# delta; that section is where the two flows are separated.
|
||||||
|
#
|
||||||
|
# The disjointness is structural, not a sampling artifact. Both arms pair a
|
||||||
|
# gap test with a coverage test, and it is the COVERAGE test that differs:
|
||||||
|
# - coarse: gap = the WHOLE category result has zero rows that year;
|
||||||
|
# covered = the recipe's generic join has ANY row that year
|
||||||
|
# (unioned across all components -- a component's own
|
||||||
|
# aggregate row counts).
|
||||||
|
# - percode: gap = one specific component has no non-aggregate row that
|
||||||
|
# year; covered = a DIFFERENT component of the SAME recipe has
|
||||||
|
# a row that year (self-coverage explicitly excluded, per
|
||||||
|
# review-doc S: 0.3's "...but OTHER components do").
|
||||||
|
# So when a whole category is empty in a year -- exactly coarse's trigger --
|
||||||
|
# and the only covering evidence is the gapped component's own wide-era
|
||||||
|
# aggregate row, coarse fires and per-code CANNOT: there is by construction
|
||||||
|
# no other component to supply the evidence. That case is already pinned as
|
||||||
|
# intended behaviour in tests/testthat/test-recipes.R ("per-code gap does
|
||||||
|
# NOT fire when a code's only coverage is its own aggregate row"). This
|
||||||
|
# script's job is to say how often it costs coverage, not to relitigate it.
|
||||||
|
#
|
||||||
|
# This script MEASURES the deltas; it does not decide whether the resulting
|
||||||
|
# signposting trade -- added "noise" in one direction, lost whole-category
|
||||||
|
# gap coverage in the other -- is acceptable. That is Jared's ruling at
|
||||||
|
# Checkpoint R3 (see
|
||||||
|
# cog_pipeline/.superpowers/sdd/phase-r-task-19c-brief.md). The selfcov arm
|
||||||
|
# exists purely to answer a narrower, mechanical question for that ruling:
|
||||||
|
# how much of the coarse -> percode delta was ever attributable to the
|
||||||
|
# self-coverage bug (selfcov -> percode), as opposed to genuine
|
||||||
|
# other-component coverage (coarse -> percode directly)?
|
||||||
|
#
|
||||||
|
# All three arms no longer coexist in R/suggestions.R (each superseded the
|
||||||
|
# last in place), so this script pulls each VERBATIM from git history and
|
||||||
|
# evaluates it in an isolated environment parented on the uscogdata
|
||||||
|
# namespace, so each still resolves the unchanged sibling helpers it
|
||||||
|
# depends on (.sql_lit_chr()) exactly as the live package did at that
|
||||||
|
# commit. This guarantees every non-live arm is the actual shipped code at
|
||||||
|
# that point, not a hand-reconstruction that could silently drift from what
|
||||||
|
# really shipped.
|
||||||
|
#
|
||||||
|
# Usage (from the uscogdata package root; a git checkout, not a tarball):
|
||||||
|
# Rscript data-raw/measure_signposting_rate.R
|
||||||
|
# USCOGDATA_URL=<staged-corpus-url> Rscript data-raw/measure_signposting_rate.R
|
||||||
|
#
|
||||||
|
# Or from R:
|
||||||
|
# source("data-raw/measure_signposting_rate.R")
|
||||||
|
# res <- measure_signposting_rate(corpus_url = "<url>")
|
||||||
|
# res$summary; res$by_category
|
||||||
|
|
||||||
|
#' Pull a historical `.build_suggestions()` (and whatever helpers it uses)
|
||||||
|
#' verbatim from git history and evaluate it in an isolated environment
|
||||||
|
#' parented on the uscogdata namespace, so it resolves unchanged sibling
|
||||||
|
#' helpers (`.sql_lit_chr()`) the same way the live package does.
|
||||||
|
#' @noRd
|
||||||
|
.measure_load_git_impl <- function(git_ref, git_path = "R/suggestions.R") {
|
||||||
|
old_src <- tryCatch(
|
||||||
|
system2("git", c("show", sprintf("%s:%s", git_ref, git_path)),
|
||||||
|
stdout = TRUE, stderr = TRUE),
|
||||||
|
error = function(e) NULL
|
||||||
|
)
|
||||||
|
status <- attr(old_src, "status")
|
||||||
|
if (is.null(old_src) || (!is.null(status) && status != 0L) ||
|
||||||
|
!any(grepl("^\\.build_suggestions", old_src))) {
|
||||||
|
stop(
|
||||||
|
"Could not retrieve the .build_suggestions() implementation from ",
|
||||||
|
"git ref '", git_ref, "' at '", git_path, "'. Run this script from ",
|
||||||
|
"inside the uscogdata git checkout (not a tarball/installed copy).",
|
||||||
|
call. = FALSE
|
||||||
|
)
|
||||||
|
}
|
||||||
|
env <- new.env(parent = asNamespace("uscogdata"))
|
||||||
|
# eval(parse()) here is safe: `old_src` is not external/untrusted input --
|
||||||
|
# it is this repo's OWN historical R/suggestions.R, fetched via `git show`
|
||||||
|
# from a fixed, hardcoded internal commit ref (overridable only by a
|
||||||
|
# caller who already has R-level code execution in this dev-only
|
||||||
|
# measurement script). No network or user-supplied data reaches this call.
|
||||||
|
eval(parse(text = old_src), envir = env)
|
||||||
|
stopifnot(is.function(env$.build_suggestions))
|
||||||
|
env
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Resolve the query battery's year span: a `pre_n`-year window immediately
|
||||||
|
#' before `boundary_year` unioned with a `post_n`-year window starting at
|
||||||
|
#' `boundary_year` (default 3+3 around 2012, the wide-aggregate ->
|
||||||
|
#' modern-leaf format boundary). Falls back to every distinct year the
|
||||||
|
#' corpus actually has in `long` when it can't fill that full design, and
|
||||||
|
#' says so explicitly in `$note` rather than silently padding or
|
||||||
|
#' fabricating years.
|
||||||
|
#' @noRd
|
||||||
|
.measure_year_span <- function(con, boundary_year = 2012L,
|
||||||
|
pre_n = 3L, post_n = 3L) {
|
||||||
|
available <- sort(as.integer(
|
||||||
|
DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long")$year
|
||||||
|
))
|
||||||
|
desired_pre <- (boundary_year - pre_n):(boundary_year - 1L)
|
||||||
|
desired_post <- boundary_year:(boundary_year + post_n - 1L)
|
||||||
|
actual_pre <- intersect(desired_pre, available)
|
||||||
|
actual_post <- intersect(desired_post, available)
|
||||||
|
full_design <- length(actual_pre) == pre_n && length(actual_post) == post_n
|
||||||
|
|
||||||
|
if (full_design) {
|
||||||
|
years <- sort(c(actual_pre, actual_post))
|
||||||
|
note <- sprintf(
|
||||||
|
"Full %d-year pre/%d-year post-%d design available -- using years: %s.",
|
||||||
|
pre_n, post_n, boundary_year, paste(years, collapse = ", ")
|
||||||
|
)
|
||||||
|
} else {
|
||||||
|
years <- available
|
||||||
|
note <- sprintf(paste(
|
||||||
|
"Corpus does NOT support a full %d-year pre/%d-year post-%d span",
|
||||||
|
"(desired pre-window %s -> only %s present; desired post-window %s",
|
||||||
|
"-> only %s present). Falling back to the WIDEST span this corpus",
|
||||||
|
"supports: all %d distinct year(s) actually in `long`: %s.",
|
||||||
|
"This is NOT a 3-year pre/post-%d design -- reported as measured,",
|
||||||
|
"not padded or fabricated."
|
||||||
|
),
|
||||||
|
pre_n, post_n, boundary_year,
|
||||||
|
paste(desired_pre, collapse = ","),
|
||||||
|
if (length(actual_pre)) paste(actual_pre, collapse = ",") else "none",
|
||||||
|
paste(desired_post, collapse = ","),
|
||||||
|
if (length(actual_post)) paste(actual_post, collapse = ",") else "none",
|
||||||
|
length(available), paste(available, collapse = ", "),
|
||||||
|
boundary_year)
|
||||||
|
}
|
||||||
|
list(years = years, full_design = full_design, note = note,
|
||||||
|
available = available)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Deterministically sample up to `n` distinct governments that actually
|
||||||
|
#' report *something* in the battery's year span (querying a government
|
||||||
|
#' with zero presence in every measured year isn't a realistic query).
|
||||||
|
#' @noRd
|
||||||
|
.measure_sample_govids <- function(con, years, n = 20L, seed = 19L) {
|
||||||
|
pool <- DBI::dbGetQuery(con, sprintf(
|
||||||
|
"SELECT DISTINCT canonical_govid FROM long WHERE year IN (%s)
|
||||||
|
ORDER BY canonical_govid",
|
||||||
|
paste(years, collapse = ",")
|
||||||
|
))$canonical_govid
|
||||||
|
if (length(pool) <= n) return(sort(pool))
|
||||||
|
set.seed(seed)
|
||||||
|
sort(sample(pool, n))
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Comma-join a suggestion list's recipe ids (stable order) for the detail
|
||||||
|
#' frame's audit columns; `""` when nothing fired.
|
||||||
|
#' @noRd
|
||||||
|
.measure_recipe_ids <- function(suggestions) {
|
||||||
|
if (length(suggestions) == 0L) return("")
|
||||||
|
paste(sort(vapply(suggestions, function(s) s$recipe_id, character(1))),
|
||||||
|
collapse = ",")
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Run one (category, government) query through the coarse (`coarse_env`),
|
||||||
|
#' self-coverage-allowed (`selfcov_env`), and live per-code
|
||||||
|
#' `.build_suggestions()` and return a one-row summary of what each fired.
|
||||||
|
#'
|
||||||
|
#' Also records the coarse arm's OWN trigger evidence -- `coarse_gap_years`,
|
||||||
|
#' the requested years in which the whole category result has zero rows --
|
||||||
|
#' so a coarse-fired/per-code-silent disagreement can be named down to
|
||||||
|
#' (category, government, year) instead of just counted. Per-code's gap
|
||||||
|
#' years are deliberately NOT re-derived here: that would mean
|
||||||
|
#' reimplementing `.recipe_component_gapped()`'s set arithmetic in the
|
||||||
|
#' measurement harness, where it could silently drift from the code under
|
||||||
|
#' measurement. Per-code rows are identified by the recipe ids they fired.
|
||||||
|
#' @noRd
|
||||||
|
.measure_one_query <- function(con, coarse_env, selfcov_env, category,
|
||||||
|
category_type, govid, years) {
|
||||||
|
view <- if (identical(category_type, "revenue")) {
|
||||||
|
"revenue_annotated_harmonized"
|
||||||
|
} else {
|
||||||
|
"spending_annotated_harmonized"
|
||||||
|
}
|
||||||
|
subtype_col <- if (identical(category_type, "revenue")) {
|
||||||
|
"revenue_subtype"
|
||||||
|
} else {
|
||||||
|
"spend_subtype"
|
||||||
|
}
|
||||||
|
sql <- .build_verb_sql(view, subtype_col, govid, years, category)
|
||||||
|
result <- tibble::as_tibble(DBI::dbGetQuery(con, sql))
|
||||||
|
|
||||||
|
coarse_sugg <- coarse_env$.build_suggestions(
|
||||||
|
con, govid, years, category, result, "harmonized"
|
||||||
|
)
|
||||||
|
selfcov_sugg <- selfcov_env$.build_suggestions(
|
||||||
|
con, govid, years, category, "harmonized"
|
||||||
|
)
|
||||||
|
percode_sugg <- .build_suggestions(con, govid, years, category, "harmonized")
|
||||||
|
|
||||||
|
result_years <- if (nrow(result) == 0L) integer(0) else unique(as.integer(result$year))
|
||||||
|
gap_years <- sort(setdiff(as.integer(years), result_years))
|
||||||
|
|
||||||
|
data.frame(
|
||||||
|
category = category,
|
||||||
|
category_type = category_type,
|
||||||
|
canonical_govid = govid,
|
||||||
|
n_result_rows = nrow(result),
|
||||||
|
coarse_gap_years = paste(gap_years, collapse = ","),
|
||||||
|
n_coarse = length(coarse_sugg),
|
||||||
|
n_selfcov = length(selfcov_sugg),
|
||||||
|
n_percode = length(percode_sugg),
|
||||||
|
fired_coarse = length(coarse_sugg) > 0L,
|
||||||
|
fired_selfcov = length(selfcov_sugg) > 0L,
|
||||||
|
fired_percode = length(percode_sugg) > 0L,
|
||||||
|
coarse_recipes = .measure_recipe_ids(coarse_sugg),
|
||||||
|
percode_recipes = .measure_recipe_ids(percode_sugg),
|
||||||
|
stringsAsFactors = FALSE
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Columns that identify a disagreeing query well enough for a human to go
|
||||||
|
#' and inspect it, in print order. Intersected with what `detail` actually
|
||||||
|
#' has, so this works on a minimal hand-built frame too.
|
||||||
|
#' @noRd
|
||||||
|
.MEASURE_IDENTITY_COLS <- c(
|
||||||
|
"category", "category_type", "canonical_govid", "coarse_gap_years",
|
||||||
|
"n_result_rows", "coarse_recipes", "percode_recipes"
|
||||||
|
)
|
||||||
|
|
||||||
|
#' Separate the two flows the `*_delta_pp` figures net together.
|
||||||
|
#'
|
||||||
|
#' Per-code is NOT a widening of coarse (see this file's header): the two
|
||||||
|
#' checks pair different gap tests with different coverage tests, so moving
|
||||||
|
#' coarse -> percode both adds and removes firings. This splits the
|
||||||
|
#' disagreement into:
|
||||||
|
#' - `violations`: coarse fired, per-code did NOT -- signposting coverage
|
||||||
|
#' LOST. These are what a net delta hides. `holds` is FALSE whenever
|
||||||
|
#' this is non-empty, i.e. whenever coarse is not a subset of per-code.
|
||||||
|
#' - `additions`: per-code fired, coarse did NOT -- the expected gain.
|
||||||
|
#'
|
||||||
|
#' Deliberately returns the offending rows, not just counts, so the
|
||||||
|
#' Checkpoint R3 ruling can be made against named (category, government,
|
||||||
|
#' year) cases. Deliberately does NOT assert -- the violation set is really
|
||||||
|
#' non-empty on the staged corpus, and a hard assertion here would only
|
||||||
|
#' break the harness that is supposed to report it.
|
||||||
|
#' @noRd
|
||||||
|
.measure_subset_relation <- function(detail) {
|
||||||
|
required <- c("category", "canonical_govid", "fired_coarse", "fired_percode")
|
||||||
|
absent <- if (is.data.frame(detail)) setdiff(required, names(detail)) else required
|
||||||
|
if (!is.data.frame(detail) || length(absent) > 0L) {
|
||||||
|
stop("`detail` must be a data frame with columns ",
|
||||||
|
paste(required, collapse = ", "), " (missing: ",
|
||||||
|
paste(absent, collapse = ", "), ").", call. = FALSE)
|
||||||
|
}
|
||||||
|
fired_coarse <- as.logical(detail$fired_coarse)
|
||||||
|
fired_percode <- as.logical(detail$fired_percode)
|
||||||
|
if (anyNA(fired_coarse) || anyNA(fired_percode)) {
|
||||||
|
stop("`fired_coarse`/`fired_percode` must be non-NA logicals.", call. = FALSE)
|
||||||
|
}
|
||||||
|
|
||||||
|
keep <- intersect(.MEASURE_IDENTITY_COLS, names(detail))
|
||||||
|
viol_idx <- which(fired_coarse & !fired_percode)
|
||||||
|
add_idx <- which(fired_percode & !fired_coarse)
|
||||||
|
|
||||||
|
list(
|
||||||
|
holds = length(viol_idx) == 0L,
|
||||||
|
n_queries = nrow(detail),
|
||||||
|
n_coarse_fired = sum(fired_coarse),
|
||||||
|
n_percode_fired = sum(fired_percode),
|
||||||
|
n_both = sum(fired_coarse & fired_percode),
|
||||||
|
n_violations = length(viol_idx),
|
||||||
|
n_additions = length(add_idx),
|
||||||
|
violations = detail[viol_idx, keep, drop = FALSE],
|
||||||
|
additions = detail[add_idx, keep, drop = FALSE]
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Render a data frame of disagreeing queries as indented report lines,
|
||||||
|
#' capped at `max_rows` with an explicit note about what was withheld (the
|
||||||
|
#' full set is always in the returned `$subset_relation`).
|
||||||
|
#' @noRd
|
||||||
|
.measure_fmt_rows <- function(df, max_rows = 50L) {
|
||||||
|
if (nrow(df) == 0L) return(" (none)")
|
||||||
|
shown <- utils::head(df, max_rows)
|
||||||
|
out <- paste0(" ", utils::capture.output(print(shown, row.names = FALSE)))
|
||||||
|
if (nrow(df) > max_rows) {
|
||||||
|
out <- c(out, sprintf(" ... %d more row(s) not shown; full set in $subset_relation.",
|
||||||
|
nrow(df) - max_rows))
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Format `.measure_subset_relation()` as a prominent, clearly-labelled
|
||||||
|
#' report section. States plainly whether coarse is a subset of per-code
|
||||||
|
#' and, when it is not, exactly where it breaks.
|
||||||
|
#' @noRd
|
||||||
|
.measure_format_subset_report <- function(rel, max_rows = 50L) {
|
||||||
|
lines <- c(
|
||||||
|
"==== Subset relation: is COARSE a subset of PER-CODE? ====",
|
||||||
|
sprintf("Queries: %d | coarse fired: %d | per-code fired: %d | both: %d",
|
||||||
|
rel$n_queries, rel$n_coarse_fired, rel$n_percode_fired, rel$n_both)
|
||||||
|
)
|
||||||
|
if (rel$holds) {
|
||||||
|
lines <- c(lines, sprintf(paste(
|
||||||
|
"HOLDS: coarse IS a subset of per-code -- 0 of %d queries fire under",
|
||||||
|
"coarse but not per-code. On THIS battery the delta is a pure",
|
||||||
|
"addition of %d query/queries, with no coverage lost."
|
||||||
|
), rel$n_queries, rel$n_additions))
|
||||||
|
} else {
|
||||||
|
lines <- c(lines,
|
||||||
|
"*** VIOLATED: coarse is NOT a subset of per-code. ***",
|
||||||
|
sprintf(paste(
|
||||||
|
"%d of %d queries fire under COARSE but NOT under PER-CODE:",
|
||||||
|
"signposting coverage the move LOSES."
|
||||||
|
), rel$n_violations, rel$n_queries),
|
||||||
|
sprintf(paste(
|
||||||
|
"Every delta reported above is therefore a NET of %d addition(s)",
|
||||||
|
"MINUS %d loss(es), and understates both. Do not read it as",
|
||||||
|
"'per-code fires wherever coarse did, plus more'."
|
||||||
|
), rel$n_additions, rel$n_violations),
|
||||||
|
"",
|
||||||
|
paste(" COVERAGE LOST -- coarse fired, per-code silent.",
|
||||||
|
"`coarse_gap_years` is the requested year(s) in which the whole",
|
||||||
|
"category result was empty (coarse's own trigger evidence):"),
|
||||||
|
.measure_fmt_rows(rel$violations, max_rows)
|
||||||
|
)
|
||||||
|
}
|
||||||
|
c(lines, "",
|
||||||
|
sprintf(" COVERAGE ADDED -- per-code fired, coarse silent (%d query/queries):",
|
||||||
|
rel$n_additions),
|
||||||
|
.measure_fmt_rows(rel$additions, max_rows))
|
||||||
|
}
|
||||||
|
|
||||||
|
#' Measure the coarse-vs-per-code signposting suggestion rate over a
|
||||||
|
#' realistic query battery (every category x a pre/post-boundary_year span
|
||||||
|
#' x up to n_gov sampled governments).
|
||||||
|
#'
|
||||||
|
#' @param corpus_url Corpus to measure against. Defaults to
|
||||||
|
#' `Sys.getenv("USCOGDATA_URL")`; if that's unset, falls back to the
|
||||||
|
#' bundled v5 fixture (so the script runs out of the box). Re-run with
|
||||||
|
#' `USCOGDATA_URL` pointed at the staged/full corpus later.
|
||||||
|
#' @param n_gov Governments to sample (deterministically). "Up to" -- if
|
||||||
|
#' the corpus has fewer distinct governments in the measured years than
|
||||||
|
#' this, every one of them is used.
|
||||||
|
#' @param seed Sampling seed (fixed for reproducibility).
|
||||||
|
#' @param boundary_year,pre_years_n,post_years_n Define the desired query
|
||||||
|
#' span: `pre_years_n` years immediately before `boundary_year`, unioned
|
||||||
|
#' with `post_years_n` years starting at `boundary_year`. Falls back to
|
||||||
|
#' the corpus's widest actually-available span when this can't be filled
|
||||||
|
#' (see `.measure_year_span()`).
|
||||||
|
#' @param coarse_ref Git ref to pull the R2 coarse `.build_suggestions()`
|
||||||
|
#' from.
|
||||||
|
#' @param selfcov_ref Git ref to pull Task 19c's first, self-coverage-
|
||||||
|
#' allowed per-code `.build_suggestions()` from (amended after review).
|
||||||
|
#' @param verbose Print progress/notes as the battery runs.
|
||||||
|
#' @return Invisibly, a list with `corpus_url`, `years`, `span_note`,
|
||||||
|
#' `full_design`, `govids`, `seed`, `detail` (one row per query),
|
||||||
|
#' `by_category`, and `summary`.
|
||||||
|
#' @noRd
|
||||||
|
measure_signposting_rate <- function(corpus_url = Sys.getenv("USCOGDATA_URL", unset = NA),
|
||||||
|
n_gov = 20L,
|
||||||
|
seed = 19L,
|
||||||
|
boundary_year = 2012L,
|
||||||
|
pre_years_n = 3L,
|
||||||
|
post_years_n = 3L,
|
||||||
|
coarse_ref = "b0df1ec",
|
||||||
|
selfcov_ref = "da72bf3",
|
||||||
|
verbose = TRUE) {
|
||||||
|
pkgload::load_all(".", quiet = TRUE)
|
||||||
|
|
||||||
|
if (is.na(corpus_url) || !nzchar(corpus_url)) {
|
||||||
|
corpus_url <- paste0(
|
||||||
|
system.file("extdata/fixture_corpus", package = "uscogdata"), "/"
|
||||||
|
)
|
||||||
|
if (verbose) {
|
||||||
|
message("No USCOGDATA_URL set; defaulting to the bundled v5 fixture: ",
|
||||||
|
corpus_url)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
||||||
|
cog_close()
|
||||||
|
Sys.setenv(USCOGDATA_URL = corpus_url)
|
||||||
|
on.exit({
|
||||||
|
cog_close()
|
||||||
|
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
||||||
|
}, add = TRUE)
|
||||||
|
|
||||||
|
con <- cog_open()
|
||||||
|
coarse_env <- .measure_load_git_impl(git_ref = coarse_ref)
|
||||||
|
selfcov_env <- .measure_load_git_impl(git_ref = selfcov_ref)
|
||||||
|
|
||||||
|
span <- .measure_year_span(con, boundary_year, pre_years_n, post_years_n)
|
||||||
|
years <- span$years
|
||||||
|
if (verbose) message(span$note)
|
||||||
|
|
||||||
|
govids <- .measure_sample_govids(con, years, n = n_gov, seed = seed)
|
||||||
|
if (verbose) {
|
||||||
|
message(sprintf("Sampled %d government(s) (seed = %d) from %d present in years %s.",
|
||||||
|
length(govids), seed,
|
||||||
|
length(DBI::dbGetQuery(con, sprintf(
|
||||||
|
"SELECT DISTINCT canonical_govid FROM long WHERE year IN (%s)",
|
||||||
|
paste(years, collapse = ",")))$canonical_govid),
|
||||||
|
paste(years, collapse = ", ")))
|
||||||
|
}
|
||||||
|
|
||||||
|
categories <- DBI::dbGetQuery(con,
|
||||||
|
"SELECT DISTINCT category, category_type FROM summary_categories
|
||||||
|
WHERE category IS NOT NULL ORDER BY category_type, category")
|
||||||
|
if (verbose) {
|
||||||
|
message(sprintf("Battery: %d categories x %d governments = %d queries.",
|
||||||
|
nrow(categories), length(govids),
|
||||||
|
nrow(categories) * length(govids)))
|
||||||
|
}
|
||||||
|
|
||||||
|
rows <- vector("list", nrow(categories) * length(govids))
|
||||||
|
k <- 0L
|
||||||
|
for (ci in seq_len(nrow(categories))) {
|
||||||
|
for (gv in govids) {
|
||||||
|
k <- k + 1L
|
||||||
|
rows[[k]] <- .measure_one_query(
|
||||||
|
con, coarse_env, selfcov_env,
|
||||||
|
category = categories$category[ci],
|
||||||
|
category_type = categories$category_type[ci],
|
||||||
|
govid = gv, years = years
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
detail <- do.call(rbind, rows)
|
||||||
|
# percode only ever fires where selfcov also fires (percode is a strict
|
||||||
|
# narrowing of selfcov: same gap detection, plus the self-coverage path
|
||||||
|
# removed) -- this is what makes the decomposition below exact rather
|
||||||
|
# than approximate. Checked, not assumed.
|
||||||
|
stopifnot(all(detail$fired_percode <= detail$fired_selfcov))
|
||||||
|
detail$fired_selfcov_only <- detail$fired_selfcov & !detail$fired_percode
|
||||||
|
|
||||||
|
# coarse vs percode is NOT a subset relation the way percode vs selfcov
|
||||||
|
# is (see header). Measured and REPORTED, never asserted: the violation
|
||||||
|
# set is genuinely non-empty on the staged corpus, and a stopifnot() here
|
||||||
|
# would break the harness whose whole job is to surface it.
|
||||||
|
subset_relation <- .measure_subset_relation(detail)
|
||||||
|
detail$coarse_only <- detail$fired_coarse & !detail$fired_percode
|
||||||
|
detail$percode_only <- detail$fired_percode & !detail$fired_coarse
|
||||||
|
|
||||||
|
by_category <- dplyr::summarise(
|
||||||
|
dplyr::group_by(detail, category, category_type),
|
||||||
|
n_queries = dplyr::n(),
|
||||||
|
coarse_rate = mean(fired_coarse),
|
||||||
|
selfcov_rate = mean(fired_selfcov),
|
||||||
|
percode_rate = mean(fired_percode),
|
||||||
|
original_delta_pp = (mean(fired_selfcov) - mean(fired_coarse)) * 100,
|
||||||
|
corrected_delta_pp = (mean(fired_percode) - mean(fired_coarse)) * 100,
|
||||||
|
selfcov_share_pp = mean(fired_selfcov_only) * 100,
|
||||||
|
# the two flows corrected_delta_pp nets together, per category
|
||||||
|
n_coarse_only = sum(coarse_only),
|
||||||
|
n_percode_only = sum(percode_only),
|
||||||
|
.groups = "drop"
|
||||||
|
)
|
||||||
|
by_category <- dplyr::arrange(by_category, dplyr::desc(corrected_delta_pp))
|
||||||
|
|
||||||
|
summary_overall <- data.frame(
|
||||||
|
n_queries = nrow(detail),
|
||||||
|
n_categories = nrow(categories),
|
||||||
|
n_governments = length(govids),
|
||||||
|
coarse_fired = sum(detail$fired_coarse),
|
||||||
|
selfcov_fired = sum(detail$fired_selfcov),
|
||||||
|
percode_fired = sum(detail$fired_percode),
|
||||||
|
coarse_rate = mean(detail$fired_coarse),
|
||||||
|
selfcov_rate = mean(detail$fired_selfcov),
|
||||||
|
percode_rate = mean(detail$fired_percode)
|
||||||
|
)
|
||||||
|
summary_overall$original_delta_pp <- (summary_overall$selfcov_rate - summary_overall$coarse_rate) * 100
|
||||||
|
summary_overall$corrected_delta_pp <- (summary_overall$percode_rate - summary_overall$coarse_rate) * 100
|
||||||
|
summary_overall$selfcov_share_pp <- mean(detail$fired_selfcov_only) * 100
|
||||||
|
summary_overall$relative_increase <- if (summary_overall$coarse_rate > 0) {
|
||||||
|
summary_overall$percode_rate / summary_overall$coarse_rate - 1
|
||||||
|
} else {
|
||||||
|
NA_real_
|
||||||
|
}
|
||||||
|
|
||||||
|
if (verbose) {
|
||||||
|
message(sprintf(
|
||||||
|
"Coarse rate: %.4f (%d/%d) | Self-cov-allowed rate: %.4f (%d/%d) | Corrected per-code rate: %.4f (%d/%d)",
|
||||||
|
summary_overall$coarse_rate, summary_overall$coarse_fired, summary_overall$n_queries,
|
||||||
|
summary_overall$selfcov_rate, summary_overall$selfcov_fired, summary_overall$n_queries,
|
||||||
|
summary_overall$percode_rate, summary_overall$percode_fired, summary_overall$n_queries
|
||||||
|
))
|
||||||
|
message(sprintf(
|
||||||
|
"Original delta (selfcov - coarse): %+.2f pp | Corrected delta (percode - coarse): %+.2f pp | Self-coverage share of original delta: %.2f pp (%d/%d queries fired ONLY via self-coverage)",
|
||||||
|
summary_overall$original_delta_pp, summary_overall$corrected_delta_pp,
|
||||||
|
summary_overall$selfcov_share_pp,
|
||||||
|
sum(detail$fired_selfcov_only), summary_overall$n_queries
|
||||||
|
))
|
||||||
|
message(if (subset_relation$holds) {
|
||||||
|
sprintf("Subset relation coarse <= percode HOLDS (0 coarse-only firings); delta is a pure addition of %d.",
|
||||||
|
subset_relation$n_additions)
|
||||||
|
} else {
|
||||||
|
sprintf("*** Subset relation coarse <= percode VIOLATED: %d coarse-only firing(s) LOST vs %d percode-only added. Deltas above are NETS. ***",
|
||||||
|
subset_relation$n_violations, subset_relation$n_additions)
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
invisible(list(
|
||||||
|
corpus_url = corpus_url,
|
||||||
|
years = years,
|
||||||
|
span_note = span$note,
|
||||||
|
full_design = span$full_design,
|
||||||
|
govids = govids,
|
||||||
|
seed = seed,
|
||||||
|
n_categories = nrow(categories),
|
||||||
|
detail = detail,
|
||||||
|
by_category = by_category,
|
||||||
|
summary = summary_overall,
|
||||||
|
subset_relation = subset_relation
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
if (identical(environment(), globalenv()) && sys.nframe() == 0L) {
|
||||||
|
res <- measure_signposting_rate()
|
||||||
|
cat("\n==== Query battery ====\n")
|
||||||
|
cat("Corpus:", res$corpus_url, "\n")
|
||||||
|
cat("Years:", paste(res$years, collapse = ", "), "\n")
|
||||||
|
cat("Full 3-year pre/post-2012 design achieved:", res$full_design, "\n")
|
||||||
|
cat(res$span_note, "\n")
|
||||||
|
cat("Governments sampled:", length(res$govids), "\n\n")
|
||||||
|
|
||||||
|
cat("==== Overall summary (coarse / self-coverage-allowed / corrected per-code) ====\n")
|
||||||
|
print(res$summary)
|
||||||
|
|
||||||
|
cat("\n==== By category (sorted by corrected delta, descending) ====\n")
|
||||||
|
cat("NOTE: corrected_delta_pp is a NET. n_coarse_only = firings LOST going\n")
|
||||||
|
cat("coarse -> percode; n_percode_only = firings ADDED. A category can be\n")
|
||||||
|
cat("negative (net coverage loss) even when the overall figure is positive.\n")
|
||||||
|
print(as.data.frame(res$by_category), row.names = FALSE)
|
||||||
|
|
||||||
|
cat("\n")
|
||||||
|
cat(paste(.measure_format_subset_report(res$subset_relation), collapse = "\n"), "\n")
|
||||||
|
}
|
||||||
@@ -11,14 +11,12 @@
|
|||||||
# Each partition is a full year (all states/govs) as published, so
|
# Each partition is a full year (all states/govs) as published, so
|
||||||
# Broward County FL and every other previously-pinned government stay
|
# Broward County FL and every other previously-pinned government stay
|
||||||
# covered without any per-gov slicing logic.
|
# covered without any per-gov slicing logic.
|
||||||
# 2. Copies every metadata parquet the publish tree ships (see
|
# 2. Copies the full canonical_fips_xwalk.parquet, canonical_alias.parquet,
|
||||||
# .FIXTURE_METADATA_FILES) as-is. These are small cross-vintage
|
# summary_categories.parquet, harmonization_map.parquet,
|
||||||
# registries, not partitioned by year, so the fixture ships the complete
|
# harmonization_recipes.parquet, and series_breaks.parquet metadata
|
||||||
# tables rather than a year-scoped subset. representation.parquet and
|
# tables as-is (these are small cross-vintage registries, not
|
||||||
# code_set.parquet are what make the sparse wide era interpretable --
|
# partitioned by year, so the fixture ships the complete tables rather
|
||||||
# absence means "$0" in a dense_source year and "not reported" in a
|
# than a year-scoped subset).
|
||||||
# sparse_source one -- so a fixture without them cannot represent the
|
|
||||||
# published corpus.
|
|
||||||
# 3. Resyncs the four reference docs (data_dictionary.md,
|
# 3. Resyncs the four reference docs (data_dictionary.md,
|
||||||
# reader-specification.md, README.md, series_breaks.md) from the
|
# reader-specification.md, README.md, series_breaks.md) from the
|
||||||
# publish tree's docs/.
|
# publish tree's docs/.
|
||||||
@@ -40,22 +38,6 @@
|
|||||||
# source("data-raw/regenerate_fixture_corpus.R")
|
# source("data-raw/regenerate_fixture_corpus.R")
|
||||||
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
|
# regenerate_fixture_corpus(publish_cache_dir = "/path/to/publish_cache")
|
||||||
|
|
||||||
# Every metadata parquet the publish tree ships, in the order they appear in
|
|
||||||
# the corpus manifest. Single source of truth for both the copy step and the
|
|
||||||
# fixture manifest, so the two can never drift apart.
|
|
||||||
.FIXTURE_METADATA_FILES <- c(
|
|
||||||
"canonical_alias.parquet",
|
|
||||||
"canonical_fips_xwalk.parquet",
|
|
||||||
"census_collection_coverage.parquet",
|
|
||||||
"code_set.parquet",
|
|
||||||
"harmonization_map.parquet",
|
|
||||||
"harmonization_recipes.parquet",
|
|
||||||
"lineage_events.parquet",
|
|
||||||
"representation.parquet",
|
|
||||||
"series_breaks.parquet",
|
|
||||||
"summary_categories.parquet"
|
|
||||||
)
|
|
||||||
|
|
||||||
regenerate_fixture_corpus <- function(
|
regenerate_fixture_corpus <- function(
|
||||||
publish_cache_dir = file.path(
|
publish_cache_dir = file.path(
|
||||||
"..", "cog_pipeline", "_targets", "publish_cache"
|
"..", "cog_pipeline", "_targets", "publish_cache"
|
||||||
@@ -118,11 +100,20 @@ regenerate_fixture_corpus <- function(
|
|||||||
invisible(NULL)
|
invisible(NULL)
|
||||||
}
|
}
|
||||||
|
|
||||||
# Copy the full (not year-scoped) metadata tables listed in
|
# Copy the full (not year-scoped) canonical_fips_xwalk, canonical_alias,
|
||||||
# .FIXTURE_METADATA_FILES.
|
# summary_categories, and (schema v5+) harmonization_map/
|
||||||
|
# harmonization_recipes/series_breaks parquet tables.
|
||||||
#' @noRd
|
#' @noRd
|
||||||
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
|
.copy_metadata_parquets <- function(publish_cache_dir, fixture_dir) {
|
||||||
for (f in .FIXTURE_METADATA_FILES) {
|
files <- c(
|
||||||
|
"canonical_fips_xwalk.parquet",
|
||||||
|
"canonical_alias.parquet",
|
||||||
|
"summary_categories.parquet",
|
||||||
|
"harmonization_map.parquet",
|
||||||
|
"harmonization_recipes.parquet",
|
||||||
|
"series_breaks.parquet"
|
||||||
|
)
|
||||||
|
for (f in files) {
|
||||||
src <- file.path(publish_cache_dir, "data", f)
|
src <- file.path(publish_cache_dir, "data", f)
|
||||||
dst <- file.path(fixture_dir, "data", f)
|
dst <- file.path(fixture_dir, "data", f)
|
||||||
if (!file.exists(src)) {
|
if (!file.exists(src)) {
|
||||||
@@ -188,7 +179,15 @@ regenerate_fixture_corpus <- function(
|
|||||||
)
|
)
|
||||||
})
|
})
|
||||||
|
|
||||||
metadata <- lapply(.FIXTURE_METADATA_FILES, function(f) {
|
metadata_files <- c(
|
||||||
|
"canonical_alias.parquet",
|
||||||
|
"canonical_fips_xwalk.parquet",
|
||||||
|
"summary_categories.parquet",
|
||||||
|
"harmonization_map.parquet",
|
||||||
|
"harmonization_recipes.parquet",
|
||||||
|
"series_breaks.parquet"
|
||||||
|
)
|
||||||
|
metadata <- lapply(metadata_files, function(f) {
|
||||||
rel <- file.path("data", f)
|
rel <- file.path("data", f)
|
||||||
path <- file.path(fixture_dir, rel)
|
path <- file.path(fixture_dir, rel)
|
||||||
list(
|
list(
|
||||||
@@ -204,16 +203,13 @@ regenerate_fixture_corpus <- function(
|
|||||||
pipeline_commit = source_manifest$pipeline_commit,
|
pipeline_commit = source_manifest$pipeline_commit,
|
||||||
fixture_note = paste(
|
fixture_note = paste(
|
||||||
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
|
"Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full",
|
||||||
"corpus available via USCOGDATA_URL. Regenerated from the sparsified",
|
"corpus available via USCOGDATA_URL. Regenerated for Phase R2",
|
||||||
"schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit",
|
"(schema_version 5, harmonization_map/harmonization_recipes/",
|
||||||
"zeros, so FY2011 absence means Census published $0 while FY2012+",
|
"series_breaks parquet tables added). 2011/2012 straddle the",
|
||||||
"absence means not reported. representation.parquet and",
|
"wide-aggregate -> modern-leaf format boundary exercised by basis=",
|
||||||
"code_set.parquet carry that rule and ship in full, as do every other",
|
"\"harmonized\" and recipe= queries; 2019/2020 retain the prior",
|
||||||
"metadata table in the publish tree. 2011/2012 straddle both the",
|
"per-capita/CPI regression anchors. Full canonical_fips_xwalk master",
|
||||||
"wide-aggregate -> modern-leaf format boundary (exercised by",
|
"and canonical_alias lookup table included via",
|
||||||
"basis=\"harmonized\" and recipe= queries) and the dense -> sparse",
|
|
||||||
"representation boundary (SB194); 2019/2020 retain the prior",
|
|
||||||
"per-capita/CPI regression anchors. Regenerated via",
|
|
||||||
"data-raw/regenerate_fixture_corpus.R."
|
"data-raw/regenerate_fixture_corpus.R."
|
||||||
),
|
),
|
||||||
data_vintage = source_manifest$data_vintage,
|
data_vintage = source_manifest$data_vintage,
|
||||||
|
|||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+10
-30
@@ -1,8 +1,8 @@
|
|||||||
{
|
{
|
||||||
"schema_version": 6,
|
"schema_version": 6,
|
||||||
"built_at": "2026-07-31T00:47:27Z",
|
"built_at": "2026-07-23T16:14:30Z",
|
||||||
"pipeline_commit": "aadb46b",
|
"pipeline_commit": "4f992a0",
|
||||||
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated from the sparsified schema-v6 corpus: the wide era (<= FY2011) no longer stores explicit zeros, so FY2011 absence means Census published $0 while FY2012+ absence means not reported. representation.parquet and code_set.parquet carry that rule and ship in full, as do every other metadata table in the publish tree. 2011/2012 straddle both the wide-aggregate -> modern-leaf format boundary (exercised by basis=\"harmonized\" and recipe= queries) and the dense -> sparse representation boundary (SB194); 2019/2020 retain the prior per-capita/CPI regression anchors. Regenerated via data-raw/regenerate_fixture_corpus.R.",
|
"fixture_note": "Four-year (2011, 2012, 2019, 2020) fixture for uscogdata tests. Full corpus available via USCOGDATA_URL. Regenerated for Phase R2 (schema_version 5, harmonization_map/harmonization_recipes/ series_breaks parquet tables added). 2011/2012 straddle the wide-aggregate -> modern-leaf format boundary exercised by basis= \"harmonized\" and recipe= queries; 2019/2020 retain the prior per-capita/CPI regression anchors. Full canonical_fips_xwalk master and canonical_alias lookup table included via data-raw/regenerate_fixture_corpus.R.",
|
||||||
"data_vintage": {
|
"data_vintage": {
|
||||||
"source_vintages": {
|
"source_vintages": {
|
||||||
"2012": "10162019",
|
"2012": "10162019",
|
||||||
@@ -36,9 +36,9 @@
|
|||||||
{
|
{
|
||||||
"year": 2011,
|
"year": 2011,
|
||||||
"path": "data/long/year=2011/part-0.parquet",
|
"path": "data/long/year=2011/part-0.parquet",
|
||||||
"sha256": "7848e18497080c8980a4f89c5b386205b2c5bc90db6773827ea01ab3943d16b1",
|
"sha256": "84302ab364dc9fc3b3fbbc3c3f8b826e3508b4d73ff7c42d094d3863cd1e37b5",
|
||||||
"row_count": 496004,
|
"row_count": 2864212,
|
||||||
"size_bytes": 2202455
|
"size_bytes": 3845911
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"year": 2012,
|
"year": 2012,
|
||||||
@@ -74,14 +74,9 @@
|
|||||||
"description": "canonical_fips_xwalk.parquet"
|
"description": "canonical_fips_xwalk.parquet"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"path": "data/census_collection_coverage.parquet",
|
"path": "data/summary_categories.parquet",
|
||||||
"sha256": "143e025616cde684da7c4442bc00d07fbd1556fabb0ea96223931b737e5d10a4",
|
"sha256": "8e6fcd4dd9bb4723841a67233b19388c9762dfc23b4479501183cebf7ea3c1b5",
|
||||||
"description": "census_collection_coverage.parquet"
|
"description": "summary_categories.parquet"
|
||||||
},
|
|
||||||
{
|
|
||||||
"path": "data/code_set.parquet",
|
|
||||||
"sha256": "4cffcb0198dd51e4ff2b694050bb371a5f9965cdac12f25521cb628fb8e118a9",
|
|
||||||
"description": "code_set.parquet"
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"path": "data/harmonization_map.parquet",
|
"path": "data/harmonization_map.parquet",
|
||||||
@@ -93,25 +88,10 @@
|
|||||||
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
|
"sha256": "1133e9a0b02f8f34f5f936e55c5ecd596bb8a55d8425dcce76767f0f3203581c",
|
||||||
"description": "harmonization_recipes.parquet"
|
"description": "harmonization_recipes.parquet"
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"path": "data/lineage_events.parquet",
|
|
||||||
"sha256": "36c16acfbe621d61010984767f1c566993b8a5f481a2c1e134c4c0a600e4502f",
|
|
||||||
"description": "lineage_events.parquet"
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"path": "data/representation.parquet",
|
|
||||||
"sha256": "31ec328a7dd505a321b45f97aafff12e53d68a1a986f63509863035b22a4360d",
|
|
||||||
"description": "representation.parquet"
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"path": "data/series_breaks.parquet",
|
"path": "data/series_breaks.parquet",
|
||||||
"sha256": "06dcc995ff533e57cc65fa25086cc9bf83ba592c58bf7cc99269dc2576f69944",
|
"sha256": "b0b6794b6887a4f300079adfa10029c2a77109faa4952fbff1c5a270793cc02b",
|
||||||
"description": "series_breaks.parquet"
|
"description": "series_breaks.parquet"
|
||||||
},
|
|
||||||
{
|
|
||||||
"path": "data/summary_categories.parquet",
|
|
||||||
"sha256": "e3b0efa00ce713b8f45829b89cfde24b55333f26101f0495df82d85997d18d8e",
|
|
||||||
"description": "summary_categories.parquet"
|
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -12,25 +12,6 @@
|
|||||||
"category": { "type": ["string", "array", "null"] },
|
"category": { "type": ["string", "array", "null"] },
|
||||||
"basis": { "type": ["string", "null"] },
|
"basis": { "type": ["string", "null"] },
|
||||||
"basis_note": { "type": ["string", "null"] },
|
"basis_note": { "type": ["string", "null"] },
|
||||||
"expenditure_concept": {
|
|
||||||
"type": "string",
|
|
||||||
"enum": ["primary", "direct", "total"],
|
|
||||||
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
|
|
||||||
},
|
|
||||||
"expenditure_concept_note": {
|
|
||||||
"type": ["string", "null"],
|
|
||||||
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
|
|
||||||
},
|
|
||||||
"expenditure_concept_direct_suppressed": {
|
|
||||||
"type": "boolean",
|
|
||||||
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
|
|
||||||
},
|
|
||||||
"revenue_concept": {
|
|
||||||
"type": "string",
|
|
||||||
"enum": ["general", "total"],
|
|
||||||
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
|
|
||||||
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
|
|
||||||
},
|
|
||||||
"harmonization": { "type": "object" },
|
"harmonization": { "type": "object" },
|
||||||
"recipe": { "type": ["object", "null"] },
|
"recipe": { "type": ["object", "null"] },
|
||||||
"suggestions": { "type": "array" },
|
"suggestions": { "type": "array" },
|
||||||
@@ -39,30 +20,6 @@
|
|||||||
"aggregate_fallback": { "type": ["object", "null"] },
|
"aggregate_fallback": { "type": ["object", "null"] },
|
||||||
"transformations":{ "type": "object" },
|
"transformations":{ "type": "object" },
|
||||||
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
||||||
"completion": {
|
|
||||||
"type": "object",
|
|
||||||
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
|
|
||||||
"properties": {
|
|
||||||
"applied": { "type": "boolean" },
|
|
||||||
"rows_filled": { "type": "integer" },
|
|
||||||
"absence_means": { "type": "object" }
|
|
||||||
}
|
|
||||||
},
|
|
||||||
"corpus_break_refs": {
|
|
||||||
"type": "array",
|
|
||||||
"items": { "type": "string" },
|
|
||||||
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
|
||||||
},
|
|
||||||
"balance_caveats": {
|
|
||||||
"type": ["object", "null"],
|
|
||||||
"description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.",
|
|
||||||
"properties": {
|
|
||||||
"not_gaap": { "type": "boolean" },
|
|
||||||
"not_gaap_note": { "type": "string" },
|
|
||||||
"coverage_window": { "type": "object" },
|
|
||||||
"truncated": { "type": "array", "items": { "type": "string" } }
|
|
||||||
}
|
|
||||||
},
|
|
||||||
"manifest": { "type": "object" },
|
"manifest": { "type": "object" },
|
||||||
"sql_query": { "type": "string" }
|
"sql_query": { "type": "string" }
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,7 +0,0 @@
|
|||||||
-- Category crosswalk. Numbered 11 (not with the other reference tables at
|
|
||||||
-- 30+) because the flow views (20-25) classify by MEMBERSHIP in this table
|
|
||||||
-- and DuckDB binds a view's sources eagerly at CREATE VIEW time, so it must
|
|
||||||
-- already exist when they register.
|
|
||||||
CREATE OR REPLACE VIEW summary_categories AS
|
|
||||||
SELECT *
|
|
||||||
FROM read_parquet('{url}data/summary_categories.parquet');
|
|
||||||
@@ -1,22 +1,5 @@
|
|||||||
-- Direct-side expenditure rows, classified by crosswalk MEMBERSHIP
|
|
||||||
-- (summary_categories.category_type = 'expenditure'), never by item-code
|
|
||||||
-- first letter: prefix Y alone spans revenue (Y01/Y02), expenditure
|
|
||||||
-- (Y05/Y06) and balance codes, so no first-letter allowlist can route it
|
|
||||||
-- (uscogdata#11, finding F-018). Which subtypes a query actually returns is
|
|
||||||
-- decided per expenditure_concept in R (.verb_spendrev); this view carries
|
|
||||||
-- every non-intergovernmental expenditure subtype: operations, capital,
|
|
||||||
-- assistance, interest, insurance_benefits.
|
|
||||||
--
|
|
||||||
-- The intergovernmental subtype (M/L/Q codes) is deliberately carved out
|
|
||||||
-- into ig_long: its legacy-era rows are published ONLY as aggregate-flagged
|
|
||||||
-- rows, so it cannot live behind this view's NOT is_aggregate filter (see
|
|
||||||
-- 24-ig_long.sql).
|
|
||||||
CREATE OR REPLACE VIEW spending_long AS
|
CREATE OR REPLACE VIEW spending_long AS
|
||||||
SELECT *
|
SELECT *
|
||||||
FROM long
|
FROM long
|
||||||
WHERE item_code IN (
|
WHERE LEFT(item_code, 1) IN ('E', 'F', 'G', 'K')
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE category_type = 'expenditure'
|
|
||||||
AND spend_subtype <> 'intergovernmental'
|
|
||||||
)
|
|
||||||
AND NOT is_aggregate;
|
AND NOT is_aggregate;
|
||||||
|
|||||||
@@ -1,18 +1,5 @@
|
|||||||
-- Revenue rows, classified by crosswalk MEMBERSHIP rather than item-code
|
|
||||||
-- first letter (see 20-spending_long.sql for why prefixes cannot work).
|
|
||||||
--
|
|
||||||
-- Carries EVERY revenue subtype. Which of Census's two published concepts a
|
|
||||||
-- query actually returns is decided per revenue_concept in R
|
|
||||||
-- (.verb_spendrev), exactly as expenditure_concept narrows spending_long:
|
|
||||||
-- general = own_source + federal + state + local_aid (the default)
|
|
||||||
-- total = general + utility + liquor_store + insurance_trust
|
|
||||||
-- Census defines the first by subtracting the other three from the second
|
|
||||||
-- (manual section 4.3), so both concepts need all four families present here.
|
|
||||||
CREATE OR REPLACE VIEW revenue_long AS
|
CREATE OR REPLACE VIEW revenue_long AS
|
||||||
SELECT *
|
SELECT *
|
||||||
FROM long
|
FROM long
|
||||||
WHERE item_code IN (
|
WHERE LEFT(item_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D')
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE category_type = 'revenue'
|
|
||||||
)
|
|
||||||
AND NOT is_aggregate;
|
AND NOT is_aggregate;
|
||||||
|
|||||||
@@ -1,16 +1,6 @@
|
|||||||
-- Harmonized-basis twin of 20-spending_long.sql: same crosswalk-membership
|
|
||||||
-- classification, applied to harmonized_code (the code the row is folded
|
|
||||||
-- onto) rather than the published item_code. Safe because the harmonized
|
|
||||||
-- space is leaf-only and every harmonized_code in the corpus is a
|
|
||||||
-- summary_categories member (verified at fixture regen; a code the
|
|
||||||
-- crosswalk cannot classify would be silently dropped here).
|
|
||||||
CREATE OR REPLACE VIEW spending_long_harmonized AS
|
CREATE OR REPLACE VIEW spending_long_harmonized AS
|
||||||
SELECT * REPLACE (harmonized_code AS item_code)
|
SELECT * REPLACE (harmonized_code AS item_code)
|
||||||
FROM long
|
FROM long
|
||||||
WHERE NOT is_aggregate
|
WHERE NOT is_aggregate
|
||||||
AND harmonized_code IS NOT NULL
|
AND harmonized_code IS NOT NULL
|
||||||
AND harmonized_code IN (
|
AND LEFT(harmonized_code, 1) IN ('E', 'F', 'G', 'K');
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE category_type = 'expenditure'
|
|
||||||
AND spend_subtype <> 'intergovernmental'
|
|
||||||
);
|
|
||||||
|
|||||||
@@ -1,12 +1,6 @@
|
|||||||
-- Harmonized-basis twin of 21-revenue_long.sql: same crosswalk-membership
|
|
||||||
-- classification (every revenue subtype; the concept narrows in R), applied
|
|
||||||
-- to harmonized_code rather than the published item_code.
|
|
||||||
CREATE OR REPLACE VIEW revenue_long_harmonized AS
|
CREATE OR REPLACE VIEW revenue_long_harmonized AS
|
||||||
SELECT * REPLACE (harmonized_code AS item_code)
|
SELECT * REPLACE (harmonized_code AS item_code)
|
||||||
FROM long
|
FROM long
|
||||||
WHERE NOT is_aggregate
|
WHERE NOT is_aggregate
|
||||||
AND harmonized_code IS NOT NULL
|
AND harmonized_code IS NOT NULL
|
||||||
AND harmonized_code IN (
|
AND LEFT(harmonized_code, 1) IN ('T', 'A', 'U', 'B', 'C', 'D');
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE category_type = 'revenue'
|
|
||||||
);
|
|
||||||
|
|||||||
@@ -1,25 +0,0 @@
|
|||||||
-- Intergovernmental expenditure rows: crosswalk spend_subtype =
|
|
||||||
-- 'intergovernmental' (M = to local govts, L = to state govts, Q11/Q12/Q18
|
|
||||||
-- = state payments to school systems -- uscogdata#11, finding F-017).
|
|
||||||
--
|
|
||||||
-- Deliberately does NOT filter `NOT is_aggregate`, unlike spending_long. In the
|
|
||||||
-- wide era (<= FY2011) the IG families M05/M12/M47/M89/L47/L89 are published
|
|
||||||
-- ONLY as aggregate-flagged rows -- filtering them would hide ~70% of legacy IG
|
|
||||||
-- dollars and make Total silently collapse to Direct. This is safe because the
|
|
||||||
-- aggregate codes and their modern leaf components are strictly year-disjoint
|
|
||||||
-- (M47 ends 2011 / M94 starts 2012; M89 is aggregate only <= 2011 and a leaf
|
|
||||||
-- from 2012 alongside M91-93), so no row is ever counted twice. Same argument
|
|
||||||
-- the pipeline's recipe joins use.
|
|
||||||
--
|
|
||||||
-- `L--` stays excluded: it is the IG-to-state FAMILY TOTAL and genuinely
|
|
||||||
-- rolls up the L-NN codes, so including it would double-count. The crosswalk
|
|
||||||
-- deliberately carries no `--` family-total codes, so membership excludes it
|
|
||||||
-- (guarded by "the IG leg never includes the L-- family total" in
|
|
||||||
-- tests/testthat/test-expenditure-concept.R).
|
|
||||||
CREATE OR REPLACE VIEW ig_long AS
|
|
||||||
SELECT *
|
|
||||||
FROM long
|
|
||||||
WHERE item_code IN (
|
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE spend_subtype = 'intergovernmental'
|
|
||||||
);
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
-- Harmonized-basis IG rows. Uses COALESCE(harmonized_code, item_code) rather
|
|
||||||
-- than harmonized_code alone: aggregate rows carry NO harmonized_code by
|
|
||||||
-- construction (harmonized space is leaf-only), so a plain
|
|
||||||
-- `harmonized_code IS NOT NULL` filter would drop every legacy IG aggregate --
|
|
||||||
-- in the bundled fixture corpus (year 2011; 2012+ all carry a harmonized_code)
|
|
||||||
-- that is $379,016,063k across 25,688 M rows and $2,277,458k across 19,266 L
|
|
||||||
-- rows (`SELECT year, LEFT(item_code,1), SUM(amt), COUNT(*) FROM ig_long
|
|
||||||
-- WHERE harmonized_code IS NULL GROUP BY 1, 2`). COALESCE keeps the one real
|
|
||||||
-- IG collapse rule (M38 -> M36, SB012, year-disjoint 1967-2011 vs 2012+)
|
|
||||||
-- while never dropping a row.
|
|
||||||
--
|
|
||||||
-- Membership is checked on the published item_code (mirroring 24-ig_long.sql)
|
|
||||||
-- rather than the COALESCEd code: every IG harmonization target (M36) is
|
|
||||||
-- itself an IG crosswalk member, so the two are equivalent, and item_code is
|
|
||||||
-- the column that exists on every row.
|
|
||||||
CREATE OR REPLACE VIEW ig_long_harmonized AS
|
|
||||||
SELECT * REPLACE (COALESCE(harmonized_code, item_code) AS item_code)
|
|
||||||
FROM long
|
|
||||||
WHERE item_code IN (
|
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE spend_subtype = 'intergovernmental'
|
|
||||||
);
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
-- Cash and security holdings, classified by crosswalk MEMBERSHIP on
|
|
||||||
-- category_type (see 21-revenue_long.sql for why first-letter prefixes cannot
|
|
||||||
-- do this job -- the X and Y families each span revenue, expenditure AND
|
|
||||||
-- balance).
|
|
||||||
--
|
|
||||||
-- These rows are STOCKS: a balance at a point in time, not a flow over a
|
|
||||||
-- fiscal year. Summing a stock with a flow is meaningless, which is why they
|
|
||||||
-- live behind a third view rather than as a subtype of either money view, and
|
|
||||||
-- why neither spending_long nor revenue_long can reach them.
|
|
||||||
--
|
|
||||||
-- `NOT is_aggregate` mirrors spending_long / revenue_long. The wide-era
|
|
||||||
-- aggregate-only holdings codes (X40/X41) are deliberately outside this view;
|
|
||||||
-- they are reachable only through the recipe path, which bypasses this filter
|
|
||||||
-- by design (cog_pipeline/docs/phase_r_harmonization_review.md § 0.2).
|
|
||||||
CREATE OR REPLACE VIEW balance_long AS
|
|
||||||
SELECT *
|
|
||||||
FROM long
|
|
||||||
WHERE item_code IN (
|
|
||||||
SELECT item_code FROM summary_categories
|
|
||||||
WHERE category_type = 'balance'
|
|
||||||
)
|
|
||||||
AND NOT is_aggregate;
|
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
CREATE OR REPLACE VIEW summary_categories AS
|
||||||
|
SELECT *
|
||||||
|
FROM read_parquet('{url}data/summary_categories.parquet');
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
CREATE OR REPLACE VIEW representation AS
|
|
||||||
SELECT *
|
|
||||||
FROM read_parquet('{url}data/representation.parquet');
|
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
CREATE OR REPLACE VIEW code_set AS
|
|
||||||
SELECT *
|
|
||||||
FROM read_parquet('{url}data/code_set.parquet');
|
|
||||||
@@ -1,16 +0,0 @@
|
|||||||
CREATE OR REPLACE VIEW ig_annotated AS
|
|
||||||
SELECT
|
|
||||||
s.*,
|
|
||||||
x.gov_name AS xwalk_gov_name,
|
|
||||||
x.govs_type,
|
|
||||||
x.type_label,
|
|
||||||
x.fips_state AS xwalk_fips_state,
|
|
||||||
x.fips_county AS xwalk_fips_county,
|
|
||||||
x.fips_place,
|
|
||||||
x.population_acs,
|
|
||||||
c.category,
|
|
||||||
c.category_type,
|
|
||||||
c.spend_subtype
|
|
||||||
FROM ig_long s
|
|
||||||
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
|
|
||||||
LEFT JOIN summary_categories c USING (item_code);
|
|
||||||
@@ -1,16 +0,0 @@
|
|||||||
CREATE OR REPLACE VIEW ig_annotated_harmonized AS
|
|
||||||
SELECT
|
|
||||||
s.*,
|
|
||||||
x.gov_name AS xwalk_gov_name,
|
|
||||||
x.govs_type,
|
|
||||||
x.type_label,
|
|
||||||
x.fips_state AS xwalk_fips_state,
|
|
||||||
x.fips_county AS xwalk_fips_county,
|
|
||||||
x.fips_place,
|
|
||||||
x.population_acs,
|
|
||||||
c.category,
|
|
||||||
c.category_type,
|
|
||||||
c.spend_subtype
|
|
||||||
FROM ig_long_harmonized s
|
|
||||||
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
|
|
||||||
LEFT JOIN summary_categories c USING (item_code);
|
|
||||||
@@ -1,16 +0,0 @@
|
|||||||
CREATE OR REPLACE VIEW balance_annotated AS
|
|
||||||
SELECT
|
|
||||||
s.*,
|
|
||||||
x.gov_name AS xwalk_gov_name,
|
|
||||||
x.govs_type,
|
|
||||||
x.type_label,
|
|
||||||
x.fips_state AS xwalk_fips_state,
|
|
||||||
x.fips_county AS xwalk_fips_county,
|
|
||||||
x.fips_place,
|
|
||||||
x.population_acs,
|
|
||||||
c.category,
|
|
||||||
c.category_type,
|
|
||||||
c.balance_subtype
|
|
||||||
FROM balance_long s
|
|
||||||
LEFT JOIN canonical_fips_xwalk x USING (canonical_govid)
|
|
||||||
LEFT JOIN summary_categories c USING (item_code);
|
|
||||||
@@ -1,76 +0,0 @@
|
|||||||
% Generated by roxygen2: do not edit by hand
|
|
||||||
% Please edit documentation in R/balances.R
|
|
||||||
\name{cog_balances}
|
|
||||||
\alias{cog_balances}
|
|
||||||
\title{Cash and security holdings for one or more governments}
|
|
||||||
\usage{
|
|
||||||
cog_balances(
|
|
||||||
govid,
|
|
||||||
years,
|
|
||||||
category = NULL,
|
|
||||||
per_capita = FALSE,
|
|
||||||
adjust_to_year = NULL,
|
|
||||||
basis = c("harmonized", "raw"),
|
|
||||||
recipe = NULL
|
|
||||||
)
|
|
||||||
}
|
|
||||||
\arguments{
|
|
||||||
\item{govid}{Canonical govid(s): a character vector, or a data frame with a
|
|
||||||
`canonical_govid` column (e.g. from [cog_gov_search()]).}
|
|
||||||
|
|
||||||
\item{years}{Integer vector of fiscal years.}
|
|
||||||
|
|
||||||
\item{category}{Optional character vector of categories to keep. One of
|
|
||||||
`"Fund Balances"`, `"Insurance Trust Balances"`,
|
|
||||||
`"Retirement System Holdings"`. There is deliberately no `subtype`
|
|
||||||
argument: for holdings, `category` is a strict coarsening of
|
|
||||||
`balance_subtype` (unlike the money verbs, where the two axes cross), so
|
|
||||||
every combination would be either redundant or empty.
|
|
||||||
`category = "Fund Balances"` is exactly the `general` family
|
|
||||||
(`W01`/`W31`/`W61`). `balance_subtype` is returned, so a finer split is
|
|
||||||
one `dplyr::filter()` away.}
|
|
||||||
|
|
||||||
\item{per_capita}{Divide holdings by population. Note this is a **stock per
|
|
||||||
resident** (reserves per person), which is *not* comparable to
|
|
||||||
[cog_spending()]'s per-capita figures -- those are a flow per person.}
|
|
||||||
|
|
||||||
\item{adjust_to_year}{Deflate to this year's dollars (CPI-U).}
|
|
||||||
|
|
||||||
\item{basis}{Accepted for uniformity with the money verbs, but currently a
|
|
||||||
**no-op**: `harmonization_map` carries no balance-code rows, so harmonized
|
|
||||||
and raw space are identical for holdings. Reported in
|
|
||||||
`provenance$basis_note`.}
|
|
||||||
|
|
||||||
\item{recipe}{Optional harmonization recipe id (see [cog_recipes()]).
|
|
||||||
`"cash_securities_z77_wide"` and `"cash_securities_z78_wide"` bridge the
|
|
||||||
wide era to the modern one.}
|
|
||||||
}
|
|
||||||
\value{
|
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
|
||||||
`balance_subtype`, `category`, `amt_nominal`, `codes_included`,
|
|
||||||
`aggregate_fallback`, plus optional `amt_per_capita_nominal` and
|
|
||||||
`pop_source` (when `per_capita = TRUE`), optional `amt_real` (when
|
|
||||||
`adjust_to_year` is set), and optional `amt_per_capita_real` (only when
|
|
||||||
**both** `per_capita = TRUE` and `adjust_to_year` are set -- there is no
|
|
||||||
nominal per-capita column to deflate otherwise). Amounts are full US
|
|
||||||
dollars.
|
|
||||||
|
|
||||||
Carries a `provenance` attribute matching
|
|
||||||
`inst/schemas/provenance-v1.json`, whose `balance_caveats` block reports
|
|
||||||
`not_gaap`, `not_gaap_note`, `coverage_window` (measured year extents for
|
|
||||||
every balance subtype in the mounted corpus, not only the observed ones)
|
|
||||||
and `truncated` (the observed subtypes whose coverage falls short of the
|
|
||||||
requested years). `expenditure_concept`/`revenue_concept` are `NA` --
|
|
||||||
holdings are a stock, not a flow, so neither concept vocabulary applies.
|
|
||||||
}
|
|
||||||
\description{
|
|
||||||
Returns Census cash-and-security holdings (`category_type = "balance"`):
|
|
||||||
fund balances, retirement system holdings and insurance trust balances.
|
|
||||||
}
|
|
||||||
\section{Holdings are not GAAP fund balance}{
|
|
||||||
|
|
||||||
Census holdings are **gross** -- no liabilities are netted -- so a reserve
|
|
||||||
ratio built from them overstates what is actually available. They are not
|
|
||||||
comparable to a GAAP fund balance from an ACFR.
|
|
||||||
}
|
|
||||||
|
|
||||||
+1
-10
@@ -11,8 +11,7 @@ cog_find_peers(
|
|||||||
same_state = FALSE,
|
same_state = FALSE,
|
||||||
pop_range = c(0.7, 1.3),
|
pop_range = c(0.7, 1.3),
|
||||||
is_ratio = TRUE,
|
is_ratio = TRUE,
|
||||||
max_peers = 10L,
|
max_peers = 10L
|
||||||
coverage = c("all", "census", "consistent")
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -35,14 +34,6 @@ target's population at `year` to produce absolute bounds. If `FALSE`,
|
|||||||
`pop_range` is interpreted as absolute population counts.}
|
`pop_range` is interpreted as absolute population counts.}
|
||||||
|
|
||||||
\item{max_peers}{Integer cap on the number of peers returned.}
|
\item{max_peers}{Integer cap on the number of peers returned.}
|
||||||
|
|
||||||
\item{coverage}{Survey-cycle handling; see [cog_peer_compare()]. Here it
|
|
||||||
governs the cohort VINTAGE when `year` is `NULL`: `"census"` snaps to the
|
|
||||||
most recent census year with an observed population, so a cohort is not
|
|
||||||
built from a sample year in which most of the candidate universe is
|
|
||||||
absent. `"consistent"` needs a year range, which cohort selection does not
|
|
||||||
have, so it selects like `"all"` and is carried on the result as
|
|
||||||
`attr(x, "coverage")` for [cog_peer_compare()].}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
Tibble with columns `canonical_govid`, `gov_name`, `fips_state`,
|
||||||
|
|||||||
@@ -9,9 +9,7 @@ cog_geographic_rollup(
|
|||||||
category,
|
category,
|
||||||
years,
|
years,
|
||||||
per_capita = FALSE,
|
per_capita = FALSE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
coverage = c("all", "census", "consistent")
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -29,33 +27,6 @@ population from `gov_population_yearly`. Govs with missing population
|
|||||||
are excluded from the result.}
|
are excluded from the result.}
|
||||||
|
|
||||||
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
|
\item{adjust_to_year}{Integer base year for CPI-U conversion, or `NULL`.}
|
||||||
|
|
||||||
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
|
|
||||||
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
|
||||||
refused here because combining Total across multiple layers of
|
|
||||||
government double-counts intergovernmental transfers (a state's payment
|
|
||||||
to a school district is the same dollar the district reports as its own
|
|
||||||
Direct spending); `"primary"` and `"direct"` combine safely.}
|
|
||||||
|
|
||||||
\item{coverage}{How to handle the Census of Governments survey cycle,
|
|
||||||
which is a **complete census only in years ending in 2 and 7** -- every
|
|
||||||
other year is a sample, and the sample varies enormously (on the bundled
|
|
||||||
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
|
||||||
and 112 in FY2019).
|
|
||||||
|
|
||||||
* `"all"` (default) -- every unit that reported that year. Unchanged
|
|
||||||
behaviour, so existing code keeps working.
|
|
||||||
* `"census"` -- census years only. Aborts if the requested range holds
|
|
||||||
none, rather than silently returning nothing.
|
|
||||||
* `"consistent"` -- only units reporting in *every* requested year, giving
|
|
||||||
a balanced panel.
|
|
||||||
|
|
||||||
Regardless of mode, `provenance$coverage` always carries per-year
|
|
||||||
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
|
||||||
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
|
||||||
statement about the **survey calendar**, never a claim of completeness:
|
|
||||||
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
|
||||||
report. `n_units_reporting` is the number that tells the truth.}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`,
|
||||||
|
|||||||
@@ -32,11 +32,8 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
|||||||
}
|
}
|
||||||
\details{
|
\details{
|
||||||
* **Utility mode** (single `name`, the original behavior): returns all
|
* **Utility mode** (single `name`, the original behavior): returns all
|
||||||
rows whose `gov_name` contains `name` as a **literal, case-insensitive
|
rows whose `gov_name` matches the regex case-insensitively, sorted by
|
||||||
substring**, sorted by `population_acs` descending. Useful for
|
`population_acs` descending. Useful for exploratory lookups.
|
||||||
exploratory lookups. Regex metacharacters in `name` are escaped, so a
|
|
||||||
government is findable by its own complete name even when that name
|
|
||||||
contains parentheses or a period.
|
|
||||||
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
* **Basket mode** (`length(name) > 1`): resolves each input row to a
|
||||||
single canonical govid and returns a tibble in input order, suitable
|
single canonical govid and returns a tibble in input order, suitable
|
||||||
for piping straight into [cog_spending()] / [cog_revenue()] /
|
for piping straight into [cog_spending()] / [cog_revenue()] /
|
||||||
@@ -48,8 +45,7 @@ the cross-vintage canonical-government registry. Operates in two modes:
|
|||||||
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
|
||||||
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
2. **Exact pass:** case-insensitive equality against `gov_name`.
|
||||||
Single hit -> resolved. Multiple -> step 4.
|
Single hit -> resolved. Multiple -> step 4.
|
||||||
3. **Substring fallback:** case-insensitive literal substring against
|
3. **Substring fallback:** case-insensitive regex against `gov_name`.
|
||||||
`gov_name` (metacharacters escaped).
|
|
||||||
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
|
||||||
`status = "no_match"`. Multiple hits -> step 4.
|
`status = "no_match"`. Multiple hits -> step 4.
|
||||||
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
4. **Disambiguation:** if matches share one `govs_type`, pick the
|
||||||
@@ -62,7 +58,7 @@ inputs (`ambiguous` / `no_match`) appear only in the sidecar.
|
|||||||
}
|
}
|
||||||
\examples{
|
\examples{
|
||||||
\dontrun{
|
\dontrun{
|
||||||
# Utility mode — exploratory substring lookup
|
# Utility mode — exploratory regex lookup
|
||||||
cog_gov_search("broward", state = "FL")
|
cog_gov_search("broward", state = "FL")
|
||||||
|
|
||||||
# Basket mode — resolve a known cohort
|
# Basket mode — resolve a known cohort
|
||||||
|
|||||||
+2
-64
@@ -10,9 +10,7 @@ cog_peer_compare(
|
|||||||
category,
|
category,
|
||||||
years,
|
years,
|
||||||
per_capita = TRUE,
|
per_capita = TRUE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
coverage = c("all", "census", "consistent")
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -29,38 +27,6 @@ cog_peer_compare(
|
|||||||
population.}
|
population.}
|
||||||
|
|
||||||
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
|
\item{adjust_to_year}{Integer base year for CPI-U conversion or `NULL`.}
|
||||||
|
|
||||||
\item{expenditure_concept}{`"primary"` (default), `"direct"`, or
|
|
||||||
`"total"` -- see [cog_spending()] for the three concepts. `"total"` is
|
|
||||||
refused here because combining Total across peer sets counts
|
|
||||||
intergovernmental transfers twice; `"primary"` and `"direct"` combine
|
|
||||||
safely.}
|
|
||||||
|
|
||||||
\item{coverage}{How to handle the Census of Governments survey cycle,
|
|
||||||
which is a **complete census only in years ending in 2 and 7** -- every
|
|
||||||
other year is a sample, and the sample varies enormously (on the bundled
|
|
||||||
fixture, Wisconsin's 608-city universe reports 597 governments in FY2012
|
|
||||||
and 112 in FY2019).
|
|
||||||
|
|
||||||
* `"all"` (default) -- every unit that reported that year. Unchanged
|
|
||||||
behaviour, so existing code keeps working.
|
|
||||||
* `"census"` -- census years only. Aborts if the requested range holds
|
|
||||||
none, rather than silently returning nothing.
|
|
||||||
* `"consistent"` -- only units reporting in *every* requested year, giving
|
|
||||||
a balanced panel.
|
|
||||||
|
|
||||||
Regardless of mode, `provenance$coverage` always carries per-year
|
|
||||||
`n_units_reporting`, `n_units_expected` and `is_census_year`, and
|
|
||||||
`provenance$coverage_mode` records the mode. `is_census_year` is a
|
|
||||||
statement about the **survey calendar**, never a claim of completeness:
|
|
||||||
FY1967 is a census year in which only 97 of Wisconsin's 608 cities
|
|
||||||
report. `n_units_reporting` is the number that tells the truth.
|
|
||||||
|
|
||||||
The comparison target is exempt from `"consistent"` balancing -- it is the
|
|
||||||
subject of the comparison, not a member of the cohort -- and the
|
|
||||||
`summary_*` quantiles are computed AFTER the filter, so they describe the
|
|
||||||
cohort actually returned. `n_units_reporting` counts peers only, against
|
|
||||||
the cohort size: "3 of your 15 peers reported in FY2019".}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble matching [cog_spending()]'s columns, plus a `role`
|
Tibble matching [cog_spending()]'s columns, plus a `role`
|
||||||
@@ -71,39 +37,11 @@ Tibble matching [cog_spending()]'s columns, plus a `role`
|
|||||||
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
`attr(peers, "cohort_year")`; `NA` when `peers` was a bare character
|
||||||
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
vector). Provenance reports `verb = "cog_peer_compare"`, `peer_count`,
|
||||||
`cohort_year`, and `cohort_govids`.
|
`cohort_year`, and `cohort_govids`.
|
||||||
|
|
||||||
**The `summary_*` rows are per-category quantiles: they are not additive.**
|
|
||||||
Each one is computed **within each `(year, spend_subtype,
|
|
||||||
category)` cell** across the peer set, so a `summary_p50` row is *the
|
|
||||||
median peer's value in that one category*, not *the value of the median
|
|
||||||
peer's total*. The median peer for Police and the median peer for Fire
|
|
||||||
are usually different governments, so summing `summary_*` rows across
|
|
||||||
categories does not give any peer's total and misstates the band it
|
|
||||||
appears to describe — measured at −32.7% to +251.0% across 24 years on
|
|
||||||
one cohort, with a sign flip at FY2012.
|
|
||||||
|
|
||||||
Facet by `role` **and** `category` (the documented use, and what the
|
|
||||||
rows are built for). For a genuine "median peer's total spending" line,
|
|
||||||
sum each peer's own categories first and take the quantile of those
|
|
||||||
per-government totals:
|
|
||||||
|
|
||||||
```r
|
|
||||||
library(dplyr)
|
|
||||||
cmp |>
|
|
||||||
filter(role %in% c("target", "peer")) |>
|
|
||||||
group_by(year, role, canonical_govid) |>
|
|
||||||
summarise(total = sum(amt_per_capita_real, na.rm = TRUE), .groups = "drop") |>
|
|
||||||
filter(role == "peer") |>
|
|
||||||
group_by(year) |>
|
|
||||||
summarise(p50 = quantile(total, 0.5, na.rm = TRUE))
|
|
||||||
```
|
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
Pulls spending for the target plus a peer set (either a
|
Pulls spending for the target plus a peer set (either a
|
||||||
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
[cog_find_peers()] result or a character vector of `canonical_govid`) and
|
||||||
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
appends peer-distribution summary rows (`summary_p25`, `summary_p50`,
|
||||||
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
`summary_p75`) so the result can be faceted by `role` in a single ggplot
|
||||||
call. Those summary rows are quantiles **within each category**, not
|
call.
|
||||||
quantiles of each peer's total — see the `@return` section before summing
|
|
||||||
them.
|
|
||||||
}
|
}
|
||||||
|
|||||||
+2
-54
@@ -11,9 +11,7 @@ cog_revenue(
|
|||||||
per_capita = FALSE,
|
per_capita = FALSE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL,
|
recipe = NULL
|
||||||
revenue_concept = c("general", "total"),
|
|
||||||
complete = FALSE
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -57,62 +55,12 @@ argument is ignored and the result's provenance reports
|
|||||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||||
|
|
||||||
\item{revenue_concept}{Which of Census's two published revenue concepts to
|
|
||||||
return. Concepts are defined as sets of the crosswalk's `revenue_subtype`
|
|
||||||
values -- never as item-code first letters, which cannot classify
|
|
||||||
correctly (prefix `Y` spans revenue, expenditure and balance codes, and
|
|
||||||
prefix `X` does the same):
|
|
||||||
|
|
||||||
* `"general"` (default) -- Census General Revenue: `own_source` +
|
|
||||||
`federal` + `state` + `local_aid`. The manual defines this concept by
|
|
||||||
subtraction (section 4.3: *"General revenue comprises all revenue
|
|
||||||
except that classified as liquor store, utility, or insurance trust
|
|
||||||
revenue"*), so utility (`A91`-`A94`), liquor store (`A90`) and
|
|
||||||
insurance trust revenue are all excluded.
|
|
||||||
* `"total"` -- Census Total Revenue: every revenue subtype, i.e.
|
|
||||||
`general` plus utility, liquor store, and insurance trust revenue
|
|
||||||
(`Y01`/`Y02`/`Y04`/`Y11`/`Y12`/`Y51`/`Y52` and the employee-retirement
|
|
||||||
`X01`/`X02`/`X05`/`X08`).
|
|
||||||
|
|
||||||
The two are related by Census's own identity, `Total Revenue = General +
|
|
||||||
Utility + Liquor Store + Insurance Trust`.
|
|
||||||
|
|
||||||
Note that the employee-retirement (`X`) codes stop at FY2016, when those
|
|
||||||
systems moved out of the annual finance file into the separate Annual
|
|
||||||
Survey of Public Pensions, so a `"total"` series steps down at the
|
|
||||||
FY2016/FY2017 seam for reasons that are about collection scope rather
|
|
||||||
than revenue (series breaks `SB197`-`SB202`).}
|
|
||||||
|
|
||||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
|
||||||
corpus does not carry still appears, labelled with **why** it is
|
|
||||||
missing, and add a `value_source` column to every row:
|
|
||||||
|
|
||||||
* `"reported"` — the corpus carries this cell.
|
|
||||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
|
||||||
Census published `$0`. `amt_nominal` is `0`.
|
|
||||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
|
||||||
government did not report, and the value is unknown. `amt_nominal` is
|
|
||||||
`NA`, **not** `0` — writing a zero there would invent data.
|
|
||||||
|
|
||||||
The grid comes from the corpus's `code_set` table, scoped to each
|
|
||||||
government's own type, so a county is never filled with cells only a
|
|
||||||
state can report. Reported rows are passed through untouched.
|
|
||||||
|
|
||||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
|
||||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
|
||||||
when `representation`/`code_set` began shipping; aborts with class
|
|
||||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
|
||||||
`recipe` or with `expenditure_concept = "total"` (class
|
|
||||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
|
||||||
`code_set`.}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
`revenue_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||||
and `value_source` when `complete = TRUE`.
|
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
Mirror of [cog_spending()] for revenue categories. One row per
|
Mirror of [cog_spending()] for revenue categories. One row per
|
||||||
|
|||||||
+3
-72
@@ -11,9 +11,7 @@ cog_spending(
|
|||||||
per_capita = FALSE,
|
per_capita = FALSE,
|
||||||
adjust_to_year = NULL,
|
adjust_to_year = NULL,
|
||||||
basis = c("harmonized", "raw"),
|
basis = c("harmonized", "raw"),
|
||||||
recipe = NULL,
|
recipe = NULL
|
||||||
expenditure_concept = c("primary", "direct", "total"),
|
|
||||||
complete = FALSE
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
\arguments{
|
\arguments{
|
||||||
@@ -57,80 +55,13 @@ argument is ignored and the result's provenance reports
|
|||||||
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
`basis = "recipe"` with an inert `harmonization` block (`applied =
|
||||||
FALSE`, pointing at the `recipe` block instead) rather than a
|
FALSE`, pointing at the `recipe` block instead) rather than a
|
||||||
possibly-misleading `"harmonized"`/`"raw"` value.}
|
possibly-misleading `"harmonized"`/`"raw"` value.}
|
||||||
|
|
||||||
\item{expenditure_concept}{Which spending concept to return. Concepts are
|
|
||||||
defined as sets of the crosswalk's `spend_subtype` values -- never as
|
|
||||||
item-code first letters, which cannot classify correctly (prefix `Y`
|
|
||||||
alone spans revenue, expenditure, and balance codes):
|
|
||||||
|
|
||||||
* `"primary"` (default) -- the government's own service provision:
|
|
||||||
`operations` + `capital` + `assistance` subtypes.
|
|
||||||
* `"direct"` -- Census's published Direct Expenditure: `primary` plus
|
|
||||||
`interest` (interest on debt) and `insurance_benefits` (insurance
|
|
||||||
trust benefit payments, e.g. pensions -- Census manual section
|
|
||||||
5.2.2.1 includes payments to retirees in Direct).
|
|
||||||
* `"total"` -- `direct` plus the intergovernmental leg: payments to
|
|
||||||
local governments (`M` codes), to the state government (`L` codes,
|
|
||||||
excluding the `L--` family-total rollup), and state payments to
|
|
||||||
school systems (`Q11`/`Q12`/`Q18`), so results gain rows with
|
|
||||||
`spend_subtype == "intergovernmental"`. Requires the active corpus's
|
|
||||||
`summary_categories` to carry M/L rows (added by cog_pipeline PR
|
|
||||||
#59); aborts with class `uscogdata_ig_categories_unsupported` on an
|
|
||||||
older corpus rather than silently under-reporting. Mutually
|
|
||||||
exclusive with `recipe` (a recipe already defines its own component
|
|
||||||
codes).
|
|
||||||
|
|
||||||
**Do not sum `"total"` results across levels of government** (e.g.
|
|
||||||
state + county + city): a state's `M12` payment to a school district is
|
|
||||||
the same dollar the district reports as its own direct `E12`, so
|
|
||||||
summing both double-counts it. This matters in particular with
|
|
||||||
[cog_geographic_rollup()], which sums across exactly that kind of
|
|
||||||
multi-layer government set.
|
|
||||||
|
|
||||||
In the legacy wide era (<= FY2011), some functions are published ONLY
|
|
||||||
as an aggregate-flagged family total (e.g. Corrections' `E04`/`E05`
|
|
||||||
split), which the Direct leg excludes by construction but the IG leg
|
|
||||||
deliberately keeps (see `inst/sql/24-ig_long.sql`). For a `"total"`
|
|
||||||
query, any (year, category) where this leaves intergovernmental rows
|
|
||||||
with NO Direct counterpart is flagged: the affected rows' `notes`
|
|
||||||
name the harmonization recipe that recovers the missing Direct
|
|
||||||
component (when one exists), and
|
|
||||||
`provenance$expenditure_concept_direct_suppressed` is `TRUE` -- the
|
|
||||||
figure in those rows is the intergovernmental leg alone, not Direct +
|
|
||||||
IG.}
|
|
||||||
|
|
||||||
\item{complete}{If `TRUE`, fill the requested grid so that a cell the
|
|
||||||
corpus does not carry still appears, labelled with **why** it is
|
|
||||||
missing, and add a `value_source` column to every row:
|
|
||||||
|
|
||||||
* `"reported"` — the corpus carries this cell.
|
|
||||||
* `"census_zero"` — dense-source year (`<= FY2011`), cell absent:
|
|
||||||
Census published `$0`. `amt_nominal` is `0`.
|
|
||||||
* `"not_reported"` — sparse-source year (`>= FY2012`), cell absent: the
|
|
||||||
government did not report, and the value is unknown. `amt_nominal` is
|
|
||||||
`NA`, **not** `0` — writing a zero there would invent data.
|
|
||||||
|
|
||||||
The grid comes from the corpus's `code_set` table, scoped to each
|
|
||||||
government's own type, so a county is never filled with cells only a
|
|
||||||
state can report. Reported rows are passed through untouched.
|
|
||||||
|
|
||||||
Defaults to `FALSE` (the historical behaviour: absent cells simply do
|
|
||||||
not appear). Needs a corpus published from 2026-07-29 onward, which is
|
|
||||||
when `representation`/`code_set` began shipping; aborts with class
|
|
||||||
`uscogdata_representation_unavailable` otherwise. Not available with
|
|
||||||
`recipe` or with `expenditure_concept = "total"` (class
|
|
||||||
`uscogdata_complete_unsupported`) — neither draws its cells from
|
|
||||||
`code_set`.}
|
|
||||||
}
|
}
|
||||||
\value{
|
\value{
|
||||||
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
Tibble with columns `year`, `canonical_govid`, `gov_name`,
|
||||||
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
`spend_subtype`, `category`, `amt_nominal`, optional `amt_real`,
|
||||||
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
optional `amt_per_capita_nominal`, optional `amt_per_capita_real`,
|
||||||
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`,
|
optional `pop_source`, `codes_included`, `aggregate_fallback`, `notes`.
|
||||||
and `value_source` when `complete = TRUE`.
|
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`.
|
||||||
Carries a `provenance` attribute matching `inst/schemas/provenance-v1.json`,
|
|
||||||
whose `completion` block reports `applied`, `rows_filled`, and the
|
|
||||||
per-year `absence_means` rule that was applied.
|
|
||||||
}
|
}
|
||||||
\description{
|
\description{
|
||||||
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
One row per `(year, canonical_govid, spend_subtype, category)`. Amounts are
|
||||||
|
|||||||
@@ -1,281 +0,0 @@
|
|||||||
# `cog_balances()` — a reader surface for cash and security holdings
|
|
||||||
|
|
||||||
**Issue:** `uscogdata#25` requirement 2 · **Downstream:** `cog-api#26`
|
|
||||||
**Date:** 2026-08-03 · **Status:** design, awaiting approval
|
|
||||||
|
|
||||||
Requirement 1 of `uscogdata#25` (no `balance` row may reach a money verb) shipped
|
|
||||||
with `#11`/`#12` and is asserted at both view and verb level. This spec covers
|
|
||||||
requirement 2 only: a way to query holdings.
|
|
||||||
|
|
||||||
## Decision: a verb, not an argument
|
|
||||||
|
|
||||||
`cog_balances()`, parallel to `cog_spending()` / `cog_revenue()`.
|
|
||||||
|
|
||||||
Holdings are a **stock** — a balance at a point in time — while the money verbs
|
|
||||||
return **flows** over a fiscal year. The flow verbs' whole argument vocabulary
|
|
||||||
is meaningless for a stock: `expenditure_concept` / `revenue_concept` describe
|
|
||||||
which flows Census aggregates into a published total, and `complete=` fills a
|
|
||||||
grid of fiscal-year cells. Overloading a money verb would put a stock behind
|
|
||||||
arguments that all assume a flow.
|
|
||||||
|
|
||||||
## The 14 codes
|
|
||||||
|
|
||||||
Measured against the published corpus 2026-08-03, not transcribed from the
|
|
||||||
issue. `year_min`/`year_max` are observed row extents.
|
|
||||||
|
|
||||||
| `balance_subtype` | `category` | codes | observed years |
|
|
||||||
|---|---|---|---|
|
|
||||||
| `general` | Fund Balances | `W01`, `W31`, `W61` | 2012–2021 |
|
|
||||||
| `employee_retirement` | Retirement System Holdings | `X21`, `X42`, `X44` | 1967–2016 |
|
|
||||||
| | | `X47` | 1988–2016 |
|
|
||||||
| | | `X30`, `Z77`, `Z78` | 2012–2016 |
|
|
||||||
| `unemployment_trust` | Insurance Trust Balances | `Y07`, `Y08` | 1967–2023 |
|
|
||||||
| `workers_comp_trust` | Insurance Trust Balances | `Y21` | 2012–2023 |
|
|
||||||
| `other_insurance_trust` | Insurance Trust Balances | `Y61` | 2012–2023 |
|
|
||||||
|
|
||||||
## Architecture
|
|
||||||
|
|
||||||
### Two new views
|
|
||||||
|
|
||||||
Mirroring the `revenue_long` / `revenue_annotated` pair exactly:
|
|
||||||
|
|
||||||
- `inst/sql/26-balance_long.sql` — `category_type = 'balance' AND NOT is_aggregate`
|
|
||||||
- `inst/sql/46-balance_annotated.sql` — joins `canonical_fips_xwalk` and
|
|
||||||
`summary_categories`, exposing `category`, `category_type`, `balance_subtype`
|
|
||||||
|
|
||||||
`.register_views()` globs `inst/sql/*.sql` in sorted order, so both register
|
|
||||||
with no new registration code.
|
|
||||||
|
|
||||||
### A third gate list in `R/views.R`
|
|
||||||
|
|
||||||
`CREATE VIEW` resolves its source schema eagerly, so a missing **column** fails
|
|
||||||
at registration time, not at query time. `46-balance_annotated.sql` selects
|
|
||||||
`c.balance_subtype`, which exists only on corpora built after pipeline `#76`/`#77`.
|
|
||||||
That arrived without a `schema_version` bump, so neither existing gate applies:
|
|
||||||
`.harmonization_view_files` keys on `schema_version`, `.representation_view_files`
|
|
||||||
on the presence of a *file*. The discriminator here is a **column on an existing
|
|
||||||
table**.
|
|
||||||
|
|
||||||
```r
|
|
||||||
.balance_view_files <- c("26-balance_long.sql", "46-balance_annotated.sql")
|
|
||||||
```
|
|
||||||
|
|
||||||
gated by probing `summary_categories` for `balance_subtype`, with
|
|
||||||
`cog_balances()` erroring cleanly via `.require_balance_support()` on an older
|
|
||||||
corpus — mirroring how `.require_schema_v5()` gates the harmonized views.
|
|
||||||
|
|
||||||
### `R/balances.R` — a dedicated path, not `.verb_spendrev()`
|
|
||||||
|
|
||||||
`.verb_spendrev()` is 825 lines whose concept scoping, intergovernmental leg and
|
|
||||||
`complete=` grid are all flow-specific, and four verbs depend on it. Threading a
|
|
||||||
third mode through it adds branching to shared code for no reuse benefit.
|
|
||||||
|
|
||||||
Reused unchanged: `.build_provenance()`, `.build_series_break_refs()`,
|
|
||||||
`.build_corpus_break_refs()`, the population join, `.inflate()`, and
|
|
||||||
`.coerce_govid_input()`.
|
|
||||||
|
|
||||||
Following the package's real two-layer convention: **view definitions** live in
|
|
||||||
`inst/sql/`; **query construction** is inline `sprintf()` in R, as in
|
|
||||||
`.verb_spendrev()`. (`CLAUDE.md` currently states "never inline SQL strings in R
|
|
||||||
files", which the verb layer has never obeyed. Corrected in a separate commit —
|
|
||||||
see Out of scope.)
|
|
||||||
|
|
||||||
## Signature
|
|
||||||
|
|
||||||
```r
|
|
||||||
cog_balances(govid, years,
|
|
||||||
category = NULL, # Fund Balances | Insurance Trust Balances |
|
|
||||||
# Retirement System Holdings
|
|
||||||
per_capita = FALSE,
|
|
||||||
adjust_to_year = NULL,
|
|
||||||
basis = c("harmonized", "raw"),
|
|
||||||
recipe = NULL)
|
|
||||||
```
|
|
||||||
|
|
||||||
Returns a `tbl_df` with a `provenance` attribute, like every other verb.
|
|
||||||
|
|
||||||
**Absent by design:** `expenditure_concept`, `revenue_concept`, `complete`,
|
|
||||||
and `subtype` — see below.
|
|
||||||
|
|
||||||
**`per_capita` is offered.** Holdings per resident is a real measure (pension
|
|
||||||
assets per capita, fund balance per resident). The roxygen `@param` states
|
|
||||||
plainly that this is a *stock per resident* and is **not** comparable to
|
|
||||||
`cog_spending()`'s per-capita figures.
|
|
||||||
|
|
||||||
**`basis` is currently a no-op** — `harmonization_map` has zero balance-code
|
|
||||||
rows, so harmonized and raw are identical for holdings. Kept for uniformity
|
|
||||||
with the money verbs (the API would otherwise special-case), and
|
|
||||||
`provenance$basis_note` says so outright rather than letting it look meaningful.
|
|
||||||
|
|
||||||
**`recipe` ships in v1 and works.** The two holdings recipes bridge the wide era
|
|
||||||
to the modern one:
|
|
||||||
|
|
||||||
```
|
|
||||||
cash_securities_z77_wide = X40 (1967-2011) + Z77 (2012-2023)
|
|
||||||
cash_securities_z78_wide = X41 (1967-2011) + Z78 (2012-2023)
|
|
||||||
```
|
|
||||||
|
|
||||||
`X40`/`X41` carry ~42,700 rows that are **100% `is_aggregate = TRUE`**, so they
|
|
||||||
are invisible to `balance_long`, which filters `NOT is_aggregate` like every
|
|
||||||
other basis view. That is by design, not a defect:
|
|
||||||
`cog_pipeline/docs/phase_r_harmonization_review.md` § 0.2 records that the wide
|
|
||||||
era exposes these split families *only* as aggregates, and that the recipe join
|
|
||||||
must therefore **not** filter `is_aggregate` — safe by construction, because
|
|
||||||
wide rows (≤2011) are aggregate-only, modern rows (2012+) are leaf-only, and
|
|
||||||
every component is year-scoped, so no double-count is possible. § 1 records the
|
|
||||||
matching decision that the planned `X40→Z77` harmonization *map* rows were
|
|
||||||
dropped and the continuity ships as recipes instead, which is why
|
|
||||||
`harmonization_map` has no balance-code rows.
|
|
||||||
|
|
||||||
The reader already implements this (`R/recipes.R`, `R/spending.R`), and it is
|
|
||||||
verified rather than assumed: `corrections_combined` for FY2007 — a recipe whose
|
|
||||||
wide leg `E05` is likewise aggregate-only — returns $906,743,000 against the
|
|
||||||
live corpus. So a recipe query reaches rows the verb's own view cannot, exactly
|
|
||||||
as intended.
|
|
||||||
|
|
||||||
### No `subtype` argument: `category` is a strict coarsening
|
|
||||||
|
|
||||||
`balance` is the only `category_type` in which `category` and the subtype column
|
|
||||||
are **not** orthogonal. Measured against the published crosswalk:
|
|
||||||
|
|
||||||
| `category_type` | subtypes spanning more than one category |
|
|
||||||
|---|---|
|
|
||||||
| expenditure | 5 of 6 (`operations`, `capital`, `interest`, `assistance`, `intergovernmental`) |
|
|
||||||
| revenue | 1 of 7 (`own_source`) |
|
|
||||||
| **balance** | **0 of 5** |
|
|
||||||
|
|
||||||
For expenditure the two axes are a genuine cross-tab — *function* (Police, Fire)
|
|
||||||
× *economic character* (operations, capital) — so both earn their place. For
|
|
||||||
balance the relation is a strict tree:
|
|
||||||
|
|
||||||
```
|
|
||||||
Fund Balances = {general} W01 W31 W61
|
|
||||||
Retirement System Holdings = {employee_retirement} X21 X30 X42 X44 X47 Z77 Z78
|
|
||||||
Insurance Trust Balances = {unemployment_trust,
|
|
||||||
workers_comp_trust,
|
|
||||||
other_insurance_trust} Y07 Y08 Y21 Y61
|
|
||||||
```
|
|
||||||
|
|
||||||
Exposing both would therefore admit no useful combination. Of the 15 possible
|
|
||||||
pairs, 3 are redundant (the subtype already implies its category) and **12 are
|
|
||||||
guaranteed empty for every government in every year** — and an impossible query
|
|
||||||
would fail by returning an empty tibble, which reads as "this government holds
|
|
||||||
none" rather than "you asked a contradiction."
|
|
||||||
|
|
||||||
Dropping `subtype` also keeps the verb aligned with the rest of the package: no
|
|
||||||
uscogdata verb exposes a subtype argument. `subtype_col` is internal plumbing in
|
|
||||||
`.verb_spendrev()`, and the API layers its own `subtype` row filter on top
|
|
||||||
(`api/R/handlers_governments.R`). `cog-api#26` can do exactly that for
|
|
||||||
`/balances`.
|
|
||||||
|
|
||||||
`#25`'s hard requirement is still met — `category = "Fund Balances"` *is* the
|
|
||||||
`general` family, precisely `W01`/`W31`/`W61`, in one filter. The only loss is
|
|
||||||
isolating one of the three insurance funds in a single argument;
|
|
||||||
`balance_subtype` remains a returned column, so that is one `dplyr::filter()`
|
|
||||||
away.
|
|
||||||
|
|
||||||
## Caveat surfacing
|
|
||||||
|
|
||||||
`provenance$balance_caveats`, always present, plus one `cli_inform()` per
|
|
||||||
session per caveat class when a query actually touches an affected family or
|
|
||||||
year. Structured so `cog-api#26` can forward the fields verbatim.
|
|
||||||
|
|
||||||
Verified against `series_breaks.csv`, not assumed:
|
|
||||||
|
|
||||||
| # | Caveat | Covered by existing machinery? |
|
|
||||||
|---|---|---|
|
|
||||||
| 1 | Gross holdings, **not GAAP fund balance**; no liabilities netted | No — a constant, new field `not_gaap = TRUE` |
|
|
||||||
| 2 | `W` is FY2012–2021 only | No — new `coverage_window`, **computed** from the corpus |
|
|
||||||
| 3 | `X`/`Z` holdings end FY2016 | **Not yet.** No `series_breaks` row exists at 2016/2017 for `Z77`/`Z78`/`X30`. Reader surfaces it via `coverage_window`; flows through `series_break_refs` once the upstream entry lands (see Out of scope) |
|
|
||||||
| 4 | `X40`/`X41` book → market at FY2002 | **Yes**, via `SB195`/`SB196` on `fin_code` `X40`/`X41`, under **two** conditions: a `recipe` query (the only path that observes those codes) **and** a year span that crosses FY2002. Asserted in the tests rather than assumed |
|
|
||||||
|
|
||||||
On caveat 4's second condition: `.build_series_break_refs()` matches
|
|
||||||
`break_year BETWEEN min(years) AND max(years)`, so a request spanning only
|
|
||||||
2011–2012 does **not** surface `SB195`. That is correct, not a gap — such a
|
|
||||||
series sits entirely after the change, on one consistent basis, and flagging a
|
|
||||||
break it never crosses would be noise. The same rule is applied deliberately in
|
|
||||||
`.build_corpus_break_refs()`. An earlier draft of this row omitted the span
|
|
||||||
condition and overclaimed.
|
|
||||||
|
|
||||||
`coverage_window` is derived per observed subtype family from the corpus, never
|
|
||||||
hardcoded, so it stays correct as the corpus grows.
|
|
||||||
|
|
||||||
`series_break_refs` and `corpus_break_refs` are otherwise populated by the
|
|
||||||
existing code-driven builders and need no change.
|
|
||||||
|
|
||||||
## Testing
|
|
||||||
|
|
||||||
New `tests/testthat/test-balances.R`. The bundled fixture covers all four
|
|
||||||
fixture years — `W` in 2012/2019/2020, the `X`/`Z` family in 2011/2012, `Y`
|
|
||||||
throughout — so every test below runs offline.
|
|
||||||
|
|
||||||
- **Inverse guard.** No flow code ever appears in `cog_balances()`, complementing
|
|
||||||
the already-asserted forward guard. Absence is verified against the raw corpus
|
|
||||||
via `read_parquet` on `data/long`, never through the verb that creates it.
|
|
||||||
- **FY2016 seam.** The `X`/`Z` family is present in 2012 and absent in 2019;
|
|
||||||
`coverage_window` reports the termination and the console message fires once.
|
|
||||||
- **Caveats.** `not_gaap` is always `TRUE`; `coverage_window` matches the
|
|
||||||
measured table above; the FY2002 valuation caveat fires only when the year
|
|
||||||
range crosses 2002 *and* touches `employee_retirement`.
|
|
||||||
- **`per_capita`.** `amt_per_capita_nominal == amt_nominal / population`.
|
|
||||||
- **`category = "Fund Balances"` is the `general` family.** Returns exactly
|
|
||||||
`W01`/`W31`/`W61` and nothing else — `#25`'s one-filter requirement, asserted
|
|
||||||
rather than assumed.
|
|
||||||
- **The hierarchy holds.** Every `balance_subtype` in the crosswalk maps to
|
|
||||||
exactly one `category`. Asserted against the crosswalk so that an upstream
|
|
||||||
change breaking the tree — which would silently make `category` lossy —
|
|
||||||
fails here rather than in a user's analysis.
|
|
||||||
- **`recipe` bridges the wide era.** `cash_securities_z77_wide` returns the
|
|
||||||
`X40` leg for a pre-2012 year, proving the aggregate-only wide rows are
|
|
||||||
reached — the property `phase_r_harmonization_review.md` § 0.2 depends on. A
|
|
||||||
regression here would silently truncate a 45-year series to five.
|
|
||||||
- **`SB195`/`SB196` reach the user on that path.** A `recipe` query spanning
|
|
||||||
FY2002 carries both in `provenance$series_break_refs`, so the book → market
|
|
||||||
basis change is disclosed wherever `X40`/`X41` are actually observed.
|
|
||||||
- **Gating.** `.require_balance_support()` errors cleanly on a corpus whose
|
|
||||||
`summary_categories` lacks `balance_subtype`.
|
|
||||||
|
|
||||||
## Out of scope, tracked separately
|
|
||||||
|
|
||||||
1. **Pipeline issue (new), non-blocking.** Catalogue the FY2016 termination of
|
|
||||||
the seven holdings codes in `series_breaks.csv`. There is currently **no**
|
|
||||||
entry at 2016/2017 for `Z77`/`Z78`/`X30`, although
|
|
||||||
`docs/phase_r_harmonization_review.md` § 2 identified the gap and recommended
|
|
||||||
exactly this — *"candidate new `series_breaks.csv` entries (recommend
|
|
||||||
`with_caution` documentation rows, no map action)"*. The follow-through never
|
|
||||||
happened. `SB197`–`SB202` set the precedent, giving the analogous X-flow
|
|
||||||
codes `coverage_restricted` + `with_caution` at 2017; `with_caution` is also
|
|
||||||
what keeps this out of the `joinable = "no"` identity-change rule, which
|
|
||||||
would otherwise oblige a harmonization-map row.
|
|
||||||
|
|
||||||
Verify the break corpus-wide and census-to-census before writing the rows.
|
|
||||||
`cog_balances()` does not wait on this — caveat 3 is covered reader-side by
|
|
||||||
`coverage_window` meanwhile, and the entry simply adds a second, catalogued
|
|
||||||
signpost when it lands.
|
|
||||||
|
|
||||||
**Superseded:** an earlier draft of this spec proposed adding
|
|
||||||
`summary_categories` rows for `X40`/`X41` and treated `recipe=` as blocked.
|
|
||||||
Both were wrong. `X40`/`X41` are deliberately aggregate-only per
|
|
||||||
`phase_r_harmonization_review.md` § 0.2, the dropped harmonization-map rows
|
|
||||||
are the documented § 1 decision, and the recipe path reaches them by design.
|
|
||||||
2. **`cog-api#26`.** Adds `/balances` in all three required places — handler,
|
|
||||||
`param_contract`, and the `plumber.R` route signature. Lands after this.
|
|
||||||
|
|
||||||
**Two contract facts the API must carry forward**, both settled during
|
|
||||||
implementation and easy to get wrong from the outside:
|
|
||||||
|
|
||||||
- `provenance$balance_caveats$coverage_window` is **corpus-scoped, not
|
|
||||||
result-scoped**. It reports the observed year extent of *every* balance
|
|
||||||
subtype in the corpus, not only the subtypes a given query returned — so a
|
|
||||||
`category = "Fund Balances"` query still returns all five windows. That is
|
|
||||||
deliberate: the windows describe what the corpus holds, which is what a
|
|
||||||
consumer needs in order to know what it did *not* ask for. The sibling
|
|
||||||
field `truncated` is the result-scoped one. Documented in
|
|
||||||
`inst/schemas/provenance-v1.json` and mutation-guarded against silent
|
|
||||||
inversion.
|
|
||||||
- `balance_caveats` appears **only** on `cog_balances()` results. It is
|
|
||||||
absent from `cog_spending()`/`cog_revenue()` provenance, and the schema
|
|
||||||
says so — an API layer that assumes it is universal will read `NULL`.
|
|
||||||
3. **`uscogdata/CLAUDE.md` refresh.** Separate commit. It is stale: it claims 7
|
|
||||||
SQL views (there are 21), 181 tests (716), a two-year fixture (four years),
|
|
||||||
and a "never inline SQL" rule the verb layer does not follow.
|
|
||||||
@@ -6,33 +6,6 @@ fixture_corpus_path <- function() {
|
|||||||
if (nzchar(p)) paste0(p, "/") else ""
|
if (nzchar(p)) paste0(p, "/") else ""
|
||||||
}
|
}
|
||||||
|
|
||||||
# Path to a file in the SOURCE tree (README.md, man/*.Rd, vignettes/*.Rmd),
|
|
||||||
# or "" when it isn't there.
|
|
||||||
#
|
|
||||||
# Tests that assert on documentation content have to read the sources, and the
|
|
||||||
# sources only exist when the suite runs from a checkout. Under R CMD check the
|
|
||||||
# suite runs from the INSTALLED package, where man/ and vignettes/ are not
|
|
||||||
# shipped and `../../README.md` does not resolve -- so those tests must skip
|
|
||||||
# rather than error. CI runs testthat::test_local() from the checkout BEFORE
|
|
||||||
# rcmdcheck, so the assertions are still enforced on every push; this only
|
|
||||||
# stops them from failing a context that structurally cannot satisfy them.
|
|
||||||
source_tree_path <- function(...) {
|
|
||||||
p <- testthat::test_path("..", "..", ...)
|
|
||||||
if (file.exists(p)) p else ""
|
|
||||||
}
|
|
||||||
|
|
||||||
# Skip unless every named source file is present (see source_tree_path()).
|
|
||||||
skip_if_no_source_tree <- function(...) {
|
|
||||||
paths <- vapply(list(...), function(rel) do.call(source_tree_path, as.list(rel)),
|
|
||||||
character(1))
|
|
||||||
missing <- vapply(paths, function(p) !nzchar(p), logical(1))
|
|
||||||
testthat::skip_if(
|
|
||||||
any(missing),
|
|
||||||
"package source tree not available (running against the installed package)"
|
|
||||||
)
|
|
||||||
invisible(paths)
|
|
||||||
}
|
|
||||||
|
|
||||||
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
# Skip a test if no corpus is reachable (bundled fixture or explicit remote URL).
|
||||||
skip_if_no_corpus <- function() {
|
skip_if_no_corpus <- function() {
|
||||||
p <- fixture_corpus_path()
|
p <- fixture_corpus_path()
|
||||||
@@ -84,106 +57,3 @@ with_doctored_schema_version <- function(version, code) {
|
|||||||
}, add = TRUE)
|
}, add = TRUE)
|
||||||
force(code)
|
force(code)
|
||||||
}
|
}
|
||||||
|
|
||||||
# Copy the bundled fixture to a temp dir with representation.parquet and
|
|
||||||
# code_set.parquet removed (and dropped from the manifest's metadata list),
|
|
||||||
# then run `code` against it. Models a corpus published BEFORE sparsification:
|
|
||||||
# schema_version is left alone deliberately, because it was never bumped for
|
|
||||||
# that change -- the pre-sparsification fixture this package shipped until
|
|
||||||
# 2026-07-30 was schema v6 and carried neither table. Presence in the manifest
|
|
||||||
# is therefore the only honest signal, and this helper is what proves the
|
|
||||||
# package keys off it rather than off the version number.
|
|
||||||
with_corpus_missing_representation <- function(code) {
|
|
||||||
src <- fixture_corpus_path()
|
|
||||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
|
||||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
|
||||||
|
|
||||||
dropped <- c("representation.parquet", "code_set.parquet")
|
|
||||||
file.remove(file.path(tmp, "data", dropped))
|
|
||||||
|
|
||||||
manifest_path <- file.path(tmp, "manifest.json")
|
|
||||||
m <- jsonlite::fromJSON(manifest_path, simplifyVector = FALSE)
|
|
||||||
m$files$metadata <- Filter(
|
|
||||||
function(f) !basename(f$path) %in% dropped, m$files$metadata
|
|
||||||
)
|
|
||||||
writeLines(
|
|
||||||
jsonlite::toJSON(m, auto_unbox = TRUE, pretty = TRUE, null = "null"),
|
|
||||||
manifest_path
|
|
||||||
)
|
|
||||||
|
|
||||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
|
||||||
on.exit({
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
|
||||||
}, add = TRUE)
|
|
||||||
force(code)
|
|
||||||
}
|
|
||||||
|
|
||||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
|
||||||
# rewritten to drop every M/L (intergovernmental) row, then run `code`
|
|
||||||
# against it with a clean session (mirrors with_fixture_corpus()/
|
|
||||||
# with_doctored_schema_version()). Models a real pre-cog_pipeline-PR#59
|
|
||||||
# corpus: the 66 M/L category rows shipped with NO schema_version bump (see
|
|
||||||
# C2 in the expenditure-concept review), so schema_version is left
|
|
||||||
# untouched here -- only the category data itself is rolled back.
|
|
||||||
with_corpus_missing_ig_categories <- function(code) {
|
|
||||||
src <- fixture_corpus_path()
|
|
||||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
|
||||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
|
||||||
|
|
||||||
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
|
|
||||||
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
|
|
||||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(write_con, sprintf(
|
|
||||||
"COPY (SELECT * FROM read_parquet(%s) WHERE LEFT(item_code, 1) NOT IN ('M', 'L'))
|
|
||||||
TO %s (FORMAT PARQUET)",
|
|
||||||
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
|
|
||||||
))
|
|
||||||
file.remove(cats_path)
|
|
||||||
file.rename(filtered_path, cats_path)
|
|
||||||
|
|
||||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
|
||||||
on.exit({
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
|
||||||
}, add = TRUE)
|
|
||||||
force(code)
|
|
||||||
}
|
|
||||||
|
|
||||||
# Copy the bundled fixture to a temp dir with summary_categories.parquet
|
|
||||||
# rewritten to DROP the balance_subtype column, then run `code` against it.
|
|
||||||
# Models a corpus published before cog_pipeline #76/#77. schema_version is
|
|
||||||
# left untouched deliberately: that change shipped without a version bump, so
|
|
||||||
# column presence is the only honest signal -- this helper is what proves the
|
|
||||||
# package keys off it. Mirrors with_corpus_missing_ig_categories().
|
|
||||||
with_corpus_missing_balance_subtype <- function(code) {
|
|
||||||
src <- fixture_corpus_path()
|
|
||||||
tmp <- withr::local_tempdir(.local_envir = parent.frame())
|
|
||||||
file.copy(list.files(src, full.names = TRUE), tmp, recursive = TRUE)
|
|
||||||
|
|
||||||
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
|
|
||||||
filtered_path <- file.path(tmp, "data", "summary_categories_filtered.parquet")
|
|
||||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(write_con, sprintf(
|
|
||||||
"COPY (SELECT * EXCLUDE (balance_subtype) FROM read_parquet(%s))
|
|
||||||
TO %s (FORMAT PARQUET)",
|
|
||||||
uscogdata:::.sql_lit_chr(cats_path), uscogdata:::.sql_lit_chr(filtered_path)
|
|
||||||
))
|
|
||||||
file.remove(cats_path)
|
|
||||||
file.rename(filtered_path, cats_path)
|
|
||||||
|
|
||||||
old_url <- Sys.getenv("USCOGDATA_URL", unset = NA)
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
Sys.setenv(USCOGDATA_URL = paste0(tmp, "/"))
|
|
||||||
on.exit({
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
if (is.na(old_url)) Sys.unsetenv("USCOGDATA_URL") else Sys.setenv(USCOGDATA_URL = old_url)
|
|
||||||
}, add = TRUE)
|
|
||||||
force(code)
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -1,44 +0,0 @@
|
|||||||
# Helper for the Madison-walkthrough finding tests (uscogdata #11-#16).
|
|
||||||
#
|
|
||||||
# Those tests all assert something about what a `cog_*` verb includes or
|
|
||||||
# excludes. The expected amounts must therefore come from the RAW corpus, never
|
|
||||||
# from the verb under test: verifying an absence through the filter that creates
|
|
||||||
# it proves nothing. `wt_raw_*()` opens its own DuckDB connection straight onto
|
|
||||||
# the corpus's `long` parquet partitions, bypassing uscogdata's SQL views (and
|
|
||||||
# therefore its `flow_prefixes` filtering) entirely.
|
|
||||||
|
|
||||||
wt_corpus_glob <- function() {
|
|
||||||
url <- Sys.getenv("USCOGDATA_URL")
|
|
||||||
if (!nzchar(url)) testthat::skip("USCOGDATA_URL is not set")
|
|
||||||
paste0(sub("/$", "", url), "/data/long/**/*.parquet")
|
|
||||||
}
|
|
||||||
|
|
||||||
wt_raw_query <- function(sql) {
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbGetQuery(con, sql)
|
|
||||||
}
|
|
||||||
|
|
||||||
# Sum of `amt` (in $1,000s, as the corpus stores it) for one government-year,
|
|
||||||
# restricted either to an explicit set of item codes or to a set of first-letter
|
|
||||||
# prefixes. Aggregate rows are excluded, matching every published verb.
|
|
||||||
wt_raw_amt <- function(govid, year, codes = NULL, prefixes = NULL) {
|
|
||||||
stopifnot(xor(is.null(codes), is.null(prefixes)))
|
|
||||||
filter_sql <- if (!is.null(codes)) {
|
|
||||||
paste0("item_code IN (", paste0("'", codes, "'", collapse = ", "), ")")
|
|
||||||
} else {
|
|
||||||
paste0("LEFT(item_code, 1) IN (", paste0("'", prefixes, "'", collapse = ", "), ")")
|
|
||||||
}
|
|
||||||
out <- wt_raw_query(paste0(
|
|
||||||
"SELECT COALESCE(SUM(amt), 0) AS amt FROM read_parquet('", wt_corpus_glob(), "') ",
|
|
||||||
"WHERE canonical_govid = '", govid, "' AND year = ", year,
|
|
||||||
" AND NOT is_aggregate AND ", filter_sql
|
|
||||||
))
|
|
||||||
out$amt[[1]]
|
|
||||||
}
|
|
||||||
|
|
||||||
# The item codes a verb reports having summed, flattened out of the
|
|
||||||
# comma-separated `codes_included` column.
|
|
||||||
wt_codes_included <- function(df) {
|
|
||||||
sort(unique(trimws(unlist(strsplit(stats::na.omit(df$codes_included), ",")))))
|
|
||||||
}
|
|
||||||
@@ -1,52 +0,0 @@
|
|||||||
# Madison walkthrough audit -- finding F-004. Tracked as uscogdata#15.
|
|
||||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# The raw Census files report thousands of dollars; this package multiplies by
|
|
||||||
# 1000 and returns full US dollars. That is the friendlier choice and is not
|
|
||||||
# wrong -- but cog_explorer's CLAUDE.md states "All raw `amt` values are in
|
|
||||||
# $1,000s", so a reader who applies that rule to amt_nominal overstates every
|
|
||||||
# figure by 1000x, and gets a plausible-looking number rather than an obvious
|
|
||||||
# error. The audit rates this the highest-consequence definitional gap it found.
|
|
||||||
#
|
|
||||||
# Deliberately NOT asserted here: man/cog_spending.Rd and man/cog_revenue.Rd,
|
|
||||||
# which ALREADY carry the statement in their @return sections (verified
|
|
||||||
# 2026-07-29), as does cog-api's data-dictionary.md (since 2b71b41). The gap is
|
|
||||||
# in the surfaces a reader meets first and in cog_explorer's own conventions
|
|
||||||
# doc -- see uscogdata#15 for the full surface-by-surface table and for the two
|
|
||||||
# secondary tasks (cog_explorer/CLAUDE.md, which has no git remote, and
|
|
||||||
# cog-api's llms.txt, which is silent on units).
|
|
||||||
|
|
||||||
test_that("returned amounts are documented as full US dollars where readers meet the package", {
|
|
||||||
|
|
||||||
# README and vignettes ship only in the source tree, not in the installed
|
|
||||||
# package, so these assertions cannot run under R CMD check -- CI's earlier
|
|
||||||
# testthat::test_local() step is what enforces them. See
|
|
||||||
# skip_if_no_source_tree() in helper-fixture.R.
|
|
||||||
docs <- skip_if_no_source_tree(
|
|
||||||
"README.md",
|
|
||||||
c("vignettes", "total-spending.Rmd"),
|
|
||||||
c("vignettes", "population-denominators.Rmd")
|
|
||||||
)
|
|
||||||
|
|
||||||
says_units <- function(path) {
|
|
||||||
txt <- paste(readLines(path, warn = FALSE), collapse = " ")
|
|
||||||
grepl("full US dollars|full U\\.S\\. dollars", txt, ignore.case = TRUE) &&
|
|
||||||
grepl("\\$1,000s|thousands of dollars", txt, ignore.case = TRUE)
|
|
||||||
}
|
|
||||||
|
|
||||||
for (path in docs) expect_true(says_units(path))
|
|
||||||
|
|
||||||
# Pin the documented claim to the actual behaviour, so the two cannot drift.
|
|
||||||
# The expected raw amount is read straight from the corpus's parquet
|
|
||||||
# partitions -- never through cog_spending(), which is the thing being
|
|
||||||
# described. Madison FY2020: E/F/G = 623,347 ($1,000s) -> $623,347,000.
|
|
||||||
raw_thousands <- wt_raw_amt("552025209777", 2020L, prefixes = c("E", "F", "G"))
|
|
||||||
expect_equal(raw_thousands, 623347)
|
|
||||||
|
|
||||||
returned <- cog_spending(govid = "552025209777", years = 2020L)
|
|
||||||
expect_equal(sum(returned$amt_nominal), raw_thousands * 1000)
|
|
||||||
|
|
||||||
units <- attr(returned, "provenance")$transformations$units_conversion
|
|
||||||
expect_true(units$applied)
|
|
||||||
expect_equal(units$multiplier, 1000)
|
|
||||||
})
|
|
||||||
@@ -1,539 +0,0 @@
|
|||||||
test_that("balance views register and carry only balance codes", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- cog_open()
|
|
||||||
on.exit(cog_close())
|
|
||||||
|
|
||||||
views <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT table_name FROM information_schema.tables
|
|
||||||
WHERE table_schema = 'main' AND table_type = 'VIEW'"
|
|
||||||
)$table_name
|
|
||||||
expect_true(all(c("balance_long", "balance_annotated") %in% views))
|
|
||||||
|
|
||||||
# Every item_code in balance_long is a category_type = 'balance' member.
|
|
||||||
leak <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n FROM balance_long
|
|
||||||
WHERE item_code NOT IN (
|
|
||||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance')"
|
|
||||||
)$n
|
|
||||||
expect_identical(as.integer(leak), 0L)
|
|
||||||
|
|
||||||
# balance_annotated exposes the subtype column the verb groups on.
|
|
||||||
cols <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT column_name FROM information_schema.columns
|
|
||||||
WHERE table_name = 'balance_annotated'"
|
|
||||||
)$column_name
|
|
||||||
expect_true(all(c("category", "category_type", "balance_subtype") %in% cols))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("inst/sql/26-balance_long.sql enforces NOT is_aggregate (real SQL text, synthetic parquet)", {
|
|
||||||
# Every category_type = 'balance' item_code in the bundled fixture has
|
|
||||||
# is_aggregate = FALSE for every row of every year -- there is no real row
|
|
||||||
# that would be excluded ONLY by the `AND NOT is_aggregate` predicate. An
|
|
||||||
# assertion against the live fixture (`WHERE is_aggregate` returns 0) is
|
|
||||||
# therefore vacuous: it passes identically whether or not the view's
|
|
||||||
# predicate is present. As with the 22-/23- and 24-/25- tests above, this
|
|
||||||
# reads the real inst/sql/26-balance_long.sql text off disk and executes it
|
|
||||||
# -- plus its 10-long.sql / 11-summary_categories.sql dependencies -- against
|
|
||||||
# a synthetic hive-partitioned parquet tree that DOES contain an aggregate
|
|
||||||
# row under a real balance item_code (W01), so a regression that drops the
|
|
||||||
# predicate changes which rows survive.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
|
|
||||||
tmp <- withr::local_tempdir()
|
|
||||||
part_dir <- file.path(tmp, "data", "long", "year=2004")
|
|
||||||
dir.create(part_dir, recursive = TRUE)
|
|
||||||
part_path <- file.path(part_dir, "part-0.parquet")
|
|
||||||
|
|
||||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('bal-A', 'W01', 100, false), -- control: ordinary balance row, survives
|
|
||||||
('bal-B', 'W01', 999999, true) -- excluded ONLY by `NOT is_aggregate`
|
|
||||||
) AS t(canonical_govid, item_code, amt, is_aggregate)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(part_path)))
|
|
||||||
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('W01', 'Fund Balances', 'balance', NULL, NULL, 'general')
|
|
||||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype, balance_subtype)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
|
||||||
|
|
||||||
sql_dir <- system.file("sql", package = "uscogdata")
|
|
||||||
.read_view_sql <- function(filename) {
|
|
||||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
|
||||||
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
|
|
||||||
}
|
|
||||||
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("26-balance_long.sql"))
|
|
||||||
|
|
||||||
rows <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT canonical_govid, item_code, amt FROM balance_long ORDER BY canonical_govid"
|
|
||||||
)
|
|
||||||
expect_equal(nrow(rows), 1L)
|
|
||||||
expect_equal(rows$canonical_govid, "bal-A")
|
|
||||||
expect_equal(rows$amt, 100)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("balance views are skipped on a corpus without balance_subtype", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_corpus_missing_balance_subtype({
|
|
||||||
con <- cog_open()
|
|
||||||
on.exit(cog_close())
|
|
||||||
views <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT table_name FROM information_schema.tables
|
|
||||||
WHERE table_schema = 'main' AND table_type = 'VIEW'"
|
|
||||||
)$table_name
|
|
||||||
# Registration must SKIP them, not error -- an older corpus stays usable.
|
|
||||||
expect_false(any(c("balance_long", "balance_annotated") %in% views))
|
|
||||||
expect_true("revenue_long" %in% views)
|
|
||||||
|
|
||||||
# ...and calling the verb on such a corpus must hit
|
|
||||||
# .require_balance_support()'s curated abort (spec § Testing: "Gating"),
|
|
||||||
# not a DuckDB binder error naming a view that was never registered.
|
|
||||||
# Asserted on the CLASS: removing the guard still errors, so a bare
|
|
||||||
# expect_error() would pass on the regression.
|
|
||||||
expect_error(
|
|
||||||
cog_balances("550000227544", 2019),
|
|
||||||
class = "uscogdata_no_balance_support"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_balances returns holdings for a government that has them", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", 2019)
|
|
||||||
expect_s3_class(r, "tbl_df")
|
|
||||||
expect_true(nrow(r) > 0L)
|
|
||||||
expect_true(all(c("year", "canonical_govid", "gov_name", "balance_subtype",
|
|
||||||
"category", "amt_nominal") %in% names(r)))
|
|
||||||
expect_identical(sort(unique(r$category)),
|
|
||||||
c("Fund Balances", "Insurance Trust Balances"))
|
|
||||||
expect_false(is.null(attr(r, "provenance")))
|
|
||||||
expect_identical(attr(r, "provenance")$verb, "cog_balances")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that('category = "Fund Balances" is exactly the general family', {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
|
||||||
expect_identical(unique(r$balance_subtype), "general")
|
|
||||||
codes <- sort(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
|
|
||||||
expect_identical(codes, c("W01", "W31", "W61"))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("no flow code can reach cog_balances", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
|
|
||||||
got <- unique(unlist(strsplit(paste(r$codes_included, collapse = ","), ",")))
|
|
||||||
|
|
||||||
# The expected set is read from the RAW corpus, never from the verb --
|
|
||||||
# verifying an absence through the filter that creates it proves nothing.
|
|
||||||
# A fresh, direct DuckDB connection against the raw parquet files (never
|
|
||||||
# cog_open()'s session, never balance_long/balance_annotated) reads
|
|
||||||
# parquet natively -- no arrow dependency needed (see CLAUDE.md).
|
|
||||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
|
||||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
|
||||||
balance_codes <- DBI::dbGetQuery(con2, sprintf(
|
|
||||||
"SELECT item_code FROM read_parquet(%s) WHERE category_type = 'balance'",
|
|
||||||
uscogdata:::.sql_lit_chr(cats_path)
|
|
||||||
))$item_code
|
|
||||||
|
|
||||||
expect_true(length(got) > 0L)
|
|
||||||
expect_true(all(got %in% balance_codes))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("every balance_subtype maps to exactly one category", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
# Dropping the `subtype` argument is only safe while this tree holds. If the
|
|
||||||
# pipeline ever gives a balance subtype a second category, `category` becomes
|
|
||||||
# a lossy filter -- fail HERE rather than in a user's analysis. Read via a
|
|
||||||
# fresh direct DuckDB connection against the raw parquet file, not through
|
|
||||||
# any registered view.
|
|
||||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
|
||||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
|
||||||
b <- DBI::dbGetQuery(con2, sprintf(
|
|
||||||
"SELECT category, balance_subtype FROM read_parquet(%s) WHERE category_type = 'balance'",
|
|
||||||
uscogdata:::.sql_lit_chr(cats_path)
|
|
||||||
))
|
|
||||||
per_subtype <- tapply(b$category, b$balance_subtype,
|
|
||||||
function(x) length(unique(x)))
|
|
||||||
expect_true(all(per_subtype == 1L))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_balances records found + missing govids in provenance", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
suppressMessages(
|
|
||||||
r <- cog_balances(c("550000227544", "XXXINVALID"), 2019)
|
|
||||||
)
|
|
||||||
prov <- attr(r, "provenance")
|
|
||||||
expect_equal(sort(prov$scope$govids_found), "550000227544")
|
|
||||||
expect_equal(sort(prov$scope$govids_missing), "XXXINVALID")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("per_capita divides holdings by population", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
plain <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
|
||||||
pc <- cog_balances("550000227544", 2019, category = "Fund Balances",
|
|
||||||
per_capita = TRUE)
|
|
||||||
expect_true("amt_per_capita_nominal" %in% names(pc))
|
|
||||||
expect_true("pop_source" %in% names(pc))
|
|
||||||
expect_identical(pc$amt_nominal, plain$amt_nominal)
|
|
||||||
|
|
||||||
# Assert against the denominator read from the corpus, NOT against a
|
|
||||||
# quantity derived from amt_per_capita_nominal itself -- dividing the
|
|
||||||
# column back out would be tautological and would pass on any value.
|
|
||||||
pop <- DBI::dbGetQuery(cog_open(), sprintf(
|
|
||||||
"SELECT population FROM gov_population_yearly
|
|
||||||
WHERE canonical_govid = %s AND year = 2019",
|
|
||||||
uscogdata:::.sql_lit_chr("550000227544")
|
|
||||||
))$population
|
|
||||||
expect_length(pop, 1L)
|
|
||||||
expect_equal(pc$amt_per_capita_nominal, pc$amt_nominal / pop,
|
|
||||||
tolerance = 1e-8)
|
|
||||||
|
|
||||||
prov <- attr(pc, "provenance")
|
|
||||||
expect_true(prov$transformations$per_capita$applied)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("adjust_to_year adds real dollars", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
|
||||||
adjust_to_year = 2020)
|
|
||||||
expect_true("amt_real" %in% names(r))
|
|
||||||
# 2012 dollars inflated to 2020 must exceed nominal.
|
|
||||||
expect_true(all(r$amt_real > r$amt_nominal))
|
|
||||||
prov <- attr(r, "provenance")
|
|
||||||
expect_true(prov$transformations$inflation$applied)
|
|
||||||
expect_identical(prov$transformations$inflation$base_year, 2020L)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("per_capita and adjust_to_year compose", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
|
||||||
per_capita = TRUE, adjust_to_year = 2020)
|
|
||||||
expect_true("amt_per_capita_real" %in% names(r))
|
|
||||||
# The per-capita column must be deflated by the SAME factor as the level
|
|
||||||
# column -- this is what the ordering at R/balances.R:101-103 guarantees.
|
|
||||||
# .attach_real_dollars() silently no-ops on the per-capita leg when
|
|
||||||
# amt_per_capita_nominal does not exist yet (R/spending.R:664), so
|
|
||||||
# reversing those two calls drops this column with no error at all.
|
|
||||||
expect_equal(r$amt_per_capita_real / r$amt_per_capita_nominal,
|
|
||||||
r$amt_real / r$amt_nominal, tolerance = 1e-8)
|
|
||||||
|
|
||||||
# And the documented condition is a conjunction: adjust_to_year ALONE
|
|
||||||
# must not produce amt_per_capita_real (pins the @return wording).
|
|
||||||
r2 <- cog_balances("550000227544", 2012, category = "Fund Balances",
|
|
||||||
adjust_to_year = 2020)
|
|
||||||
expect_true("amt_real" %in% names(r2))
|
|
||||||
expect_false("amt_per_capita_real" %in% names(r2))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
# --- input validation ------------------------------------------------------
|
|
||||||
|
|
||||||
test_that("cog_balances validates its inputs like the money verbs", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
G <- "550000227544"
|
|
||||||
# Pinned to the message, not bare expect_error(): every one of these
|
|
||||||
# already produces *some* error or *some* quiet wrong answer today --
|
|
||||||
# years = integer(0) leaks `Parser Error ... AND year IN ()` with the
|
|
||||||
# generated SQL, recipe = c("a","b") throws "the condition has length > 1",
|
|
||||||
# and the govid/category cases return 0 rows with no error at all.
|
|
||||||
expect_error(cog_balances(G, integer(0)), "non-empty integer vector")
|
|
||||||
expect_error(cog_balances(character(0), 2019), "non-empty character vector")
|
|
||||||
expect_error(cog_balances(G, 2019, category = 5), "must be character or NULL")
|
|
||||||
expect_error(cog_balances(G, 2019, recipe = c("a", "b")),
|
|
||||||
"length-1 character string")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("validation runs after govid coercion, so a data frame still works", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# .validate_verb_inputs() asserts is.character(govid); it must therefore
|
|
||||||
# run AFTER .coerce_govid_input(), never before, or the documented
|
|
||||||
# data-frame input (cog_gov_search() output) would abort.
|
|
||||||
df <- data.frame(canonical_govid = "550000227544", stringsAsFactors = FALSE)
|
|
||||||
r <- suppressMessages(cog_balances(df, 2019))
|
|
||||||
expect_true(nrow(r) > 0L)
|
|
||||||
expect_identical(unique(r$canonical_govid), "550000227544")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("recipe and category are mutually exclusive", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
expect_error(
|
|
||||||
cog_balances("550000227544", c(2011, 2012),
|
|
||||||
category = "Fund Balances",
|
|
||||||
recipe = "cash_securities_z77_wide"),
|
|
||||||
class = "uscogdata_recipe_category_conflict"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
# --- recipe = : the wide-era holdings bridge -------------------------------
|
|
||||||
|
|
||||||
test_that("recipe bridges the wide era into the modern one", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", c(2011, 2012),
|
|
||||||
recipe = "cash_securities_z77_wide")
|
|
||||||
# .run_recipe()'s SQL returns `long.year` as a DOUBLE (a corpus-wide trait,
|
|
||||||
# not specific to this recipe -- see the money-verb recipe tests, which
|
|
||||||
# only ever assert on it with expect_equal), so compare numerically rather
|
|
||||||
# than with expect_identical()'s type-strict comparison.
|
|
||||||
expect_equal(sort(r$year), c(2011, 2012))
|
|
||||||
|
|
||||||
# The 2011 leg can ONLY come from X40, which is 100% is_aggregate = TRUE
|
|
||||||
# and therefore invisible to balance_long. If the recipe path ever starts
|
|
||||||
# filtering aggregates, a 45-year series silently truncates to five --
|
|
||||||
# this is the regression guard for phase_r_harmonization_review.md § 0.2.
|
|
||||||
codes <- attr(r, "provenance")$codes_summed$observed
|
|
||||||
expect_true("X40" %in% codes)
|
|
||||||
expect_true("Z77" %in% codes)
|
|
||||||
expect_true(all(r$amt_nominal > 0))
|
|
||||||
|
|
||||||
prov <- attr(r, "provenance")
|
|
||||||
expect_identical(prov$recipe$recipe_id, "cash_securities_z77_wide")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the FY2002 book-to-market basis change is disclosed on the recipe path", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# 2002 is in the year vector deliberately, and must stay -- do not
|
|
||||||
# "simplify" this back to c(2011, 2012).
|
|
||||||
#
|
|
||||||
# .build_series_break_refs() (R/series_breaks.R, shared with every verb)
|
|
||||||
# matches breaks with `break_year BETWEEN min(years) AND max(years)`, and
|
|
||||||
# SB195's break_year is 2002. A c(2011, 2012) span never crosses the
|
|
||||||
# FY2002 book -> market change -- that whole span sits after it, on one
|
|
||||||
# consistent basis -- so NOT disclosing SB195 there is correct behaviour,
|
|
||||||
# not a gap (same reasoning as the "a request that never crosses the
|
|
||||||
# boundary is not affected by it" comment on .build_corpus_break_refs()).
|
|
||||||
#
|
|
||||||
# The property actually worth testing is: a recipe query that observes
|
|
||||||
# X40 AND spans FY2002 discloses SB195. This fixture has no 2002
|
|
||||||
# partition data for X40/Z77 (confirmed: only 2011/2012/2019/2020
|
|
||||||
# partitions exist), so including 2002 in `years` widens the
|
|
||||||
# break-matching window without changing which rows the recipe join
|
|
||||||
# returns -- verified empirically: r$year below is exactly {2011, 2012}
|
|
||||||
# whether or not 2002 is in the request (see task-4-report.md).
|
|
||||||
# Removing 2002 would silently turn this back into the non-crossing case
|
|
||||||
# above and destroy the test's purpose.
|
|
||||||
r <- cog_balances("550000227544", c(2002, 2011, 2012),
|
|
||||||
recipe = "cash_securities_z77_wide")
|
|
||||||
expect_equal(sort(r$year), c(2011, 2012))
|
|
||||||
refs <- attr(r, "provenance")$series_break_refs
|
|
||||||
# SB195 sits on fin_code X40; it can only fire where X40 is observed,
|
|
||||||
# which is exactly the recipe path.
|
|
||||||
expect_true("SB195" %in% refs)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the second holdings bridge works too", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# X41 -> Z78, the securities counterpart. Wisconsin carries X41 in 2011
|
|
||||||
# and Z78 in 2012, so both legs are exercised.
|
|
||||||
r <- cog_balances("550000227544", c(2011, 2012),
|
|
||||||
recipe = "cash_securities_z78_wide")
|
|
||||||
codes <- attr(r, "provenance")$codes_summed$observed
|
|
||||||
expect_true(all(c("X41", "Z78") %in% codes))
|
|
||||||
expect_equal(sort(r$year), c(2011, 2012))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("an unknown recipe id is rejected", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# Asserted on the CLASS .validate_recipe_id() sets (R/recipes.R:84).
|
|
||||||
# Without it the test is non-discriminating: deleting the validation call
|
|
||||||
# leaves .recipe_components() returning 0 rows and comps$label[[1]]
|
|
||||||
# throwing "subscript out of bounds", which a bare expect_error() accepts
|
|
||||||
# while the user loses the curated "valid recipe ids are ..." message.
|
|
||||||
expect_error(
|
|
||||||
cog_balances("550000227544", 2019, recipe = "no_such_recipe"),
|
|
||||||
class = "uscogdata_unknown_recipe"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
# --- balance_caveats: GAAP disclosure + measured coverage windows ----------
|
|
||||||
|
|
||||||
test_that("balance_caveats is always present and flags the GAAP distinction", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", 2019)
|
|
||||||
cav <- attr(r, "provenance")$balance_caveats
|
|
||||||
expect_false(is.null(cav))
|
|
||||||
expect_true(cav$not_gaap)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("coverage_window is computed from the corpus, not hardcoded", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_balances("550000227544", c(2011, 2012, 2019, 2020))
|
|
||||||
cav <- attr(r, "provenance")$balance_caveats
|
|
||||||
|
|
||||||
# Read the "general" family's true year extent independently, via a
|
|
||||||
# fresh DuckDB connection against the raw parquet files (never through
|
|
||||||
# balance_long/.balance_caveats() itself, and never via arrow -- this
|
|
||||||
# package reads parquet through DuckDB only, see CLAUDE.md). Replicates
|
|
||||||
# the same predicates 26-balance_long.sql applies (category_type =
|
|
||||||
# 'balance', NOT is_aggregate) so this is a faithful, independent
|
|
||||||
# measurement rather than a re-statement of the view under test.
|
|
||||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
|
||||||
long_glob <- file.path(fixture_corpus_path(), "data", "long", "**", "*.parquet")
|
|
||||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
|
||||||
obs <- DBI::dbGetQuery(con2, sprintf(
|
|
||||||
"SELECT MIN(l.year) AS y0, MAX(l.year) AS y1
|
|
||||||
FROM read_parquet(%s, hive_partitioning = true) l
|
|
||||||
JOIN read_parquet(%s) c USING (item_code)
|
|
||||||
WHERE c.balance_subtype = 'general' AND NOT l.is_aggregate",
|
|
||||||
uscogdata:::.sql_lit_chr(long_glob), uscogdata:::.sql_lit_chr(cats_path)
|
|
||||||
))
|
|
||||||
|
|
||||||
expect_identical(as.integer(cav$coverage_window$general),
|
|
||||||
c(as.integer(obs$y0), as.integer(obs$y1)))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("coverage_window covers every corpus subtype, not just observed ones", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# Deliberate contract (provenance-v1.json): the window block is corpus-
|
|
||||||
# scoped so a caller can ask "is there a family I missed?", while
|
|
||||||
# `truncated` is the observed-scoped field. A single-category query must
|
|
||||||
# therefore still report every balance family in the mounted corpus.
|
|
||||||
r <- cog_balances("550000227544", 2019, category = "Fund Balances")
|
|
||||||
expect_identical(unique(r$balance_subtype), "general")
|
|
||||||
|
|
||||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
|
||||||
cats_path <- file.path(fixture_corpus_path(), "data", "summary_categories.parquet")
|
|
||||||
all_subtypes <- DBI::dbGetQuery(con2, sprintf(
|
|
||||||
"SELECT DISTINCT balance_subtype FROM read_parquet(%s)
|
|
||||||
WHERE balance_subtype IS NOT NULL",
|
|
||||||
uscogdata:::.sql_lit_chr(cats_path)
|
|
||||||
))$balance_subtype
|
|
||||||
|
|
||||||
cav <- attr(r, "provenance")$balance_caveats
|
|
||||||
expect_setequal(names(cav$coverage_window), all_subtypes)
|
|
||||||
expect_true(length(all_subtypes) > 1L)
|
|
||||||
# ...while `truncated` stays scoped to what this query actually observed.
|
|
||||||
expect_true(all(cav$truncated %in% unique(r$balance_subtype)))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the corpus-constant coverage windows are memoised per session", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# The windows query has no govid/year predicate: its answer depends only
|
|
||||||
# on which corpus is mounted, so re-running the full balance_long scan on
|
|
||||||
# every call is pure waste (35% of verb runtime on the fixture). Same
|
|
||||||
# memoise-and-invalidate pattern as .uscogdata_env$manifest.
|
|
||||||
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
|
|
||||||
suppressMessages(cog_balances("550000227544", 2019))
|
|
||||||
memo <- uscogdata:::.uscogdata_env$balance_coverage_windows
|
|
||||||
expect_false(is.null(memo))
|
|
||||||
expect_true("general" %in% names(memo))
|
|
||||||
|
|
||||||
uscogdata:::cog_close()
|
|
||||||
expect_null(uscogdata:::.uscogdata_env$balance_coverage_windows)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("a request past a family's coverage window is flagged", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# employee_retirement (X21/X30/X47/Z77/Z78) genuinely ends at FY2016 in
|
|
||||||
# the LIVE corpus -- Census moved employee retirement reporting to the
|
|
||||||
# Annual Survey of Public Pensions after that year. This bundled FIXTURE
|
|
||||||
# doesn't carry 2013-2016 at all (only 2011/2012/2019/2020 are present),
|
|
||||||
# so the family's *observed* max here is 2012, not 2016. Either way the
|
|
||||||
# requested span (2012, 2019) reaches past what the family covers in
|
|
||||||
# THIS corpus, which is what makes .balance_caveats() flag it -- the
|
|
||||||
# assertion below is about the fixture's measured window, not the FY2016
|
|
||||||
# live-corpus cutoff.
|
|
||||||
r <- cog_balances("550000227544", c(2012, 2019))
|
|
||||||
cav <- attr(r, "provenance")$balance_caveats
|
|
||||||
expect_true("employee_retirement" %in% cav$truncated)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the provenance schema documents balance_caveats", {
|
|
||||||
sch <- jsonlite::fromJSON(
|
|
||||||
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
|
|
||||||
simplifyVector = FALSE
|
|
||||||
)
|
|
||||||
expect_true("balance_caveats" %in% names(sch$properties))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_explain surfaces the balance caveats", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# Asserted on the RENDERED text, not on prov$balance_caveats: the field
|
|
||||||
# is already covered above, and the once-per-session cli_inform() means
|
|
||||||
# cog_explain() is the only surface a caller who missed (or suppressed)
|
|
||||||
# the first message can still audit.
|
|
||||||
r <- suppressMessages(cog_balances("550000227544", c(2012, 2019)))
|
|
||||||
# Both streams: cli routes most of its output through conditions that
|
|
||||||
# land on stderr, so a stdout-only capture would be empty (the pattern
|
|
||||||
# used throughout test-explain.R).
|
|
||||||
out <- paste(c(capture.output(cog_explain(r)),
|
|
||||||
capture.output(cog_explain(r), type = "message")),
|
|
||||||
collapse = "\n")
|
|
||||||
expect_match(out, "GAAP")
|
|
||||||
expect_match(out, "employee_retirement")
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_explain on a money-verb result has no balance caveat section", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- suppressMessages(cog_spending("550000227544", 2019))
|
|
||||||
out <- paste(c(capture.output(cog_explain(r)),
|
|
||||||
capture.output(cog_explain(r), type = "message")),
|
|
||||||
collapse = "\n")
|
|
||||||
# Guard against the capture itself being vacuous: the section must be
|
|
||||||
# absent from output that demonstrably contains the rest of the report.
|
|
||||||
expect_match(out, "Data vintage")
|
|
||||||
expect_false(grepl("GAAP", out))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the caveat message fires once per session", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
expect_message(cog_balances("550000227544", 2019), "not.*GAAP")
|
|
||||||
expect_no_message(cog_balances("550000227544", 2020))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
@@ -8,55 +8,22 @@ test_that("cog_categories returns all categories grouped by subtype", {
|
|||||||
expect_gt(nrow(r), 10L)
|
expect_gt(nrow(r), 10L)
|
||||||
# corpus preserves Census-native "expenditure" vocabulary; the API takes
|
# corpus preserves Census-native "expenditure" vocabulary; the API takes
|
||||||
# "spending" as a friendlier alias.
|
# "spending" as a friendlier alias.
|
||||||
#
|
expect_setequal(unique(r$category_type), c("expenditure", "revenue"))
|
||||||
# `balance` joined as a third category_type with the cash-and-security
|
|
||||||
# holding codes (pipeline#76). `cog_categories()` is a CATALOGUE verb, not a
|
|
||||||
# money verb, so it surfaces every category_type the corpus carries -- the
|
|
||||||
# stock/flow guard belongs on cog_spending()/cog_revenue(), which must never
|
|
||||||
# return a balance row.
|
|
||||||
expect_setequal(unique(r$category_type),
|
|
||||||
c("expenditure", "revenue", "balance"))
|
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("cog_categories(type = 'spending') returns only expenditure rows", {
|
test_that("cog_categories(type = 'spending') returns only expenditure rows", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
r <- cog_categories(type = "spending")
|
r <- cog_categories(type = "spending")
|
||||||
expect_true(all(r$category_type == "expenditure"))
|
expect_true(all(r$category_type == "expenditure"))
|
||||||
# "assistance" (the J-prefix aid/benefit codes) joined the vocabulary with
|
expect_true(all(r$subtype %in% c("operations", "capital")))
|
||||||
# the crosswalk completion in cog_pipeline#60/#65 -- every flow code
|
|
||||||
# carrying dollars now maps to a category.
|
|
||||||
# `interest` (I89, I91-I94) and `insurance_benefits` (Y05/Y06/Y14/Y53)
|
|
||||||
# joined with the I/Q/Y flow batch -- the last two characters of Census's
|
|
||||||
# expenditure taxonomy. `interest` is what makes the three-concept model
|
|
||||||
# computable: primary = direct minus debt service.
|
|
||||||
expect_true(all(r$subtype %in%
|
|
||||||
c("operations", "capital", "intergovernmental", "assistance",
|
|
||||||
"interest", "insurance_benefits")))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_categories surfaces the intergovernmental spending subtype", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
r <- cog_categories(type = "spending")
|
|
||||||
expect_true("intergovernmental" %in% r$subtype)
|
|
||||||
# IG rows reuse the existing functional categories -- they add a subtype,
|
|
||||||
# not new category values.
|
|
||||||
ig_cats <- sort(unique(r$category[r$subtype == "intergovernmental"]))
|
|
||||||
direct_cats <- sort(unique(r$category[r$subtype != "intergovernmental"]))
|
|
||||||
expect_true(all(ig_cats %in% c(direct_cats, "Other Education")))
|
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("cog_categories(type = 'revenue') returns only revenue rows", {
|
test_that("cog_categories(type = 'revenue') returns only revenue rows", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
r <- cog_categories(type = "revenue")
|
r <- cog_categories(type = "revenue")
|
||||||
expect_true(all(r$category_type == "revenue"))
|
expect_true(all(r$category_type == "revenue"))
|
||||||
# The four non-general subtypes are deliberately NOT own_source: Census's
|
|
||||||
# General Revenue excludes insurance trust (Y01 alone is $1.31T corpus-wide,
|
|
||||||
# plus the employee-retirement X codes), utility (A91-A94) and liquor store
|
|
||||||
# (A90) revenue by definition, which is what makes both of its published
|
|
||||||
# revenue concepts computable -- see `revenue_concept` in `?cog_revenue`.
|
|
||||||
expect_true(all(r$subtype %in%
|
expect_true(all(r$subtype %in%
|
||||||
c("own_source", "federal", "state", "local_aid",
|
c("own_source", "federal", "state", "local_aid")))
|
||||||
"insurance_trust", "utility", "liquor_store")))
|
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("cog_categories(pattern = ...) filters case-insensitively", {
|
test_that("cog_categories(pattern = ...) filters case-insensitively", {
|
||||||
|
|||||||
@@ -1,193 +0,0 @@
|
|||||||
# tests/testthat/test-complete.R
|
|
||||||
#
|
|
||||||
# uscogdata#18. The published corpus no longer stores the wide era's explicit
|
|
||||||
# zeros (cog_pipeline#64, series break SB194), so absence means two different
|
|
||||||
# things:
|
|
||||||
#
|
|
||||||
# <= FY2011 (dense_source) : cell absent => Census published $0
|
|
||||||
# >= FY2012 (sparse_source): cell absent => not reported, unknown
|
|
||||||
#
|
|
||||||
# `complete = TRUE` fills the requested grid from `code_set` and stamps every
|
|
||||||
# row's `value_source` so the two are distinguishable. Expected row sets here
|
|
||||||
# are built from the corpus parquet directly, never from the verb under test --
|
|
||||||
# verifying what a filter does through that same filter proves nothing.
|
|
||||||
|
|
||||||
# The (subtype, category) cells that SHOULD exist for one government-year:
|
|
||||||
# every code in force for that government's type, mapped through
|
|
||||||
# summary_categories, matching the verb's crosswalk subtype scope (the
|
|
||||||
# default concept, `primary`, is operations/capital/assistance -- see
|
|
||||||
# uscogdata#11) and excluding aggregate-flagged codes (which
|
|
||||||
# spending_long/revenue_long drop).
|
|
||||||
raw_expected_cells <- function(govid, year, subtypes, subtype_col) {
|
|
||||||
fx <- sub("/$", "", Sys.getenv("USCOGDATA_URL"))
|
|
||||||
q <- function(f) sprintf("read_parquet('%s/data/%s')", fx, f)
|
|
||||||
wt_raw_query(sprintf(
|
|
||||||
"SELECT DISTINCT c.%s AS subtype, c.category
|
|
||||||
FROM %s cs
|
|
||||||
JOIN %s x ON x.govs_type = cs.type
|
|
||||||
JOIN %s c ON c.item_code = cs.item_code
|
|
||||||
WHERE x.canonical_govid = '%s'
|
|
||||||
AND cs.year = %d
|
|
||||||
AND NOT cs.is_aggregate
|
|
||||||
AND c.category IS NOT NULL
|
|
||||||
AND c.%s IN (%s)",
|
|
||||||
subtype_col, q("code_set.parquet"), q("canonical_fips_xwalk.parquet"),
|
|
||||||
q("summary_categories.parquet"), govid, year,
|
|
||||||
subtype_col, paste0("'", subtypes, "'", collapse = ",")
|
|
||||||
))
|
|
||||||
}
|
|
||||||
|
|
||||||
# The default expenditure concept's subtype scope, mirrored from
|
|
||||||
# R/spending.R's .spend_subtypes_primary.
|
|
||||||
primary_subtypes <- c("operations", "capital", "assistance")
|
|
||||||
|
|
||||||
test_that("complete = FALSE is the default and changes nothing", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
plain <- cog_spending("121011212191", 2011L)
|
|
||||||
explicit <- cog_spending("121011212191", 2011L, complete = FALSE)
|
|
||||||
expect_equal(nrow(plain), nrow(explicit))
|
|
||||||
expect_false("value_source" %in% names(plain))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE round-trips a dense-source year to the pre-sparsification cells", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# FY2011 is dense_source: before sparsification this government carried a
|
|
||||||
# row for every code in force, most of them $0. complete = TRUE must
|
|
||||||
# reproduce that cell set exactly.
|
|
||||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
|
||||||
expected <- raw_expected_cells("121011212191", 2011L,
|
|
||||||
primary_subtypes, "spend_subtype")
|
|
||||||
|
|
||||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
|
||||||
expect_setequal(key(r$spend_subtype, r$category),
|
|
||||||
key(expected$subtype, expected$category))
|
|
||||||
expect_gt(nrow(expected), 0L)
|
|
||||||
|
|
||||||
# Every filled cell in a dense-source year is a Census-published $0 --
|
|
||||||
# never "unknown", which is what the modern era's absences mean.
|
|
||||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
|
||||||
expect_true(all(r$amt_nominal[r$value_source == "census_zero"] == 0))
|
|
||||||
expect_true(all(r$amt_nominal[r$value_source == "reported"] != 0))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE preserves the reported rows and their amounts exactly", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
plain <- cog_spending("121011212191", 2011L)
|
|
||||||
full <- cog_spending("121011212191", 2011L, complete = TRUE)
|
|
||||||
|
|
||||||
# Filling adds rows; it must never alter or drop one.
|
|
||||||
expect_gt(nrow(full), nrow(plain))
|
|
||||||
reported <- full[full$value_source == "reported", ]
|
|
||||||
expect_equal(nrow(reported), nrow(plain))
|
|
||||||
expect_equal(sum(reported$amt_nominal), sum(plain$amt_nominal))
|
|
||||||
# ... and the total is unchanged, because every added cell is $0.
|
|
||||||
expect_equal(sum(full$amt_nominal, na.rm = TRUE), sum(plain$amt_nominal))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("a sparse-source year's absences are unknown, not zero", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# FY2019 is sparse_source: an absent cell means the government did not
|
|
||||||
# report, which is NOT a zero. Filling those with 0 would invent data --
|
|
||||||
# the exact error the representation contract exists to prevent.
|
|
||||||
r <- cog_spending("121011212191", 2019L, complete = TRUE)
|
|
||||||
filled <- r[r$value_source != "reported", ]
|
|
||||||
expect_gt(nrow(filled), 0L)
|
|
||||||
expect_true(all(filled$value_source == "not_reported"))
|
|
||||||
expect_true(all(is.na(filled$amt_nominal)))
|
|
||||||
expect_false(any(r$value_source == "census_zero"))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the fill is scoped to each government's own type", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# Filling against the union of all types would invent cells for codes a
|
|
||||||
# county can never report. Every filled category must be one that
|
|
||||||
# code_set puts in force for type 1 (county) specifically.
|
|
||||||
r <- cog_spending("121011212191", 2011L, complete = TRUE)
|
|
||||||
county_cells <- raw_expected_cells("121011212191", 2011L,
|
|
||||||
primary_subtypes, "spend_subtype")
|
|
||||||
expect_true(all(r$category %in% county_cells$category))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE respects the category filter", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_spending("121011212191", 2011L, category = "Police",
|
|
||||||
complete = TRUE)
|
|
||||||
expect_true(all(r$category == "Police"))
|
|
||||||
expect_true("value_source" %in% names(r))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_revenue() completes on its own flow", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_revenue("121011212191", 2011L, complete = TRUE)
|
|
||||||
expected <- raw_expected_cells("121011212191", 2011L,
|
|
||||||
c("own_source", "federal", "state", "local_aid"),
|
|
||||||
"revenue_subtype")
|
|
||||||
key <- function(sub, cat) paste(sub, cat, sep = "|")
|
|
||||||
expect_setequal(key(r$revenue_subtype, r$category),
|
|
||||||
key(expected$subtype, expected$category))
|
|
||||||
expect_setequal(unique(r$value_source), c("reported", "census_zero"))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("provenance records the completion and its absence rule", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
prov <- attr(cog_spending("121011212191", 2011L, complete = TRUE),
|
|
||||||
"provenance")
|
|
||||||
expect_true(prov$completion$applied)
|
|
||||||
expect_equal(prov$completion$absence_means$`2011`, "census_zero")
|
|
||||||
expect_gt(prov$completion$rows_filled, 0L)
|
|
||||||
|
|
||||||
off <- attr(cog_spending("121011212191", 2011L), "provenance")
|
|
||||||
expect_false(off$completion$applied)
|
|
||||||
expect_equal(off$completion$rows_filled, 0L)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE is refused where the fill would be guesswork", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# A recipe defines its own component codes and does not go through
|
|
||||||
# summary_categories at all, so there is no grid to fill from.
|
|
||||||
expect_error(
|
|
||||||
cog_spending("121011212191", 2011L, recipe = "corrections_combined",
|
|
||||||
complete = TRUE),
|
|
||||||
class = "uscogdata_complete_unsupported"
|
|
||||||
)
|
|
||||||
# The intergovernmental leg keeps aggregate rows by design
|
|
||||||
# (inst/sql/24-ig_long.sql), so its grid is not code_set's grid.
|
|
||||||
expect_error(
|
|
||||||
cog_spending("121011212191", 2011L, expenditure_concept = "total",
|
|
||||||
complete = TRUE),
|
|
||||||
class = "uscogdata_complete_unsupported"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("complete = TRUE aborts on a corpus with no representation contract", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
# A corpus published before sparsification carries neither table, so there
|
|
||||||
# is nothing to fill from and no rule saying what an absence means. That
|
|
||||||
# must abort rather than guess.
|
|
||||||
with_corpus_missing_representation({
|
|
||||||
expect_error(
|
|
||||||
cog_spending("121011212191", 2011L, complete = TRUE),
|
|
||||||
class = "uscogdata_representation_unavailable"
|
|
||||||
)
|
|
||||||
# ... while an ordinary query on the same corpus still works.
|
|
||||||
expect_gt(nrow(cog_spending("121011212191", 2011L)), 0L)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
@@ -26,41 +26,3 @@ test_that(".resolve_cache_dir falls back to R_user_dir", {
|
|||||||
})
|
})
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
|
||||||
# Trailing-slash normalization (uscogdata #3 follow-up).
|
|
||||||
#
|
|
||||||
# EVERY consumer builds paths by concatenation: paste0(url, "manifest.json")
|
|
||||||
# (manifest.R), paste0(url, e$path) (mirror.R), and the parquet glob in
|
|
||||||
# views.R. mirror.R:104 even comments 'url ends in "/"' -- an assumption the
|
|
||||||
# package documents and relies on but never enforced.
|
|
||||||
#
|
|
||||||
# A URL missing its trailing slash therefore fails SILENTLY and confusingly:
|
|
||||||
# HTTPS -> ".../downloadmanifest.json" -> the host answers with an HTML 404
|
|
||||||
# page -> the jsonlite lexical error that issue #3 reported;
|
|
||||||
# local -> ".../corpusdata/long/**/*.parquet" -> DuckDB "No files found".
|
|
||||||
# Neither message points at the real cause. Normalize once, at resolution.
|
|
||||||
# ---------------------------------------------------------------------------
|
|
||||||
|
|
||||||
test_that(".resolve_url appends a missing trailing slash", {
|
|
||||||
withr::local_envvar(USCOGDATA_URL = "https://example.org/s/TOKEN/download")
|
|
||||||
expect_equal(.resolve_url(), "https://example.org/s/TOKEN/download/")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".resolve_url leaves an existing trailing slash alone", {
|
|
||||||
withr::local_envvar(USCOGDATA_URL = "https://example.org/s/TOKEN/download/")
|
|
||||||
expect_equal(.resolve_url(), "https://example.org/s/TOKEN/download/")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".resolve_url normalizes a local path without a trailing slash", {
|
|
||||||
withr::local_envvar(USCOGDATA_URL = "/tmp/corpus")
|
|
||||||
expect_equal(.resolve_url(), "/tmp/corpus/")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".resolve_url does not invent a slash for an empty setting", {
|
|
||||||
# An unset/empty URL must stay empty so the "not configured" guard in
|
|
||||||
# manifest.R still fires, rather than degrading into a bare "/" root.
|
|
||||||
withr::local_envvar(USCOGDATA_URL = "")
|
|
||||||
withr::local_options(uscogdata.url = "")
|
|
||||||
expect_equal(.resolve_url(), "")
|
|
||||||
})
|
|
||||||
|
|||||||
@@ -1,94 +0,0 @@
|
|||||||
# tests/testthat/test-corpus-breaks.R
|
|
||||||
#
|
|
||||||
# uscogdata#19. Four catalogued series breaks carry fin_code = "ALL" -- they
|
|
||||||
# are caveats about the corpus itself rather than about one item code:
|
|
||||||
#
|
|
||||||
# SB085 1977 dollar precision across the 1976/1977 boundary
|
|
||||||
# SB087 2002 imputation exclusion FY2002-2006
|
|
||||||
# SB194 2012 dense -> sparse representation change
|
|
||||||
# SB086 2017 government ID scheme change
|
|
||||||
#
|
|
||||||
# .build_series_break_refs() matches `fin_code IN (<codes in the result>)`,
|
|
||||||
# and no row's item_code is ever the literal "ALL", so none of them could
|
|
||||||
# ever reach a user. They now travel in their own provenance field,
|
|
||||||
# `corpus_break_refs`, which keeps them distinguishable from the
|
|
||||||
# code-specific `series_break_refs` (an ALL caveat qualifies the whole
|
|
||||||
# result, not one series).
|
|
||||||
|
|
||||||
test_that("corpus_break_refs surfaces an ALL-scoped break the year range spans", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# SB194 sits at FY2012 -- the dense/sparse boundary. A query spanning
|
|
||||||
# 2011 -> 2012 straddles it, and this is the case cog_pipeline#64's
|
|
||||||
# DoD 4 intended to reach users.
|
|
||||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
|
||||||
prov <- attr(r, "provenance")
|
|
||||||
expect_true("SB194" %in% prov$corpus_break_refs)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("corpus_break_refs stays empty when no ALL break falls in the range", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# 2019-2020 spans no catalogued corpus-wide break.
|
|
||||||
r <- cog_spending("121011212191", 2019:2020, "Police")
|
|
||||||
expect_equal(attr(r, "provenance")$corpus_break_refs, character(0))
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("corpus_break_refs and series_break_refs stay disjoint", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
|
||||||
prov <- attr(r, "provenance")
|
|
||||||
expect_type(prov$series_break_refs, "character")
|
|
||||||
expect_type(prov$corpus_break_refs, "character")
|
|
||||||
# An ALL caveat must never masquerade as a break in a specific series.
|
|
||||||
expect_length(intersect(prov$series_break_refs, prov$corpus_break_refs), 0L)
|
|
||||||
expect_false("SB194" %in% prov$series_break_refs)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".build_corpus_break_refs matches on the break_year window alone", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
con <- cog_open()
|
|
||||||
on.exit(cog_close())
|
|
||||||
|
|
||||||
# SB085's boundary is 1976/1977, outside the fixture's partitions -- the
|
|
||||||
# series_breaks table is a full cross-vintage registry, so the matching
|
|
||||||
# logic is testable there even though no long partition covers it.
|
|
||||||
expect_true("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
|
||||||
con, years = 1975:1980, schema_version = 6L
|
|
||||||
))
|
|
||||||
# ... and does not fire for a range that misses it, unlike a filter keyed
|
|
||||||
# on the era rather than the boundary.
|
|
||||||
expect_false("SB085" %in% uscogdata:::.build_corpus_break_refs(
|
|
||||||
con, years = 1978:1980, schema_version = 6L
|
|
||||||
))
|
|
||||||
|
|
||||||
# Unlike code-specific refs, these do not depend on which codes a result
|
|
||||||
# happens to contain -- that dependency is the whole defect.
|
|
||||||
expect_setequal(
|
|
||||||
uscogdata:::.build_corpus_break_refs(con, years = 2001:2003, schema_version = 6L),
|
|
||||||
"SB087"
|
|
||||||
)
|
|
||||||
|
|
||||||
# Gated on schema_version >= 5: series_breaks_pq is not registered below it.
|
|
||||||
expect_equal(
|
|
||||||
uscogdata:::.build_corpus_break_refs(con, years = 2011:2012, schema_version = 4L),
|
|
||||||
character(0)
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_explain() prints corpus-wide caveats under their own heading", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
r <- cog_spending("121011212191", 2011:2012, "Police")
|
|
||||||
out <- paste(c(
|
|
||||||
capture.output(cog_explain(r)),
|
|
||||||
capture.output(cog_explain(r), type = "message")
|
|
||||||
), collapse = "\n")
|
|
||||||
expect_match(out, "Corpus-wide caveats", fixed = TRUE)
|
|
||||||
expect_match(out, "SB194", fixed = TRUE)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
@@ -1,103 +0,0 @@
|
|||||||
# Madison walkthrough audit -- findings F-020 and F-023. Tracked as uscogdata#13.
|
|
||||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# The owner's settled design (2026-07-28): a `coverage` argument on
|
|
||||||
# cog_geographic_rollup(), cog_find_peers()/cog_peer_compare() and their
|
|
||||||
# cog-api equivalents --
|
|
||||||
# "all" every unit that reported that year (today's behaviour, DEFAULT)
|
|
||||||
# "census" census years only (years ending 2 or 7)
|
|
||||||
# "consistent" only units reporting in every requested year (balanced panel)
|
|
||||||
# -- PLUS always-on coverage metadata on every result regardless of mode:
|
|
||||||
# n_units_reporting, n_units_expected, is_census_year.
|
|
||||||
#
|
|
||||||
# Motivating principle: using these verbs correctly must not require the user to
|
|
||||||
# know that the Census of Governments is a complete census only in years ending
|
|
||||||
# in 2 and 7.
|
|
||||||
#
|
|
||||||
# The helper below accepts that metadata either as columns on the returned
|
|
||||||
# tibble or as a per-year table in provenance$coverage -- the design fixes the
|
|
||||||
# three field names and that they reach the caller, not the container.
|
|
||||||
|
|
||||||
wt_coverage <- function(x) {
|
|
||||||
prov <- attr(x, "provenance")
|
|
||||||
cov <- prov$coverage
|
|
||||||
if (is.null(cov)) {
|
|
||||||
needed <- c("year", "n_units_reporting", "n_units_expected", "is_census_year")
|
|
||||||
expect_true(all(needed %in% names(x)))
|
|
||||||
cov <- unique(x[, needed])
|
|
||||||
}
|
|
||||||
cov[order(cov$year), ]
|
|
||||||
}
|
|
||||||
|
|
||||||
test_that("multi-government aggregates disclose reporting coverage on every result", {
|
|
||||||
|
|
||||||
# -- F-020: geographic rollups -------------------------------------------
|
|
||||||
# Wisconsin's city/village universe is 608 governments. On the bundled
|
|
||||||
# fixture, FY2012 (a census year) has 597 of them reporting while FY2019 and
|
|
||||||
# FY2020 (sample years) have 112 and 114 -- an 18%-98% swing that today's
|
|
||||||
# return value says nothing about. Counts cross-checked against the raw
|
|
||||||
# corpus, not through cog_geographic_rollup(), which is under test.
|
|
||||||
wi <- cog_gov_search(name = NULL, state = "WI", type = "city")
|
|
||||||
expect_equal(nrow(wi), 608L)
|
|
||||||
|
|
||||||
roll <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
|
|
||||||
category = NULL, years = c(2011L, 2012L, 2019L, 2020L))
|
|
||||||
cov <- wt_coverage(roll)
|
|
||||||
|
|
||||||
expect_equal(cov$n_units_expected, rep(608L, 4L))
|
|
||||||
expect_equal(cov$n_units_reporting, c(152L, 597L, 112L, 114L))
|
|
||||||
expect_equal(cov$is_census_year, c(FALSE, TRUE, FALSE, FALSE))
|
|
||||||
|
|
||||||
# Cross-check against the raw partitions, scoped to the SAME universe the
|
|
||||||
# rollup was given -- the 608 govids above. Scoping instead on the long
|
|
||||||
# table's own `type`/`fips_state` asks a different question and answers 595:
|
|
||||||
# VERNON VILLAGE and WAUKESHA VILLAGE carry type = 3 there (their as-of-year
|
|
||||||
# identity, when they were townships) while the xwalk lists them as
|
|
||||||
# govs_type = 2 (their present identity, as villages). Schema v6 made the
|
|
||||||
# long table's geography present-harmonized and moved as-of-year to the
|
|
||||||
# *_asof columns, but `type` still reads as-of-year -- see .validate_schema()
|
|
||||||
# in R/manifest.R. n_units_reporting counts against the requested universe,
|
|
||||||
# so 597 is the number that answers "how many of the governments I asked
|
|
||||||
# about reported".
|
|
||||||
raw_2012 <- wt_raw_query(paste0(
|
|
||||||
"SELECT COUNT(DISTINCT canonical_govid) n FROM read_parquet('", wt_corpus_glob(), "') ",
|
|
||||||
"WHERE year = 2012 AND LEFT(item_code, 1) IN ('E','F','G') AND NOT is_aggregate ",
|
|
||||||
"AND canonical_govid IN (",
|
|
||||||
paste0("'", wi$canonical_govid, "'", collapse = ","), ")"))
|
|
||||||
expect_equal(cov$n_units_reporting[cov$year == 2012], as.integer(raw_2012$n[[1]]))
|
|
||||||
|
|
||||||
# -- F-023: peer cohorts --------------------------------------------------
|
|
||||||
# CHILTON CITY, WI (ACS population 4,017): a 15-peer cohort fixed at FY2012
|
|
||||||
# reports 15 of 15 in FY2012 and only 3 of 15 in FY2019 and FY2020. Nothing
|
|
||||||
# in cog_peer_compare()'s return distinguishes those years today.
|
|
||||||
chilton <- "552015177095"
|
|
||||||
peers <- cog_find_peers(chilton, year = 2012L, max_peers = 15L)
|
|
||||||
expect_equal(nrow(peers), 15L)
|
|
||||||
|
|
||||||
cmp <- cog_peer_compare(target_govid = chilton, peers = peers, category = NULL,
|
|
||||||
years = c(2012L, 2019L, 2020L), per_capita = TRUE)
|
|
||||||
cov_peers <- wt_coverage(cmp)
|
|
||||||
expect_equal(cov_peers$n_units_expected, rep(15L, 3L))
|
|
||||||
expect_equal(cov_peers$n_units_reporting, c(15L, 3L, 3L))
|
|
||||||
expect_equal(cov_peers$is_census_year, c(TRUE, FALSE, FALSE))
|
|
||||||
|
|
||||||
# -- the three coverage modes --------------------------------------------
|
|
||||||
expect_equal(attr(cog_peer_compare(target_govid = chilton, peers = peers,
|
|
||||||
category = NULL, years = c(2012L, 2019L, 2020L),
|
|
||||||
per_capita = TRUE),
|
|
||||||
"provenance")$coverage_mode, "all") # unchanged default
|
|
||||||
|
|
||||||
consistent <- cog_peer_compare(target_govid = chilton, peers = peers,
|
|
||||||
category = NULL, years = c(2012L, 2019L, 2020L),
|
|
||||||
per_capita = TRUE, coverage = "consistent")
|
|
||||||
n_by_year <- tapply(consistent$canonical_govid[consistent$role == "peer"],
|
|
||||||
consistent$year[consistent$role == "peer"],
|
|
||||||
function(g) length(unique(g)))
|
|
||||||
expect_equal(unname(as.integer(n_by_year)), c(3L, 3L, 3L)) # balanced panel
|
|
||||||
|
|
||||||
census_only <- cog_geographic_rollup(govids = list(city = wi$canonical_govid),
|
|
||||||
category = NULL,
|
|
||||||
years = c(2011L, 2012L, 2019L, 2020L),
|
|
||||||
coverage = "census")
|
|
||||||
expect_equal(sort(unique(census_only$year)), 2012)
|
|
||||||
})
|
|
||||||
@@ -1,553 +0,0 @@
|
|||||||
test_that("the corpus contains no K-prefix rows, so the Direct leg omits K", {
|
|
||||||
con <- .ensure_session()
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n FROM long WHERE LEFT(item_code, 1) = 'K'")$n
|
|
||||||
expect_equal(n, 0)
|
|
||||||
|
|
||||||
sql_files <- c("20-spending_long.sql", "22-spending_long_harmonized.sql")
|
|
||||||
for (f in sql_files) {
|
|
||||||
txt <- paste(readLines(system.file("sql", f, package = "uscogdata")),
|
|
||||||
collapse = " ")
|
|
||||||
expect_false(grepl("'K'", txt, fixed = TRUE),
|
|
||||||
label = paste(f, "must not reference the inert K prefix"))
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("expenditure_concept defaults to primary; direct matches it on a pure operations/capital category", {
|
|
||||||
gov <- "010000226085" # Alabama state government
|
|
||||||
base <- cog_spending(gov, years = 2019, category = "Police")
|
|
||||||
expect_equal(attr(base, "provenance")$expenditure_concept, "primary")
|
|
||||||
# Police maps only to operations/capital codes (E62/F62/G62), so the
|
|
||||||
# direct concept's extra subtypes (interest, insurance_benefits) cannot
|
|
||||||
# contribute and the two concepts must agree exactly here.
|
|
||||||
expl <- cog_spending(gov, years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
expect_equal(base$amt_nominal, expl$amt_nominal)
|
|
||||||
expect_false("intergovernmental" %in% base$spend_subtype)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("expenditure_concept = 'total' adds an intergovernmental subtype", {
|
|
||||||
gov <- "010000226085"
|
|
||||||
d <- cog_spending(gov, years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
t <- cog_spending(gov, years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
expect_true("intergovernmental" %in% t$spend_subtype)
|
|
||||||
# Direct rows are untouched; Total only ever ADDS. Use %in% rather than
|
|
||||||
# != : a category = NULL result can contain a NULL-subtype group (codes
|
|
||||||
# with no summary_categories row, e.g. E16/E21/E85/F16/F85/G16/G21/G85),
|
|
||||||
# and `NA != "intergovernmental"` is NA, not TRUE, which would silently
|
|
||||||
# smuggle an all-NA phantom row into dt.
|
|
||||||
dt <- t[!(t$spend_subtype %in% "intergovernmental"), ]
|
|
||||||
expect_equal(sort(dt$amt_nominal), sort(d$amt_nominal))
|
|
||||||
expect_gt(sum(t$amt_nominal), sum(d$amt_nominal))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("legacy-era Total does not collapse to Direct (the is_aggregate trap)", {
|
|
||||||
# In the wide era the IG dollars live almost entirely on aggregate-flagged
|
|
||||||
# rows. A Total leg that inherited the Direct leg's NOT is_aggregate filter
|
|
||||||
# would silently return Total == Direct here.
|
|
||||||
gov <- "010000226085"
|
|
||||||
d <- cog_spending(gov, years = 2011, category = "Education K-12",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
t <- cog_spending(gov, years = 2011, category = "Education K-12",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
expect_true("intergovernmental" %in% t$spend_subtype)
|
|
||||||
ig <- sum(t$amt_nominal[t$spend_subtype == "intergovernmental"])
|
|
||||||
expect_gt(ig, 0)
|
|
||||||
expect_gt(sum(t$amt_nominal), sum(d$amt_nominal))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the IG leg never includes the L-- family total", {
|
|
||||||
con <- .ensure_session()
|
|
||||||
codes <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT item_code FROM ig_long")$item_code
|
|
||||||
expect_false(any(grepl("--$", codes)))
|
|
||||||
# Q joined the IG family with the crosswalk-membership rewrite
|
|
||||||
# (uscogdata#11 / F-017: Q11/Q12/Q18 are state payments to school systems).
|
|
||||||
expect_true(all(substr(codes, 1, 1) %in% c("M", "L", "Q")))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("expenditure_concept rejects unknown values", {
|
|
||||||
expect_error(
|
|
||||||
cog_spending("010000226085", years = 2019, expenditure_concept = "gross"),
|
|
||||||
class = "rlang_error"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("total composes with basis = 'raw' and basis = 'harmonized'", {
|
|
||||||
gov <- "010000226085"
|
|
||||||
h <- cog_spending(gov, years = 2011, category = "Education K-12",
|
|
||||||
expenditure_concept = "total", basis = "harmonized")
|
|
||||||
r <- cog_spending(gov, years = 2011, category = "Education K-12",
|
|
||||||
expenditure_concept = "total", basis = "raw")
|
|
||||||
ig_h <- sum(h$amt_nominal[h$spend_subtype == "intergovernmental"])
|
|
||||||
ig_r <- sum(r$amt_nominal[r$spend_subtype == "intergovernmental"])
|
|
||||||
# The only IG harmonization rule is M38 -> M36 (year-disjoint), so the IG
|
|
||||||
# total must agree between bases even though the code labels may differ.
|
|
||||||
expect_equal(ig_h, ig_r)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("recipe = and expenditure_concept = 'total' together aborts", {
|
|
||||||
expect_error(
|
|
||||||
cog_spending("121011212191", 2020L, recipe = "corrections_combined",
|
|
||||||
expenditure_concept = "total"),
|
|
||||||
class = "uscogdata_recipe_concept_conflict"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("aggregate-sourced IG dollars are flagged aggregate_fallback = TRUE (bool_or, not bool_and)", {
|
|
||||||
# Regression test: .build_verb_sql() originally used bool_and(is_aggregate)
|
|
||||||
# for aggregate_fallback, which is correct for the Direct leg (a group can
|
|
||||||
# never mix aggregate and non-aggregate rows there -- spending_long filters
|
|
||||||
# NOT is_aggregate) but wrong for the IG leg. The wide era is dense -- every
|
|
||||||
# government has a $0 row for every code in a family -- so a $0 leaf sits in
|
|
||||||
# the same (year, gov, subtype, category) group as the real aggregate row
|
|
||||||
# and flips bool_and() to FALSE. Measured: AL state 2011 had $5,740,775,000
|
|
||||||
# of aggregate-sourced IG dollars (Corrections $31,358,000 + Education K-12
|
|
||||||
# $5,152,385,000 + General Government $557,032,000) reporting
|
|
||||||
# aggregate_fallback = FALSE under bool_and(), with the only TRUE row being
|
|
||||||
# Transit Utilities at $0. bool_or() reports all of them correctly.
|
|
||||||
gov <- "010000226085"
|
|
||||||
t <- cog_spending(gov, years = 2011, category = "Education K-12",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
ig <- t[t$spend_subtype == "intergovernmental", ]
|
|
||||||
expect_equal(nrow(ig), 1L)
|
|
||||||
expect_true(ig$aggregate_fallback)
|
|
||||||
expect_true(nzchar(ig$notes))
|
|
||||||
expect_match(ig$notes, "Aggregate fallback applied", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("legacy aggregate IG codes are year-disjoint from their modern leaf components", {
|
|
||||||
# The safety of ig_long's deliberate omission of `NOT is_aggregate` (see
|
|
||||||
# inst/sql/24-ig_long.sql) rests entirely on each legacy code's AGGREGATE
|
|
||||||
# instance being year-disjoint from the modern leaf codes it rolls up --
|
|
||||||
# if a future corpus rebuild ever back-filled a leaf into a year where the
|
|
||||||
# code is still flagged aggregate, `total` would silently double-count and
|
|
||||||
# this suite would still pass. This test fails loudly if that ever
|
|
||||||
# happens.
|
|
||||||
#
|
|
||||||
# Note the invariant is scoped to the AGGREGATE flag, not bare code
|
|
||||||
# presence: M89/L89 do NOT disappear after the wide era the way M47/L47
|
|
||||||
# do -- they continue past 2011 as their OWN independent leaf line item
|
|
||||||
# (is_aggregate = FALSE) alongside M91-93/L91-93, which is fine because a
|
|
||||||
# non-aggregate M89/L89 no longer represents a rollup of those codes.
|
|
||||||
# (Verified in the fixture: M89/L89 are is_aggregate = TRUE only in 2011,
|
|
||||||
# when M91-93/L91-93 don't exist yet; from 2012 on M89/L89 are
|
|
||||||
# is_aggregate = FALSE leaves coexisting with M91-93/L91-93.)
|
|
||||||
#
|
|
||||||
# Pairs are the M/L-prefixed components (this package's ig_long only
|
|
||||||
# covers M/L; other prefixes in the same rollup, e.g. N/O/P/Q/R, fall
|
|
||||||
# outside its domain and are irrelevant here) enumerated in
|
|
||||||
# cog_pipeline's data/wide_to_long_xwalk.csv `full_desc` column (read
|
|
||||||
# once at authoring time, not at test time -- this test stays offline):
|
|
||||||
# M47 "To local governments, total (includes N47, O47, P47, R47, and M94)"
|
|
||||||
# M89 "To local governments, total (incl N89, O89, P89, R89, M91, M92, and M93)"
|
|
||||||
# L47 "To state government (includes L94)"
|
|
||||||
# L89 "To state government (includes L91, L92, and L93)"
|
|
||||||
con <- .ensure_session()
|
|
||||||
pairs <- list(
|
|
||||||
list(aggregate = "M47", components = "M94"),
|
|
||||||
list(aggregate = "M89", components = c("M91", "M92", "M93")),
|
|
||||||
list(aggregate = "L47", components = "L94"),
|
|
||||||
list(aggregate = "L89", components = c("L91", "L92", "L93"))
|
|
||||||
)
|
|
||||||
agg_years_by_code <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT year, item_code FROM ig_long WHERE is_aggregate")
|
|
||||||
codes_by_year <- DBI::dbGetQuery(con, "SELECT DISTINCT year, item_code FROM ig_long")
|
|
||||||
|
|
||||||
for (p in pairs) {
|
|
||||||
agg_years <- agg_years_by_code$year[agg_years_by_code$item_code == p$aggregate]
|
|
||||||
for (yr in agg_years) {
|
|
||||||
codes_yr <- codes_by_year$item_code[codes_by_year$year == yr]
|
|
||||||
has_component <- any(p$components %in% codes_yr)
|
|
||||||
expect_false(
|
|
||||||
has_component,
|
|
||||||
label = sprintf(
|
|
||||||
"year %s has aggregate-flagged %s co-occurring with a modern component (%s)",
|
|
||||||
yr, p$aggregate, paste(p$components, collapse = ",")
|
|
||||||
)
|
|
||||||
)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".verb_spendrev rejects expenditure_concept = 'total' for a non-spending view_base", {
|
|
||||||
# cog_revenue() never exposes expenditure_concept and always resolves it
|
|
||||||
# to the "direct" default, so there is no revenue codepath that reaches
|
|
||||||
# this today -- but .verb_spendrev() is shared, and nothing else stops a
|
|
||||||
# future caller from passing expenditure_concept = "total" alongside
|
|
||||||
# view_base = "revenue_annotated", which would UNION expenditure M/L rows
|
|
||||||
# into a revenue result. Exercise the internal helper directly.
|
|
||||||
expect_error(
|
|
||||||
uscogdata:::.verb_spendrev(
|
|
||||||
verb = "cog_revenue_test", view_base = "revenue_annotated",
|
|
||||||
subtype_col = "revenue_subtype",
|
|
||||||
flow_prefixes = c("T", "A", "U", "B", "C", "D"),
|
|
||||||
call = quote(cog_revenue_test()),
|
|
||||||
govid = "010000226085", years = 2019L, category = NULL,
|
|
||||||
per_capita = FALSE, adjust_to_year = NULL, basis = "raw",
|
|
||||||
recipe = NULL, expenditure_concept = "total"
|
|
||||||
),
|
|
||||||
class = "uscogdata_expenditure_concept_unsupported"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_geographic_rollup refuses expenditure_concept = 'total'", {
|
|
||||||
expect_error(
|
|
||||||
cog_geographic_rollup(
|
|
||||||
govids = list(state = "010000226085"),
|
|
||||||
category = "Police", years = 2019,
|
|
||||||
expenditure_concept = "total"
|
|
||||||
),
|
|
||||||
class = "uscogdata_concept_not_aggregatable"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_peer_compare refuses expenditure_concept = 'total'", {
|
|
||||||
expect_error(
|
|
||||||
cog_peer_compare(
|
|
||||||
target_govid = "010000226085", peers = "010000226085",
|
|
||||||
category = "Police", years = 2019,
|
|
||||||
expenditure_concept = "total"
|
|
||||||
),
|
|
||||||
class = "uscogdata_concept_not_aggregatable"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the refusal message names the fix and the reason", {
|
|
||||||
err <- tryCatch(
|
|
||||||
cog_geographic_rollup(govids = list(state = "010000226085"),
|
|
||||||
category = "Police", years = 2019,
|
|
||||||
expenditure_concept = "total"),
|
|
||||||
condition = function(e) e
|
|
||||||
)
|
|
||||||
msg <- paste(conditionMessage(err), collapse = " ")
|
|
||||||
expect_match(msg, "direct")
|
|
||||||
expect_match(msg, "double-count|double count")
|
|
||||||
expect_match(msg, "cog_geographic_rollup")
|
|
||||||
|
|
||||||
# Test that cog_peer_compare's message names its own function
|
|
||||||
err2 <- tryCatch(
|
|
||||||
cog_peer_compare(target_govid = "010000226085", peers = "010000226085",
|
|
||||||
category = "Police", years = 2019,
|
|
||||||
expenditure_concept = "total"),
|
|
||||||
condition = function(e) e
|
|
||||||
)
|
|
||||||
msg2 <- paste(conditionMessage(err2), collapse = " ")
|
|
||||||
expect_match(msg2, "direct")
|
|
||||||
expect_match(msg2, "double-count|double count")
|
|
||||||
expect_match(msg2, "cog_peer_compare")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("both cross-government verbs still accept the direct default", {
|
|
||||||
expect_no_error(
|
|
||||||
cog_geographic_rollup(govids = list(state = "010000226085"),
|
|
||||||
category = "Police", years = 2019)
|
|
||||||
)
|
|
||||||
expect_no_error(
|
|
||||||
cog_peer_compare(target_govid = "010000226085", peers = "010000226085",
|
|
||||||
category = "Police", years = 2019)
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("provenance always records the expenditure concept", {
|
|
||||||
p <- cog_spending("010000226085", years = 2019, category = "Police")
|
|
||||||
d <- cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
t <- cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
expect_equal(attr(p, "provenance")$expenditure_concept, "primary")
|
|
||||||
expect_equal(attr(d, "provenance")$expenditure_concept, "direct")
|
|
||||||
expect_equal(attr(t, "provenance")$expenditure_concept, "total")
|
|
||||||
# The note explains the non-obvious part: how legacy IG was assembled.
|
|
||||||
expect_true(nzchar(attr(t, "provenance")$expenditure_concept_note))
|
|
||||||
expect_true(is.na(attr(d, "provenance")$expenditure_concept_note) ||
|
|
||||||
!nzchar(attr(d, "provenance")$expenditure_concept_note))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the provenance schema documents expenditure_concept", {
|
|
||||||
sch <- jsonlite::fromJSON(
|
|
||||||
system.file("schemas", "provenance-v1.json", package = "uscogdata"),
|
|
||||||
simplifyVector = FALSE
|
|
||||||
)
|
|
||||||
expect_true("expenditure_concept" %in% names(sch$properties))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("a firing suggestion names the intergovernmental counterpart recipe", {
|
|
||||||
# Corrections has no legacy leaf rows, so the coverage-gap suggestion fires;
|
|
||||||
# corrections_ig_local_combined is its IG counterpart.
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("010000226085", years = c(2005, 2011), category = "Corrections")
|
|
||||||
)
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
expect_gt(length(sugg), 0L)
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
expect_true("corrections_combined" %in% ids)
|
|
||||||
ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id))
|
|
||||||
expect_true("corrections_ig_local_combined" %in% ig)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("no suggestion fires for a healthy query", {
|
|
||||||
r <- cog_spending("010000226085", years = 2019, category = "Police")
|
|
||||||
expect_length(attr(r, "provenance")$suggestions, 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("a mis-scoped cog_spending() call never attaches an M/L counterpart to a revenue-flavored recipe", {
|
|
||||||
# "IG Federal" is a revenue-only category (summary_categories maps it to
|
|
||||||
# B-prefixed component codes only; its recipes are ig_federal_b47_wide /
|
|
||||||
# ig_federal_b89_wide). A cog_spending() call scoped to it returns zero
|
|
||||||
# spending rows for every requested year -- there is no spending
|
|
||||||
# component in this category at all -- so the coverage-gap machinery
|
|
||||||
# fires for real (not hypothetically) even though this isn't the kind of
|
|
||||||
# format-boundary gap the recipe catalog is meant to signpost. This is
|
|
||||||
# exactly the live-corpus risk flagged in review: ig_federal_b47_wide's
|
|
||||||
# own component codes (B47/B94, suffixes {"47","94"}) are an EXACT
|
|
||||||
# suffix-set match for the expenditure recipe ige_local_m47_wide
|
|
||||||
# (M47/M94, same suffixes) -- a coincidence of reused digits, not a real
|
|
||||||
# Direct/Total pairing. The flow-family gate in
|
|
||||||
# .attach_ig_counterparts() must keep ig_recipe_id NULL here.
|
|
||||||
#
|
|
||||||
# Anchored on FL state government, not AL. Coverage is presence-based: a
|
|
||||||
# recipe is only suggested when its component codes have rows for the
|
|
||||||
# requested government-year. AL state's only FY2011 B47 cell was an
|
|
||||||
# explicit zero, which the corpus no longer stores after sparsification
|
|
||||||
# (SB194, cog_pipeline#64), so the recipe stopped being a candidate there.
|
|
||||||
# FL state carries a real FY2011 B47 amount, so this exercises the guard
|
|
||||||
# against a suggestion that genuinely fires.
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("120000226351", years = c(2005, 2011), category = "IG Federal")
|
|
||||||
)
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
expect_gt(length(sugg), 0L)
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
expect_true("ig_federal_b47_wide" %in% ids)
|
|
||||||
ig <- unlist(lapply(sugg, function(s) s$ig_recipe_id))
|
|
||||||
expect_length(ig, 0L)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("C1: 'total' on a legacy aggregate-only family reports the IG-only figure honestly, not as Direct + IG", {
|
|
||||||
# AL state government, Corrections, 2011. Measured pre-fix: 'total'
|
|
||||||
# returned $31,358,000 (the IG leg alone, on an aggregate-flagged M04/M05
|
|
||||||
# row) with 0 suggestions (the surviving IG row made the gap-detection
|
|
||||||
# machinery think the Direct leg was covered) and a note asserting
|
|
||||||
# "Total = Direct + intergovernmental" with no caveat. True Direct (via
|
|
||||||
# recipe = "corrections_combined") is $521,651,000 -- the IG-only figure
|
|
||||||
# is ~6% of it.
|
|
||||||
gov <- "010000226085"
|
|
||||||
|
|
||||||
d <- cog_spending(gov, years = 2011, category = "Corrections",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
expect_equal(nrow(d), 0L)
|
|
||||||
|
|
||||||
t <- suppressMessages(cog_spending(
|
|
||||||
gov, years = 2011, category = "Corrections", expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
expect_equal(nrow(t), 1L)
|
|
||||||
expect_equal(t$spend_subtype, "intergovernmental")
|
|
||||||
expect_equal(t$amt_nominal, 31358000)
|
|
||||||
|
|
||||||
r <- cog_spending(gov, years = 2011, recipe = "corrections_combined")
|
|
||||||
expect_equal(r$amt_nominal, 521651000)
|
|
||||||
|
|
||||||
# C1(a): the recipe hints must fire for "total" exactly as they do for
|
|
||||||
# "direct" -- the surviving IG row must not be mistaken for Direct
|
|
||||||
# coverage.
|
|
||||||
prov <- attr(t, "provenance")
|
|
||||||
expect_gt(length(prov$suggestions), 0L)
|
|
||||||
ids <- vapply(prov$suggestions, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
expect_true("corrections_combined" %in% ids)
|
|
||||||
|
|
||||||
# C1(b): the affected row's notes name a recovering recipe rather than
|
|
||||||
# staying silent, and the provenance carries a flag a downstream consumer
|
|
||||||
# (e.g. cog-api, which passes provenance through verbatim) can test.
|
|
||||||
expect_true(nzchar(t$notes))
|
|
||||||
expect_match(t$notes, "unavailable", fixed = TRUE)
|
|
||||||
expect_match(t$notes, "corrections_combined", fixed = TRUE)
|
|
||||||
expect_true(prov$expenditure_concept_direct_suppressed)
|
|
||||||
|
|
||||||
# The base "Total = Direct + IG" note must NOT stand unqualified when that
|
|
||||||
# arithmetic didn't actually happen for this row.
|
|
||||||
expect_match(prov$expenditure_concept_note, "NOTE", fixed = TRUE)
|
|
||||||
expect_match(prov$expenditure_concept_note,
|
|
||||||
"expenditure_concept_direct_suppressed", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("C1(b): expenditure_concept_direct_suppressed is FALSE when the Direct leg is present", {
|
|
||||||
d <- cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
t <- cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
expect_false(isTRUE(attr(d, "provenance")$expenditure_concept_direct_suppressed))
|
|
||||||
expect_false(isTRUE(attr(t, "provenance")$expenditure_concept_direct_suppressed))
|
|
||||||
expect_false(any(nzchar(t$notes[t$spend_subtype == "intergovernmental"]) &
|
|
||||||
grepl("unavailable", t$notes[t$spend_subtype == "intergovernmental"])))
|
|
||||||
})
|
|
||||||
|
|
||||||
# M/I fix: .detect_direct_suppressed() was equating "no Direct sibling row"
|
|
||||||
# with "Direct was suppressed", but the dominant real cause is a government
|
|
||||||
# that simply has no direct spending in that category -- correct, ordinary
|
|
||||||
# data. The fix gates the flag (and its row note) on a harmonization recipe
|
|
||||||
# ACTUALLY covering that exact (year, canonical_govid, category) triple.
|
|
||||||
|
|
||||||
test_that("M/I: true positive, category supplied explicitly (unchanged behavior)", {
|
|
||||||
al <- "010000226085"
|
|
||||||
t_cat <- suppressMessages(cog_spending(
|
|
||||||
al, years = 2011, category = "Corrections", expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
expect_true(attr(t_cat, "provenance")$expenditure_concept_direct_suppressed)
|
|
||||||
expect_match(t_cat$notes, "corrections_combined", fixed = TRUE)
|
|
||||||
expect_match(t_cat$notes, "unavailable", fixed = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("M/I: true positive, category = NULL now also names the recipe (was the fallback bug)", {
|
|
||||||
# Root bug: .build_suggestions() short-circuits to list() when category is
|
|
||||||
# NULL, so the note previously always hit its "no covering recipe found"
|
|
||||||
# fallback here even though corrections_combined genuinely covers this row.
|
|
||||||
al <- "010000226085"
|
|
||||||
t_null <- suppressMessages(cog_spending(
|
|
||||||
al, years = 2011, category = NULL, expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
corr_row <- t_null[t_null$category %in% "Corrections", ]
|
|
||||||
expect_equal(nrow(corr_row), 1L)
|
|
||||||
expect_true(attr(t_null, "provenance")$expenditure_concept_direct_suppressed)
|
|
||||||
expect_match(corr_row$notes, "corrections_combined", fixed = TRUE)
|
|
||||||
expect_match(corr_row$notes, "unavailable", fixed = TRUE)
|
|
||||||
expect_false(grepl("no covering recipe found", corr_row$notes, fixed = TRUE))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("M/I: false positive -- Virginia Education K-12 FY2019 total is NOT flagged", {
|
|
||||||
# States fund K-12 through school districts, so the Direct leg (E12/F12/
|
|
||||||
# G12) is genuinely, correctly zero -- not suppressed. Must not be flagged
|
|
||||||
# and must carry no suppression note.
|
|
||||||
va <- "510000227542"
|
|
||||||
t_va <- suppressMessages(cog_spending(
|
|
||||||
va, years = 2019, category = "Education K-12", expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
expect_equal(nrow(t_va), 1L)
|
|
||||||
expect_equal(t_va$spend_subtype, "intergovernmental")
|
|
||||||
expect_equal(t_va$amt_nominal, 8028179000)
|
|
||||||
expect_false(isTRUE(attr(t_va, "provenance")$expenditure_concept_direct_suppressed))
|
|
||||||
expect_false(nzchar(t_va$notes) && grepl("unavailable", t_va$notes))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("M/I: false positive by construction -- 'Other Education' has no E/F/G code, never flagged", {
|
|
||||||
# "Other Education" maps only to M21/L21 in summary_categories -- there is
|
|
||||||
# no E/F/G code for it in this corpus at all, so no Direct-recovering
|
|
||||||
# recipe can exist and it must never be flagged, in any fixture year.
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
years_all <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
|
|
||||||
states <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
|
|
||||||
oe <- suppressMessages(cog_spending(
|
|
||||||
states, years = years_all, category = "Other Education",
|
|
||||||
expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
expect_false(isTRUE(attr(oe, "provenance")$expenditure_concept_direct_suppressed))
|
|
||||||
expect_false(any(nzchar(oe$notes) & grepl("unavailable", oe$notes)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("M/I: a clean FY2019 category = NULL total query flags far fewer than the pre-fix 32/50 states", {
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
states <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT canonical_govid FROM long WHERE type = 0")$canonical_govid
|
|
||||||
r <- suppressMessages(cog_spending(
|
|
||||||
states, years = 2019, category = NULL, expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
ig <- r[r$spend_subtype == "intergovernmental", ]
|
|
||||||
flagged <- ig[nzchar(ig$notes) & grepl("unavailable", ig$notes), ]
|
|
||||||
expect_lt(length(unique(flagged$canonical_govid)), 32L)
|
|
||||||
# Every remaining flagged row must actually name a covering recipe --
|
|
||||||
# never the old no-recipe-found fallback.
|
|
||||||
expect_true(all(grepl("recipe = '", flagged$notes, fixed = TRUE)))
|
|
||||||
expect_false(any(grepl("no covering recipe found", flagged$notes, fixed = TRUE)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("C2: expenditure_concept = 'total' aborts on a corpus with no intergovernmental category rows", {
|
|
||||||
with_corpus_missing_ig_categories({
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT COUNT(*) AS n FROM summary_categories WHERE LEFT(item_code, 1) IN ('M', 'L')"
|
|
||||||
)$n
|
|
||||||
expect_equal(n, 0)
|
|
||||||
|
|
||||||
err <- tryCatch(
|
|
||||||
cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total"),
|
|
||||||
condition = function(e) e
|
|
||||||
)
|
|
||||||
expect_s3_class(err, "uscogdata_ig_categories_unsupported")
|
|
||||||
msg <- conditionMessage(err)
|
|
||||||
expect_match(msg, "PR #59|predates", perl = TRUE)
|
|
||||||
})
|
|
||||||
|
|
||||||
# 'direct' is unaffected on the same corpus -- the guard is scoped to
|
|
||||||
# expenditure_concept = "total" only.
|
|
||||||
with_corpus_missing_ig_categories({
|
|
||||||
expect_no_error(
|
|
||||||
cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("C2: expenditure_concept = 'total' still works on a corpus that DOES carry M/L category rows", {
|
|
||||||
expect_no_error(
|
|
||||||
cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("I2: an intergovernmental (M/L) recipe never appears as its own top-level suggestion", {
|
|
||||||
# Task 1's M04/M05 category rows share the "Corrections" summary_categories
|
|
||||||
# category with the Direct-flavored E04/E05, so `corrections_ig_local_
|
|
||||||
# combined` (entirely M-prefixed) becomes a raw *candidate* in
|
|
||||||
# .build_suggestions()'s component_code-driven query. Following a
|
|
||||||
# "re-run with recipe = 'corrections_ig_local_combined'" hint on a plain
|
|
||||||
# cog_spending() call would silently return intergovernmental dollars
|
|
||||||
# under provenance$expenditure_concept = "direct". Task 6's gate
|
|
||||||
# (.attach_ig_counterparts()) already protects the *counterpart* lookup;
|
|
||||||
# this exercises that the candidate list itself is filtered too.
|
|
||||||
r <- suppressMessages(
|
|
||||||
cog_spending("010000226085", years = c(2005, 2011), category = "Corrections")
|
|
||||||
)
|
|
||||||
sugg <- attr(r, "provenance")$suggestions
|
|
||||||
ids <- vapply(sugg, function(s) s$recipe_id %||% "", character(1))
|
|
||||||
expect_true("corrections_combined" %in% ids)
|
|
||||||
expect_false("corrections_ig_local_combined" %in% ids)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".attach_ig_counterparts() never pairs a revenue-side recipe with its coincidental M/L suffix twin", {
|
|
||||||
# Broader version of the case above, run at the matching-helper level
|
|
||||||
# (the same level code review's pairwise enumeration was done at) rather
|
|
||||||
# than end-to-end: the fixture has no (govid, year) combination where
|
|
||||||
# cog_revenue() itself produces a covered gap for any B/C/D recipe, so an
|
|
||||||
# end-to-end repro for THIS specific set of recipes isn't reachable
|
|
||||||
# today. Each of these six recipes shares an exact suffix set with an
|
|
||||||
# M/L expenditure recipe purely by reused-digit coincidence:
|
|
||||||
# ig_federal_b47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
|
|
||||||
# ig_federal_b89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
|
|
||||||
# ig_state_c47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
|
|
||||||
# ig_state_c89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
|
|
||||||
# ig_local_d47_wide {"47","94"} == ige_local_m47_wide / ige_state_l47_wide
|
|
||||||
# ig_local_d89_wide {"89","91","92","93"} == ige_local_m89_wide / ige_state_l89_wide
|
|
||||||
# None of them may receive an ig_recipe_id under cog_revenue()'s own
|
|
||||||
# flow_prefixes, since M/L only ever pairs with the direct-expenditure
|
|
||||||
# (E/F/G) family.
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
fake_suggestion <- function(rid) {
|
|
||||||
list(recipe_id = rid, label = "x", available_years = c(1967L, 2023L),
|
|
||||||
hint = "h")
|
|
||||||
}
|
|
||||||
fake_suggestions <- lapply(
|
|
||||||
c("ig_federal_b47_wide", "ig_federal_b89_wide",
|
|
||||||
"ig_state_c47_wide", "ig_state_c89_wide",
|
|
||||||
"ig_local_d47_wide", "ig_local_d89_wide"),
|
|
||||||
fake_suggestion
|
|
||||||
)
|
|
||||||
out <- uscogdata:::.attach_ig_counterparts(
|
|
||||||
con, fake_suggestions, c("T", "A", "U", "B", "C", "D")
|
|
||||||
)
|
|
||||||
ig <- unlist(lapply(out, function(s) s$ig_recipe_id))
|
|
||||||
expect_length(ig, 0L)
|
|
||||||
})
|
|
||||||
@@ -1,112 +0,0 @@
|
|||||||
# Madison walkthrough audit -- findings F-012, F-017, F-018.
|
|
||||||
# Tracked as uscogdata#11. See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# The owner's settled three-concept model (2026-07-28):
|
|
||||||
# total = primary + interest + intergovernmental transfers
|
|
||||||
# direct = primary + interest (Census's published Direct Expenditure)
|
|
||||||
# primary = direct minus debt service (the NEW DEFAULT)
|
|
||||||
# implemented by reclassifying on the crosswalk's `spend_type` column, NOT on
|
|
||||||
# item-code first letters -- F-018 shows prefix `Y` carries both revenue
|
|
||||||
# (Y01/Y02) and expenditure (Y05/Y06) codes, so no first-letter allowlist can
|
|
||||||
# route them correctly.
|
|
||||||
#
|
|
||||||
# Fixture reproducibility: the finding's headline reconciliation is Madison
|
|
||||||
# FY2022, where the corpus carries I89 = 46,609 (thousands) and Census's
|
|
||||||
# published Direct Expenditure is $654,893,000 against cog_spending()'s
|
|
||||||
# $608,284,000 (-7.1%). FY2022 is outside the bundled fixture's year window
|
|
||||||
# (2011/2012/2019/2020), so the same invariant is asserted on FY2020, where the
|
|
||||||
# fixture carries I89 = 27,704. Anyone running against the full corpus should
|
|
||||||
# also check the FY2022 numbers above.
|
|
||||||
|
|
||||||
test_that("expenditure concepts classify on spend_type, not item-code prefix", {
|
|
||||||
mad <- "552025209777" # MADISON CITY, WI
|
|
||||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
|
||||||
|
|
||||||
# -- F-012: `primary` is the new default, and equals today's E/F/G figure ---
|
|
||||||
primary <- cog_spending(govid = mad, years = 2020L)
|
|
||||||
expect_equal(attr(primary, "provenance")$expenditure_concept, "primary")
|
|
||||||
expect_equal(sum(primary$amt_nominal), 623347000)
|
|
||||||
|
|
||||||
# -- F-012: `direct` adds interest on long-term debt ------------------------
|
|
||||||
# Expected interest read from the RAW corpus, never through cog_spending(),
|
|
||||||
# which is the filter under test.
|
|
||||||
interest <- wt_raw_amt(mad, 2020L, prefixes = "I")
|
|
||||||
expect_equal(interest, 27704) # I89, in $1,000s
|
|
||||||
|
|
||||||
direct <- cog_spending(govid = mad, years = 2020L, expenditure_concept = "direct")
|
|
||||||
expect_equal(sum(direct$amt_nominal), 651051000) # 623,347 + 27,704 thousands
|
|
||||||
expect_equal(sum(direct$amt_nominal) - sum(primary$amt_nominal), interest * 1000)
|
|
||||||
expect_true("I89" %in% wt_codes_included(direct))
|
|
||||||
|
|
||||||
# -- F-017: `total` carries Q12/Q18, state IG transfers to school districts --
|
|
||||||
# Wisconsin FY2019: Q12 = 6,431,530 and Q18 = 533,391 (thousands). Today
|
|
||||||
# neither verb's flow_prefixes contains "Q", so both are dropped from the one
|
|
||||||
# concept that is supposed to include intergovernmental transfers.
|
|
||||||
ig_expected <- wt_raw_amt(wi_state, 2019L, prefixes = c("M", "L", "Q"))
|
|
||||||
expect_equal(ig_expected, 11609814) # M 4,644,893 + Q 6,964,921
|
|
||||||
|
|
||||||
wi_direct <- cog_spending(govid = wi_state, years = 2019L,
|
|
||||||
expenditure_concept = "direct")
|
|
||||||
wi_total <- cog_spending(govid = wi_state, years = 2019L,
|
|
||||||
expenditure_concept = "total")
|
|
||||||
|
|
||||||
# total - direct is exactly the intergovernmental component. Asserted as a
|
|
||||||
# delta rather than a grand total so this stays correct however the J and Y
|
|
||||||
# families land inside `primary`.
|
|
||||||
expect_equal(sum(wi_total$amt_nominal) - sum(wi_direct$amt_nominal),
|
|
||||||
ig_expected * 1000)
|
|
||||||
expect_true(all(c("Q12", "Q18") %in% wt_codes_included(wi_total)))
|
|
||||||
|
|
||||||
# -- F-018: prefix Y splits revenue from expenditure, by spend_type ---------
|
|
||||||
# Y01/Y02 are Insurance Trust revenue; Y05/Y06 are Insurance Trust benefit
|
|
||||||
# payments. All four share the first letter `Y`, so no first-letter allowlist
|
|
||||||
# can route them. The proof that classification is crosswalk-keyed:
|
|
||||||
# Y05 lands in `total` spending (insurance_benefits is inside `direct`),
|
|
||||||
# while Y01 -- same prefix -- is classified `revenue` by the crosswalk and
|
|
||||||
# therefore can never appear in a spending result.
|
|
||||||
#
|
|
||||||
# Per the owner's 2026-07-30 ruling (#11 DoD item 4 vs #12), cog_revenue()'s
|
|
||||||
# DEFAULT stays Census General Revenue and so excludes insurance-trust
|
|
||||||
# revenue; Y01's revenue-side classification is asserted against the
|
|
||||||
# crosswalk itself, not the default call. Surfacing Y01 through an explicit
|
|
||||||
# revenue concept argument is uscogdata#12.
|
|
||||||
wi_revenue <- cog_revenue(govid = wi_state, years = 2019L)
|
|
||||||
spend_codes <- wt_codes_included(wi_total)
|
|
||||||
rev_codes <- wt_codes_included(wi_revenue)
|
|
||||||
|
|
||||||
expect_true("Y05" %in% spend_codes)
|
|
||||||
expect_false("Y05" %in% rev_codes)
|
|
||||||
expect_false("Y01" %in% spend_codes)
|
|
||||||
expect_false("Y01" %in% rev_codes) # default = general revenue (#12 ruling)
|
|
||||||
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
y_class <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT item_code, category_type, spend_subtype, revenue_subtype
|
|
||||||
FROM summary_categories WHERE item_code IN ('Y01', 'Y05')")
|
|
||||||
expect_equal(y_class$category_type[y_class$item_code == "Y01"], "revenue")
|
|
||||||
expect_equal(y_class$revenue_subtype[y_class$item_code == "Y01"], "insurance_trust")
|
|
||||||
expect_equal(y_class$category_type[y_class$item_code == "Y05"], "expenditure")
|
|
||||||
expect_equal(y_class$spend_subtype[y_class$item_code == "Y05"], "insurance_benefits")
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("no balance code or category ever reaches a spending or revenue result (uscogdata#25)", {
|
|
||||||
# Stocks are not flows. The crosswalk's balance codes (W/X/Y/Z fund
|
|
||||||
# balances) share first letters with flow codes, so this could never be
|
|
||||||
# guaranteed under prefix classification; under crosswalk membership it
|
|
||||||
# falls out structurally -- asserted here at the verb level, on a
|
|
||||||
# government-year the fixture gives real balance rows (Wisconsin carries
|
|
||||||
# Y07/Y08/Y21/Y61-type balances in FY2019).
|
|
||||||
wi_state <- "550000227544"
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
balance <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
|
|
||||||
expect_gt(nrow(balance), 0L)
|
|
||||||
|
|
||||||
spend <- cog_spending(wi_state, 2019L, expenditure_concept = "total")
|
|
||||||
rev <- cog_revenue(wi_state, 2019L)
|
|
||||||
|
|
||||||
expect_false(any(spend$category %in% balance$category))
|
|
||||||
expect_false(any(rev$category %in% balance$category))
|
|
||||||
expect_length(intersect(wt_codes_included(spend), balance$item_code), 0L)
|
|
||||||
expect_length(intersect(wt_codes_included(rev), balance$item_code), 0L)
|
|
||||||
})
|
|
||||||
@@ -70,37 +70,6 @@ test_that("cog_explain prints a Suggestions section when the provenance has one"
|
|||||||
expect_true(grepl("re-run with recipe", txt))
|
expect_true(grepl("re-run with recipe", txt))
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("cog_explain prints the expenditure concept (I1)", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
d <- cog_spending("010000226085", years = 2019, category = "Police")
|
|
||||||
t <- cog_spending("010000226085", years = 2019, category = "Police",
|
|
||||||
expenditure_concept = "total")
|
|
||||||
txt_d <- paste(c(
|
|
||||||
capture.output(cog_explain(d)),
|
|
||||||
capture.output(cog_explain(d), type = "message")
|
|
||||||
), collapse = "\n")
|
|
||||||
txt_t <- paste(c(
|
|
||||||
capture.output(cog_explain(t)),
|
|
||||||
capture.output(cog_explain(t), type = "message")
|
|
||||||
), collapse = "\n")
|
|
||||||
expect_true(grepl("Concept: primary", txt_d))
|
|
||||||
expect_true(grepl("Concept: total", txt_t))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_explain surfaces the C1(b) direct-suppressed flag as a warning", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
t <- suppressMessages(cog_spending(
|
|
||||||
"010000226085", years = 2011, category = "Corrections",
|
|
||||||
expenditure_concept = "total"
|
|
||||||
))
|
|
||||||
expect_true(attr(t, "provenance")$expenditure_concept_direct_suppressed)
|
|
||||||
txt <- paste(c(
|
|
||||||
capture.output(cog_explain(t)),
|
|
||||||
capture.output(cog_explain(t), type = "message")
|
|
||||||
), collapse = "\n")
|
|
||||||
expect_true(grepl("Direct leg unavailable", txt))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("cog_explain prints denominator + popyear_range + counts", {
|
test_that("cog_explain prints denominator + popyear_range + counts", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
with_fixture_corpus({
|
with_fixture_corpus({
|
||||||
|
|||||||
@@ -1,112 +0,0 @@
|
|||||||
# tests/testthat/test-fixture-vintage.R
|
|
||||||
#
|
|
||||||
# The bundled fixture is a slice of a real cog_pipeline publish tree, and
|
|
||||||
# every test in this package -- plus the whole cog-api suite -- runs against
|
|
||||||
# it. When the published corpus changes shape and the fixture does not, both
|
|
||||||
# suites stay green against a corpus that no longer exists (uscogdata#18).
|
|
||||||
#
|
|
||||||
# These tests pin the structural facts that distinguish the current published
|
|
||||||
# vintage from its predecessor, so a stale fixture fails loudly instead of
|
|
||||||
# passing quietly. They assert shape, never dollar values: re-running
|
|
||||||
# data-raw/regenerate_fixture_corpus.R against a newer publish tree should
|
|
||||||
# keep them green.
|
|
||||||
|
|
||||||
# Open a bare DuckDB connection on the fixture's parquet files. Deliberately
|
|
||||||
# not the package session: these assertions are about what the fixture
|
|
||||||
# CONTAINS, and routing them through the reader's own views would let a
|
|
||||||
# filter hide the very absence being checked.
|
|
||||||
fixture_query <- function(sql, ...) {
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
path <- function(rel) {
|
|
||||||
sprintf("read_parquet(%s)",
|
|
||||||
DBI::dbQuoteString(con, file.path(fixture_corpus_path(), rel)))
|
|
||||||
}
|
|
||||||
DBI::dbGetQuery(con, do.call(sprintf, c(list(sql), lapply(c(...), path))))
|
|
||||||
}
|
|
||||||
|
|
||||||
test_that("fixture ships every metadata table the publish tree does", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
# representation/code_set are what make a sparse corpus interpretable; a
|
|
||||||
# fixture without them predates sparsification (cog_pipeline#64).
|
|
||||||
expected <- c(
|
|
||||||
"canonical_alias.parquet", "canonical_fips_xwalk.parquet",
|
|
||||||
"census_collection_coverage.parquet", "code_set.parquet",
|
|
||||||
"harmonization_map.parquet", "harmonization_recipes.parquet",
|
|
||||||
"lineage_events.parquet", "representation.parquet",
|
|
||||||
"series_breaks.parquet", "summary_categories.parquet"
|
|
||||||
)
|
|
||||||
on_disk <- basename(list.files(
|
|
||||||
file.path(fixture_corpus_path(), "data"), pattern = "\\.parquet$"
|
|
||||||
))
|
|
||||||
expect_true(all(expected %in% on_disk))
|
|
||||||
|
|
||||||
# The manifest must list them too -- consumers read the manifest, not ls().
|
|
||||||
in_manifest <- with_fixture_corpus(
|
|
||||||
basename(vapply(cog_manifest()$files$metadata, function(f) f$path, character(1)))
|
|
||||||
)
|
|
||||||
expect_true(all(expected %in% in_manifest))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("fixture carries the dense/sparse representation contract", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
rep <- fixture_query(
|
|
||||||
"SELECT year, representation, absence_means FROM %s
|
|
||||||
WHERE year IN (2011, 2012, 2019, 2020) ORDER BY year",
|
|
||||||
"data/representation.parquet"
|
|
||||||
)
|
|
||||||
expect_equal(nrow(rep), 4L)
|
|
||||||
expect_equal(rep$representation, c("dense_source", rep("sparse_source", 3L)))
|
|
||||||
expect_equal(rep$absence_means, c("census_zero", rep("not_reported", 3L)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("the fixture's wide era is sparse, not zero-padded", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
# FY2011 is a dense_source year: the corpus publishes only the cells Census
|
|
||||||
# reported non-zero, and an absent cell means Census published $0. Before
|
|
||||||
# sparsification this partition was 2,864,212 rows, ~83% of them explicit
|
|
||||||
# zeros. A single explicit zero here means the fixture predates the change.
|
|
||||||
zeros_2011 <- fixture_query(
|
|
||||||
"SELECT COUNT(*) AS n FROM %s WHERE amt = 0",
|
|
||||||
"data/long/year=2011/part-0.parquet"
|
|
||||||
)$n
|
|
||||||
expect_equal(zeros_2011, 0L)
|
|
||||||
|
|
||||||
# The modern era is a different regime: a reported zero there is real data
|
|
||||||
# (the government filed $0), so zeros legitimately survive and must not be
|
|
||||||
# asserted away.
|
|
||||||
expect_gt(
|
|
||||||
fixture_query("SELECT COUNT(*) AS n FROM %s", "data/long/year=2012/part-0.parquet")$n,
|
|
||||||
0L
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("code_set covers every fixture year with the reader-spec columns", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
cs <- fixture_query(
|
|
||||||
"SELECT * FROM %s WHERE year IN (2011, 2012, 2019, 2020)",
|
|
||||||
"data/code_set.parquet"
|
|
||||||
)
|
|
||||||
expect_true(all(
|
|
||||||
c("code_set_id", "year", "type", "item_code", "is_aggregate", "n_units")
|
|
||||||
%in% names(cs)
|
|
||||||
))
|
|
||||||
expect_setequal(unique(cs$year), c(2011L, 2012L, 2019L, 2020L))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("every flow code carrying dollars has a category, J-prefix included", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
# The J (assistance/benefit) codes were uncategorised until the crosswalk
|
|
||||||
# completion shipped (cog_pipeline#60/#65, J19 held back until #64's
|
|
||||||
# duplication fix landed). Their absence is how a pre-crosswalk fixture
|
|
||||||
# gives itself away.
|
|
||||||
j <- fixture_query(
|
|
||||||
"SELECT item_code, category, category_type, spend_subtype FROM %s
|
|
||||||
WHERE LEFT(item_code, 1) = 'J' ORDER BY item_code",
|
|
||||||
"data/summary_categories.parquet"
|
|
||||||
)
|
|
||||||
expect_true("J19" %in% j$item_code)
|
|
||||||
expect_true(all(j$category_type == "expenditure"))
|
|
||||||
expect_true(all(j$spend_subtype == "assistance"))
|
|
||||||
expect_false(any(is.na(j$category)))
|
|
||||||
})
|
|
||||||
@@ -1,57 +0,0 @@
|
|||||||
# Madison walkthrough audit -- finding F-025. Tracked as uscogdata#16.
|
|
||||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# cog_gov_search()'s UTILITY mode interpolates `name` into
|
|
||||||
# regexp_matches(gov_name, <name>, 'i')
|
|
||||||
# unescaped (R/search.R:102), while BASKET mode in the same file already routes
|
|
||||||
# it through .escape_regex() (R/search.R:307) with the comment "so `name` is
|
|
||||||
# treated as a literal substring". Two failure modes result:
|
|
||||||
# correctness -- a real government cannot be found by its own exact name, and
|
|
||||||
# a single "." matches everything (HTTP 200 both ways via the API);
|
|
||||||
# robustness -- malformed regex reaches the engine and errors, which cog-api
|
|
||||||
# surfaces as a 500, reachable by typing a real name one
|
|
||||||
# character at a time.
|
|
||||||
#
|
|
||||||
# NOT asserted here: the finding's `q=St. Louis` example. Under correct literal
|
|
||||||
# matching that search still returns 0 rows, because the stored name is
|
|
||||||
# "ST LOUIS CITY" with no period -- it demonstrates today's over-matching
|
|
||||||
# semantics, not a row the fix makes findable.
|
|
||||||
|
|
||||||
test_that("cog_gov_search() matches name literally, not as an unescaped regex", {
|
|
||||||
|
|
||||||
# -- correctness (1): a government must be findable by its own exact name ---
|
|
||||||
# FREDONIA (BRISCOE) CITY is real; today the parentheses are read as regex
|
|
||||||
# grouping, so its own complete name matches nothing.
|
|
||||||
fredonia <- cog_gov_search(name = "FREDONIA (BRISCOE) CITY")
|
|
||||||
expect_equal(nrow(fredonia), 1L)
|
|
||||||
expect_equal(fredonia$canonical_govid, "052117184386")
|
|
||||||
expect_equal(cog_gov_search(name = "FREDONIA (BRISCOE)")$canonical_govid,
|
|
||||||
"052117184386")
|
|
||||||
|
|
||||||
# -- correctness (2): a metacharacter must not become a wildcard ------------
|
|
||||||
# No Wisconsin city or village name contains a literal period -- established
|
|
||||||
# against the raw registry below, NOT through the verb under test. A literal
|
|
||||||
# search for "." must therefore return nothing; today it returns all 608.
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
xwalk <- paste0(sub("/$", "", Sys.getenv("USCOGDATA_URL")),
|
|
||||||
"/data/canonical_fips_xwalk.parquet")
|
|
||||||
with_dot <- DBI::dbGetQuery(con, paste0(
|
|
||||||
"SELECT COUNT(*) n FROM read_parquet('", xwalk, "') ",
|
|
||||||
"WHERE fips_state = '55' AND govs_type = 2 AND gov_name LIKE '%.%'"))
|
|
||||||
expect_equal(as.integer(with_dot$n[[1]]), 0L)
|
|
||||||
|
|
||||||
expect_equal(nrow(cog_gov_search(name = ".", state = "WI", type = "city")), 0L)
|
|
||||||
expect_equal(nrow(cog_gov_search(name = "M.dison", state = "WI", type = "city")), 0L)
|
|
||||||
expect_equal(nrow(cog_gov_search(name = "Mad(i|o)son", state = "WI", type = "city")), 0L)
|
|
||||||
|
|
||||||
# A metacharacter-free name still resolves exactly as before.
|
|
||||||
expect_equal(nrow(cog_gov_search(name = "Madison", state = "WI", type = "city")), 1L)
|
|
||||||
|
|
||||||
# -- robustness: malformed pattern text returns no rows, and does not error --
|
|
||||||
# "[" alone, and "Athens-Clarke County (bal" -- an in-progress substring of
|
|
||||||
# ATHENS-CLARKE COUNTY (BALANCE), a real government -- both currently raise
|
|
||||||
# (DuckDB: "Invalid Input Error: missing ]").
|
|
||||||
expect_equal(nrow(cog_gov_search(name = "[")), 0L)
|
|
||||||
expect_equal(nrow(cog_gov_search(name = "Athens-Clarke County (bal")), 0L)
|
|
||||||
})
|
|
||||||
@@ -1,51 +0,0 @@
|
|||||||
# Madison walkthrough audit -- finding F-021. Tracked as uscogdata#14.
|
|
||||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# .peer_summary_rows() computes stats::quantile() separately INSIDE each
|
|
||||||
# (year, spend_subtype, category) cell. A summary_p50 row is therefore "the
|
|
||||||
# median peer's value in that one category", not "the value of the median
|
|
||||||
# peer's total". Summing those rows across categories -- the obvious move for a
|
|
||||||
# caller who wants one peer-median total line and reads only the column names --
|
|
||||||
# misstated a total-spending band by -32.7% to +251.0% across the 24 years the
|
|
||||||
# audit tested, with a sign flip at FY2012.
|
|
||||||
#
|
|
||||||
# The verb is not wrong and its documented use (faceting by role AND category)
|
|
||||||
# is unaffected, so the fix is documentation: one sentence in @return.
|
|
||||||
|
|
||||||
test_that("cog_peer_compare() documents that summary_* rows are per-category quantiles", {
|
|
||||||
|
|
||||||
# man/ ships only in the source tree (the installed package carries a
|
|
||||||
# compiled help database instead), so the prose assertions below cannot run
|
|
||||||
# under R CMD check -- CI's earlier testthat::test_local() step enforces
|
|
||||||
# them. The numeric pin further down needs only the corpus, but it lives in
|
|
||||||
# the same test_that() as the sentence it protects, deliberately: they are
|
|
||||||
# one claim, and splitting them would let the prose drift while a separate
|
|
||||||
# test kept passing.
|
|
||||||
rd_path <- skip_if_no_source_tree(c("man", "cog_peer_compare.Rd"))
|
|
||||||
rd <- paste(readLines(rd_path, warn = FALSE), collapse = " ")
|
|
||||||
|
|
||||||
# The @return section must say the quantile is computed within each cell...
|
|
||||||
expect_match(rd, "within each|per-category|per category", ignore.case = TRUE)
|
|
||||||
# ...and must warn that the rows are not additive across category.
|
|
||||||
expect_match(rd, "not additive|do(es)? not sum|cannot be summed", ignore.case = TRUE)
|
|
||||||
# ...naming the grouping explicitly.
|
|
||||||
expect_match(rd, "spend_subtype", fixed = TRUE)
|
|
||||||
|
|
||||||
# Pin the mechanism numerically so a future refactor that quietly changes the
|
|
||||||
# quantile grouping fails here rather than silently invalidating the sentence
|
|
||||||
# above. Fixture: Madison, 10 peers found at FY2020, category = NULL.
|
|
||||||
peers <- cog_find_peers("552025209777", year = 2020L, max_peers = 10L)
|
|
||||||
cmp <- cog_peer_compare(target_govid = "552025209777", peers = peers,
|
|
||||||
category = NULL, years = 2020L, per_capita = TRUE)
|
|
||||||
|
|
||||||
naive <- sum(cmp$amt_per_capita_nominal[cmp$role == "summary_p50"], na.rm = TRUE)
|
|
||||||
|
|
||||||
peer_rows <- cmp[cmp$role == "peer", ]
|
|
||||||
per_gov <- tapply(peer_rows$amt_per_capita_nominal, peer_rows$canonical_govid,
|
|
||||||
sum, na.rm = TRUE)
|
|
||||||
correct <- unname(stats::quantile(per_gov, 0.5, na.rm = TRUE))
|
|
||||||
|
|
||||||
expect_equal(round(naive), 6180) # summing the built-in summary rows
|
|
||||||
expect_equal(round(correct), 2043) # quantile of each peer's OWN total
|
|
||||||
expect_gt(naive / correct, 2) # a +200% misstatement on this cohort
|
|
||||||
})
|
|
||||||
@@ -176,6 +176,16 @@ test_that("recipe = requires schema_version >= 5", {
|
|||||||
})
|
})
|
||||||
|
|
||||||
# --- signposting -------------------------------------------------------
|
# --- signposting -------------------------------------------------------
|
||||||
|
#
|
||||||
|
# Phase R3 / Task 19c: .build_suggestions() was narrowed from a whole-result
|
||||||
|
# gap check (R2: does the ENTIRE category result have zero rows in a
|
||||||
|
# requested year) to per-code gap detection (does a specific recipe
|
||||||
|
# component -- itself a member of the requested category -- have zero rows
|
||||||
|
# in a year the recipe's own generic join otherwise covers). See
|
||||||
|
# R/suggestions.R's header comment and docs/phase_r_harmonization_review.md
|
||||||
|
# § 0.3. The R2 test below ("...across the 2011->2012 gap") is unaffected
|
||||||
|
# by the refinement (it already passed under both the coarse and per-code
|
||||||
|
# rule). The next few pin cases the coarse rule specifically could NOT see.
|
||||||
|
|
||||||
test_that("signposting suggests corrections_combined across the 2011->2012 gap", {
|
test_that("signposting suggests corrections_combined across the 2011->2012 gap", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
@@ -193,10 +203,96 @@ test_that("signposting suggests corrections_combined across the 2011->2012 gap",
|
|||||||
expect_equal(hit$available_years, c(1967L, 2023L))
|
expect_equal(hit$available_years, c(1967L, 2023L))
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("no signposting when the result already has full year coverage", {
|
test_that("per-code gap does NOT fire when a code's only coverage is its own aggregate row (self-coverage is not \"other components\")", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
|
# Broward, FY2011 ONLY (isolating the 2011 half of the query above): E05,
|
||||||
|
# F05, and G05 each report SOLELY as a wide-era AGGREGATE row that year
|
||||||
|
# (216088, 1453, 270 respectively); E04/F04/G04 -- their modern-only
|
||||||
|
# siblings -- don't exist as codes at all before 2012, corpus-wide (zero
|
||||||
|
# rows for any government). Each component's own aggregate row would
|
||||||
|
# trivially satisfy a same-component "covered" check, but review-doc
|
||||||
|
# § 0.3's criterion is explicit that a gap must be covered by "OTHER
|
||||||
|
# components", not the gapped component's own aggregate form. With no
|
||||||
|
# OTHER component present for any of the three Corrections recipes in
|
||||||
|
# 2011, none of them should fire -- this is what the combined
|
||||||
|
# 2011-2012 test above actually relies on 2012 (E05 gapped, E04 -- a
|
||||||
|
# genuinely different component -- covers) to fire, not 2011.
|
||||||
|
r <- cog_spending("121011212191", years = 2011L, category = "Corrections")
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
|
expect_length(prov$suggestions, 0L)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("per-code gap fires even when a sibling code masks the whole-result check (Cleburne County, FY2012)", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# Cleburne County, AL (canonical_govid 011029122489), FY2012: E04 ($854)
|
||||||
|
# and E05 ($1) both report ("operations" subtype), and G04 ($14,000,
|
||||||
|
# corrections_other_capital_combined's modern-only leg) also reports
|
||||||
|
# ("capital" subtype) -- so the WHOLE category result is non-empty for
|
||||||
|
# 2012 (2 rows) and the R2 whole-result check would never look further.
|
||||||
|
# But G05 -- G04's OWN recipe sibling, the 1967-2023 wide leg -- has
|
||||||
|
# ZERO rows at all that year: a genuine, per-code gap the recipe exists
|
||||||
|
# to bridge, invisible at the category-result grain because it's masked
|
||||||
|
# by G04's own data, let alone the unrelated E04/E05 pair.
|
||||||
|
r <- cog_spending("011029122489", years = 2012L, category = "Corrections")
|
||||||
|
expect_equal(nrow(r), 2L) # operations + capital rows: a non-empty result
|
||||||
|
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
|
ids <- vapply(prov$suggestions, function(s) s$recipe_id, character(1))
|
||||||
|
expect_true("corrections_other_capital_combined" %in% ids)
|
||||||
|
hit <- prov$suggestions[[which(ids == "corrections_other_capital_combined")]]
|
||||||
|
expect_equal(hit$hint, "re-run with recipe = 'corrections_other_capital_combined'")
|
||||||
|
expect_equal(hit$available_years, c(1967L, 2023L))
|
||||||
|
|
||||||
|
# corrections_combined must NOT fire: E04 AND E05 both have real 2012
|
||||||
|
# data for this government, so neither of ITS OWN components is gapped.
|
||||||
|
expect_false("corrections_combined" %in% ids)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("per-code gap does not fire when no recipe component has any data at all (ordinary reporting variance, not a format-boundary gap)", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# Same government/year as above: F04 and F05 (corrections_capital_combined)
|
||||||
|
# are BOTH completely absent -- Cleburne simply never reported capital
|
||||||
|
# corrections spending under that code family in 2012, wide-era or
|
||||||
|
# modern. The recipe's own generic join (aggregate-inclusive, either
|
||||||
|
# component) has nothing to offer either, so this must stay silent --
|
||||||
|
# the per-government `covered` guard the header comment describes is
|
||||||
|
# unchanged and still does this filtering.
|
||||||
|
r <- cog_spending("011029122489", years = 2012L, category = "Corrections")
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
|
ids <- vapply(prov$suggestions, function(s) s$recipe_id, character(1))
|
||||||
|
expect_false("corrections_capital_combined" %in% ids)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("per-code gap fires for Broward 2019-2020 even though the category result looks complete", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# Broward reports E04 + G04 (modern leaf codes) in BOTH 2019 and 2020 but
|
||||||
|
# never reports E05 or G05 (their own recipe siblings) in either year --
|
||||||
|
# a real per-code gap in two of the three Corrections recipes, invisible
|
||||||
|
# under the R2 coarse check because the category *result* is non-empty
|
||||||
|
# both years (this replaces the old R2-era "full year coverage" test,
|
||||||
|
# whose premise -- that a non-empty result implies nothing to signpost --
|
||||||
|
# is exactly what this refinement narrows; see data-raw/
|
||||||
|
# measure_signposting_rate.R for the measured rate change this causes).
|
||||||
|
# corrections_capital_combined correctly stays silent: Broward reports
|
||||||
|
# neither F04 nor F05 in 2019 or 2020, so that recipe's own join has
|
||||||
|
# nothing to offer either (ordinary non-reporting, not a format-boundary
|
||||||
|
# gap) -- the per-government `covered` guard still does its job here too.
|
||||||
r <- cog_spending("121011212191", years = 2019:2020, category = "Corrections")
|
r <- cog_spending("121011212191", years = 2019:2020, category = "Corrections")
|
||||||
prov <- attr(r, "provenance")
|
prov <- attr(r, "provenance")
|
||||||
|
ids <- vapply(prov$suggestions, function(s) s$recipe_id, character(1))
|
||||||
|
expect_true("corrections_combined" %in% ids)
|
||||||
|
expect_true("corrections_other_capital_combined" %in% ids)
|
||||||
|
expect_false("corrections_capital_combined" %in% ids)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("no signposting when every recipe component genuinely has data (true full per-code coverage)", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# Maricopa County, AZ (canonical_govid 041013160815): all six Corrections
|
||||||
|
# codes (E04, E05, F04, F05, G04, G05) report real, nonzero, non-aggregate
|
||||||
|
# amounts in BOTH 2019 and 2020 -- genuinely nothing for any recipe to
|
||||||
|
# fill, even at the finer per-code grain this refinement now checks.
|
||||||
|
r <- cog_spending("041013160815", years = 2019:2020, category = "Corrections")
|
||||||
|
prov <- attr(r, "provenance")
|
||||||
expect_length(prov$suggestions, 0L)
|
expect_length(prov$suggestions, 0L)
|
||||||
})
|
})
|
||||||
|
|
||||||
|
|||||||
@@ -1,139 +0,0 @@
|
|||||||
# Madison walkthrough audit -- finding F-014. Tracked as uscogdata#12.
|
|
||||||
# See docs/walkthroughs/FINDINGS.md in cog_explorer.
|
|
||||||
#
|
|
||||||
# cog_revenue()'s flow_prefixes = c("T","A","U","B","C","D") never returns
|
|
||||||
# item-code prefix X (Employee Retirement) or Y (other Insurance Trust). Per
|
|
||||||
# Census's standard identity, Total Revenue = General + Utility + Liquor Store +
|
|
||||||
# Insurance Trust Revenue, and Employee Retirement System contributions and
|
|
||||||
# earnings ARE the Insurance Trust Revenue component -- so prefix X sits inside
|
|
||||||
# a published Census revenue concept exactly the way I89 sits inside Census's
|
|
||||||
# Direct Expenditure concept (finding F-012).
|
|
||||||
#
|
|
||||||
# RULED 2026-07-30. `revenue_concept = c("general", "total")` mirrors
|
|
||||||
# `expenditure_concept`, and the two values are Census's two published revenue
|
|
||||||
# concepts, related by the manual's own identity (section 4.3, which defines
|
|
||||||
# the first by SUBTRACTING from the second):
|
|
||||||
#
|
|
||||||
# Total Revenue = General + Utility + Liquor Store + Insurance Trust
|
|
||||||
#
|
|
||||||
# so `general` is the four general subtypes (own_source/federal/state/
|
|
||||||
# local_aid) and `total` is every revenue subtype. Naming utility (A91-A94)
|
|
||||||
# and liquor store (A90) separately is what makes BOTH computable -- before
|
|
||||||
# cog_pipeline#79 they sat in own_source, so the default was really
|
|
||||||
# "General + Utility + Liquor", a concept Census does not publish.
|
|
||||||
#
|
|
||||||
# Fixture reproducibility: Madison's own X-prefix revenue (FY1970-FY1986,
|
|
||||||
# $15,098,000 nominal, $0 thereafter) is outside the bundled fixture's year
|
|
||||||
# window (2011/2012/2019/2020), so the same invariant is asserted on Wisconsin
|
|
||||||
# state government FY2012, where the fixture carries nonzero X01/X02/X05/X08.
|
|
||||||
|
|
||||||
test_that("cog_revenue() can return Census Total Revenue including Insurance Trust (prefix X)", {
|
|
||||||
wi_state <- "550000227544" # WISCONSIN (state government)
|
|
||||||
|
|
||||||
# Revenue-shaped Employee Retirement codes, read from the RAW corpus rather
|
|
||||||
# than through cog_revenue(), which is the filter under test:
|
|
||||||
# X01/X02 employee contributions, X05 contributions from other governments,
|
|
||||||
# X08 total earnings on investments.
|
|
||||||
#
|
|
||||||
# X04 and X06 are deliberately NOT in this set, though an earlier draft of
|
|
||||||
# this test included X04. Both are exhibit codes for INTRAgovernmental
|
|
||||||
# transfers (the administering government paying into its own fund), which
|
|
||||||
# X05's own definition excludes by name. Census agrees: its computed "Total
|
|
||||||
# Emp Ret Rev" for this government-year is exactly the four codes below.
|
|
||||||
x_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("X01", "X02", "X05", "X08"))
|
|
||||||
expect_equal(x_revenue, 2283883) # 615,835 + 245,083 + 560,382 + 862,583
|
|
||||||
|
|
||||||
# The Y-prefix insurance trust revenue (unemployment + workers comp), which
|
|
||||||
# is the other half of the same Census concept.
|
|
||||||
y_revenue <- wt_raw_amt(wi_state, 2012L, codes = c("Y01", "Y11"))
|
|
||||||
expect_equal(y_revenue, 1259785)
|
|
||||||
|
|
||||||
general <- cog_revenue(govid = wi_state, years = 2012L)
|
|
||||||
expect_equal(attr(general, "provenance")$revenue_concept, "general")
|
|
||||||
expect_equal(sum(general$amt_nominal), 31338293000)
|
|
||||||
|
|
||||||
total <- cog_revenue(govid = wi_state, years = 2012L, revenue_concept = "total")
|
|
||||||
expect_equal(attr(total, "provenance")$revenue_concept, "total")
|
|
||||||
|
|
||||||
# total - general is the whole insurance trust leg, X and Y together.
|
|
||||||
# Asserted as a delta as well as a level so this stays correct however the
|
|
||||||
# utility/liquor families land (both are $0 for WI state in FY2012).
|
|
||||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
|
|
||||||
(x_revenue + y_revenue) * 1000)
|
|
||||||
expect_equal(sum(total$amt_nominal), 34881961000)
|
|
||||||
expect_true(all(c("X01", "X02", "X05", "X08") %in% wt_codes_included(total)))
|
|
||||||
|
|
||||||
# Sibling codes under the SAME first letter must stay out: X11/X12 are
|
|
||||||
# benefit payments (an expenditure) and X21/X30/X47 are cash and securities
|
|
||||||
# holdings (a balance-sheet stock). This is the F-018 point restated on the
|
|
||||||
# revenue side -- the split comes from the crosswalk, not from the letter X.
|
|
||||||
expect_false(any(c("X11", "X12", "X21", "X30", "X47") %in% wt_codes_included(total)))
|
|
||||||
|
|
||||||
# Every returned row still resolves to a category (cog_pipeline#79 added the
|
|
||||||
# X crosswalk rows; relaxing a prefix filter alone would have produced
|
|
||||||
# category = NA rows).
|
|
||||||
expect_false(any(is.na(total$category)))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("revenue_concept = 'general' is the default and is strict Census General Revenue", {
|
|
||||||
wi_state <- "550000227544"
|
|
||||||
default <- cog_revenue(govid = wi_state, years = 2012L)
|
|
||||||
explicit <- cog_revenue(govid = wi_state, years = 2012L,
|
|
||||||
revenue_concept = "general")
|
|
||||||
expect_equal(sum(default$amt_nominal), sum(explicit$amt_nominal))
|
|
||||||
|
|
||||||
# General Revenue excludes utility, liquor store AND insurance trust
|
|
||||||
# revenue. WI state carries $0 of utility/liquor in FY2012, so the level
|
|
||||||
# assertion above cannot see those two -- assert the subtype scope directly.
|
|
||||||
#
|
|
||||||
# A subset, not setequal: `state` means "intergovernmental revenue FROM the
|
|
||||||
# state government" (the C codes), which a STATE government does not receive
|
|
||||||
# from itself, so it is legitimately absent here.
|
|
||||||
expect_true(all(default$revenue_subtype %in%
|
|
||||||
c("own_source", "federal", "state", "local_aid")))
|
|
||||||
expect_false(any(c("utility", "liquor_store", "insurance_trust") %in%
|
|
||||||
default$revenue_subtype))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("utility and liquor store revenue are inside `total` and outside `general`", {
|
|
||||||
# A city, where utility revenue is material: this is the case the WI state
|
|
||||||
# baseline structurally cannot exercise. Measured on the fixture, utility +
|
|
||||||
# liquor is 15.9% of what cog_revenue() returned for type-2 governments
|
|
||||||
# before the general/total split, so this is the largest behaviour change
|
|
||||||
# the concept split introduces.
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
gov <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT canonical_govid, SUM(amt) amt FROM long
|
|
||||||
WHERE year = 2012 AND type = 2 AND NOT is_aggregate
|
|
||||||
AND item_code IN ('A91','A92','A93','A94')
|
|
||||||
GROUP BY 1 ORDER BY amt DESC LIMIT 1")$canonical_govid
|
|
||||||
|
|
||||||
util_raw <- wt_raw_amt(gov, 2012L, codes = c("A90", "A91", "A92", "A93", "A94"))
|
|
||||||
expect_gt(util_raw, 0)
|
|
||||||
|
|
||||||
general <- cog_revenue(govid = gov, years = 2012L)
|
|
||||||
total <- cog_revenue(govid = gov, years = 2012L, revenue_concept = "total")
|
|
||||||
|
|
||||||
expect_false(any(c("utility", "liquor_store") %in% general$revenue_subtype))
|
|
||||||
expect_true("utility" %in% total$revenue_subtype)
|
|
||||||
expect_equal(sum(total$amt_nominal) - sum(general$amt_nominal),
|
|
||||||
util_raw * 1000 +
|
|
||||||
wt_raw_amt(gov, 2012L, codes = c("Y01", "Y11", "X01", "X02",
|
|
||||||
"X05", "X08")) * 1000)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("revenue_concept rejects unknown values and never returns a balance row", {
|
|
||||||
expect_error(
|
|
||||||
cog_revenue("550000227544", years = 2012L, revenue_concept = "gross"),
|
|
||||||
class = "uscogdata_invalid_revenue_concept"
|
|
||||||
)
|
|
||||||
|
|
||||||
# uscogdata#25 restated for the widest revenue concept: stocks are not
|
|
||||||
# flows, and `total` must not quietly admit the X/Y/W/Z balance families.
|
|
||||||
con <- uscogdata:::.ensure_session()
|
|
||||||
balance <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT item_code, category FROM summary_categories WHERE category_type = 'balance'")
|
|
||||||
total <- cog_revenue("550000227544", years = 2012L, revenue_concept = "total")
|
|
||||||
expect_false(any(total$category %in% balance$category))
|
|
||||||
expect_length(intersect(wt_codes_included(total), balance$item_code), 0L)
|
|
||||||
})
|
|
||||||
@@ -80,11 +80,8 @@ test_that("cog_geographic_rollup provenance reports the outer verb", {
|
|||||||
|
|
||||||
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
test_that("cog_geographic_rollup accepts data.frames per layer", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
# Unanchored: utility mode matches literally now, so "^...$" would be
|
fl_state <- cog_gov_search("^FLORIDA$", type = "state")
|
||||||
# searched for as characters rather than read as anchors (uscogdata#16).
|
broward <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||||
# Both still resolve to exactly one row once scoped by type/state.
|
|
||||||
fl_state <- cog_gov_search("FLORIDA", type = "state")
|
|
||||||
broward <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
|
||||||
r <- cog_geographic_rollup(
|
r <- cog_geographic_rollup(
|
||||||
govids = list(state = fl_state, county = broward),
|
govids = list(state = fl_state, county = broward),
|
||||||
category = "Police", years = 2020L
|
category = "Police", years = 2020L
|
||||||
|
|||||||
@@ -0,0 +1,164 @@
|
|||||||
|
# tests/testthat/test-signposting-harness.R
|
||||||
|
#
|
||||||
|
# Pins the subset-relation REPORTING in data-raw/measure_signposting_rate.R.
|
||||||
|
#
|
||||||
|
# Phase R3 Task 19c narrowed signposting from a coarse whole-result gap
|
||||||
|
# check to per-code gap detection. Those two checks are partly DISJOINT,
|
||||||
|
# not nested: a query can fire under coarse and stay silent under per-code,
|
||||||
|
# so the harness's `*_delta_pp` figures are NETS that can hide a coverage
|
||||||
|
# loss. `.measure_subset_relation()` is what separates the two flows, and
|
||||||
|
# `.measure_format_subset_report()` is what puts the loss in front of a
|
||||||
|
# human. Both are load-bearing for the Checkpoint R3 ruling, so both are
|
||||||
|
# pinned here: if the violation detection is deleted, inverted, or quietly
|
||||||
|
# downgraded to a count with no identities, these tests fail.
|
||||||
|
#
|
||||||
|
# These tests do NOT assert that the violation set is empty -- it is
|
||||||
|
# genuinely non-empty, and asserting otherwise would be pinning a bug as a
|
||||||
|
# contract. They assert only that a real violation is DETECTED and NAMED.
|
||||||
|
|
||||||
|
# The harness lives in data-raw/, which is .Rbuildignore'd, so it is absent
|
||||||
|
# from an installed/checked tarball. Source it into an env parented on the
|
||||||
|
# namespace so it resolves the package internals it calls (.build_verb_sql,
|
||||||
|
# .build_suggestions) exactly as it does when run for real.
|
||||||
|
harness_env <- function() {
|
||||||
|
path <- testthat::test_path("..", "..", "data-raw", "measure_signposting_rate.R")
|
||||||
|
skip_if_not(file.exists(path),
|
||||||
|
"data-raw/ is .Rbuildignore'd; harness not present in this tree")
|
||||||
|
env <- new.env(parent = asNamespace("uscogdata"))
|
||||||
|
source(path, local = env)
|
||||||
|
env
|
||||||
|
}
|
||||||
|
|
||||||
|
# A detail frame in exactly the shape .measure_one_query() emits, covering
|
||||||
|
# all four quadrants of the coarse x percode cross-tab.
|
||||||
|
fake_detail <- function() {
|
||||||
|
data.frame(
|
||||||
|
category = c("Corrections", "Other Taxes", "Police", "Fire"),
|
||||||
|
category_type = c("expenditure", "revenue", "expenditure", "expenditure"),
|
||||||
|
canonical_govid = c("121011212191", "472155175824", "011029122489",
|
||||||
|
"041013160815"),
|
||||||
|
n_result_rows = c(0L, 1L, 2L, 6L),
|
||||||
|
coarse_gap_years = c("2011", "", "", ""),
|
||||||
|
fired_coarse = c(TRUE, FALSE, TRUE, FALSE),
|
||||||
|
fired_percode = c(FALSE, TRUE, TRUE, FALSE),
|
||||||
|
coarse_recipes = c("corrections_combined", "", "police_combined", ""),
|
||||||
|
percode_recipes = c("", "t29_license_wide", "police_combined", ""),
|
||||||
|
stringsAsFactors = FALSE
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
test_that(".measure_subset_relation() separates coverage LOST from coverage ADDED", {
|
||||||
|
e <- harness_env()
|
||||||
|
rel <- e$.measure_subset_relation(fake_detail())
|
||||||
|
|
||||||
|
# Row 1 (coarse fired, per-code silent) is the violation; row 2 is the
|
||||||
|
# addition; row 3 agrees; row 4 is silent.
|
||||||
|
expect_false(rel$holds)
|
||||||
|
expect_equal(rel$n_violations, 1L)
|
||||||
|
expect_equal(rel$n_additions, 1L)
|
||||||
|
expect_equal(rel$n_coarse_fired, 2L)
|
||||||
|
expect_equal(rel$n_percode_fired, 2L)
|
||||||
|
expect_equal(rel$n_both, 1L)
|
||||||
|
expect_equal(rel$n_queries, 4L)
|
||||||
|
|
||||||
|
# The violation must be NAMED down to (category, government, year), not
|
||||||
|
# merely counted -- that is what makes it inspectable at Checkpoint R3.
|
||||||
|
expect_equal(rel$violations$category, "Corrections")
|
||||||
|
expect_equal(rel$violations$canonical_govid, "121011212191")
|
||||||
|
expect_equal(rel$violations$coarse_gap_years, "2011")
|
||||||
|
expect_equal(rel$violations$coarse_recipes, "corrections_combined")
|
||||||
|
|
||||||
|
# Inversion guard: an implementation that swapped the two directions
|
||||||
|
# would report the addition as a violation and vice versa.
|
||||||
|
expect_false("Other Taxes" %in% rel$violations$category)
|
||||||
|
expect_equal(rel$additions$category, "Other Taxes")
|
||||||
|
expect_false("Corrections" %in% rel$additions$category)
|
||||||
|
|
||||||
|
# Agreeing and silent queries belong to neither set.
|
||||||
|
expect_false("Police" %in% c(rel$violations$category, rel$additions$category))
|
||||||
|
expect_false("Fire" %in% c(rel$violations$category, rel$additions$category))
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that(".measure_subset_relation() reports holds = TRUE only when nothing fires coarse-only", {
|
||||||
|
e <- harness_env()
|
||||||
|
# Drop the violating row: coarse is now genuinely a subset of per-code.
|
||||||
|
clean <- fake_detail()[-1L, , drop = FALSE]
|
||||||
|
rel <- e$.measure_subset_relation(clean)
|
||||||
|
|
||||||
|
expect_true(rel$holds)
|
||||||
|
expect_equal(rel$n_violations, 0L)
|
||||||
|
expect_equal(nrow(rel$violations), 0L)
|
||||||
|
expect_equal(rel$n_additions, 1L)
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that(".measure_subset_relation() validates its input rather than silently mis-reporting", {
|
||||||
|
e <- harness_env()
|
||||||
|
expect_error(e$.measure_subset_relation("not a data frame"), "must be a data frame")
|
||||||
|
expect_error(e$.measure_subset_relation(fake_detail()[, c("category", "canonical_govid")]),
|
||||||
|
"fired_coarse")
|
||||||
|
bad <- fake_detail()
|
||||||
|
bad$fired_percode[1] <- NA
|
||||||
|
expect_error(e$.measure_subset_relation(bad), "non-NA logicals")
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("the subset report NAMES a coarse-only firing as a violation", {
|
||||||
|
e <- harness_env()
|
||||||
|
txt <- paste(e$.measure_format_subset_report(e$.measure_subset_relation(fake_detail())),
|
||||||
|
collapse = "\n")
|
||||||
|
|
||||||
|
# Stated plainly as a violation, not buried.
|
||||||
|
expect_match(txt, "VIOLATED")
|
||||||
|
expect_match(txt, "COVERAGE LOST")
|
||||||
|
expect_no_match(txt, "HOLDS")
|
||||||
|
# ...and the offending query named, so a human can go look at it.
|
||||||
|
expect_match(txt, "Corrections")
|
||||||
|
expect_match(txt, "121011212191")
|
||||||
|
expect_match(txt, "2011")
|
||||||
|
# ...and the delta explicitly flagged as a net of both directions.
|
||||||
|
expect_match(txt, "NET")
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("the subset report says HOLDS when coarse really is a subset", {
|
||||||
|
e <- harness_env()
|
||||||
|
rel <- e$.measure_subset_relation(fake_detail()[-1L, , drop = FALSE])
|
||||||
|
txt <- paste(e$.measure_format_subset_report(rel), collapse = "\n")
|
||||||
|
|
||||||
|
expect_match(txt, "HOLDS")
|
||||||
|
expect_no_match(txt, "VIOLATED")
|
||||||
|
expect_no_match(txt, "COVERAGE LOST")
|
||||||
|
})
|
||||||
|
|
||||||
|
test_that("a REAL coarse-fires/per-code-silent query is measured and reported as a violation", {
|
||||||
|
skip_if_no_corpus()
|
||||||
|
# Broward County FY2011, Corrections: E05/F05/G05 report SOLELY as
|
||||||
|
# wide-era aggregate rows, which basis = "harmonized" excludes, so the
|
||||||
|
# whole category result is empty -- coarse's trigger. Their modern-only
|
||||||
|
# siblings E04/F04/G04 do not exist as codes at all before 2012, so no
|
||||||
|
# OTHER component can supply per-code's covering evidence and per-code
|
||||||
|
# is structurally unable to fire. This is the disjointness the harness
|
||||||
|
# exists to surface, measured end-to-end through the real git-loaded
|
||||||
|
# coarse arm and the live per-code arm (not a hand-built frame).
|
||||||
|
e <- harness_env()
|
||||||
|
con <- uscogdata:::cog_open()
|
||||||
|
row <- e$.measure_one_query(
|
||||||
|
con,
|
||||||
|
coarse_env = e$.measure_load_git_impl("b0df1ec"),
|
||||||
|
selfcov_env = e$.measure_load_git_impl("da72bf3"),
|
||||||
|
category = "Corrections", category_type = "expenditure",
|
||||||
|
govid = "121011212191", years = 2011L
|
||||||
|
)
|
||||||
|
|
||||||
|
expect_true(row$fired_coarse)
|
||||||
|
expect_false(row$fired_percode)
|
||||||
|
expect_equal(row$n_result_rows, 0L)
|
||||||
|
expect_equal(row$coarse_gap_years, "2011")
|
||||||
|
|
||||||
|
rel <- e$.measure_subset_relation(row)
|
||||||
|
expect_false(rel$holds)
|
||||||
|
expect_equal(rel$n_violations, 1L)
|
||||||
|
expect_equal(rel$violations$canonical_govid, "121011212191")
|
||||||
|
|
||||||
|
txt <- paste(e$.measure_format_subset_report(rel), collapse = "\n")
|
||||||
|
expect_match(txt, "VIOLATED")
|
||||||
|
expect_match(txt, "121011212191")
|
||||||
|
})
|
||||||
@@ -98,11 +98,7 @@ test_that("cog_spending rejects invalid inputs", {
|
|||||||
|
|
||||||
test_that("cog_spending accepts a cog_gov_search result directly", {
|
test_that("cog_spending accepts a cog_gov_search result directly", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
# Unanchored: utility mode matches `name` as a literal substring now, so
|
picks <- cog_gov_search("^BROWARD COUNTY$", state = "FL", type = "county")
|
||||||
# "^...$" would be searched for as those characters rather than read as
|
|
||||||
# anchors (uscogdata#16). Scoped by state and type, the bare name still
|
|
||||||
# resolves to exactly one row.
|
|
||||||
picks <- cog_gov_search("BROWARD COUNTY", state = "FL", type = "county")
|
|
||||||
expect_gt(nrow(picks), 0L)
|
expect_gt(nrow(picks), 0L)
|
||||||
r <- cog_spending(picks, 2020L, "Corrections")
|
r <- cog_spending(picks, 2020L, "Corrections")
|
||||||
expect_equal(unique(r$canonical_govid), "121011212191")
|
expect_equal(unique(r$canonical_govid), "121011212191")
|
||||||
@@ -248,42 +244,24 @@ test_that("basis defaults to 'harmonized' when not passed", {
|
|||||||
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
|
test_that("provenance carries basis + harmonization block with na_rows_excluded", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
with_fixture_corpus({
|
with_fixture_corpus({
|
||||||
# FL state government. The harmonization block is scoped by government,
|
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
||||||
# year and flow prefix -- NOT by category -- so a Corrections query still
|
|
||||||
# counts every E/F/G-prefixed row the harmonized basis drops for having
|
|
||||||
# no harmonized_code. The three that apply here are E21/F21/G21
|
|
||||||
# (Education NEC, SB184-186, "discontinued_na", wide-era window ending
|
|
||||||
# FY2011); the other discontinued_na rulings live outside E/F/G.
|
|
||||||
# See docs/phase_r_harmonization_review.md § 1.3/1.4 and cog_pipeline
|
|
||||||
# data/harmonization_map.csv.
|
|
||||||
r <- cog_spending("120000226351", 2011:2012, "Corrections")
|
|
||||||
prov <- attr(r, "provenance")
|
prov <- attr(r, "provenance")
|
||||||
expect_equal(prov$basis, "harmonized")
|
expect_equal(prov$basis, "harmonized")
|
||||||
expect_true(prov$harmonization$applied)
|
expect_true(prov$harmonization$applied)
|
||||||
|
expect_true(prov$harmonization$na_rows_excluded >= 0L)
|
||||||
|
expect_true(prov$harmonization$na_amount_excluded >= 0)
|
||||||
|
# Data-verified for the v6 fixture (corpus 2026-07-22). The Task 18 map
|
||||||
|
# extension added E/F/G-prefix discontinued_na rulings the earlier pin's
|
||||||
|
# comment predated: E21/F21/G21 (Education NEC local, SB184-186,
|
||||||
|
# "trivial; explicit-NA, full wide-era window"). Broward's 2011 legacy
|
||||||
|
# partition zero-pads exactly those three codes, so this query now
|
||||||
|
# excludes 3 NA-harmonized rows -- all with amt = 0, hence the excluded
|
||||||
|
# AMOUNT stays exactly zero. (The other discontinued_na rulings -- S74,
|
||||||
|
# Z61, X04, X06, the debt-detail family, L24 -- remain outside the
|
||||||
|
# E/F/G/K prefixes.) See docs/phase_r_harmonization_review.md § 1.3/1.4
|
||||||
|
# and cog_pipeline data/harmonization_map.csv E21/F21/G21 rows.
|
||||||
expect_equal(prov$harmonization$na_rows_excluded, 3L)
|
expect_equal(prov$harmonization$na_rows_excluded, 3L)
|
||||||
# $2,825,439 thousands of FY2011 E21 + F21 + G21, reported in full USD.
|
expect_equal(prov$harmonization$na_amount_excluded, 0)
|
||||||
# Pinning a non-zero amount is the point: the earlier Broward anchor's
|
|
||||||
# three rows were all explicit zeros, so the AMOUNT accounting was
|
|
||||||
# asserted only against 0 and could not have caught a bug.
|
|
||||||
expect_equal(prov$harmonization$na_amount_excluded, 2825439 * 1000)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("sparsification removed the wide era's zero-pads from the exclusion count", {
|
|
||||||
skip_if_no_corpus()
|
|
||||||
with_fixture_corpus({
|
|
||||||
# Broward County FY2011 used to carry E21/F21/G21 rows of exactly $0 --
|
|
||||||
# the wide era stored every government x every code, zeros included. The
|
|
||||||
# published corpus no longer does (SB194, cog_pipeline#64), so there is
|
|
||||||
# now nothing for the harmonized basis to exclude. Absence in a
|
|
||||||
# dense_source year means Census published $0; it does not mean the
|
|
||||||
# exclusion machinery stopped working, which the FL state anchor above
|
|
||||||
# proves independently.
|
|
||||||
r <- cog_spending("121011212191", 2011:2012, "Corrections")
|
|
||||||
h <- attr(r, "provenance")$harmonization
|
|
||||||
expect_true(h$applied)
|
|
||||||
expect_equal(h$na_rows_excluded, 0L)
|
|
||||||
expect_equal(h$na_amount_excluded, 0)
|
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
@@ -325,13 +303,11 @@ test_that("provenance$series_break_refs is a populated-when-applicable character
|
|||||||
r <- cog_spending("121011212191", 2020L, "Corrections")
|
r <- cog_spending("121011212191", 2020L, "Corrections")
|
||||||
refs <- attr(r, "provenance")$series_break_refs
|
refs <- attr(r, "provenance")$series_break_refs
|
||||||
expect_type(refs, "character")
|
expect_type(refs, "character")
|
||||||
# No catalogued code-specific series_breaks_pq row falls inside this
|
# No catalogued series_breaks_pq row falls inside this fixture's
|
||||||
# fixture's 2011/2012/2019/2020 window for the codes this query touches
|
# 2011/2012/2019/2020 window for the codes this query touches (E04/G04)
|
||||||
# (E04/G04) -- data-verified; the mechanism itself is what's under test
|
# -- data-verified; the mechanism itself is what's under test here, via
|
||||||
# here, via a query-shaped unit test in test-views.R since the fixture
|
# a query-shaped unit test in test-views.R since the fixture has no
|
||||||
# has no positive case to pin against. Corpus-wide ("ALL") entries never
|
# positive case to pin against.
|
||||||
# appear in this field by construction -- they travel in
|
|
||||||
# corpus_break_refs; see test-corpus-breaks.R.
|
|
||||||
expect_equal(refs, character(0))
|
expect_equal(refs, character(0))
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|||||||
+29
-270
@@ -9,9 +9,7 @@ test_that("all expected views register on session open", {
|
|||||||
expected <- c(
|
expected <- c(
|
||||||
"long", "spending_long", "revenue_long",
|
"long", "spending_long", "revenue_long",
|
||||||
"canonical_fips_xwalk", "summary_categories",
|
"canonical_fips_xwalk", "summary_categories",
|
||||||
"spending_annotated", "revenue_annotated",
|
"spending_annotated", "revenue_annotated"
|
||||||
"ig_long", "ig_annotated",
|
|
||||||
"ig_long_harmonized", "ig_annotated_harmonized"
|
|
||||||
)
|
)
|
||||||
expect_true(all(expected %in% views$table_name))
|
expect_true(all(expected %in% views$table_name))
|
||||||
})
|
})
|
||||||
@@ -36,8 +34,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
|||||||
# {url} exactly as .register_views() does, and executes them -- plus
|
# {url} exactly as .register_views() does, and executes them -- plus
|
||||||
# their 10-long.sql dependency -- against a synthetic hive-partitioned
|
# their 10-long.sql dependency -- against a synthetic hive-partitioned
|
||||||
# parquet tree written to a temp dir. A regression in any predicate (e.g.
|
# parquet tree written to a temp dir. A regression in any predicate (e.g.
|
||||||
# `NOT is_aggregate` dropped, the crosswalk-membership subquery changed,
|
# `NOT is_aggregate` dropped, the prefix list changed, the NULL guard
|
||||||
# the NULL guard removed) would change which of the rows below survive.
|
# removed) would change which of the rows below survive.
|
||||||
#
|
#
|
||||||
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
|
# The synthetic parquet is written with DuckDB's own COPY ... TO (FORMAT
|
||||||
# PARQUET) rather than the arrow package: this package has no arrow
|
# PARQUET) rather than the arrow package: this package has no arrow
|
||||||
@@ -61,45 +59,25 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
|||||||
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
|
('spend-B', 'E38', 50, false, 'E36'), -- collapse-fold: passes every predicate, renamed to E36
|
||||||
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
|
('spend-C', 'E05', 999999, true, 'E05'), -- excluded ONLY by `NOT is_aggregate`
|
||||||
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
('spend-D', 'E99', 888888, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
||||||
-- 'S74' and 'Z61' are classified `balance` in the synthetic
|
-- 'S74' is outside BOTH flow families (E/F/G/K spending and
|
||||||
-- crosswalk below (mirroring the real corpus's own non-flow codes),
|
-- T/A/U/B/C/D revenue -- it mirrors the real corpus's own
|
||||||
-- so each is excluded from its view ONLY by the crosswalk-membership
|
-- non-flow-type codes like S74/Z61), so it can only leak into
|
||||||
-- subquery -- the mechanism that replaced the prefix allowlists
|
-- EITHER view via the E/F/G/K or T/A/U/B/C/D prefix filter, never
|
||||||
-- (uscogdata#11) and keeps balance stocks out of both flows
|
-- both at once -- a prefix drawn from the other view's own family
|
||||||
-- (uscogdata#25).
|
-- (e.g. a real T-code for the spending row) would incorrectly
|
||||||
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by crosswalk membership (balance)
|
-- leak into the other view's assertion below and not discriminate
|
||||||
-- Revenue rows, exercised against revenue_long_harmonized:
|
-- the predicate under test.
|
||||||
|
('spend-E', 'S74', 777777, false, 'S74'), -- excluded ONLY by the E/F/G/K prefix filter
|
||||||
|
-- Revenue (T/A/U/B/C/D) rows, exercised against revenue_long_harmonized:
|
||||||
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
|
('rev-A', 'U11', 200, false, 'U11'), -- control: passes every predicate as-is
|
||||||
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
|
('rev-B', 'U10', 25, false, 'U11'), -- collapse-fold: passes every predicate, renamed to U11
|
||||||
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
|
('rev-C', 'T29', 555555, true, 'T29'), -- excluded ONLY by `NOT is_aggregate`
|
||||||
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
('rev-D', 'T88', 444444, false, NULL), -- excluded by `harmonized_code IS NOT NULL`
|
||||||
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by crosswalk membership (balance)
|
('rev-E', 'Z61', 333333, false, 'Z61') -- excluded ONLY by the T/A/U/B/C/D prefix filter
|
||||||
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
|
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
|
||||||
) TO %s (FORMAT PARQUET)
|
) TO %s (FORMAT PARQUET)
|
||||||
", uscogdata:::.sql_lit_chr(part_path)))
|
", uscogdata:::.sql_lit_chr(part_path)))
|
||||||
|
|
||||||
# The flow views classify by membership in summary_categories, so the
|
|
||||||
# synthetic corpus needs one too. Every flow code above is a member of its
|
|
||||||
# own flow (so is_aggregate / NULL-harmonized exclusions stay the SOLE
|
|
||||||
# excluder for those rows); S74/Z61 are members but classified balance, so
|
|
||||||
# membership itself is what excludes them.
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('E36', 'Water Utilities', 'expenditure', 'operations', NULL),
|
|
||||||
('E38', 'Water Utilities', 'expenditure', 'operations', NULL),
|
|
||||||
('E05', 'Corrections', 'expenditure', 'operations', NULL),
|
|
||||||
('E99', 'Other', 'expenditure', 'operations', NULL),
|
|
||||||
('S74', 'Fund Balances', 'balance', NULL, NULL),
|
|
||||||
('U11', 'Interest Earnings','revenue', NULL, 'own_source'),
|
|
||||||
('U10', 'Interest Earnings','revenue', NULL, 'own_source'),
|
|
||||||
('T29', 'Other Taxes', 'revenue', NULL, 'own_source'),
|
|
||||||
('T88', 'Other Taxes', 'revenue', NULL, 'own_source'),
|
|
||||||
('Z61', 'Fund Balances', 'balance', NULL, NULL)
|
|
||||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
|
||||||
|
|
||||||
sql_dir <- system.file("sql", package = "uscogdata")
|
sql_dir <- system.file("sql", package = "uscogdata")
|
||||||
.read_view_sql <- function(filename) {
|
.read_view_sql <- function(filename) {
|
||||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
||||||
@@ -109,7 +87,6 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
|||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
con <- DBI::dbConnect(duckdb::duckdb())
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
||||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
||||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
|
DBI::dbExecute(con, .read_view_sql("22-spending_long_harmonized.sql"))
|
||||||
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
|
DBI::dbExecute(con, .read_view_sql("23-revenue_long_harmonized.sql"))
|
||||||
|
|
||||||
@@ -118,8 +95,8 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
|||||||
GROUP BY item_code ORDER BY item_code"
|
GROUP BY item_code ORDER BY item_code"
|
||||||
)
|
)
|
||||||
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
|
# Exactly one surviving row: spend-C (aggregate), spend-D (NULL
|
||||||
# harmonized_code), and spend-E (balance, not an expenditure member) must
|
# harmonized_code), and spend-E (wrong prefix family) must all be gone,
|
||||||
# all be gone, and spend-A + spend-B must be folded together under E36.
|
# and spend-A + spend-B must be folded together under E36.
|
||||||
expect_equal(nrow(spend), 1L)
|
expect_equal(nrow(spend), 1L)
|
||||||
expect_equal(spend$item_code, "E36")
|
expect_equal(spend$item_code, "E36")
|
||||||
expect_equal(spend$amt, 150)
|
expect_equal(spend$amt, 150)
|
||||||
@@ -133,95 +110,12 @@ test_that("inst/sql/22- and 23- harmonized views enforce every WHERE predicate (
|
|||||||
expect_equal(rev$amt, 225)
|
expect_equal(rev$amt, 225)
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("inst/sql/24- and 25- IG views retain aggregates, COALESCE NULL harmonized_code, and exclude the L-- family total (real SQL text, synthetic parquet)", {
|
|
||||||
# ig_long / ig_long_harmonized have the subtlest predicates in the package:
|
|
||||||
# a deliberately ABSENT `NOT is_aggregate` (unlike every other *_long view),
|
|
||||||
# and COALESCE(harmonized_code, item_code) instead of a plain
|
|
||||||
# `harmonized_code IS NOT NULL` filter. The only end-to-end guard on this
|
|
||||||
# today is bound to AL state / 2011 / Education K-12, where M12 happens to
|
|
||||||
# be the sole IG code present -- regenerate the fixture without that one
|
|
||||||
# row and the guard would die silently while staying green. As with the
|
|
||||||
# 22-/23- test above, this reads the real inst/sql/24-/25- text off disk
|
|
||||||
# and executes it against a synthetic hive-partitioned parquet tree, so a
|
|
||||||
# regression in either predicate changes which rows survive.
|
|
||||||
skip_if_no_corpus()
|
|
||||||
|
|
||||||
tmp <- withr::local_tempdir()
|
|
||||||
part_dir <- file.path(tmp, "data", "long", "year=2004")
|
|
||||||
dir.create(part_dir, recursive = TRUE)
|
|
||||||
part_path <- file.path(part_dir, "part-0.parquet")
|
|
||||||
|
|
||||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('ig-A', 'M04', 100, false, 'M04'), -- control: passes through as-is
|
|
||||||
('ig-B', 'M38', 50, false, 'M36'), -- fold control: real SB012 rule, renamed to M36 under harmonized basis
|
|
||||||
('ig-C', 'M47', 99999, true, NULL), -- legacy aggregate, NO harmonized_code: must survive BOTH views
|
|
||||||
('ig-D', 'L--', 55555, false, 'L--'), -- family total: deliberately NOT a crosswalk member, excluded from BOTH views
|
|
||||||
('ig-E', 'T29', 44444, false, 'T29') -- revenue member, not intergovernmental: excluded from BOTH views
|
|
||||||
) AS t(canonical_govid, item_code, amt, is_aggregate, harmonized_code)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(part_path)))
|
|
||||||
|
|
||||||
# The IG views classify by summary_categories membership
|
|
||||||
# (spend_subtype = 'intergovernmental'). L-- is deliberately absent --
|
|
||||||
# exactly as it is from the real crosswalk -- which is what excludes it.
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('M04', 'Corrections', 'expenditure', 'intergovernmental', NULL),
|
|
||||||
('M38', 'Health', 'expenditure', 'intergovernmental', NULL),
|
|
||||||
('M36', 'Health', 'expenditure', 'intergovernmental', NULL),
|
|
||||||
('M47', 'IG Other', 'expenditure', 'intergovernmental', NULL),
|
|
||||||
('T29', 'Other Taxes', 'revenue', NULL, 'own_source')
|
|
||||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(file.path(tmp, "data", "summary_categories.parquet"))))
|
|
||||||
|
|
||||||
sql_dir <- system.file("sql", package = "uscogdata")
|
|
||||||
.read_view_sql <- function(filename) {
|
|
||||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
|
||||||
gsub("\\{url\\}", paste0(tmp, "/"), txt, fixed = FALSE)
|
|
||||||
}
|
|
||||||
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(con, .read_view_sql("10-long.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("11-summary_categories.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("24-ig_long.sql"))
|
|
||||||
DBI::dbExecute(con, .read_view_sql("25-ig_long_harmonized.sql"))
|
|
||||||
|
|
||||||
raw <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT item_code, SUM(amt) AS amt FROM ig_long
|
|
||||||
GROUP BY item_code ORDER BY item_code"
|
|
||||||
)
|
|
||||||
# L-- (family total, not a member) and T29 (revenue, not IG) are gone; the
|
|
||||||
# aggregate row M47 survives -- proof `NOT is_aggregate` is absent from
|
|
||||||
# ig_long.
|
|
||||||
expect_equal(raw$item_code, c("M04", "M38", "M47"))
|
|
||||||
expect_equal(raw$amt, c(100, 50, 99999))
|
|
||||||
|
|
||||||
harmonized <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT item_code, SUM(amt) AS amt FROM ig_long_harmonized
|
|
||||||
GROUP BY item_code ORDER BY item_code"
|
|
||||||
)
|
|
||||||
# M38 folds to M36 (real harmonized_code present); M47 keeps its raw code
|
|
||||||
# via COALESCE(NULL, 'M47') -- proof the aggregate row is NOT dropped by
|
|
||||||
# a plain `harmonized_code IS NOT NULL` filter. L-- and T29 stay excluded.
|
|
||||||
expect_equal(harmonized$item_code, c("M04", "M36", "M47"))
|
|
||||||
expect_equal(harmonized$amt, c(100, 50, 99999))
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
test_that(".build_series_break_refs matches fin_code + break_year window", {
|
||||||
# No CODE-SPECIFIC series_breaks_pq row falls inside the bundled fixture's
|
# No series_breaks_pq row falls inside the bundled fixture's 2011-2020
|
||||||
# 2011-2020 window (data-verified; see the "series_break_refs" test in
|
# window (data-verified; see the "series_break_refs" test in
|
||||||
# test-spending.R), so this proves the matching logic itself against the
|
# test-spending.R), so this proves the matching logic itself against the
|
||||||
# live view + a synthetic year window that DOES hit a cataloged break
|
# live view + a synthetic year window that DOES hit a cataloged break
|
||||||
# (SB075, fin_code E62, break_year 2005). The corpus-wide entries are a
|
# (SB075, fin_code E62, break_year 2005).
|
||||||
# separate path with its own coverage -- SB194 does sit at 2012, inside
|
|
||||||
# the fixture window; see test-corpus-breaks.R.
|
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
con <- cog_open()
|
con <- cog_open()
|
||||||
on.exit(cog_close())
|
on.exit(cog_close())
|
||||||
@@ -258,134 +152,14 @@ test_that("schema v5 harmonization views register when the corpus supports them"
|
|||||||
expect_true(all(expected_v5 %in% views$table_name))
|
expect_true(all(expected_v5 %in% views$table_name))
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that(".harmonization_view_files guard is necessary: registration against a v4-shaped corpus (no harmonized_code column at all) succeeds only because the harmonized views are skipped", {
|
test_that("spending_long filters to E/F/G/K prefixes and excludes aggregates", {
|
||||||
# with_doctored_schema_version() (used elsewhere in this suite) only
|
|
||||||
# rewrites manifest.json's schema_version -- the underlying `long` parquet
|
|
||||||
# is still the bundled v6 fixture, which DOES have a harmonized_code
|
|
||||||
# column, so it only proves the skip *happens*, not that it is *required*.
|
|
||||||
# This test builds a genuinely v4-shaped corpus: `long` has no
|
|
||||||
# harmonized_code column at all, matching a real pre-Phase-R2 publish
|
|
||||||
# tree, and then shows two things: (1) the real .register_views(), gated
|
|
||||||
# on manifest$schema_version, registers cleanly against it; (2) the exact
|
|
||||||
# SQL text of a gated file (25-ig_long_harmonized.sql), executed directly
|
|
||||||
# against the same corpus without the gate, fails -- proving the gate is
|
|
||||||
# load-bearing, not incidental.
|
|
||||||
tmp <- withr::local_tempdir()
|
|
||||||
part_dir <- file.path(tmp, "data", "long", "year=2004")
|
|
||||||
dir.create(part_dir, recursive = TRUE)
|
|
||||||
part_path <- file.path(part_dir, "part-0.parquet")
|
|
||||||
|
|
||||||
write_con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(write_con, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('gov-1', 'E36', 100, false, 500000, 2020)
|
|
||||||
) AS t(canonical_govid, item_code, amt, is_aggregate, population, popyear)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(part_path)))
|
|
||||||
|
|
||||||
xwalk_path <- file.path(tmp, "data", "canonical_fips_xwalk.parquet")
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('gov-1', 'Test Gov', 1, 'County', '01', '001', NULL, 500000)
|
|
||||||
) AS t(canonical_govid, gov_name, govs_type, type_label, fips_state,
|
|
||||||
fips_county, fips_place, population_acs)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(xwalk_path)))
|
|
||||||
|
|
||||||
cats_path <- file.path(tmp, "data", "summary_categories.parquet")
|
|
||||||
DBI::dbExecute(write_con, sprintf("
|
|
||||||
COPY (
|
|
||||||
SELECT * FROM (VALUES
|
|
||||||
('E36', 'Test Category', 'expenditure', 'direct', NULL)
|
|
||||||
) AS t(item_code, category, category_type, spend_subtype, revenue_subtype)
|
|
||||||
) TO %s (FORMAT PARQUET)
|
|
||||||
", uscogdata:::.sql_lit_chr(cats_path)))
|
|
||||||
|
|
||||||
# Confirm the synthetic `long` genuinely lacks harmonized_code (not just
|
|
||||||
# NULL values -- the column itself must be absent) before trusting the
|
|
||||||
# rest of this test.
|
|
||||||
cols <- DBI::dbGetQuery(write_con, sprintf(
|
|
||||||
"DESCRIBE SELECT * FROM read_parquet(%s)", uscogdata:::.sql_lit_chr(part_path)
|
|
||||||
))$column_name
|
|
||||||
expect_false("harmonized_code" %in% cols)
|
|
||||||
|
|
||||||
url <- paste0(tmp, "/")
|
|
||||||
|
|
||||||
# (1) Full .register_views() against this v4-shaped corpus must succeed --
|
|
||||||
# this is the behavior the guard exists to protect.
|
|
||||||
con <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con, shutdown = TRUE), add = TRUE)
|
|
||||||
expect_no_error(
|
|
||||||
uscogdata:::.register_views(con, url, manifest = list(schema_version = 4L))
|
|
||||||
)
|
|
||||||
views <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT table_name FROM information_schema.tables
|
|
||||||
WHERE table_schema = 'main' AND table_type = 'VIEW'")$table_name
|
|
||||||
expect_true(all(c("ig_long", "ig_annotated", "spending_annotated") %in% views))
|
|
||||||
expect_false(any(c("ig_long_harmonized", "ig_annotated_harmonized",
|
|
||||||
"spending_long_harmonized") %in% views))
|
|
||||||
|
|
||||||
# (2) Prove the gate is load-bearing: the exact SQL text of the skipped
|
|
||||||
# file, executed directly (bypassing .register_views()'s schema_version
|
|
||||||
# check) against the SAME corpus, fails because it references
|
|
||||||
# long.harmonized_code, a column this corpus's `long` does not have.
|
|
||||||
sql_dir <- system.file("sql", package = "uscogdata")
|
|
||||||
.read_view_sql <- function(filename) {
|
|
||||||
txt <- paste(readLines(file.path(sql_dir, filename), warn = FALSE), collapse = "\n")
|
|
||||||
gsub("\\{url\\}", url, txt, fixed = FALSE)
|
|
||||||
}
|
|
||||||
con2 <- DBI::dbConnect(duckdb::duckdb())
|
|
||||||
on.exit(DBI::dbDisconnect(con2, shutdown = TRUE), add = TRUE)
|
|
||||||
DBI::dbExecute(con2, .read_view_sql("10-long.sql"))
|
|
||||||
expect_error(DBI::dbExecute(con2, .read_view_sql("25-ig_long_harmonized.sql")))
|
|
||||||
|
|
||||||
# Reconciling this test with the C2 guard (expenditure-concept review):
|
|
||||||
# `ig_annotated`/`spending_annotated` registering cleanly above proves
|
|
||||||
# only that CREATE VIEW binds against a `summary_categories` with no M/L
|
|
||||||
# rows at all (this synthetic corpus's own summary_categories has a
|
|
||||||
# single E36 row, see the COPY above) -- a LEFT JOIN never fails to
|
|
||||||
# resolve regardless of what the joined-to table contains. It does NOT
|
|
||||||
# mean querying expenditure_concept = "total" against this shape is safe:
|
|
||||||
# exactly this corpus (schema_version reported as supported, but
|
|
||||||
# summary_categories predates the M/L rows cog_pipeline PR #59 added) is
|
|
||||||
# what .require_ig_categories() exists to catch at the *verb* level,
|
|
||||||
# since PR #59 shipped those rows with no schema_version bump. Confirm
|
|
||||||
# the new runtime guard actually fires against this same `con`.
|
|
||||||
expect_error(
|
|
||||||
uscogdata:::.require_ig_categories(con),
|
|
||||||
class = "uscogdata_ig_categories_unsupported"
|
|
||||||
)
|
|
||||||
})
|
|
||||||
|
|
||||||
test_that("spending_long carries exactly the non-IG expenditure crosswalk codes and excludes aggregates", {
|
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
con <- cog_open()
|
con <- cog_open()
|
||||||
on.exit(cog_close())
|
on.exit(cog_close())
|
||||||
|
prefixes <- DBI::dbGetQuery(con,
|
||||||
# Classification is crosswalk membership, not prefixes (uscogdata#11):
|
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM spending_long"
|
||||||
# every row's code must classify as expenditure and never as
|
)$pfx
|
||||||
# intergovernmental (which lives in ig_long).
|
expect_true(all(prefixes %in% c("E", "F", "G", "K")))
|
||||||
stray <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT s.item_code
|
|
||||||
FROM spending_long s
|
|
||||||
LEFT JOIN summary_categories c USING (item_code)
|
|
||||||
WHERE c.category_type IS DISTINCT FROM 'expenditure'
|
|
||||||
OR c.spend_subtype = 'intergovernmental'"
|
|
||||||
)$item_code
|
|
||||||
expect_length(stray, 0L)
|
|
||||||
|
|
||||||
# Balance codes are stocks, not flows -- they must never appear in a
|
|
||||||
# spending result (uscogdata#25). Prefix filtering could not guarantee
|
|
||||||
# this (W/X/Y/Z balance codes share letters with flow codes).
|
|
||||||
balance_n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT count(*) AS n FROM spending_long WHERE item_code IN (
|
|
||||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
|
|
||||||
)"
|
|
||||||
)$n
|
|
||||||
expect_equal(balance_n, 0)
|
|
||||||
|
|
||||||
agg_count <- DBI::dbGetQuery(con,
|
agg_count <- DBI::dbGetQuery(con,
|
||||||
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
|
"SELECT count(*) AS n FROM spending_long WHERE is_aggregate"
|
||||||
@@ -393,29 +167,14 @@ test_that("spending_long carries exactly the non-IG expenditure crosswalk codes
|
|||||||
expect_equal(agg_count, 0)
|
expect_equal(agg_count, 0)
|
||||||
})
|
})
|
||||||
|
|
||||||
test_that("revenue_long carries exactly the revenue crosswalk codes and excludes aggregates", {
|
test_that("revenue_long filters to T/A/U/B/C/D prefixes and excludes aggregates", {
|
||||||
skip_if_no_corpus()
|
skip_if_no_corpus()
|
||||||
con <- cog_open()
|
con <- cog_open()
|
||||||
on.exit(cog_close())
|
on.exit(cog_close())
|
||||||
|
prefixes <- DBI::dbGetQuery(con,
|
||||||
# The view carries EVERY revenue subtype; which of Census's two published
|
"SELECT DISTINCT LEFT(item_code, 1) AS pfx FROM revenue_long"
|
||||||
# concepts a query returns is decided per `revenue_concept` in R
|
)$pfx
|
||||||
# (uscogdata#12), exactly as `expenditure_concept` narrows spending_long.
|
expect_true(all(prefixes %in% c("T", "A", "U", "B", "C", "D")))
|
||||||
stray <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT DISTINCT s.item_code
|
|
||||||
FROM revenue_long s
|
|
||||||
LEFT JOIN summary_categories c USING (item_code)
|
|
||||||
WHERE c.category_type IS DISTINCT FROM 'revenue'"
|
|
||||||
)$item_code
|
|
||||||
expect_length(stray, 0L)
|
|
||||||
|
|
||||||
# No balance stock ever appears in a revenue result (uscogdata#25).
|
|
||||||
balance_n <- DBI::dbGetQuery(con,
|
|
||||||
"SELECT count(*) AS n FROM revenue_long WHERE item_code IN (
|
|
||||||
SELECT item_code FROM summary_categories WHERE category_type = 'balance'
|
|
||||||
)"
|
|
||||||
)$n
|
|
||||||
expect_equal(balance_n, 0)
|
|
||||||
|
|
||||||
agg_count <- DBI::dbGetQuery(con,
|
agg_count <- DBI::dbGetQuery(con,
|
||||||
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
|
"SELECT count(*) AS n FROM revenue_long WHERE is_aggregate"
|
||||||
|
|||||||
@@ -13,8 +13,6 @@ knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>")
|
|||||||
|
|
||||||
# Why per-year population matters
|
# Why per-year population matters
|
||||||
|
|
||||||
A note on units first, since every figure below is a rate: the numerator is in **full US dollars**. The raw Census files report **thousands of dollars** and the corpus keeps them that way in its own `amt` column, but `cog_spending()` and `cog_revenue()` multiply by 1000 on the way out, so `amt_per_capita_nominal` is already dollars per person. Do not scale it again.
|
|
||||||
|
|
||||||
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change.
|
||||||
|
|
||||||
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%.
|
||||||
|
|||||||
@@ -1,256 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Total spending: Primary, Direct, Total, and when each is right"
|
|
||||||
output: rmarkdown::html_vignette
|
|
||||||
vignette: >
|
|
||||||
%\VignetteIndexEntry{Total spending: Primary, Direct, Total, and when each is right}
|
|
||||||
%\VignetteEngine{knitr::rmarkdown}
|
|
||||||
%\VignetteEncoding{UTF-8}
|
|
||||||
---
|
|
||||||
|
|
||||||
```{r setup, include = FALSE}
|
|
||||||
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
|
|
||||||
```
|
|
||||||
|
|
||||||
# Two questions that sound the same but aren't
|
|
||||||
|
|
||||||
"Total spending" means two different things depending on whether the question
|
|
||||||
is about one government or several:
|
|
||||||
|
|
||||||
1. **"What did my county spend in total, a decade ago vs today?"** — one
|
|
||||||
government, tracked over time. Any concept answers this correctly, as
|
|
||||||
long as the same concept is used for both years.
|
|
||||||
2. **"How do all the counties in my state compare, a decade ago vs today,
|
|
||||||
against the neighboring state?"** — several governments, summed together.
|
|
||||||
Here only a non-intergovernmental concept (`primary` or `direct`) gives
|
|
||||||
the right answer; summing `total` across governments double-counts money
|
|
||||||
that passes between them.
|
|
||||||
|
|
||||||
`cog_spending()`'s `expenditure_concept` argument controls which of these a
|
|
||||||
query answers, via three nested concepts defined as sets of the crosswalk's
|
|
||||||
`spend_subtype` values (never item-code first letters — the letter `Y` alone
|
|
||||||
spans revenue, expenditure, and balance codes):
|
|
||||||
|
|
||||||
- `"primary"` (the default) — the government's own service provision:
|
|
||||||
`operations` + `capital` + `assistance`.
|
|
||||||
- `"direct"` — Census's published Direct Expenditure: `primary` plus
|
|
||||||
`interest` on debt and `insurance_benefits` (e.g. pension payments).
|
|
||||||
- `"total"` — `direct` plus the `intergovernmental` leg.
|
|
||||||
|
|
||||||
This vignette walks through both questions with code that actually runs
|
|
||||||
against the package's bundled fixture corpus, then explains why the second
|
|
||||||
question refuses `"total"` outright.
|
|
||||||
|
|
||||||
Before any of the numbers below: every amount column here — `amt_nominal`,
|
|
||||||
`amt_real`, and their `amt_per_capita_*` counterparts — is in **full US
|
|
||||||
dollars**. The raw Census files report **thousands of dollars** and the
|
|
||||||
corpus preserves that in its own `amt` column, but the verbs multiply by 1000
|
|
||||||
on the way out. So `amt_nominal = 1317000` means $1.317 million, not $1.317
|
|
||||||
billion. Do not scale it again.
|
|
||||||
|
|
||||||
```{r}
|
|
||||||
library(uscogdata)
|
|
||||||
|
|
||||||
# Point at the bundled offline fixture (years 2011, 2012, 2019, 2020, all 50
|
|
||||||
# states) so this vignette knits without network access. In real use,
|
|
||||||
# USCOGDATA_URL is instead set to the published corpus URL -- see README.md.
|
|
||||||
Sys.setenv(USCOGDATA_URL = paste0(
|
|
||||||
system.file("extdata/fixture_corpus", package = "uscogdata"), "/"
|
|
||||||
))
|
|
||||||
```
|
|
||||||
|
|
||||||
The fixture doesn't carry 2017 or the present year, so the examples below use
|
|
||||||
the closest years it does ship -- **2012 and 2020** -- in place of "2017 vs
|
|
||||||
today" / "ten years ago vs today". Point `USCOGDATA_URL` at the published
|
|
||||||
corpus and swap in real years; the mechanics are identical.
|
|
||||||
|
|
||||||
# Archetype 1: one government's own trend
|
|
||||||
|
|
||||||
For a single government, `total` is a legitimate way to describe "everything
|
|
||||||
this government spent, including money it handed to other governments to
|
|
||||||
spend on its behalf":
|
|
||||||
|
|
||||||
```{r}
|
|
||||||
al_total <- cog_spending(
|
|
||||||
"010000226085", # Alabama, the state government
|
|
||||||
years = c(2012, 2020),
|
|
||||||
category = "Highways",
|
|
||||||
expenditure_concept = "total"
|
|
||||||
)
|
|
||||||
al_total
|
|
||||||
```
|
|
||||||
|
|
||||||
The `intergovernmental` rows are what `"total"` adds on top of the
|
|
||||||
non-intergovernmental subtypes (here `capital` + `operations`): Alabama's
|
|
||||||
own payments out to counties and cities for highway work. Because this
|
|
||||||
query only ever concerns Alabama, including that piece is safe -- there's
|
|
||||||
no other government's number it could be double-counted against.
|
|
||||||
|
|
||||||
`"primary"` (the default) answers the same trend question just as validly
|
|
||||||
(for Highways, which maps only to operations/capital codes, `"primary"` and
|
|
||||||
`"direct"` coincide -- there is no highway-specific interest or insurance
|
|
||||||
benefit to add):
|
|
||||||
|
|
||||||
```{r}
|
|
||||||
al_primary <- cog_spending(
|
|
||||||
"010000226085", years = c(2012, 2020), category = "Highways"
|
|
||||||
# expenditure_concept = "primary" is the default; shown here for contrast
|
|
||||||
)
|
|
||||||
al_primary
|
|
||||||
```
|
|
||||||
|
|
||||||
Both are internally consistent series. What breaks the comparison is
|
|
||||||
**switching concepts between the two years being compared** -- e.g. `direct`
|
|
||||||
for 2012 and `total` for 2020 -- which manufactures a trend that isn't
|
|
||||||
really there. Pick one concept for a given question and hold it fixed across
|
|
||||||
every year in the series.
|
|
||||||
|
|
||||||
# Archetype 2: a cross-government rollup
|
|
||||||
|
|
||||||
`cog_geographic_rollup()` sums spending across state/county/city layers for
|
|
||||||
a place. Its default is `"primary"`, and (as shown below) it accepts only
|
|
||||||
the non-intergovernmental concepts, `"primary"` and `"direct"`:
|
|
||||||
|
|
||||||
```{r}
|
|
||||||
fl_rollup <- cog_geographic_rollup(
|
|
||||||
govids = list(
|
|
||||||
state = "120000226351", # Florida
|
|
||||||
county = c("121011212191", "121099101897") # Broward + Palm Beach
|
|
||||||
),
|
|
||||||
category = "Highways",
|
|
||||||
years = c(2012, 2020)
|
|
||||||
)
|
|
||||||
fl_rollup
|
|
||||||
```
|
|
||||||
|
|
||||||
For the neighboring state, the comparison is a single government, so it's a
|
|
||||||
plain `cog_spending()` call rather than a rollup:
|
|
||||||
|
|
||||||
```{r}
|
|
||||||
ga_state <- cog_spending(
|
|
||||||
"130000226087", years = c(2012, 2020), category = "Highways" # Georgia
|
|
||||||
)
|
|
||||||
ga_state
|
|
||||||
```
|
|
||||||
|
|
||||||
Now the same rollup, but asking for `expenditure_concept = "total"`:
|
|
||||||
|
|
||||||
```{r, error = TRUE}
|
|
||||||
cog_geographic_rollup(
|
|
||||||
govids = list(state = "120000226351", county = "121011212191"),
|
|
||||||
category = "Highways",
|
|
||||||
years = 2020,
|
|
||||||
expenditure_concept = "total"
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
`cog_geographic_rollup()` (and `cog_peer_compare()`, for the same reason)
|
|
||||||
refuses `"total"` outright rather than silently returning an inflated
|
|
||||||
number. The next section is why.
|
|
||||||
|
|
||||||
# The mechanism
|
|
||||||
|
|
||||||
Suppose Alabama gives a county $10M toward a highway project. That $10M
|
|
||||||
shows up **twice** in the underlying corpus:
|
|
||||||
|
|
||||||
- Once on Alabama's own record, coded `M44` ("to local governments,
|
|
||||||
Highways") -- Alabama's intergovernmental leg.
|
|
||||||
- Again on the county's record, coded `E44` / `F44` ("Highways, current
|
|
||||||
operations" / "capital outlay") -- the county's direct spending, because
|
|
||||||
the county is the government that actually lets the contract and pays the
|
|
||||||
paving crew.
|
|
||||||
|
|
||||||
`primary` and `direct` (the crosswalk's non-intergovernmental expenditure
|
|
||||||
subtypes) only ever count the second of those -- the government that
|
|
||||||
actually did the spending. `total` (Direct plus the intergovernmental leg)
|
|
||||||
counts the first one *as well*, which is exactly right for describing
|
|
||||||
Alabama's own budget: Alabama's `total` genuinely includes the $10M it
|
|
||||||
committed to highways, whether it built the road itself or paid the county
|
|
||||||
to. But sum `total` across Alabama **and** the county, and that $10M is
|
|
||||||
counted twice -- once as Alabama's payment out, once as the county's
|
|
||||||
spending in -- reporting $20M of highway work for $10M actually spent.
|
|
||||||
|
|
||||||
This is exactly the shape of query `cog_geographic_rollup()` exists to run
|
|
||||||
(summing across layers of government), so it refuses `"total"` rather than
|
|
||||||
silently overstating every multi-layer figure it produces.
|
|
||||||
|
|
||||||
# How big is the risk in practice
|
|
||||||
|
|
||||||
Intergovernmental transfers aren't evenly distributed by government type.
|
|
||||||
Measured against the bundled fixture corpus (all 50 states, each of its
|
|
||||||
four years -- 2011, 2012, 2019, 2020), intergovernmental spending as a
|
|
||||||
share of a government's own Direct spending is:
|
|
||||||
|
|
||||||
| Government type | Intergovernmental / Direct |
|
|
||||||
|---|---|
|
|
||||||
| State | 33.1%-40.5% (varies by year; 36.2% pooled across all four) |
|
|
||||||
| County | 3.3%-4.8% (varies by year) |
|
|
||||||
| City | 2.4%-2.9% (varies by year) |
|
|
||||||
|
|
||||||
So the Direct/Total choice matters overwhelmingly for **state**
|
|
||||||
governments -- a state's Total genuinely differs from its Direct by more
|
|
||||||
than a third, while for a county or city the two are close. (The state
|
|
||||||
share is much larger than pre-#11 measurements suggested, because the
|
|
||||||
intergovernmental leg now correctly includes the `Q11`/`Q12`/`Q18` state
|
|
||||||
payments to school systems -- for most states the single largest transfer
|
|
||||||
they make.) That's also why the mistake this vignette warns about is easy
|
|
||||||
to make unnoticed at the county/city level and costly at the state level:
|
|
||||||
rolling up every government using `total` instead of `primary`/`direct`
|
|
||||||
overstates the FY2019 figure by 24.1% for Alabama and 23.2% nationally.
|
|
||||||
|
|
||||||
# Why Total = Direct + M + L + Q, not Direct + M
|
|
||||||
|
|
||||||
It's tempting to assume `total` only needs to add `M`. But the
|
|
||||||
intergovernmental leg has three families, all money the queried government
|
|
||||||
itself pays **out** -- they're not different accounts of a receiving
|
|
||||||
government's revenue. `M` is what it pays to other **local** governments
|
|
||||||
(e.g. a county paying a city for a shared paving contract); `L` is what it
|
|
||||||
pays **up** to its **state** government (e.g. a county's contribution to a
|
|
||||||
state-administered program); and `Q11`/`Q12`/`Q18` are a state's payments
|
|
||||||
to **school systems** (K-12 and higher-ed aid -- for most states the
|
|
||||||
single largest transfer they make, and the piece the pre-#11 prefix
|
|
||||||
allowlist silently dropped, finding F-017). A government's Total genuinely
|
|
||||||
includes every leg it pays, because each is its own spending, just routed
|
|
||||||
to a different kind of recipient. On the bundled fixture corpus (all 50
|
|
||||||
states, 2011/2012/2019/2020), `L` is 0 for state governments (a state has
|
|
||||||
no "payments to the state government" leg of its own) but is 43%-51% the
|
|
||||||
size of `M` for counties (varies by year) and 144%-189% the size of `M`
|
|
||||||
for cities (varies by year; 166% pooled across all four) -- so a `total`
|
|
||||||
that omitted `L` would silently undercount Total specifically for local
|
|
||||||
governments, and for cities `L` is often the *larger* of the two legs.
|
|
||||||
`cog_spending(expenditure_concept = "total")` includes every leg
|
|
||||||
(excluding the `L--` family-total rollup row, which would double-count its
|
|
||||||
own components).
|
|
||||||
|
|
||||||
# Composition rules
|
|
||||||
|
|
||||||
- `expenditure_concept` (whose spending counts -- Primary, Direct, or
|
|
||||||
Direct plus intergovernmental) is **orthogonal** to `basis` (which
|
|
||||||
vintage of the item-code space a query resolves against --
|
|
||||||
`"harmonized"` vs `"raw"`).
|
|
||||||
They combine freely: `expenditure_concept = "total", basis = "raw"` is a
|
|
||||||
valid, meaningful query, and so is every other pairing.
|
|
||||||
- `expenditure_concept = "total"` is **mutually exclusive** with `recipe`: a
|
|
||||||
recipe already defines its own component codes (some recipes have their
|
|
||||||
own matching intergovernmental counterpart recipe instead -- see
|
|
||||||
`cog_recipes()` and the "firing suggestion" notes surfaced in
|
|
||||||
`cog_spending()`'s provenance), so layering a second, generic `total`
|
|
||||||
union on top of a recipe query has no well-defined meaning. Passing both
|
|
||||||
together aborts with an error naming the conflict.
|
|
||||||
- `expenditure_concept` is a **spending-only** concept: `cog_revenue()`
|
|
||||||
doesn't expose it (revenue's own intergovernmental codes are a different
|
|
||||||
axis -- see `?cog_revenue`).
|
|
||||||
|
|
||||||
# Summary
|
|
||||||
|
|
||||||
- Comparing one government to itself over time: any concept works -- pick
|
|
||||||
one and hold it fixed across every year compared.
|
|
||||||
- Comparing or summing across governments -- counties within a state, a
|
|
||||||
state against its neighbor, cities against counties: use `"primary"`
|
|
||||||
(the default) or `"direct"`. `cog_geographic_rollup()` and
|
|
||||||
`cog_peer_compare()` enforce this by refusing `"total"`.
|
|
||||||
- `"primary"` = operations + capital + assistance. `"direct"` = primary +
|
|
||||||
interest on debt + insurance trust benefits (Census's published Direct
|
|
||||||
Expenditure). `"total"` = direct + intergovernmental (`M` to local
|
|
||||||
governments, `L` to the state government excluding the `L--`
|
|
||||||
family-total row, and `Q11`/`Q12`/`Q18` state payments to school
|
|
||||||
systems).
|
|
||||||
Reference in New Issue
Block a user