diff --git a/plans/2026-04-29-per-year-population-denominator.md b/plans/2026-04-29-per-year-population-denominator.md deleted file mode 100644 index 22351a8..0000000 --- a/plans/2026-04-29-per-year-population-denominator.md +++ /dev/null @@ -1,1398 +0,0 @@ -# Per-year population denominator implementation plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Replace `cog_spending(per_capita = TRUE)` / `cog_revenue(per_capita = TRUE)` / `cog_geographic_rollup()` static-ACS denominator with the per-year Census F-33 population already present in `long.population`. Switch peer matching to a user-selectable cohort year. Add `pop_source` column, multi-note concatenation, updated provenance, vignette, and pipeline data dictionary entry. - -**Architecture:** A new `gov_population_yearly` DuckDB view exposes `(year, canonical_govid, population, popyear)` from `long`. `.attach_per_capita()` joins on `(canonical_govid, year)` instead of querying the static `canonical_fips_xwalk.population_acs`. `cog_find_peers()` queries the new view at a chosen year (default = most recent observed year for the target). `cog_geographic_rollup()` sums population across observed govs only. Type-4 (special districts) and type-5 (school districts) govs return `NA` per-capita with `pop_source = "unavailable"` because the F-33 schema masks population for those types. - -**Tech Stack:** R (DuckDB via DBI), testthat 3, roxygen2, pkgdown, tibble/dplyr, cli. - -**Spec:** `specs/2026-04-29-per-year-population-denominator-design.md` - -**Coverage assumption (fixture):** The bundled `inst/extdata/fixture_corpus` covers years 2019 and 2020 across all 50 states for types 0–3. It contains no type-4 or type-5 govs, so unavailability tests rely on querying govids absent from `gov_population_yearly` rather than on type filters. A list of fixture govids used in tests below: Alabama state `010000000` (pop 4,874,747 → 4,903,185), Broward County `101006006` (1,935,878 → 1,952,778), Wayne County `231082082`, Bexar County `441015015`, Tarrant County `441220220`. Static ACS values: Alabama `5,028,092`, Broward `1,940,907`. - ---- - -## Task 1: Add `gov_population_yearly` view - -**Files:** -- Create: `inst/sql/32-gov_population_yearly.sql` -- Modify: `tests/testthat/test-views.R` (append a test) - -- [ ] **Step 1: Read existing `test-views.R` to learn the pattern** - -```bash -cat tests/testthat/test-views.R -``` - -Note the pattern: tests open `with_fixture_corpus({ ... })`, then run a `DBI::dbGetQuery()` against `.ensure_session()` to verify a view exists and returns expected columns. - -- [ ] **Step 2: Write the failing test** - -Append to `tests/testthat/test-views.R`: - -```r -test_that("gov_population_yearly exposes one row per (year, canonical_govid)", { - skip_if_no_corpus() - with_fixture_corpus({ - con <- uscogdata:::.ensure_session() - df <- DBI::dbGetQuery( - con, - "SELECT year, canonical_govid, population, popyear - FROM gov_population_yearly - WHERE canonical_govid = '101006006' - ORDER BY year" - ) - expect_setequal(df$year, c(2019L, 2020L)) - expect_equal(nrow(df), 2L) - expect_true(all(!is.na(df$population))) - expect_equal(df$population[df$year == 2019L], 1935878L) - expect_equal(df$population[df$year == 2020L], 1952778L) - # Uniqueness on (year, canonical_govid) across the whole view. - dup <- DBI::dbGetQuery( - con, - "SELECT year, canonical_govid, COUNT(*) AS n - FROM gov_population_yearly - GROUP BY year, canonical_govid HAVING n > 1" - ) - expect_equal(nrow(dup), 0L) - }) -}) -``` - -- [ ] **Step 3: Run the test to verify it fails** - -```bash -Rscript -e 'devtools::test(filter = "views")' -``` - -Expected: FAIL — `gov_population_yearly` view does not exist (DuckDB binder error). - -- [ ] **Step 4: Create the SQL view file** - -Write `inst/sql/32-gov_population_yearly.sql`: - -```sql -CREATE OR REPLACE VIEW gov_population_yearly AS -SELECT DISTINCT - year, - canonical_govid, - population, - popyear -FROM long -WHERE population IS NOT NULL; -``` - -The numeric prefix `32` slots between the existing `30-canonical_fips_xwalk.sql` and `40-spending_annotated.sql` so it loads before any annotated views that might depend on it. - -- [ ] **Step 5: Re-run the test to verify it passes** - -```bash -Rscript -e 'devtools::test(filter = "views")' -``` - -Expected: PASS for the new test plus the existing view tests. - -- [ ] **Step 6: Commit** - -```bash -git add inst/sql/32-gov_population_yearly.sql tests/testthat/test-views.R -git commit -m "feat(sql): add gov_population_yearly view - -Exposes one row per (year, canonical_govid) drawn from long.population. -Used by per-capita denominators and peer matching." -``` - ---- - -## Task 2: Per-year denominator in `.attach_per_capita` (RED) - -Write the failing test first; implement in Task 3. - -**Files:** -- Modify: `tests/testthat/test-spending.R` (append) - -- [ ] **Step 1: Read existing `test-spending.R` patterns** - -```bash -head -80 tests/testthat/test-spending.R -``` - -- [ ] **Step 2: Write the failing test** - -Append to `tests/testthat/test-spending.R`: - -```r -test_that("per_capita denominator is the per-year F-33 population", { - skip_if_no_corpus() - with_fixture_corpus({ - r <- cog_spending("101006006", years = 2019:2020, - category = "Police", per_capita = TRUE) - r_ops <- r[r$spend_subtype == "operations", ] - # Implied denominator from amt_nominal / amt_per_capita_nominal - implied_pop <- r_ops$amt_nominal / r_ops$amt_per_capita_nominal - names(implied_pop) <- r_ops$year - expect_equal(implied_pop[["2019"]], 1935878, tolerance = 1) - expect_equal(implied_pop[["2020"]], 1952778, tolerance = 1) - # And the implied denominator does NOT equal the static ACS value - expect_false(all(abs(implied_pop - 1940907) < 1)) - }) -}) -``` - -- [ ] **Step 3: Run test to verify it fails** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: FAIL — implied denominator equals 1,940,907 (static ACS) for both years. - -- [ ] **Step 4: Commit the failing test** - -```bash -git add tests/testthat/test-spending.R -git commit -m "test(spending): per-year denominator expectation (failing)" -``` - ---- - -## Task 3: Rewrite `.attach_per_capita` to join on (canonical_govid, year) (GREEN) - -**Files:** -- Modify: `R/spending.R:143-159` - -- [ ] **Step 1: Read current `.attach_per_capita`** - -```bash -sed -n '143,160p' R/spending.R -``` - -- [ ] **Step 2: Replace `.attach_per_capita` with per-year join** - -Edit `R/spending.R`. Replace the existing function body (lines 143–159) with: - -```r -#' @noRd -.attach_per_capita <- function(result, con, govid) { - if (nrow(result) == 0L) { - result$amt_per_capita_nominal <- numeric(0) - result$pop_source <- character(0) - return(result) - } - years_lit <- paste(unique(as.integer(result$year)), collapse = ",") - sql <- sprintf( - "SELECT canonical_govid, year, population - FROM gov_population_yearly - WHERE canonical_govid IN (%s) - AND year IN (%s)", - .sql_lit_chr(govid), years_lit - ) - pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) - result <- dplyr::left_join(result, pops, - by = c("canonical_govid", "year")) - result$amt_per_capita_nominal <- result$amt_nominal / result$population - result$pop_source <- ifelse(is.na(result$population), - "unavailable", "census_f33") - result$population <- NULL - result -} -``` - -Key changes vs. the prior implementation: -- Joins on `(canonical_govid, year)` instead of `canonical_govid` alone. -- Queries `gov_population_yearly` instead of `canonical_fips_xwalk`. -- Adds `pop_source` column (`"census_f33"` or `"unavailable"`). -- `amt_per_capita_nominal` is naturally NA when `population` is NA (R's `NA / x = NA`). - -- [ ] **Step 3: Run the failing test from Task 2 to verify it passes** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: PASS for the new test. Existing per-capita tests in `test-spending.R` may now fail because they were written against the static-ACS denominator. Inspect each failure — most should be updated to assert the per-year implied denominator. Tests that asserted *equality* of per-capita across years for the same gov are no longer correct expectations. - -- [ ] **Step 4: Update any existing per-capita tests in `test-spending.R` that fail** - -Read each failing test. If it merely asserted `amt_per_capita_nominal` is positive/finite, no change needed. If it asserted a specific numeric value derived from `population_acs`, recompute the expected value using the per-year `population` for that gov-year. If it asserted `amt_per_capita` is the same across two years, change it to assert that the per-year implied denominator matches `gov_population_yearly`. - -- [ ] **Step 5: Run the full spending test file** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: ALL PASS. - -- [ ] **Step 6: Commit** - -```bash -git add R/spending.R tests/testthat/test-spending.R -git commit -m "feat(per-capita): use per-year F-33 population in spending verbs - -cog_spending(per_capita = TRUE) and cog_revenue(per_capita = TRUE) now -divide each year's amount by that gov-year's population from -gov_population_yearly (drawn from long.population) instead of a single -static ACS 2018-2022 value. Adds pop_source column with values -'census_f33' or 'unavailable'." -``` - ---- - -## Task 4: Multi-note concatenation + unavailable-pop note - -**Files:** -- Modify: `R/spending.R:177-185` (`.notes_column`) -- Modify: `R/spending.R:42-80` (`.verb_spendrev`) to call `.notes_column` after per-capita is attached -- Modify: `tests/testthat/test-spending.R` (append) - -- [ ] **Step 1: Read current `.notes_column`** - -```bash -sed -n '177,190p' R/spending.R -``` - -The current implementation builds a one-element note ("Aggregate fallback applied; see cog_explain()" or "") off `aggregate_fallback`. After this task it concatenates multiple sources of notes, joined by `"; "`. - -- [ ] **Step 2: Write a failing test** - -Append to `tests/testthat/test-spending.R`: - -```r -test_that("pop_source = 'unavailable' produces a note and NA per-capita", { - skip_if_no_corpus() - with_fixture_corpus({ - # No type-4/5 govs in fixture; use a govid present in long but synthetically - # absent from gov_population_yearly by querying a year out of fixture range. - # Better: query a govid that doesn't exist anywhere — cog_spending will - # return zero rows. Use a real govid in 2019 with per_capita to confirm - # the no-NA branch works, then test the NA branch with a manual round-trip: - r <- cog_spending("101006006", years = 2019L, - category = "Police", per_capita = TRUE) - expect_true(all(r$pop_source == "census_f33")) - expect_true(all(is.na(r$notes) | r$notes == "" | - !grepl("No population denominator", r$notes))) - }) -}) - -test_that("aggregate fallback + unavailable pop produce concatenated notes", { - # Unit-level test of .notes_column with a synthetic data frame so we don't - # depend on having a type-4/5 gov in the fixture. - result <- tibble::tibble( - aggregate_fallback = c(FALSE, TRUE, TRUE), - pop_source = c("census_f33", "census_f33", "unavailable") - ) - notes <- uscogdata:::.notes_column(result) - expect_equal(notes[1], "") - expect_equal(notes[2], "Aggregate fallback applied; see cog_explain()") - expect_equal(notes[3], - "Aggregate fallback applied; see cog_explain(); No population denominator available for this gov type") -}) -``` - -- [ ] **Step 3: Run test to verify the second one fails** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: FAIL — `.notes_column` does not yet handle `pop_source`. - -- [ ] **Step 4: Rewrite `.notes_column`** - -Replace `.notes_column` in `R/spending.R` with: - -```r -#' @noRd -.notes_column <- function(result) { - n <- nrow(result) - if (n == 0L) return(character(0)) - parts <- vector("list", 2L) - agg <- result[["aggregate_fallback"]] - parts[[1]] <- ifelse( - !is.null(agg) & isTRUE(any(agg, na.rm = TRUE)) & agg %in% TRUE, - "Aggregate fallback applied; see cog_explain()", - NA_character_ - ) - ps <- result[["pop_source"]] - parts[[2]] <- if (!is.null(ps)) { - ifelse(ps == "unavailable", - "No population denominator available for this gov type", - NA_character_) - } else { - rep(NA_character_, n) - } - out <- character(n) - for (i in seq_len(n)) { - pieces <- vapply(parts, `[[`, character(1), i) - pieces <- pieces[!is.na(pieces)] - out[i] <- if (length(pieces) == 0L) "" else paste(pieces, collapse = "; ") - } - out -} -``` - -The per-row loop is unavoidable in base R for this exact join semantics; the result tibble is small (one row per year × govid × subtype × category) so this is fine. - -- [ ] **Step 5: Verify `.notes_column` is called after `.attach_per_capita`** - -Read `.verb_spendrev` (`R/spending.R:42-80`). The existing flow is: - -``` -sql -> result -if (per_capita) result <- .attach_per_capita(result, con, govid) -if (!is.null(adjust_to_year)) result <- .attach_real_dollars(...) -result$notes <- .notes_column(result) -``` - -`.notes_column` is already invoked after per-capita attachment, so no change to `.verb_spendrev` is required. Verify with: - -```bash -sed -n '54,63p' R/spending.R -``` - -Expected output: shows `.notes_column(result)` on a line after the per-capita block. - -- [ ] **Step 6: Run tests** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: ALL PASS. - -- [ ] **Step 7: Commit** - -```bash -git add R/spending.R tests/testthat/test-spending.R -git commit -m "feat(notes): concatenate notes; flag unavailable population - -.notes_column now joins multiple per-row notes with '; '. Adds the -'No population denominator available for this gov type' note when -pop_source is 'unavailable'." -``` - ---- - -## Task 5: `cog_find_peers()` per-year (RED) - -**Files:** -- Modify: `tests/testthat/test-peers.R` (append) - -- [ ] **Step 1: Write failing tests** - -Append to `tests/testthat/test-peers.R`: - -```r -test_that("cog_find_peers defaults `year` to most recent observed year for target", { - skip_if_no_corpus() - peers <- cog_find_peers("101006006") - expect_equal(attr(peers, "cohort_year"), 2020L) - # Returned column is now `population`, not `population_acs` - expect_true("population" %in% names(peers)) - expect_false("population_acs" %in% names(peers)) -}) - -test_that("cog_find_peers honors an explicit `year`", { - skip_if_no_corpus() - peers <- cog_find_peers("101006006", year = 2019L) - expect_equal(attr(peers, "cohort_year"), 2019L) -}) - -test_that("cog_find_peers errors when target has no observed pop in `year`", { - skip_if_no_corpus() - expect_error( - cog_find_peers("101006006", year = 1999L), - "no observed population" - ) -}) -``` - -- [ ] **Step 2: Run tests to verify they fail** - -```bash -Rscript -e 'devtools::test(filter = "peers")' -``` - -Expected: FAIL — function doesn't accept `year` arg, returns `population_acs` column. - -- [ ] **Step 3: Commit failing tests** - -```bash -git add tests/testthat/test-peers.R -git commit -m "test(peers): per-year cohort expectations (failing)" -``` - ---- - -## Task 6: `cog_find_peers()` per-year (GREEN) - -**Files:** -- Modify: `R/peers.R:24-88` - -- [ ] **Step 1: Replace `cog_find_peers` body** - -Replace the entire `cog_find_peers` function in `R/peers.R` with: - -```r -#' Find peer governments by similarity criteria -#' -#' Selects peer governments by combinations of government type, state, and -#' population range at a chosen `year`. Peers are ordered by `|log(pop_ratio)|` -#' ascending (closest to the target's population first). -#' -#' @param target_govid Character scalar — `canonical_govid` of the target. -#' @param year Integer scalar. Cohort vintage. When `NULL` (default), uses the -#' most recent year for which the target has an observed population in -#' `gov_population_yearly`. -#' @param same_type If `TRUE` (default) restrict peers to the target's -#' `govs_type`. -#' @param same_state If `TRUE` restrict peers to the target's `fips_state`. -#' Default `FALSE`. -#' @param pop_range Length-2 numeric vector giving lower/upper bounds. -#' @param is_ratio If `TRUE` (default) `pop_range` is multiplied by the -#' target's population at `year` to produce absolute bounds. If `FALSE`, -#' `pop_range` is interpreted as absolute population counts. -#' @param max_peers Integer cap on the number of peers returned. -#' @return Tibble with columns `canonical_govid`, `gov_name`, `fips_state`, -#' `population`, `pop_ratio`, `rank`. The cohort year is attached as -#' `attr(x, "cohort_year")`. -#' @export -cog_find_peers <- function(target_govid, - year = NULL, - same_type = TRUE, - same_state = FALSE, - pop_range = c(0.7, 1.3), - is_ratio = TRUE, - max_peers = 10L) { - if (!is.character(target_govid) || length(target_govid) != 1L) { - cli::cli_abort("`target_govid` must be a length-1 character string.") - } - if (!is.numeric(pop_range) || length(pop_range) != 2L || - pop_range[1] >= pop_range[2]) { - cli::cli_abort("`pop_range` must be a length-2 numeric with lo < hi.") - } - if (!is.null(year) && - (!(is.numeric(year) || is.integer(year)) || length(year) != 1L)) { - cli::cli_abort("`year` must be NULL or a length-1 integer.") - } - - con <- .ensure_session() - - # Confirm target exists in the xwalk and pull govs_type / fips_state. - meta_sql <- sprintf( - "SELECT canonical_govid, gov_name, govs_type, fips_state - FROM canonical_fips_xwalk - WHERE canonical_govid = %s", - .sql_lit_chr(target_govid) - ) - meta <- DBI::dbGetQuery(con, meta_sql) - if (nrow(meta) == 0L) { - cli::cli_abort(c( - "govid {target_govid} not found in corpus.", - i = "v0.1 covers types 0-3 only (state/county/city/township); see vignette('coverage-scope')." - )) - } - - cohort_year <- .resolve_cohort_year(con, target_govid, year) - - pop_sql <- sprintf( - "SELECT population FROM gov_population_yearly - WHERE canonical_govid = %s AND year = %d", - .sql_lit_chr(target_govid), as.integer(cohort_year) - ) - target_pop <- DBI::dbGetQuery(con, pop_sql)$population - if (length(target_pop) == 0L || is.na(target_pop) || target_pop <= 0) { - cli::cli_abort(c( - "Target {target_govid} has no observed population in {cohort_year}.", - i = "Use a year for which population is observed; see gov_population_yearly." - )) - } - - if (isTRUE(is_ratio)) { - lo <- target_pop * pop_range[1] - hi <- target_pop * pop_range[2] - } else { - lo <- pop_range[1]; hi <- pop_range[2] - } - - preds <- c( - sprintf("p.canonical_govid != %s", .sql_lit_chr(target_govid)), - sprintf("p.year = %d", as.integer(cohort_year)), - sprintf("p.population BETWEEN %.6f AND %.6f", lo, hi) - ) - if (isTRUE(same_type)) preds <- c(preds, sprintf("x.govs_type = %d", meta$govs_type)) - if (isTRUE(same_state)) preds <- c(preds, sprintf("x.fips_state = %s", .sql_lit_chr(meta$fips_state))) - - peers_sql <- sprintf( - "SELECT p.canonical_govid, x.gov_name, x.fips_state, p.population, - p.population / %.6f AS pop_ratio - FROM gov_population_yearly p - JOIN canonical_fips_xwalk x USING (canonical_govid) - WHERE %s - ORDER BY ABS(LN(CAST(p.population AS DOUBLE) / %.6f)) - LIMIT %d", - target_pop, - paste(preds, collapse = " AND "), - target_pop, - as.integer(max_peers) - ) - peers <- tibble::as_tibble(DBI::dbGetQuery(con, peers_sql)) - peers$rank <- if (nrow(peers) > 0L) seq_len(nrow(peers)) else integer(0) - attr(peers, "cohort_year") <- as.integer(cohort_year) - peers -} - -#' @noRd -.resolve_cohort_year <- function(con, target_govid, year) { - if (!is.null(year)) return(as.integer(year)) - sql <- sprintf( - "SELECT MAX(year) AS y FROM gov_population_yearly - WHERE canonical_govid = %s", - .sql_lit_chr(target_govid) - ) - y <- DBI::dbGetQuery(con, sql)$y - if (length(y) == 0L || is.na(y)) { - cli::cli_abort( - "Target {target_govid} has no observed population in any year." - ) - } - as.integer(y) -} -``` - -- [ ] **Step 2: Run failing tests from Task 5** - -```bash -Rscript -e 'devtools::test(filter = "peers")' -``` - -Expected: PASS for the three new tests. Existing tests in `test-peers.R` may fail because they reference `population_acs` — fix in next step. - -- [ ] **Step 3: Update existing `test-peers.R` assertions** - -The existing tests that reference `population_acs` need column name and value updates. Specifically lines that check `expected_cols` and `peers$population_acs`: - -```r -# In test "cog_find_peers returns same-type peers in the default pop band": -expected_cols <- c("canonical_govid", "gov_name", "fips_state", - "population", "pop_ratio", "rank") - -# In test "cog_find_peers absolute pop range works": -expect_true(all(peers$population >= 1.5e6 & - peers$population <= 2.5e6)) -``` - -Run `grep -n population_acs tests/testthat/test-peers.R` to find every instance and rename to `population`. - -- [ ] **Step 4: Run all peer tests** - -```bash -Rscript -e 'devtools::test(filter = "peers")' -``` - -Expected: ALL PASS. - -- [ ] **Step 5: Commit** - -```bash -git add R/peers.R tests/testthat/test-peers.R -git commit -m "feat(peers): cog_find_peers uses per-year population - -Adds optional 'year' argument (defaults to most recent observed year for -the target). Filters and ranks candidates by gov_population_yearly.population -at that year. Returned column renamed population_acs -> population. -Cohort year attached as attr(x, 'cohort_year')." -``` - ---- - -## Task 7: `cog_peer_compare()` cohort_year column - -**Files:** -- Modify: `R/peers.R:112-146` -- Modify: `tests/testthat/test-peers.R` (append) - -- [ ] **Step 1: Write failing test** - -Append to `tests/testthat/test-peers.R`: - -```r -test_that("cog_peer_compare stamps cohort_year from peers attribute", { - skip_if_no_corpus() - peers <- cog_find_peers("101006006", year = 2019L, max_peers = 4L) - r <- cog_peer_compare("101006006", peers, "Police", years = 2020L) - expect_true("cohort_year" %in% names(r)) - expect_true(all(r$cohort_year == 2019L)) - prov <- attr(r, "provenance") - expect_equal(prov$cohort_year, 2019L) -}) - -test_that("cog_peer_compare cohort_year is NA for bare character peers", { - skip_if_no_corpus() - r <- cog_peer_compare( - "101006006", - peers = c("441015015", "441220220"), - category = "Police", years = 2020L - ) - expect_true(all(is.na(r$cohort_year))) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -```bash -Rscript -e 'devtools::test(filter = "peers")' -``` - -Expected: FAIL — `cohort_year` column does not exist. - -- [ ] **Step 3: Update `cog_peer_compare`** - -In `R/peers.R`, edit `cog_peer_compare`. Just before the existing `attr(out, "provenance") <- prov` line, add the cohort_year derivation and column stamp; and inside the provenance block add the `cohort_year` field. - -Replace the section from `peer_govids <- if (...)` through the final return with: - -```r - cohort_year <- if (is.data.frame(peers)) { - ay <- attr(peers, "cohort_year") - if (is.null(ay)) NA_integer_ else as.integer(ay) - } else { - NA_integer_ - } - peer_govids <- if (is.data.frame(peers)) { - as.character(peers$canonical_govid) - } else { - as.character(peers) - } - peer_govids <- peer_govids[!is.na(peer_govids) & nzchar(peer_govids)] - all_govids <- unique(c(target_govid, peer_govids)) - - r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) - r$role <- ifelse(r$canonical_govid == target_govid, "target", "peer") - - value_col <- .peer_value_col(per_capita, adjust_to_year) - - summary_rows <- .peer_summary_rows(r, value_col) - out <- dplyr::bind_rows(r, summary_rows) - rank_val <- .peer_target_rank(r, target_govid, years, value_col) - out$target_rank <- ifelse(out$role == "target", rank_val, NA_integer_) - out$cohort_year <- cohort_year - - prov <- attr(r, "provenance") %||% list() - prov$verb <- "cog_peer_compare" - prov$call <- paste(deparse(call), collapse = " ") - prov$peer_count <- length(peer_govids) - prov$cohort_year <- cohort_year - prov$cohort_govids <- peer_govids - prov$target <- list( - canonical_govid = target_govid, - gov_name = unique(r$gov_name[r$role == "target"]) - ) - attr(out, "provenance") <- prov - out -} -``` - -- [ ] **Step 4: Run tests** - -```bash -Rscript -e 'devtools::test(filter = "peers")' -``` - -Expected: ALL PASS. - -- [ ] **Step 5: Commit** - -```bash -git add R/peers.R tests/testthat/test-peers.R -git commit -m "feat(peers): stamp cohort_year on cog_peer_compare results - -Reads attr(peers, 'cohort_year') when the caller passed a cog_find_peers() -tibble; NA when the caller passed a bare character vector. Stamped as a -constant column on the result and recorded in provenance alongside the -cohort govids." -``` - ---- - -## Task 8: Population-aware `cog_geographic_rollup()` (RED) - -**Files:** -- Modify: `tests/testthat/test-rollup.R` (append) - -- [ ] **Step 1: Read existing rollup test patterns** - -```bash -head -60 tests/testthat/test-rollup.R -``` - -- [ ] **Step 2: Write failing test** - -Append to `tests/testthat/test-rollup.R`: - -```r -test_that("cog_geographic_rollup per-capita uses summed per-year populations", { - skip_if_no_corpus() - with_fixture_corpus({ - r <- cog_geographic_rollup( - govids = list(state = "010000000", - county = "101006006"), - category = "Police", - years = 2019:2020, - per_capita = TRUE - ) - state_ops <- r[r$layer == "state" & r$spend_subtype == "operations", ] - county_ops <- r[r$layer == "county" & r$spend_subtype == "operations", ] - state_implied <- state_ops$amt_nominal / state_ops$amt_per_capita_nominal - county_implied <- county_ops$amt_nominal / - county_ops$amt_per_capita_nominal - # Per-year, per-layer denominator is the layer's own per-year population - expect_equal(state_implied[state_ops$year == 2019], 4874747, tolerance = 1) - expect_equal(state_implied[state_ops$year == 2020], 4903185, tolerance = 1) - expect_equal(county_implied[county_ops$year == 2019], 1935878, tolerance = 1) - }) -}) - -test_that("cog_geographic_rollup records included/excluded govids in provenance", { - skip_if_no_corpus() - with_fixture_corpus({ - r <- cog_geographic_rollup( - govids = list(county = "101006006"), - category = "Police", - years = 2019:2020, - per_capita = TRUE - ) - prov <- attr(r, "provenance") - expect_true("rollup" %in% names(prov)) - expect_true("101006006" %in% prov$rollup$included_govids) - expect_true(is.character(prov$rollup$excluded_govids)) - }) -}) -``` - -- [ ] **Step 3: Run to verify failure** - -```bash -Rscript -e 'devtools::test(filter = "rollup")' -``` - -Expected: the second test fails (no `rollup` block in provenance). The first may pass already because rollup currently delegates to `cog_spending` and the per-capita is per-row — verify. - -- [ ] **Step 4: Commit failing tests** - -```bash -git add tests/testthat/test-rollup.R -git commit -m "test(rollup): per-year denominator + provenance expectations (failing)" -``` - ---- - -## Task 9: Population-aware `cog_geographic_rollup()` (GREEN) - -**Files:** -- Modify: `R/rollup.R:26-57` - -The rollup currently passes through to `cog_spending()` and returns one row per `(year, canonical_govid, subtype, category)` tagged with its layer — a side-by-side comparison, not a summed total. The spec preserves that semantics. - -The GREEN step: -- After `cog_spending(...)` returns, drop rows with `pop_source == "unavailable"` *only when `per_capita = TRUE`*. -- Record included/excluded govids in provenance. -- Document the rule in roxygen. - -- [ ] **Step 1: Replace `cog_geographic_rollup`** - -Edit `R/rollup.R`. Replace the existing function with: - -```r -#' Aggregate spending across state/county/city layers for a place -#' -#' Wraps [cog_spending()], tags each row with its layer, and attaches a -#' human-readable `scope_note` documenting geographic-scope caveats (e.g. -#' "county totals include areas outside the listed city"). Useful for -#' "place portraits" that compare a city to the surrounding county and -#' containing state on one set of axes. -#' -#' When `per_capita = TRUE`, rows whose government has no observed -#' population in that year (`pop_source == "unavailable"`) are dropped from -#' the result. The dropped govids are recorded in -#' `provenance$rollup$excluded_govids`. This excludes special districts -#' (gov type 4) and school districts (gov type 5) from per-capita rollups -#' by design — see `vignette('population-denominators')`. -#' -#' @param govids Named list with any non-empty subset of elements named -#' `state`, `county`, `city`. Each element is a character vector of -#' `canonical_govid` values. At least one layer required. -#' @param category Single category name or character vector (passed through -#' to [cog_spending()]). -#' @param years Integer vector of years. -#' @param per_capita If `TRUE`, per-capita uses each gov's own per-year -#' population from `gov_population_yearly`. Govs with missing population -#' are excluded from the result. -#' @param adjust_to_year Integer base year for CPI-U conversion, or `NULL`. -#' @return Tibble with columns `year`, `layer`, `canonical_govid`, `gov_name`, -#' `spend_subtype`, `category`, `amt_nominal`, optional `amt_real` / -#' `amt_per_capita_nominal` / `amt_per_capita_real`, optional `pop_source`, -#' `codes_included`, `aggregate_fallback`, `scope_note`, `notes`. Carries a -#' `provenance` attribute with `verb = "cog_geographic_rollup"`, `layers`, -#' and `rollup$included_govids` / `rollup$excluded_govids`. -#' @export -cog_geographic_rollup <- function(govids, category, years, - per_capita = FALSE, adjust_to_year = NULL) { - call <- match.call() - .validate_rollup_layers(govids) - - govids <- lapply(govids, .coerce_govid_input, arg = "govids[[layer]]") - if (any(lengths(govids) == 0L)) { - cli::cli_abort("Each layer in `govids` must be non-empty after coercion.") - } - layer_names <- names(govids) - all_govids <- unlist(govids, use.names = FALSE) - layer_map <- tibble::tibble( - canonical_govid = all_govids, - layer = rep(layer_names, lengths(govids)) - ) - - r <- cog_spending(all_govids, years, category, per_capita, adjust_to_year) - r <- dplyr::left_join(r, layer_map, by = "canonical_govid", - relationship = "many-to-many") - r$scope_note <- .rollup_scope_note(r$layer) - - excluded <- character(0) - if (isTRUE(per_capita) && "pop_source" %in% names(r)) { - drop <- r$pop_source == "unavailable" - excluded <- unique(r$canonical_govid[drop]) - r <- r[!drop, , drop = FALSE] - } - included <- unique(r$canonical_govid) - - r <- .reorder_rollup_cols(r) - - prov <- attr(r, "provenance") - prov$verb <- "cog_geographic_rollup" - prov$call <- paste(deparse(call), collapse = " ") - prov$layers <- layer_names - prov$rollup <- list( - included_govids = included, - excluded_govids = excluded - ) - attr(r, "provenance") <- prov - - r -} -``` - -`.validate_rollup_layers`, `.rollup_scope_note`, `.reorder_rollup_cols` are unchanged. - -- [ ] **Step 2: Run tests** - -```bash -Rscript -e 'devtools::test(filter = "rollup")' -``` - -Expected: ALL PASS. - -- [ ] **Step 3: Commit** - -```bash -git add R/rollup.R -git commit -m "feat(rollup): drop unavailable-pop rows + provenance audit - -cog_geographic_rollup(per_capita = TRUE) now drops rows whose government -has no observed population for that year (pop_source == 'unavailable'), -matching the spec's exclusion rule. Records included/excluded govids in -provenance\$rollup." -``` - ---- - -## Task 10: Update provenance for new denominator metadata - -**Files:** -- Modify: `R/provenance.R:54-74` -- Modify: `R/spending.R` `.verb_spendrev` to pass result with `pop_source` to `.build_provenance` -- Modify: `tests/testthat/test-spending.R` (append) - -- [ ] **Step 1: Write failing test** - -Append to `tests/testthat/test-spending.R`: - -```r -test_that("provenance records per-year denominator metadata", { - skip_if_no_corpus() - with_fixture_corpus({ - r <- cog_spending("101006006", years = 2019:2020, - category = "Police", per_capita = TRUE) - pc <- attr(r, "provenance")$transformations$per_capita - expect_true(pc$applied) - expect_match(pc$denominator_source, "Census F-33", fixed = FALSE) - expect_match(pc$denominator_source, "per-year", fixed = TRUE) - expect_equal(pc$pop_source_counts$census_f33, nrow(r)) - expect_equal(pc$pop_source_counts$unavailable, 0L) - expect_equal(length(pc$popyear_range), 2L) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -```bash -Rscript -e 'devtools::test(filter = "spending")' -``` - -Expected: FAIL — `pop_source_counts` is NULL, `denominator_source` still reads "ACS 2018-2022". - -- [ ] **Step 3: Plumb popyear through to provenance** - -The popyear range needs to come from `gov_population_yearly`. Two options: re-query inside `.build_provenance`, or have `.attach_per_capita` stash a `popyear_range` attribute on the result. Stash on the result is simpler and avoids a duplicate query. - -Edit `.attach_per_capita` in `R/spending.R` so the SQL also pulls `popyear`, and after computing per-capita, drop the column but stash min/max as attributes: - -Replace the SQL block + assignment: - -```r - sql <- sprintf( - "SELECT canonical_govid, year, population, popyear - FROM gov_population_yearly - WHERE canonical_govid IN (%s) - AND year IN (%s)", - .sql_lit_chr(govid), years_lit - ) - pops <- tibble::as_tibble(DBI::dbGetQuery(con, sql)) - result <- dplyr::left_join(result, pops, - by = c("canonical_govid", "year")) - result$amt_per_capita_nominal <- result$amt_nominal / result$population - result$pop_source <- ifelse(is.na(result$population), - "unavailable", "census_f33") - py <- result$popyear[!is.na(result$popyear)] - attr(result, ".popyear_range") <- if (length(py) > 0L) { - as.integer(c(min(py), max(py))) - } else { - integer(0) - } - result$population <- NULL - result$popyear <- NULL - result -} -``` - -- [ ] **Step 4: Update `.build_provenance`** - -Replace the `per_capita` block in `R/provenance.R`: - -```r - per_capita = list( - applied = isTRUE(per_capita), - denominator_source = if (isTRUE(per_capita)) { - "Census F-33 population (per-year, from long.population)" - } else { - NA_character_ - }, - popyear_range = if (isTRUE(per_capita)) { - attr(result, ".popyear_range") %||% integer(0) - } else { - integer(0) - }, - pop_source_counts = if (isTRUE(per_capita)) { - ps <- result[["pop_source"]] - if (is.null(ps) || length(ps) == 0L) { - list(census_f33 = 0L, unavailable = 0L) - } else { - list( - census_f33 = sum(ps == "census_f33", na.rm = TRUE), - unavailable = sum(ps == "unavailable", na.rm = TRUE) - ) - } - } else { - NULL - } - ), -``` - -- [ ] **Step 5: Strip `.popyear_range` attr after provenance is built** - -In `.verb_spendrev` (`R/spending.R:42-80`), after the line `attr(result, "provenance") <- prov`, add: - -```r - attr(result, ".popyear_range") <- NULL -``` - -so the helper attribute does not leak into the public surface. - -- [ ] **Step 6: Run tests** - -```bash -Rscript -e 'devtools::test()' -``` - -Expected: ALL PASS across spending, rollup, peers, explain. - -- [ ] **Step 7: Commit** - -```bash -git add R/spending.R R/provenance.R tests/testthat/test-spending.R -git commit -m "feat(provenance): record per-year denominator metadata - -Updates transformations\$per_capita with the new denominator_source string, -popyear_range, and pop_source_counts. .attach_per_capita stashes -popyear_range on the result; .verb_spendrev strips the helper attr after -provenance is built." -``` - ---- - -## Task 11: Render new provenance fields in `cog_explain()` - -**Files:** -- Modify: `R/explain.R:69-82` (Transformations block) -- Modify: `tests/testthat/test-explain.R` (append) - -- [ ] **Step 1: Write failing test** - -Append to `tests/testthat/test-explain.R`: - -```r -test_that("cog_explain prints denominator + popyear_range + counts", { - skip_if_no_corpus() - with_fixture_corpus({ - r <- cog_spending("101006006", years = 2019:2020, - category = "Police", per_capita = TRUE) - out <- capture.output(cog_explain(r)) - expect_true(any(grepl("Census F-33", out))) - expect_true(any(grepl("popyear", out, ignore.case = TRUE))) - expect_true(any(grepl("census_f33", out))) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -```bash -Rscript -e 'devtools::test(filter = "explain")' -``` - -Expected: FAIL. - -- [ ] **Step 3: Update `.print_provenance`** - -In `R/explain.R`, replace the `pc <- prov$transformations$per_capita` block with: - -```r - pc <- prov$transformations$per_capita - if (isTRUE(pc$applied)) { - cli::cli_text("Per-capita denominator: {pc$denominator_source}") - if (length(pc$popyear_range) == 2L) { - cli::cli_text( - " popyear range: {pc$popyear_range[1]}-{pc$popyear_range[2]}" - ) - } - if (!is.null(pc$pop_source_counts)) { - cli::cli_text( - " pop_source counts: census_f33={pc$pop_source_counts$census_f33}, unavailable={pc$pop_source_counts$unavailable}" - ) - } - } -``` - -- [ ] **Step 4: Run tests** - -```bash -Rscript -e 'devtools::test(filter = "explain")' -``` - -Expected: PASS. - -- [ ] **Step 5: Commit** - -```bash -git add R/explain.R tests/testthat/test-explain.R -git commit -m "feat(explain): render new per-capita provenance fields - -cog_explain() now prints denominator_source, popyear_range, and -pop_source_counts under the Transformations section." -``` - ---- - -## Task 12: Population-denominators vignette - -**Files:** -- Create: `vignettes/population-denominators.Rmd` -- Modify: `DESCRIPTION` (add `Suggests: knitr, rmarkdown` if missing) - -- [ ] **Step 1: Verify vignette infrastructure** - -```bash -grep -E "VignetteBuilder|knitr|rmarkdown" DESCRIPTION -ls vignettes/ 2>/dev/null -``` - -If no other vignettes exist or `knitr`/`rmarkdown` is missing from `Suggests`, add to DESCRIPTION: - -``` -Suggests: - knitr, - rmarkdown, - testthat (>= 3.0.0), - withr -VignetteBuilder: knitr -``` - -(Only add the lines that aren't already present.) - -- [ ] **Step 2: Create the vignette** - -Write `vignettes/population-denominators.Rmd`: - -````markdown ---- -title: "Population denominators" -output: rmarkdown::html_vignette -vignette: > - %\VignetteIndexEntry{Population denominators} - %\VignetteEngine{knitr::knitr} - %\VignetteEncoding{UTF-8} ---- - -```{r setup, include = FALSE} -knitr::opts_chunk$set(eval = FALSE, collapse = TRUE, comment = "#>") -``` - -# Why per-year population matters - -Per-capita finance numbers divide each year's spending or revenue by a population denominator. The choice of denominator is a research decision, not an implementation detail: a 24-year corpus paired with a single 5-year ACS estimate produces biased per-capita values whose magnitude scales with each government's population change. - -`uscogdata` defaults to the **Census F-33 population value Census itself uses to compute its published per-capita tables.** That value is recorded on every COG row as `population`, with `popyear` indicating the vintage. For a city that grew from 200,000 to 300,000 between 2000 and 2023, this default reproduces the per-capita value Census published. A static ACS denominator would have understated 2000 per-capita by ~33%. - -# The four population sources - -| Source | What it is | Default in uscogdata? | -|---|---|---| -| Census F-33 `population` | Population value Census used on each COG row to compute its published per-capita tables. Almost always a Population Estimates Program (PEP) estimate; sometimes lagged a year for fiscal-year alignment, recorded in `popyear`. | **Yes — default for `cog_spending(per_capita = TRUE)` etc.** | -| PEP (raw) | Census Bureau's official annual intercensal estimates, distinct from F-33 because F-33 sometimes uses a lagged vintage. | No (not in corpus) | -| ACS 5-year | American Community Survey 5-year rolling average. Different methodology, has margin of error, only available 2005-2009 onward. | Used by `cog_find_peers()` historically; replaced in 0.1 by per-year F-33. Still available in `canonical_fips_xwalk.population_acs` for non-time-series uses. | -| Decennial count | Actual count, every 10 years. | No (not in corpus) | - -The F-33 denominator is preferred because it's the same value Census uses internally — so `uscogdata` per-capita numbers reconcile with Census's own published tables. - -# Coverage - -F-33 `population` is observed for gov types 0–3 (state, county, city, township). Gov types 4 (special districts) and 5 (school districts) have `population` masked to NA in the F-33 schema. uscogdata returns: - -- `pop_source = "census_f33"` and a numeric `amt_per_capita_*` for types 0–3. -- `pop_source = "unavailable"` and `NA` per-capita for types 4–5, with a corresponding entry in `notes`. - -`cog_geographic_rollup(per_capita = TRUE)` excludes unavailable-pop rows from the result; the dropped govids are listed in `provenance\$rollup\$excluded_govids`. - -# The popyear quirk - -Census sometimes uses a population estimate from one year prior to the fiscal year being reported (e.g., FY2018 paired with a 2017 PEP estimate) so the denominator is available before the fiscal year closes. `popyear` records which vintage was paired; `cog_spending()` returns the popyear range in `provenance\$transformations\$per_capita\$popyear_range` rather than as a per-row column. - -# Time-varying peer cohorts - -`cog_find_peers(target, year = Y)` builds a cohort matched on each candidate's population at year `Y`. The cohort is fixed once chosen; `cog_peer_compare()` then runs that cohort across whatever `years` you ask for. To run a moving-window comparison, build cohorts year-by-year and stitch the results: - -```r -years <- 2010:2023 -out <- purrr::map_dfr(years, function(y) { - peers <- cog_find_peers("231082082", year = y, max_peers = 10L) - cog_peer_compare("231082082", peers, - category = "Police", years = y, - per_capita = TRUE) -}) -``` - -Each row in `out` has `cohort_year == year`, so a faceted plot shows cohort drift directly. - -# Future direction - -`pop_source` is a column on the result, not a fixed value, so adding a new denominator (PEP from tidycensus, decennial counts, ACS time-series) is a join change rather than an API change. A future release may add `cog_spending(..., pop_source = "pep")` for users who need a single externally-audited series. -```` - -- [ ] **Step 3: Build vignettes locally to confirm** - -```bash -Rscript -e 'devtools::build_vignettes()' -``` - -Expected: builds without error; HTML appears in `doc/`. - -- [ ] **Step 4: Commit** - -```bash -git add vignettes/population-denominators.Rmd DESCRIPTION -git commit -m "docs(vignette): population denominators rationale + usage - -Explains the four population sources, why F-33 is the default, type-4/5 -coverage gap, the popyear quirk, and how to build moving-window peer -cohorts manually." -``` - ---- - -## Task 13: Pipeline data dictionary entry - -**Files:** -- Modify: `../cog_pipeline/docs/data_dictionary.md` (sibling repo) - -This task touches a different git repo. The cog_pipeline repo lives at `../cog_pipeline/` relative to the uscogdata working directory. - -- [ ] **Step 1: Locate the data dictionary** - -```bash -ls ../cog_pipeline/docs/data_dictionary.md -``` - -If the file does not exist, create it; otherwise append to the appropriate column-reference section. - -- [ ] **Step 2: Add the population block** - -Append (or insert under any existing column reference) the following section: - -```markdown -## `population` and `popyear` - -`long.population` and `long.popyear` are population metadata columns from the F-33 fixed-width files. Census uses them to compute the per-capita tables in its own COG publications. - -- **Source bytes:** `population` from cols 124-132 (older years) or cols 117-125 (modern format); `popyear` from cols 133-134 / 126-127. See `R/read_modern.R` for the exact mappings per fiscal-year layout. -- **Vintage:** `popyear` is a 2-digit year identifying which Population Estimates Program (PEP) value Census paired with that fiscal year. PEP estimates are sometimes lagged a year for fiscal-year alignment (e.g., FY2018 paired with 2017 PEP). -- **Coverage:** Populated for gov types 0–3 (state, county, city, township). Masked to NA for gov types 4 (special districts) and 5 (school districts) in `R/read_modern.R::.apply_phase_e_masking()`. Schools instead carry `enrollment` / `enrollyear`. -- **Relationship to PEP:** `population` is approximately the PEP estimate for `popyear` for that geography. It is *not* identical to a tidycensus `get_estimates()` pull because Census occasionally revises PEP retroactively while the F-33 value is frozen at publication. -- **Downstream use:** `uscogdata::cog_spending(per_capita = TRUE)` exposes this as the `census_f33` denominator via the `gov_population_yearly` view. -``` - -- [ ] **Step 3: Commit in cog_pipeline** - -```bash -cd ../cog_pipeline -git add docs/data_dictionary.md -git commit -m "docs(data-dict): document long.population / long.popyear - -Adds source byte ranges, vintage semantics, type-4/5 masking rule, and -downstream use by uscogdata's per-capita denominator." -cd ../uscogdata -``` - ---- - -## Task 14: NEWS.md entry - -**Files:** -- Modify: `NEWS.md` (top of file, under or above the existing latest entry) - -- [ ] **Step 1: Read current NEWS.md** - -```bash -head -30 NEWS.md -``` - -- [ ] **Step 2: Add Unreleased entry** - -Insert at the top of `NEWS.md` (above the most recent dated entry): - -```markdown -# uscogdata (development version) - -## Per-capita denominators now use per-year Census F-33 population - -`cog_spending()` and `cog_revenue()` previously divided all years' amounts by a single ACS 2018-2022 estimate (`canonical_fips_xwalk.population_acs`), producing biased per-capita values for time-series analysis. They now divide by the F-33 `population` recorded on each gov-year via the new `gov_population_yearly` view. Result tibbles gain a `pop_source` column with values `"census_f33"` or `"unavailable"`. `notes` is updated to concatenate multiple notes with `"; "`. - -## Peer cohorts can be set to a chosen year - -`cog_find_peers()` adds a `year` argument (default: most recent year for which the target has an observed population in `gov_population_yearly`). The returned column previously named `population_acs` is now `population` and reflects the cohort year's vintage. The cohort year is attached to the returned tibble as `attr(x, "cohort_year")`. - -`cog_peer_compare()` now stamps a `cohort_year` column on its result (read from the peers tibble's attribute) and records `cohort_year` plus `cohort_govids` in provenance. When the caller supplies a bare character vector instead of a `cog_find_peers()` result, `cohort_year` is `NA`. - -## Rollups exclude govs missing population - -`cog_geographic_rollup(per_capita = TRUE)` drops rows whose government has `pop_source == "unavailable"` and records the dropped govids in `provenance$rollup$excluded_govids`. This excludes special districts (type 4) and school districts (type 5) from per-capita rollups by design. - -## New: vignette and provenance metadata - -- New vignette `population-denominators` covers the four population sources, the type-4/5 coverage gap, the popyear quirk, and how to build moving-window peer cohorts manually. -- Provenance gains `transformations$per_capita$popyear_range` and `pop_source_counts`. `cog_explain()` renders both. -``` - -- [ ] **Step 3: Run the full test suite one more time** - -```bash -Rscript -e 'devtools::test()' -``` - -Expected: ALL tests pass (181+ existing + ~12 new = ~193+). - -- [ ] **Step 4: Commit** - -```bash -git add NEWS.md -git commit -m "docs(news): per-year population denominators (unreleased) - -Summarizes the per-capita and peer-cohort behavior changes for users -upgrading from earlier 0.1 snapshots." -``` - ---- - -## Task 15: Document refresh + final R CMD check - -**Files:** -- Re-generate: `man/*.Rd` from updated roxygen -- Re-generate: `NAMESPACE` (no exports change, but document() will refresh) - -- [ ] **Step 1: Run `devtools::document()`** - -```bash -Rscript -e 'devtools::document()' -``` - -Expected: `man/cog_find_peers.Rd`, `man/cog_geographic_rollup.Rd`, -`man/cog_spending.Rd`, `man/cog_revenue.Rd`, `man/cog_explain.Rd` are -regenerated. - -- [ ] **Step 2: Run R CMD check** - -```bash -Rscript -e 'devtools::check(args = c("--no-manual", "--no-build-vignettes"))' -``` - -Expected: 0 ERRORs, 0 WARNINGs, 0 NOTEs (or only pre-existing acceptable notes). - -- [ ] **Step 3: Commit doc regeneration** - -```bash -git add man/ NAMESPACE -git commit -m "docs(roxygen): regenerate man pages for per-year population work" -``` - -- [ ] **Step 4: Push** - -```bash -git push origin main -``` - -(Only when human review of the full series is complete and the engineer is authorized to push.) - ---- - -## Verification checklist - -After all tasks land, verify: - -- [ ] `devtools::test()` reports 0 failures -- [ ] `devtools::check()` reports 0 ERRORs / 0 WARNINGs -- [ ] `cog_spending("101006006", 2019:2020, "Police", per_capita = TRUE)` returns different `amt_per_capita_nominal` for 2019 vs 2020 with the implied denominators matching `gov_population_yearly` -- [ ] `cog_find_peers("101006006")` returns a tibble with `population` (not `population_acs`) and `attr(x, "cohort_year") == 2020` -- [ ] `cog_geographic_rollup(per_capita = TRUE)` for a state+county+city set returns rows for each gov tagged by layer, with per-row per-year per-capita -- [ ] `cog_explain()` output mentions "Census F-33" and "popyear range" -- [ ] `vignette("population-denominators", package = "uscogdata")` opens -- [ ] cog_pipeline's `docs/data_dictionary.md` includes the population entry - -## Out of scope (do not do) - -- Adding PEP / ACS time-series / decennial denominators (architected for, not implemented) -- Adding a `per_pupil` denominator using `long.enrollment` -- Backfilling type-4/5 population from external sources -- Updating `cog_explorer/` callers — separate follow-up -- Bumping the package version diff --git a/plans/2026-08-04-partial-coverage-signposting.md b/plans/2026-08-04-partial-coverage-signposting.md deleted file mode 100644 index 3418367..0000000 --- a/plans/2026-08-04-partial-coverage-signposting.md +++ /dev/null @@ -1,1061 +0,0 @@ -# Partial-Coverage Signposting Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Fire a harmonization-recipe suggestion when a requested category still returns rows but structurally excludes component dollars the corpus does hold — closing uscogdata#9, where `cog_spending(category = "Public Welfare")` silently drops aggregate-published `E67`/`E68` in legacy years. - -**Architecture:** `.build_suggestions()` keeps its recipe-candidate query, M/L exclusion, and IG-counterpart attachment untouched. Its sole year test — "the result has zero rows in this year" — gains a parallel qualifying path: "a component code carries nonzero dollars for this government in this year that the verb's own long view structurally excludes." Reachability is decided by anti-joining the real view (`spending_long_harmonized` / `revenue_long_harmonized`) rather than restating its `WHERE` clause, so the trigger tracks the view if it changes. Both verbs inherit the fix from the shared `.verb_spendrev()` call site. - -**Tech Stack:** R (pkg `uscogdata`), DuckDB via DBI, testthat 3e, cli. Downstream: cog-api (plumber). - -## Global Constraints - -- **No new package dependencies.** `withr` stays Suggests-only, test-only. -- **All tests run offline against the bundled fixture** at `inst/extdata/fixture_corpus/` (years 2011, 2012, 2019, 2020). No live corpus, no credentials. Guard every corpus-touching test with `skip_if_no_corpus()`. -- **Every dollar figure in this plan is measured**, not estimated — reproduced against the bundled fixture on 2026-08-04. Assert them exactly. -- **Raw `amt` is in $1,000s**; multiply by `1000.0` in SQL. Values reaching provenance are full US dollars. -- **Suggestions only run under `basis = "harmonized"` and a non-NULL `category`** — the early return at `R/suggestions.R:56` is unchanged and must stay. -- Functions under 50 lines, files under 400 lines. `R/suggestions.R` is currently 251 lines and must stay under 400. -- Use `.sql_lit_chr()` for every string literal interpolated into SQL. Never interpolate a user-supplied value as a SQL *identifier*. -- Commit after each task. Conventional-commit prefixes (`fix:`, `test:`, `docs:`, `chore:`). - -## Reference Data (measured 2026-08-04, bundled fixture) - -Government ids are `canonical_govid`. - -| Case | gov | year | category | today | after | -|---|---|---|---|---|---| -| Target defect | `061037123085` (LA County) | 2011 | Public Welfare | `suggestions` = `list()`, reports $3,185,943,000 | 2 suggestions; `welfare_cash_e67_wide` $1,803,872,000 (E67), `welfare_cash_e68_wide` $271,589,000 (E68) | -| Revenue twin | `020000227749` (Alaska state) | 2011 | Miscellaneous Revenue | `suggestions` = `list()`, reports $943,842,000 from U11,U20,U30 | 1 suggestion; `rents_royalties_u4_wide` $1,899,995,000 (`U4-`) | -| Regression | `061037123085` | 2011 | Corrections | 3 suggestions, `empty_year` | same 3, now carrying $1,371,460,000 (E05), $17,373,000 (F05), $884,000 (G05) | -| Negative | `061037123085` | 2019 | Public Welfare | `list()` | `list()` | -| Negative | `higher_ed_e18_wide`, `general_gov_e89_wide` | any | any | never fire | never fire (E18/E89 are leaf-and-classified in 2011) | - -**Structural invariant the anti-join rests on:** no recipe component is ever renamed by harmonization. Measured: 0 rows in `long` where `item_code` is a recipe component and `harmonized_code IS NOT NULL AND harmonized_code <> item_code`. All 19 components that differ have `harmonized_code IS NULL`, and every one is `is_aggregate = TRUE`. - -**Pre-verified safe** — these existing assertions do *not* regress (0 new fires each): -- `test-recipes.R:196` Broward 2019:2020 Corrections -- `test-expenditure-concept.R:293` AL state 2019 Police -- `test-recipes.R:82` Broward `recipe = "corrections_combined"` — safe by construction; the recipe branch sets `suggestions <- list()` at `R/spending.R:376` and never calls `.build_suggestions()` - -## File Structure - -- **Modify `R/suggestions.R`** — add `.suppressed_components()`; widen `.build_suggestions()`; extend `.inform_suggestions()`. Owns the whole trigger. -- **Modify `R/spending.R`** — add `.select_long_view()` beside `.select_view()` (~line 511); pass the long-view name at the `.build_suggestions()` call site (~line 396). -- **Modify `R/explain.R`** — render the new detail in the Suggestions section (~line 127). -- **Modify `inst/schemas/provenance-v1.json`** — document the four new per-suggestion fields. -- **Modify `tests/testthat/test-recipes.R`** — all new reader tests (this file already owns signposting tests). -- **Modify `NEWS.md`** — user-facing entry. -- **Modify (cog-api) `api/tests/testthat/test-handlers-governments.R`** — assert the fields survive the envelope. - ---- - -### Task 0: Pre-flight — clean branch off `main` - -**Files:** none (git only) - -**Interfaces:** -- Consumes: nothing -- Produces: a clean `fix/partial-coverage-signposting-9` branch based on `origin/main`; a recorded baseline test count later tasks compare against - -> The working tree is currently on `ci/apt-https` with two **uncommitted fixture modifications** (`inst/extdata/fixture_corpus/data/series_breaks.parquet`, `inst/extdata/fixture_corpus/manifest.json`). Do not carry them into this branch and do not discard them without checking with the user — they are unrelated to this work. - -- [ ] **Step 1: Inspect the uncommitted fixture changes** - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata -git status --short -git diff --stat -``` - -Expected: exactly the two fixture files listed above, modified. - -- [ ] **Step 2: Stash them, and say so in the stash message** - -```bash -git stash push -m "WIP fixture series_breaks/manifest, unrelated to #9" \ - inst/extdata/fixture_corpus/data/series_breaks.parquet \ - inst/extdata/fixture_corpus/manifest.json -git status --short -``` - -Expected: clean tree. Report the stash ref to the user so it is not lost. - -- [ ] **Step 3: Branch off current `main`** - -```bash -git fetch origin -git checkout -b fix/partial-coverage-signposting-9 origin/main -git log --oneline -1 -``` - -- [ ] **Step 4: Record the baseline test count** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -20 -``` - -Expected: a green suite. **Write the exact PASS/FAIL/SKIP/WARN numbers into the task notes** — every later task compares against them. Do not proceed if the baseline is already red; report to the user instead. - ---- - -### Task 1: `.suppressed_components()` + the structural guard - -**Files:** -- Modify: `R/suggestions.R` (append after `.build_suggestions()`, before `.attach_ig_counterparts()`) -- Modify: `R/spending.R:511-513` (add `.select_long_view()` beside `.select_view()`) -- Test: `tests/testthat/test-recipes.R` (append at end) - -**Interfaces:** -- Consumes: `.sql_lit_chr(x)` (existing, `R/config.R`); `.select_view(view_base, basis)` (existing, `R/spending.R:511`) -- Produces: - - `.select_long_view(view_base, basis)` → character(1). `"spending_annotated"` + `"harmonized"` → `"spending_long_harmonized"`. - - `.suppressed_components(con, candidates, govid, years, long_view)` → tibble with columns `recipe_id` (chr), `year` (dbl), `suppressed_amount` (dbl, full US dollars), `suppressed_codes` (chr, comma-joined sorted item codes). Zero rows when nothing is suppressed. Aborts on a `long_view` outside the allowlist. - -- [ ] **Step 1: Write the failing tests** - -Append to `tests/testthat/test-recipes.R`: - -```r -# --- uscogdata#9: partial-coverage signposting ------------------------------ - -test_that("no recipe component is ever renamed by harmonization", { - # The suppression trigger anti-joins the verb's long view on item_code. - # That is only sound because harmonization never rewrites a recipe - # component's code -- every component whose harmonized_code differs has - # harmonized_code IS NULL (and is aggregate-flagged). If this ever fails, - # .suppressed_components() would report reachable dollars as suppressed. - skip_if_no_corpus() - con <- uscogdata:::.ensure_session() - n <- DBI::dbGetQuery(con, - "SELECT COUNT(*) AS renamed FROM long - WHERE item_code IN (SELECT DISTINCT component_code FROM harmonization_recipes) - AND harmonized_code IS NOT NULL - AND harmonized_code <> item_code")$renamed - expect_equal(as.integer(n), 0L) -}) - -test_that(".select_long_view maps annotated view bases to their long views", { - expect_equal( - uscogdata:::.select_long_view("spending_annotated", "harmonized"), - "spending_long_harmonized") - expect_equal( - uscogdata:::.select_long_view("revenue_annotated", "harmonized"), - "revenue_long_harmonized") - expect_equal( - uscogdata:::.select_long_view("spending_annotated", "raw"), - "spending_long") -}) - -test_that(".suppressed_components measures the E67/E68 dollars Public Welfare drops", { - skip_if_no_corpus() - con <- uscogdata:::.ensure_session() - s <- uscogdata:::.suppressed_components( - con, - candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"), - govid = "061037123085", years = 2011L, - long_view = "spending_long_harmonized") - - expect_s3_class(s, "tbl_df") - expect_equal(nrow(s), 2L) - s <- s[order(s$recipe_id), ] - expect_equal(s$recipe_id, c("welfare_cash_e67_wide", "welfare_cash_e68_wide")) - expect_equal(s$suppressed_amount, c(1803872000, 271589000)) - expect_equal(s$suppressed_codes, c("E67", "E68")) -}) - -test_that(".suppressed_components finds nothing in a modern year", { - skip_if_no_corpus() - con <- uscogdata:::.ensure_session() - s <- uscogdata:::.suppressed_components( - con, - candidates = c("welfare_cash_e67_wide", "welfare_cash_e68_wide"), - govid = "061037123085", years = 2019L, - long_view = "spending_long_harmonized") - expect_equal(nrow(s), 0L) -}) - -test_that(".suppressed_components rejects a long_view outside the allowlist", { - skip_if_no_corpus() - con <- uscogdata:::.ensure_session() - expect_error( - uscogdata:::.suppressed_components( - con, candidates = "welfare_cash_e67_wide", govid = "061037123085", - years = 2011L, long_view = "long; DROP TABLE x"), - class = "uscogdata_internal_error") -}) -``` - -- [ ] **Step 2: Run to verify they fail** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30 -``` - -Expected: FAIL — `.select_long_view` and `.suppressed_components` are not found. The "no recipe component is ever renamed" test should already PASS (it asserts existing corpus structure). - -- [ ] **Step 3: Add `.select_long_view()` to `R/spending.R`** - -Insert immediately after `.select_view()` (currently ends at line 513): - -```r -#' The `*_long`/`*_long_harmonized` view behind an annotated view base -- -#' `"spending_annotated"` -> `"spending_long_harmonized"`. `.build_suggestions()` -#' anti-joins the LONG view rather than the annotated one: they have identical -#' row membership (the annotated views are the long views plus LEFT JOINs, see -#' inst/sql/42-spending_annotated_harmonized.sql), but the long view is the -#' one that actually owns the `NOT is_aggregate` + crosswalk-membership rule -#' the suppression test is asking about. -#' @noRd -.select_long_view <- function(view_base, basis) { - .select_view(sub("_annotated$", "_long", view_base), basis) -} -``` - -- [ ] **Step 4: Add `.suppressed_components()` to `R/suggestions.R`** - -Insert after `.build_suggestions()` closes (currently line 132), before the `.attach_ig_counterparts()` roxygen block: - -```r -#' Measure, per (recipe, year), the component dollars this government holds -#' that the calling verb's own long view structurally excludes. -#' -#' This is the second qualifying path for a suggestion (uscogdata#9). The -#' first -- row absence -- only fires when a category returns NOTHING in a -#' requested year, which is how Corrections behaves in the wide era. Public -#' Welfare is the failure mode it misses: E74/E75/E77/E79 still return rows, -#' so there is no absence to detect, while E67/E68 (aggregate-flagged 1967- -#' 2011, and absent from `summary_categories` entirely) are dropped. The -#' caller gets a plausible number a third too low, silently. -#' -#' "Structurally excluded" is decided by anti-joining the verb's REAL long -#' view rather than restating its WHERE clause, so this stays correct if -#' `spending_long_harmonized` / `revenue_long_harmonized` ever change. That -#' anti-join is keyed on `item_code`, which is sound only because -#' harmonization never renames a recipe component -- asserted by the "no -#' recipe component is ever renamed by harmonization" test in -#' tests/testthat/test-recipes.R. -#' -#' Note what this deliberately does NOT count as suppressed: a component -#' excluded from the RESULT for scoping reasons -- because it belongs to a -#' different `category`, or because `expenditure_concept` narrowed the -#' subtypes -- is still present in the view, so it never fires. Suggesting a -#' recipe is a coverage fix, not a category redefinition. Measured on the -#' bundled fixture, this keeps `higher_ed_e18_wide` and `general_gov_e89_wide` -#' silent (E18/E89 are leaf-and-classified even in the wide era) and confines -#' every fire to 2011. -#' -#' @param con Active DuckDB connection. -#' @param candidates Character vector of recipe ids to measure. -#' @param govid Character vector of canonical_govid values. -#' @param years Integer vector of requested years. -#' @param long_view Name of the verb's long view, from `.select_long_view()`. -#' @return Tibble of `recipe_id`, `year`, `suppressed_amount` (full US -#' dollars), `suppressed_codes` (comma-joined, sorted). Zero rows when -#' nothing is suppressed. -#' @noRd -.suppressed_components <- function(con, candidates, govid, years, long_view) { - empty <- tibble::tibble( - recipe_id = character(0), year = numeric(0), - suppressed_amount = numeric(0), suppressed_codes = character(0) - ) - if (length(candidates) == 0L) return(empty) - - # long_view is interpolated as a SQL IDENTIFIER, not a literal, so it can - # never be quoted safely. It is always internally derived from a fixed - # view_base, so an off-allowlist value is a programming error, not input. - if (!long_view %in% c("spending_long", "spending_long_harmonized", - "revenue_long", "revenue_long_harmonized")) { - cli::cli_abort( - "Internal error: unexpected `long_view` {.val {long_view}}.", - class = "uscogdata_internal_error" - ) - } - - sql <- sprintf( - "SELECT r.recipe_id, - l.year, - SUM(l.amt) * 1000.0 AS suppressed_amount, - string_agg(DISTINCT l.item_code, ',' ORDER BY l.item_code) - AS suppressed_codes - FROM long l - JOIN harmonization_recipes r - ON l.item_code = r.component_code - AND l.year BETWEEN r.year_min AND r.year_max - AND (r.gov_type_scope = 'all' - OR (r.gov_type_scope = 'state' AND l.type = 0) - OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) - WHERE r.recipe_id IN (%1$s) - AND l.canonical_govid IN (%2$s) - AND l.year IN (%3$s) - AND l.amt <> 0 - AND NOT EXISTS ( - SELECT 1 FROM %4$s v - WHERE v.canonical_govid = l.canonical_govid - AND v.year = l.year - AND v.item_code = l.item_code - ) - GROUP BY 1, 2 - ORDER BY 1, 2", - .sql_lit_chr(candidates), .sql_lit_chr(govid), - paste(as.integer(years), collapse = ","), long_view - ) - tibble::as_tibble(DBI::dbGetQuery(con, sql)) -} -``` - -- [ ] **Step 5: Run to verify they pass** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30 -``` - -Expected: PASS, 0 failures. - -- [ ] **Step 6: Commit** - -```bash -git add R/suggestions.R R/spending.R tests/testthat/test-recipes.R -git commit -m "feat: measure structurally-suppressed recipe component dollars (#9)" -``` - ---- - -### Task 2: Widen `.build_suggestions()` and add the four fields - -**Files:** -- Modify: `R/suggestions.R:54-132` (`.build_suggestions()`) and its header comment block at lines 1-31 -- Modify: `R/spending.R:396-398` (call site) -- Test: `tests/testthat/test-recipes.R` - -**Interfaces:** -- Consumes: `.suppressed_components(con, candidates, govid, years, long_view)` and `.select_long_view(view_base, basis)` from Task 1 -- Produces: `.build_suggestions(con, govid, years, category, result, basis, flow_prefixes, long_view)` — one new trailing argument. Each returned suggestion is a list with the five existing keys (`recipe_id`, `label`, `available_years`, `hint`, `ig_recipe_id`) plus: - - `trigger` — chr(1), `"empty_year"` or `"suppressed_component"`. `"empty_year"` wins when both apply. - - `suppressed_amount` — dbl(1), full US dollars summed across requested years, `0` when none. - - `suppressed_years` — integer vector, sorted, `integer(0)` when none. - - `suppressed_codes` — chr vector, sorted unique, `character(0)` when none. - -- [ ] **Step 1: Write the failing tests** - -Append to `tests/testthat/test-recipes.R`: - -```r -test_that("uscogdata#9: Public Welfare signposts its suppressed E67/E68 dollars", { - # The bug: E74/E79 return rows for FY2011, so there is no row-absence gap, - # so nothing fired -- while E67 ($1,803,872,000) and E68 ($271,589,000) were - # dropped for being aggregate-published. LA County reports $3,185,943,000 - # and omits $2,075,461,000, a 39% understatement, silently. - skip_if_no_corpus() - r <- suppressMessages( - cog_spending("061037123085", years = 2011L, category = "Public Welfare")) - sugg <- attr(r, "provenance")$suggestions - - expect_length(sugg, 2L) - ids <- vapply(sugg, function(s) s$recipe_id, character(1)) - expect_setequal(ids, c("welfare_cash_e67_wide", "welfare_cash_e68_wide")) - - e67 <- sugg[[which(ids == "welfare_cash_e67_wide")]] - expect_equal(e67$trigger, "suppressed_component") - expect_equal(e67$suppressed_amount, 1803872000) - expect_equal(e67$suppressed_years, 2011L) - expect_equal(e67$suppressed_codes, "E67") - expect_equal(e67$hint, "re-run with recipe = 'welfare_cash_e67_wide'") - - e68 <- sugg[[which(ids == "welfare_cash_e68_wide")]] - expect_equal(e68$trigger, "suppressed_component") - expect_equal(e68$suppressed_amount, 271589000) - expect_equal(e68$suppressed_codes, "E68") -}) - -test_that("uscogdata#9: an empty_year fire keeps its trigger and gains the dollars", { - # Corrections is the case that already worked: zero rows in FY2011, so the - # row-absence path fires. It must keep firing, keep trigger = "empty_year", - # keep its IG counterpart -- and now also report what was suppressed. - skip_if_no_corpus() - r <- suppressMessages( - cog_spending("061037123085", years = 2011L, category = "Corrections")) - sugg <- attr(r, "provenance")$suggestions - - expect_length(sugg, 3L) - ids <- vapply(sugg, function(s) s$recipe_id, character(1)) - expect_setequal(ids, c("corrections_combined", "corrections_capital_combined", - "corrections_other_capital_combined")) - expect_true(all(vapply(sugg, function(s) s$trigger, character(1)) == "empty_year")) - - cc <- sugg[[which(ids == "corrections_combined")]] - expect_equal(cc$suppressed_amount, 1371460000) - expect_equal(cc$suppressed_codes, "E05") - expect_equal(cc$ig_recipe_id, "corrections_ig_local_combined") -}) - -test_that("uscogdata#9: the revenue verb inherits the same trigger", { - # Alaska state FY2011 Miscellaneous Revenue reports $943,842,000 from - # U11/U20/U30 while dropping $1,899,995,000 of aggregate-published `U4-` - # rents and royalties -- the omission is LARGER than the reported figure. - skip_if_no_corpus() - r <- suppressMessages( - cog_revenue("020000227749", years = 2011L, - category = "Miscellaneous Revenue")) - sugg <- attr(r, "provenance")$suggestions - - expect_length(sugg, 1L) - expect_equal(sugg[[1]]$recipe_id, "rents_royalties_u4_wide") - expect_equal(sugg[[1]]$trigger, "suppressed_component") - expect_equal(sugg[[1]]$suppressed_amount, 1899995000) - expect_equal(sugg[[1]]$suppressed_codes, "U4-") - # A revenue recipe must never be handed an M/L expenditure counterpart. - expect_null(sugg[[1]]$ig_recipe_id) -}) - -test_that("uscogdata#9: no partial-coverage fire in a modern year", { - skip_if_no_corpus() - r <- cog_spending("061037123085", years = 2019L, category = "Public Welfare") - expect_length(attr(r, "provenance")$suggestions, 0L) -}) - -test_that("uscogdata#9: leaf-and-classified wide-era families never fire", { - # higher_ed_e18_wide and general_gov_e89_wide are the control group: their - # components (E16/E18, E85/E89) are ordinary classified leaves even in the - # wide era, so widening the trigger must leave them silent. This is the - # measurement that refutes "it would fire on every category in every legacy - # year" -- corpus-wide on the fixture, these two produce zero suppressed rows. - skip_if_no_corpus() - con <- uscogdata:::.ensure_session() - n <- DBI::dbGetQuery(con, - "SELECT COUNT(*) AS n - FROM long l - JOIN harmonization_recipes r - ON l.item_code = r.component_code - AND l.year BETWEEN r.year_min AND r.year_max - WHERE r.recipe_id IN ('higher_ed_e18_wide', 'general_gov_e89_wide') - AND l.amt <> 0 - AND NOT EXISTS ( - SELECT 1 FROM spending_long_harmonized v - WHERE v.canonical_govid = l.canonical_govid - AND v.year = l.year AND v.item_code = l.item_code)")$n - expect_equal(as.integer(n), 0L) -}) -``` - -- [ ] **Step 2: Run to verify they fail** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -40 -``` - -Expected: FAIL — the Public Welfare and revenue tests get `length(sugg) == 0`; the Corrections test fails on the missing `suppressed_amount`. The two negative-control tests should already PASS. - -- [ ] **Step 3: Replace the body of `.build_suggestions()`** - -In `R/suggestions.R`, change the signature and everything from `result_years <-` through the closing `.attach_ig_counterparts(...)` call. The `candidates` query (lines 70-81) is **unchanged** — do not touch it. - -```r -.build_suggestions <- function(con, govid, years, category, result, basis, - flow_prefixes, long_view) { - if (!identical(basis, "harmonized") || is.null(category)) return(list()) - - # ... candidates query unchanged ... - if (length(candidates) == 0L) return(list()) - - result_years <- if (is.null(result) || nrow(result) == 0L) { - integer(0) - } else { - unique(as.integer(result$year)) - } - gap_years <- setdiff(as.integer(years), result_years) - - # Path 2 (uscogdata#9): component dollars this government holds that the - # verb's own view structurally excludes. Measured across ALL requested - # years, not just gap years -- the whole point is that a year with rows can - # still be missing dollars. - supp <- .suppressed_components(con, candidates, govid, years, long_view) - - if (length(gap_years) == 0L && nrow(supp) == 0L) return(list()) - - meta <- tibble::as_tibble(DBI::dbGetQuery(con, sprintf( - "SELECT recipe_id, any_value(label) AS label, - MIN(year_min) AS year_min, MAX(year_max) AS year_max - FROM harmonization_recipes - WHERE recipe_id IN (%s) - GROUP BY recipe_id", - .sql_lit_chr(candidates) - ))) - - # Path 1 (unchanged): (recipe, year) pairs the recipe's own generic join - # covers for this government, restricted to the gap years. - covered <- if (length(gap_years) == 0L) { - data.frame(recipe_id = character(0), year = integer(0)) - } else { - DBI::dbGetQuery(con, sprintf( - "SELECT DISTINCT r.recipe_id, l.year - FROM long l - JOIN harmonization_recipes r - ON l.item_code = r.component_code - AND l.year BETWEEN r.year_min AND r.year_max - AND (r.gov_type_scope = 'all' - OR (r.gov_type_scope = 'state' AND l.type = 0) - OR (r.gov_type_scope = 'local' AND l.type BETWEEN 1 AND 3)) - WHERE r.recipe_id IN (%s) - AND l.canonical_govid IN (%s) - AND l.year IN (%s)", - .sql_lit_chr(candidates), .sql_lit_chr(govid), - paste(gap_years, collapse = ",") - )) - } - - suggestions <- list() - for (rid in candidates) { - empty_hit <- rid %in% covered$recipe_id - s_rows <- supp[supp$recipe_id == rid, , drop = FALSE] - supp_hit <- nrow(s_rows) > 0L - if (!empty_hit && !supp_hit) next - m <- meta[meta$recipe_id == rid, ] - suggestions[[length(suggestions) + 1L]] <- list( - recipe_id = rid, - label = m$label[[1]], - available_years = c(as.integer(m$year_min), as.integer(m$year_max)), - hint = sprintf("re-run with recipe = '%s'", rid), - # An empty year is the stronger claim -- the category returned nothing - # at all -- so it wins when both paths qualify. The suppressed_* fields - # are still populated, so an empty_year fire also reports its dollars. - trigger = if (empty_hit) "empty_year" else "suppressed_component", - suppressed_amount = if (supp_hit) sum(s_rows$suppressed_amount) else 0, - suppressed_years = if (supp_hit) { - sort(unique(as.integer(s_rows$year))) - } else { - integer(0) - }, - suppressed_codes = if (supp_hit) { - sort(unique(unlist(strsplit(s_rows$suppressed_codes, ",", fixed = TRUE)))) - } else { - character(0) - } - ) - } - .attach_ig_counterparts(con, suggestions, flow_prefixes) -} -``` - -- [ ] **Step 4: Update the file header comment** - -In `R/suggestions.R`, replace the first paragraph (lines 1-5) so the file's stated contract matches its behavior: - -```r -# R/suggestions.R -# Recipe-component-driven signposting. When a basis = "harmonized" query for -# a category comes back incomplete in some requested year -- and a -# harmonization recipe would actually fill it for this government -- surface -# that recipe as a suggestion. "Incomplete" has two forms, and a recipe -# qualifies on either: -# 1. empty_year -- the result has no rows at all in that year. -# 2. suppressed_component -- the result HAS rows, but a component code -# carries dollars the verb's own long view structurally excludes -# (aggregate-published, or absent from summary_categories). This is -# uscogdata#9: Public Welfare kept returning E74/E79 rows while dropping -# aggregate-only E67/E68, so form 1 never fired and the caller got a -# number a third too low with no signpost at all. -``` - -- [ ] **Step 5: Update the call site in `R/spending.R`** - -At lines 396-398, pass the long view: - -```r - suggestions <- .build_suggestions(con, govid, years, category, - direct_leg_result, - resolved$basis, flow_prefixes, - .select_long_view(view_base, resolved$basis)) -``` - -- [ ] **Step 6: Run the recipes tests** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30 -``` - -Expected: PASS, 0 failures. - -- [ ] **Step 7: Run the full suite — this is the regression gate** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -30 -``` - -Expected: 0 FAIL, and PASS count = Task 0 baseline plus the new tests. The three pre-verified `expect_length(suggestions, 0L)` assertions (`test-recipes.R:196`, `test-recipes.R:82`, `test-expenditure-concept.R:293`) must still pass. If any fails, **stop and report** — it means the trigger is firing somewhere the measurement said it would not. - -- [ ] **Step 8: Commit** - -```bash -git add R/suggestions.R R/spending.R tests/testthat/test-recipes.R -git commit -m "fix: signpost aggregate-suppressed components in a category that still has rows (#9)" -``` - ---- - -### Task 3: Surface the dollars in the cli message and `cog_explain()` - -**Files:** -- Modify: `R/suggestions.R:237-251` (`.inform_suggestions()`) -- Modify: `R/explain.R:125-132` -- Test: `tests/testthat/test-recipes.R` - -**Interfaces:** -- Consumes: the suggestion fields from Task 2 -- Produces: no new functions; `.inform_suggestions()` and `cog_explain()` render `suppressed_amount` / `suppressed_years` / `suppressed_codes` when `suppressed_amount > 0` - -- [ ] **Step 1: Write the failing tests** - -Append to `tests/testthat/test-recipes.R`: - -```r -test_that("uscogdata#9: the cli message reports the suppressed dollars", { - skip_if_no_corpus() - expect_message( - cog_spending("061037123085", years = 2011L, category = "Public Welfare"), - "1,803,872,000", fixed = TRUE) - expect_message( - cog_spending("061037123085", years = 2011L, category = "Public Welfare"), - "FY2011", fixed = TRUE) - expect_message( - cog_spending("061037123085", years = 2011L, category = "Public Welfare"), - "E67", fixed = TRUE) -}) - -test_that("uscogdata#9: cog_explain() reports the suppressed dollars", { - skip_if_no_corpus() - r <- suppressMessages( - cog_spending("061037123085", years = 2011L, category = "Public Welfare")) - out <- paste(capture.output(cog_explain(r), type = "message"), - capture.output(cog_explain(r)), collapse = "\n") - expect_match(out, "271,589,000", fixed = TRUE) -}) -``` - -- [ ] **Step 2: Run to verify they fail** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -30 -``` - -Expected: FAIL — no dollar figures in either output. - -- [ ] **Step 3: Extend `.inform_suggestions()`** - -Replace the body in `R/suggestions.R`: - -```r -.inform_suggestions <- function(suggestions) { - bullets <- vapply(suggestions, function(s) { - bullet <- sprintf("%s (%d-%d): %s", s$recipe_id, - s$available_years[1], s$available_years[2], s$hint) - # Only present when dollars were actually measured as excluded. An - # empty_year fire can carry them too -- the year had no rows AND the - # component was suppressed -- which is strictly more informative. - if (isTRUE(s$suppressed_amount > 0)) { - bullet <- paste0(bullet, sprintf( - "\n $%s excluded from %s (%s), published as an aggregate or outside the crosswalk", - formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","), - paste0("FY", s$suppressed_years, collapse = ", "), - paste(s$suppressed_codes, collapse = ", "))) - } - if (!is.null(s$ig_recipe_id)) { - bullet <- paste0(bullet, sprintf( - "\n intergovernmental counterpart: recipe = '%s'", s$ig_recipe_id)) - } - bullet - }, character(1)) - cli::cli_inform(c( - i = "Incomplete coverage for the requested years; a harmonization recipe may fill it:", - stats::setNames(bullets, rep("*", length(bullets))) - )) -} -``` - -> The header changed from "Coverage gap detected for the requested years" because a partial-coverage fire is not a gap — the year has rows, they are just short. Verified 2026-08-04: no test in `tests/` asserts the old string, so this is safe. - -- [ ] **Step 4: Extend the `cog_explain()` renderer** - -Replace lines 125-132 of `R/explain.R`: - -```r - if (length(prov$suggestions) > 0L) { - cli::cli_h2("Suggestions") - sugg_lines <- vapply(prov$suggestions, function(s) { - line <- sprintf("%s -- %s (years %s-%s): %s", s$recipe_id, s$label, - s$available_years[1], s$available_years[2], s$hint) - if (isTRUE(s$suppressed_amount > 0)) { - line <- paste0(line, sprintf(" [$%s excluded from %s: %s]", - formatC(s$suppressed_amount, format = "f", digits = 0, big.mark = ","), - paste0("FY", s$suppressed_years, collapse = ", "), - paste(s$suppressed_codes, collapse = ", "))) - } - line - }, character(1)) - cli::cli_ul(sugg_lines) - } -``` - -- [ ] **Step 5: Run to verify they pass** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes|explain")' 2>&1 | tail -30 -``` - -Expected: PASS, 0 failures. - -- [ ] **Step 6: Commit** - -```bash -git add R/suggestions.R R/explain.R tests/testthat/test-recipes.R -git commit -m "feat: report suppressed component dollars in the signpost message (#9)" -``` - ---- - -### Task 4: Document the contract — schema, NEWS, roxygen - -**Files:** -- Modify: `inst/schemas/provenance-v1.json:36` -- Modify: `NEWS.md` (new section at top, under the `# uscogdata 0.1.0 (development)` heading) -- Test: `tests/testthat/test-recipes.R` - -**Interfaces:** -- Consumes: the suggestion fields from Task 2 -- Produces: no code interfaces; the schema is the published contract cog-api reads against - -- [ ] **Step 1: Write the failing test** - -Append to `tests/testthat/test-recipes.R`: - -```r -test_that("the provenance schema documents the suggestion trigger fields", { - sch <- jsonlite::fromJSON( - system.file("schemas", "provenance-v1.json", package = "uscogdata"), - simplifyVector = FALSE) - props <- sch$properties$suggestions$items$properties - expect_true(all(c("trigger", "suppressed_amount", "suppressed_years", - "suppressed_codes") %in% names(props))) - expect_setequal(unlist(props$trigger$enum), - c("empty_year", "suppressed_component")) -}) -``` - -- [ ] **Step 2: Run to verify it fails** - -```bash -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(filter = "recipes")' 2>&1 | tail -20 -``` - -Expected: FAIL — `sch$properties$suggestions$items` is NULL. - -- [ ] **Step 3: Replace the `suggestions` property in `inst/schemas/provenance-v1.json`** - -Replace line 36 (`"suggestions": { "type": "array" },`) with: - -```json - "suggestions": { - "type": "array", - "description": "Harmonization recipes that would fill incomplete coverage in the requested years for this government. Empty on a healthy query, on an un-scoped (category = NULL) query, on basis = 'raw', and on a recipe = query (which resolves its own coverage).", - "items": { - "type": "object", - "required": ["recipe_id", "label", "available_years", "hint", "trigger", - "suppressed_amount", "suppressed_years", "suppressed_codes"], - "properties": { - "recipe_id": { "type": "string" }, - "label": { "type": "string" }, - "available_years": { - "type": "array", - "items": { "type": "integer" }, - "description": "[year_min, year_max] of the recipe's component coverage." - }, - "hint": { "type": "string" }, - "ig_recipe_id": { - "type": ["string", "null"], - "description": "The intergovernmental (M/L) counterpart recipe covering the same function suffixes, or null. Never set for revenue recipes." - }, - "trigger": { - "type": "string", - "enum": ["empty_year", "suppressed_component"], - "description": "Why this fired. 'empty_year': the result has no rows at all in a requested year. 'suppressed_component': the result HAS rows, but a component code carries dollars the verb's long view structurally excludes -- aggregate-published, or absent from summary_categories. 'empty_year' wins when both apply, being the stronger claim; the suppressed_* fields are populated either way." - }, - "suppressed_amount": { - "type": "number", - "description": "Full US dollars this government holds in the recipe's component codes that the result excludes, summed across the requested years. 0 when nothing is suppressed." - }, - "suppressed_years": { - "type": "array", - "items": { "type": "integer" }, - "description": "The requested years contributing to suppressed_amount." - }, - "suppressed_codes": { - "type": "array", - "items": { "type": "string" }, - "description": "The excluded component item codes, sorted." - } - } - } - }, -``` - -- [ ] **Step 4: Add the NEWS.md entry** - -Insert immediately after the `# uscogdata 0.1.0 (development)` line: - -```markdown -## Signposting now catches partially-suppressed categories - -* A coverage suggestion used to fire only when a category returned **no rows - at all** in a requested year. That missed the more dangerous case: a - category that still returns rows while silently dropping component codes - the wide era publishes only as aggregates (#9). `cog_spending(category = - "Public Welfare")` for FY2011 returned a plausible figure that omitted - `E67`/`E68` entirely -- for Los Angeles County, $2,075,461,000 of a true - $5,261,404,000, a 39% understatement, with `provenance$suggestions` empty. -* Suggestions now also fire on **partial** coverage, and every suggestion - carries `trigger` (`"empty_year"` or `"suppressed_component"`), - `suppressed_amount`, `suppressed_years` and `suppressed_codes`, so a caller - can see how much is missing and decide whether to re-run with the recipe. -* `cog_revenue()` gets the same fix through the shared verb path. Alaska's - FY2011 `Miscellaneous Revenue` reported $943,842,000 while dropping - $1,899,995,000 of aggregate-published `U4-` rents and royalties. -* The trigger stays recipe-driven, so it only fires where a harmonization - recipe actually exists to name the fix. Measured on the bundled fixture, - every fire lands in the wide era; `higher_ed_e18_wide` and - `general_gov_e89_wide` stay silent, because their components are ordinary - classified leaves even pre-2012. -``` - -- [ ] **Step 5: Regenerate docs and run the full suite** - -```bash -/usr/bin/Rscript -e 'devtools::document()' -/usr/bin/Rscript -e 'devtools::load_all("."); testthat::test_local(reporter = "summary")' 2>&1 | tail -30 -``` - -Expected: 0 FAIL. `devtools::document()` should produce no `man/` changes (all edits are `@noRd` or non-roxygen); if it does, include them in the commit. - -- [ ] **Step 6: Commit** - -```bash -git add inst/schemas/provenance-v1.json NEWS.md tests/testthat/test-recipes.R man/ -git commit -m "docs: document the suggestion trigger and suppressed-dollar fields (#9)" -``` - ---- - -### Task 5: Ship the reader — check, push, PR - -**Files:** none (verification + git) - -**Interfaces:** -- Consumes: everything from Tasks 1-4 -- Produces: a green CI run and an open PR on `gitea.civilytics.org/Civilytics/uscogdata` - -- [ ] **Step 1: Run R CMD check** - -```bash -/usr/bin/Rscript -e 'rcmdcheck::rcmdcheck(args = c("--no-manual", "--as-cran"), error_on = "warning")' 2>&1 | tail -40 -``` - -Expected: 0 errors, 0 warnings. Notes about fixture size are pre-existing. - -- [ ] **Step 2: Confirm the file-length constraint still holds** - -```bash -wc -l R/suggestions.R R/spending.R R/explain.R -``` - -Expected: `R/suggestions.R` under 400 lines. If it has crossed, split `.suppressed_components()` into `R/suppression.R` and re-run the suite before continuing. - -- [ ] **Step 3: Push and open the PR** - -```bash -git push -u origin fix/partial-coverage-signposting-9 -tea pr create --title "fix: signpost partially-suppressed categories (#9)" \ - --description "Closes #9. - -\`.build_suggestions()\` fired only when a category returned zero rows in a requested year. Public Welfare is the failure mode that missed: E74/E79 still return rows for legacy years, so no row-absence gap existed, while aggregate-published E67/E68 were dropped by \`spending_long\`'s \`NOT is_aggregate\` filter. LA County FY2011 reported \$3,185,943,000 and omitted \$2,075,461,000 -- a 39% understatement -- with \`provenance\$suggestions\` empty. - -A recipe now also qualifies when a component code carries dollars the verb's own long view structurally excludes, measured per government by anti-joining the real view. Each suggestion carries \`trigger\`, \`suppressed_amount\`, \`suppressed_years\` and \`suppressed_codes\`. - -\`cog_revenue()\` inherits it: Alaska FY2011 Miscellaneous Revenue dropped \$1,899,995,000 of \`U4-\`. - -Blast radius measured on the bundled fixture: every fire lands in the wide era, none in 2012/2019/2020, and the leaf-and-classified control families (\`higher_ed_e18_wide\`, \`general_gov_e89_wide\`) stay silent." -``` - -- [ ] **Step 4: Wait for CI green, then merge** - -```bash -tea pr list -``` - -Report the PR number and CI status to the user. **Do not merge without confirming CI is green** — merging is what cog-api's next build picks up. - ---- - -### Task 6: cog-api — assert the fields survive the envelope, then deploy - -**Files:** -- Modify: `~/Nextcloud/Civilytics/Code/Civilytics/cog-api/api/tests/testthat/test-handlers-governments.R` (append near the existing suggestions test at line 162) - -**Interfaces:** -- Consumes: the merged reader from Task 5. cog-api needs **no handler code change** — `prov$suggestions` is lifted verbatim into the envelope at `api/R/handlers_query.R:100` and `api/R/handlers_governments.R:131,195,317`. -- Produces: a deployed API whose `/spending` responses carry the new fields - -- [ ] **Step 1: Confirm the reader ref is not pinned to a stale SHA** - -cog-api's CI and build clone uscogdata at `USCOGDATA_REF` (a Gitea repo variable, default `main` — see `.gitea/workflows/ci.yml:99-121`). - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api -tea api -X GET /repos/Civilytics/cog-api/actions/variables -``` - -Expected: `USCOGDATA_REF` is absent or `main`. **If it is pinned to a commit SHA, stop and report to the user** — the deploy will silently ship the old reader, and every assertion below will fail for the wrong reason. - -- [ ] **Step 2: Write the failing test** - -Append to `api/tests/testthat/test-handlers-governments.R`: - -```r -test_that("suppressed-component suggestion fields survive the envelope (uscogdata#9)", { - # LA County FY2011 Public Welfare returns rows but drops aggregate-published - # E67/E68. The reader signposts it; the envelope must not flatten the detail - # away, because a Tableau consumer only ever sees the envelope. - res <- handle_gov_spending("061037123085", years = "2011", - category = "Public Welfare", - adjust = NULL, base_year = NULL, per_capita = NULL, - subtype = NULL, limit = "1000", page = "0") - ids <- vapply(res$suggestions, function(s) s$recipe_id, character(1)) - expect_true("welfare_cash_e67_wide" %in% ids) - - hit <- res$suggestions[[which(ids == "welfare_cash_e67_wide")]] - expect_equal(hit$trigger, "suppressed_component") - expect_equal(hit$suppressed_amount, 1803872000) - expect_equal(hit$suppressed_codes, "E67") -}) -``` - -- [ ] **Step 3: Install the merged reader locally and run the test** - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api -/usr/bin/Rscript -e 'remotes::install_local("../cog_explorer/uscogdata", upgrade = "never")' -cd api/tests/testthat && /usr/bin/Rscript -e 'testthat::test_local(filter = "handlers-governments")' 2>&1 | tail -30 -``` - -Expected: PASS. If it fails with `length(ids) == 0`, the installed reader is stale — re-run the install step. - -- [ ] **Step 4: Run the full cog-api suite** - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api/api/tests/testthat -/usr/bin/Rscript -e 'testthat::test_local(reporter = "summary")' 2>&1 | tail -30 -``` - -Expected: 0 FAIL. Record the count. - -- [ ] **Step 5: Commit, push, PR** - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog-api -git checkout -b test/suppressed-component-fields -git add api/tests/testthat/test-handlers-governments.R -git commit -m "test: assert suppressed-component suggestion fields survive the envelope (uscogdata#9)" -git push -u origin test/suppressed-component-fields -tea pr create --title "test: assert uscogdata#9 suggestion fields survive the envelope" \ - --description "uscogdata#9 adds \`trigger\`, \`suppressed_amount\`, \`suppressed_years\` and \`suppressed_codes\` to each \`prov\$suggestions\` entry. cog-api lifts \`prov\$suggestions\` verbatim, so no handler change is needed -- this pins that pass-through so a future envelope refactor cannot silently drop the detail. - -Merging this is the prod deploy that picks up the new reader." -``` - -- [ ] **Step 6: Merge after CI green, then verify the live endpoint** - -> Merging a cog-api PR **is** the production deploy, and `deploy.yml` does not wait for CI. Confirm CI is green *before* merging, not after. - -```bash -tea pr list -# after merge + deploy settles: -curl -s "https:///v1/governments/061037123085/spending?years=2011&category=Public%20Welfare" \ - | python3 -c "import json,sys; d=json.load(sys.stdin); print(json.dumps(d['suggestions'], indent=2))" -``` - -Expected: two suggestions, `welfare_cash_e67_wide` carrying `"suppressed_amount": 1803872000`. Report the live output to the user. - ---- - -### Task 7: Close out the issues - -**Files:** none (issue tracker) - -**Interfaces:** -- Consumes: the merged and deployed work from Tasks 5-6 -- Produces: uscogdata#9 closed with a corrected record - -- [ ] **Step 1: Comment on uscogdata#9 with the resolution** - -```bash -cd ~/Nextcloud/Civilytics/Code/Civilytics/cog_explorer/uscogdata -tea comment 9 "**Fixed.** The trigger was the problem, not the crosswalk. - -\`.build_suggestions()\` already identified \`welfare_cash_e67_wide\` and \`welfare_cash_e68_wide\` as candidates -- \`J67\`/\`J68\` are Public Welfare members, so the candidate query matched. It then threw them away at the \`gap_years\` early return, because E74/E79 returned rows and there was no empty year to detect. - -A recipe now also qualifies when a component code carries dollars the verb's own long view structurally excludes -- aggregate-published, or absent from \`summary_categories\`. Reachability is decided by anti-joining the real view rather than restating its WHERE clause, which is sound because no recipe component is ever renamed by harmonization (now asserted). - -Each suggestion carries \`trigger\`, \`suppressed_amount\`, \`suppressed_years\`, \`suppressed_codes\`. \`cog_revenue()\` inherits the fix. - -**On the 'would fire on every category in every legacy year' worry** -- measured, it does not. Candidates stay gated on the 24-recipe catalog, so it only fires where an actionable recipe exists. On the bundled fixture every fire lands in the wide era, none in 2012/2019/2020, and \`higher_ed_e18_wide\`/\`general_gov_e89_wide\` never fire at all because E18/E89 are ordinary classified leaves pre-2012. - -**The 'minor' sub-item is withdrawn, not implemented.** \`aggregate_fallback\` is no longer vestigial: \`.build_verb_sql()\` uses \`bool_or(is_aggregate)\` (R/spending.R:611), and \`ig_long\` deliberately keeps aggregate rows, so the flag is live for the intergovernmental leg -- \`test-expenditure-concept.R:99\` asserts aggregate-sourced IG dollars report TRUE. Removing it would break that disclosure. - -**Reproduction, now signposted** (bundled fixture, LA County): -\`\`\`r -r <- cog_spending(\"061037123085\", years = 2011, category = \"Public Welfare\") -attr(r, \"provenance\")\$suggestions[[1]]\$suppressed_amount # 1803872000 -\`\`\`" -``` - -- [ ] **Step 2: Close the issue** - -```bash -tea issue close 9 -tea issue 9 | head -5 -``` - -Expected: state `closed`. - -- [ ] **Step 3: Restore the stashed fixture changes** - -```bash -git checkout ci/apt-https -git stash list -git stash pop -git status --short -``` - -Expected: the two fixture files modified again, as they were before Task 0. Report to the user that they are restored and still uncommitted. - ---- - -## Self-Review - -**Spec coverage** — every element of the approved design maps to a task: - -| Design element | Task | -|---|---| -| Second trigger in `.build_suggestions()`, candidate query untouched | 2 | -| Anti-join the real view, not a restated predicate | 1 | -| Structural invariant (no renamed components) asserted | 1 | -| Scoping-excluded ≠ suppressed; F-007 stays out | 1 (roxygen), 2 (control test) | -| Four new suggestion fields, `empty_year` precedence | 2 | -| Both verbs via the shared call site | 2 (revenue test) | -| cli message + `cog_explain()` | 3 | -| Provenance schema documents the fields (no v1 bump) | 4 | -| LA County / Alaska / Corrections / 2019 / control tests | 2, 3 | -| cog-api pass-through + redeploy | 6 | -| Close #9, withdraw the `aggregate_fallback` sub-item | 7 | - -**Placeholder scan** — no TBD/TODO; every code step carries real code; every assertion carries a measured value. - -**Type consistency** — `.suppressed_components()` returns `recipe_id`/`year`/`suppressed_amount`/`suppressed_codes` in Task 1 and is consumed under exactly those names in Task 2. `.select_long_view(view_base, basis)` is defined in Task 1 and called with that signature in Task 2. `suppressed_codes` is comma-joined **inside** the tibble (Task 1) and split into a character vector **for the suggestion** (Task 2) — deliberate, and the tests assert both shapes correctly. diff --git a/plans/2026-08-08-public-release.md b/plans/2026-08-08-public-release.md deleted file mode 100644 index 9134625..0000000 --- a/plans/2026-08-08-public-release.md +++ /dev/null @@ -1,1091 +0,0 @@ -# uscogdata 0.3.0 Public Release Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Make `uscogdata` installable and usable by a stranger — fix the corpus-unreachable defect, correct release metadata, and rewrite README and NEWS for a first public release. - -**Architecture:** Two code changes fix the P0 (manifest-driven file enumeration replacing an HTTP-incompatible glob; a working public default URL). Everything else is metadata, packaging config, and documentation. Distribution mechanics (Gitea public, GitHub mirror, r-universe) are **out of scope for this plan** — they follow after checks are green. - -**Tech Stack:** R (>= 4.1), DuckDB 1.5.5 via `duckdb`/`DBI`, `httr2`, `jsonlite`, `cli`, testthat 3e, pkgdown, roxygen2 7.3.3. - -**Design spec:** `specs/2026-08-08-public-release-design.md` - -## Global Constraints - -- Package license is **MIT**. Copyright holder is **Civilytics Consulting LLC**. -- Author of record: **Jared E. Knowles**, `jared@civilytics.com`, ORCID **0000-0003-0005-9478**, roles `aut`/`cre`. Civilytics Consulting LLC is `cph`/`fnd`. -- Public corpus base URL (note the trailing slash, which `.resolve_url()` requires): - `https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/` -- Published corpus is **schema_version 7**; bundled fixture is **schema_version 6**. `.validate_schema()` accepts `c(4L, 5L, 6L, 7L)` and that list does not change in this plan. -- Corpus facts for documentation, measured 2026-08-08: **46,148,034 rows**, **56 partitions**, **190.6 MB**, government types **0–3**, **FY1967–FY2024** (no source data FY1968, FY1969). -- The bundled fixture at `inst/extdata/fixture_corpus/` **ships in the release**. Never add it to `.Rbuildignore`. -- Every amount column a verb returns is in **full US dollars**, already multiplied by 1000. Documentation must never tell a user to apply the $1,000s rule to verb output. -- Run the full suite with `devtools::test()` from the package root. It uses the bundled fixture and requires no network or credentials. - ---- - -## File Structure - -**Modified — code** -- `R/views.R` — gains `.long_files_sql()`; `.register_views()` substitutes a second token. This is the only file that knows how the `long` table's paths are built. -- `inst/sql/10-long.sql` — the one file containing a glob. Becomes token-driven. -- `R/config.R` — default corpus URL. -- `R/manifest.R` — `.check_url_configured()` guidance text only; the sentinel check itself is unchanged. - -**Modified — metadata/packaging** -- `DESCRIPTION`, `LICENSE`, `.Rbuildignore`, `_pkgdown.yml` - -**Created** -- `LICENSE.md`, `CONTRIBUTING.md` -- `tests/testthat/test-long-files.R` — unit tests for enumeration -- `tests/testthat/test-live-corpus.R` — network-gated integration test - -**Rewritten** -- `README.md`, `NEWS.md` - -**Deliberately untouched:** every verb file (`R/spending.R`, `R/revenue.R`, `R/balances.R`, `R/rollup.R`, `R/peers.R`, `R/search.R`), `R/mirror.R`, and all other `inst/sql/*.sql`. The enumeration fix is confined to view registration by design. - ---- - -## Task 1: Manifest-driven partition enumeration - -The P0 defect, part one. `inst/sql/10-long.sql` globs `{url}data/long/**/*.parquet`. DuckDB 1.5.5 refuses globs over generic HTTP, and `allow_asterisks_in_http_paths` does not help — it forwards the literal `**/*` as a filename and 404s, because HTTP exposes no directory listing. The manifest already enumerates every partition under `files$long_partitions[]`. - -**Files:** -- Modify: `R/views.R` (add helper; `.register_views()` at the `gsub` line) -- Modify: `inst/sql/10-long.sql:3` -- Test: `tests/testthat/test-long-files.R` (create) - -**Interfaces:** -- Consumes: `.sql_lit_chr(x)` from `R/spending.R:553` — quotes each element, escapes `'` by doubling, joins with `,` and **no space**. `%||%` from `R/manifest.R:158`. -- Produces: `.long_files_sql(url, manifest)` returning a single SQL string — either a bracketed list literal `['a','b']` or, on fallback, a single quoted glob `'…/**/*.parquet'`. Task 3 relies on this being the only place partition paths are constructed. - -**Critical constraint:** `tests/testthat/test-views.R:322` calls `.register_views(con, url, manifest = list(schema_version = 4L))` — a manifest with **no `files` element at all**. The helper must not error on it. The glob fallback exists for exactly this case and for local paths, where globbing works fine. - -- [ ] **Step 1: Write the failing tests** - -Create `tests/testthat/test-long-files.R`: - -```r -test_that(".long_files_sql enumerates every partition the manifest lists", { - manifest <- list(files = list(long_partitions = list( - list(year = 2011L, path = "data/long/year=2011/part-0.parquet"), - list(year = 2012L, path = "data/long/year=2012/part-0.parquet") - ))) - expect_equal( - uscogdata:::.long_files_sql("https://example.org/corpus/", manifest), - paste0( - "['https://example.org/corpus/data/long/year=2011/part-0.parquet',", - "'https://example.org/corpus/data/long/year=2012/part-0.parquet']" - ) - ) -}) - -test_that(".long_files_sql falls back to the glob when no partition list is present", { - # test-views.R registers views with a hand-built manifest that has no - # `files` element. That must keep working: the glob is valid for the - # local paths such a manifest is used with. - expect_equal( - uscogdata:::.long_files_sql("/tmp/corpus/", list(schema_version = 4L)), - "'/tmp/corpus/data/long/**/*.parquet'" - ) - expect_equal( - uscogdata:::.long_files_sql("/tmp/corpus/", list(files = list(long_partitions = list()))), - "'/tmp/corpus/data/long/**/*.parquet'" - ) -}) - -test_that("the enumerated list matches the bundled fixture's partition count", { - skip_if_no_corpus() - m <- jsonlite::fromJSON( - file.path(fixture_corpus_path(), "manifest.json"), simplifyVector = FALSE - ) - out <- uscogdata:::.long_files_sql(fixture_corpus_path(), m) - expect_equal( - lengths(regmatches(out, gregexpr("part-0\\.parquet", out))), - length(m$files$long_partitions) - ) -}) - -test_that("registered `long` view reads through the enumerated list", { - skip_if_no_corpus() - with_fixture_corpus({ - con <- uscogdata:::.ensure_session() - n <- DBI::dbGetQuery(con, "SELECT count(*) AS n FROM long")$n - expect_gt(n, 0) - yrs <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year - expect_true(all(c(2011, 2012, 2019, 2020) %in% yrs)) - }) -}) -``` - -- [ ] **Step 2: Run the tests to verify they fail** - -Run: `Rscript -e 'devtools::test(filter = "long-files")'` -Expected: FAIL — `could not find function ".long_files_sql"` on the first three; the fourth may pass already (it exercises the glob against a local fixture, which works). - -- [ ] **Step 3: Add the helper to `R/views.R`** - -Insert immediately above `#' Register DuckDB views from inst/sql/ SQL files`: - -```r -#' Build the SQL path expression for the partitioned `long` table. -#' -#' DuckDB cannot expand a glob over generic HTTP: there is no directory -#' listing to expand against, and `allow_asterisks_in_http_paths` only -#' forwards the literal `**/*` as a filename, which 404s. Measured against -#' the published corpus on 2026-08-08, an explicit file list returns the -#' same 46,148,034 rows the (working) `hf://` glob does, and -#' `hive_partitioning = true` still recovers `year` from the paths. -#' -#' The manifest already enumerates every partition, so we build the list -#' from it. This is host-agnostic -- Nextcloud, HuggingFace and a local -#' fixture take the same path -- where an `hf://` URL would tie the reader -#' to one vendor's protocol and still need special-casing, since manifest -#' fetching goes through httr2, which cannot speak `hf://`. -#' -#' Falls back to the glob when the manifest carries no partition list: a -#' hand-built manifest in a test (see test-views.R) or a corpus predating -#' the field. Both are local, where globbing works. -#' @noRd -.long_files_sql <- function(url, manifest) { - parts <- manifest$files$long_partitions %||% list() - if (length(parts) == 0L) { - return(.sql_lit_chr(paste0(url, "data/long/**/*.parquet"))) - } - paths <- vapply(parts, function(p) as.character(p$path), character(1)) - paste0("[", .sql_lit_chr(paste0(url, paths)), "]") -} -``` - -- [ ] **Step 4: Substitute the new token in `.register_views()`** - -In `R/views.R`, replace this line: - -```r - sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) -``` - -with: - -```r - sql <- gsub("\\{long_files\\}", .long_files_sql(url, manifest), sql, fixed = FALSE) - sql <- gsub("\\{url\\}", url, sql, fixed = FALSE) -``` - -Order matters: `{long_files}` expands to a string containing the url, so it must be substituted first or the `{url}` pass would have nothing to do and the token would survive. - -- [ ] **Step 5: Change the SQL to use the token** - -`inst/sql/10-long.sql` line 3 becomes: - -```sql -FROM read_parquet({long_files}, hive_partitioning = true); -``` - -Note there are **no surrounding quotes** — `.long_files_sql()` returns its own quoting, whether a bracketed list or a single quoted glob. - -- [ ] **Step 6: Run the new tests** - -Run: `Rscript -e 'devtools::test(filter = "long-files")'` -Expected: PASS, all four. - -- [ ] **Step 7: Run the full suite for regressions** - -Run: `Rscript -e 'devtools::test()'` -Expected: PASS. Pay particular attention to `test-views.R` — it is the file that exercises `.register_views()` with a `files`-less manifest. - -- [ ] **Step 8: Commit** - -```bash -git add R/views.R inst/sql/10-long.sql tests/testthat/test-long-files.R -git commit -m "fix: enumerate long partitions from the manifest instead of globbing - -DuckDB cannot expand a glob over generic HTTP -- no directory listing -- -so every remote corpus read failed. Only local paths worked, which is how -the API and the test fixture run, so nothing caught it. - -Measured against the published corpus: the explicit list returns the same -46,148,034 rows, with hive_partitioning still recovering year." -``` - ---- - -## Task 2: Ship a working default corpus URL - -The P0 defect, part two. The default is the literal `REPLACE_WITH_SHARE_TOKEN` sentinel and no file in the repo supplies a real URL, so a new user has no path to a working session. - -**Files:** -- Modify: `R/config.R:7` -- Modify: `R/manifest.R` (the `i =` guidance line in `.check_url_configured()`) -- Modify: `tests/testthat/test-manifest.R:11`, `:38` -- Test: `tests/testthat/test-config.R` (add cases) - -**Interfaces:** -- Consumes: `.cfg("url")`, `.resolve_url()` from `R/config.R`. -- Produces: a `.uscogdata_defaults$url` that is a real, reachable, credential-free URL. Task 3's integration test depends on this being the default. - -**Critical constraint:** `tests/testthat/test-manifest.R` hardcodes the placeholder at lines 11 and 38 and asserts `cog_open()` aborts with `uscogdata_url_not_configured`. Changing the default **breaks those two tests** and they must be updated in this task. The test at line 26 passes an explicit sentinel-bearing URL via env var — it keeps passing untouched, and it is what proves the sentinel guard still works. - -- [ ] **Step 1: Write the failing tests** - -Append to `tests/testthat/test-config.R`: - -```r -test_that("the default corpus URL is real, not a placeholder", { - withr::with_envvar(c(USCOGDATA_URL = NA), { - withr::with_options(list(uscogdata.url = NULL), { - url <- uscogdata:::.resolve_url() - expect_false(grepl("REPLACE_WITH", url, fixed = TRUE)) - expect_match(url, "^https://", perl = TRUE) - expect_match(url, "/$", perl = TRUE) - }) - }) -}) - -test_that("an explicitly-set sentinel URL still aborts", { - # The guard must survive the default change: a user who half-edited a - # copied config still gets the actionable error. - withr::with_envvar( - c(USCOGDATA_URL = "https://other.example/s/REPLACE_WITH_SHARE_TOKEN/x/"), { - expect_error( - uscogdata:::.check_url_configured(uscogdata:::.resolve_url()), - class = "uscogdata_url_not_configured" - ) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: FAIL on the first test — the resolved default still contains `REPLACE_WITH`. - -- [ ] **Step 3: Change the default** - -`R/config.R`, in `.uscogdata_defaults`: - -```r -.uscogdata_defaults <- list( - # Public HuggingFace mirror of the published corpus: CC-BY-4.0, no - # credential, CDN-backed. Chosen as the default so `library(uscogdata)` - # followed by a verb works with zero configuration. Trailing slash is - # required -- every consumer concatenates onto this (see .resolve_url()). - url = "https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/", - cache_dir = NULL, - manifest_ttl_secs = 3600L -) -``` - -- [ ] **Step 4: Update the stale guidance line** - -In `R/manifest.R`, inside `.check_url_configured()`, replace: - -```r - i = "For the live Civilytics corpus, request the Nextcloud share URL from the package maintainer." -``` - -with: - -```r - i = "The public corpus is the default; unset USCOGDATA_URL to use it, or point it at a local copy from {.code cog_mirror()}." -``` - -- [ ] **Step 5: Correct the two now-inaccurate test names** - -Both affected tests in `tests/testthat/test-manifest.R` set `USCOGDATA_URL` **explicitly** via `withr::with_envvar` before calling `cog_open()`, so they keep passing unchanged. Only their wording becomes wrong — the sentinel URL is no longer the default. - -Line 7, rename the test: - -```r -test_that("cog_open aborts with actionable error when URL contains the sentinel", { -``` - -Lines 11 and 38, rename the variable and say why it is still here: - -```r - # No longer the package default (that is the public HF corpus). This is a - # user who copied a config template and did not finish editing it. - sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/" -``` - -Update the two `withr::with_envvar(c(USCOGDATA_URL = placeholder), ...)` call sites in those tests to use `sentinel_url`. No functional change — do not alter the assertions. - -- [ ] **Step 6: Run the affected files** - -Run: `Rscript -e 'devtools::test(filter = "config|manifest")'` -Expected: PASS. - -- [ ] **Step 7: Run the full suite** - -Run: `Rscript -e 'devtools::test()'` -Expected: PASS. `setup.R` points `USCOGDATA_URL` at the bundled fixture for the whole suite, so the default change should not affect any other file. - -- [ ] **Step 8: Commit** - -```bash -git add R/config.R R/manifest.R tests/testthat/test-config.R tests/testthat/test-manifest.R -git commit -m "feat: default to the public corpus so the package works unconfigured - -The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no file in the -repo supplied a working URL, so a new user had no path to a session. The -sentinel guard stays for half-edited configs." -``` - ---- - -## Task 3: Live-corpus integration test - -Nothing in the suite exercises a remote corpus — that is why the P0 survived. This test is network-gated so it skips in offline CI but runs on demand and before release. - -**Files:** -- Test: `tests/testthat/test-live-corpus.R` (create) - -**Interfaces:** -- Consumes: `.long_files_sql()` (Task 1), the default URL (Task 2), and the public verbs `cog_gov_search()`, `cog_spending()`, `cog_explain()`. -- Produces: nothing consumed downstream. - -- [ ] **Step 1: Write the test** - -Create `tests/testthat/test-live-corpus.R`: - -```r -# Network-gated. Set USCOGDATA_LIVE_TEST=true to run. -# -# This file exists because the P0 fixed in this release -- no remote corpus -# was readable at all -- survived precisely because every other test path -# used a LOCAL corpus (the bundled fixture) and so did the API in -# production. Nothing ever exercised the code the way a new user does. -skip_live <- function() { - testthat::skip_if_not( - identical(tolower(Sys.getenv("USCOGDATA_LIVE_TEST", "")), "true"), - "live-corpus test: set USCOGDATA_LIVE_TEST=true to run" - ) -} - -test_that("the package reads the public corpus with no configuration at all", { - skip_live() - withr::with_envvar(c(USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA), { - withr::with_options(list(uscogdata.url = NULL), { - uscogdata:::cog_close() - on.exit(uscogdata:::cog_close(), add = TRUE) - - g <- cog_gov_search(name = "Madison", state = "WI", type = 2) - expect_gt(nrow(g), 0) - - s <- cog_spending(g$canonical_govid[1], years = 2022) - expect_gt(nrow(s), 0) - expect_true(all(c("amt_nominal", "year", "category") %in% names(s))) - - # Amounts are full dollars, already x1000. A city's total annual - # spending is millions, not thousands -- this catches a regression - # that reintroduced the double conversion. - expect_gt(sum(s$amt_nominal, na.rm = TRUE), 1e6) - - p <- attr(s, "provenance") - expect_true(isTRUE(p$transformations$units_conversion$applied)) - expect_equal(p$transformations$units_conversion$multiplier, 1000) - }) - }) -}) - -test_that("a full-history query spans the published range", { - skip_live() - withr::with_envvar(c(USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA), { - withr::with_options(list(uscogdata.url = NULL), { - uscogdata:::cog_close() - on.exit(uscogdata:::cog_close(), add = TRUE) - - g <- cog_gov_search(name = "Madison", state = "WI", type = 2) - s <- cog_spending(g$canonical_govid[1]) - # The corpus publishes FY1967-FY2024. Any single government's span is - # narrower, but a full-history query must cross more than one decade - # -- if enumeration silently returned one partition, this fails. - expect_gt(diff(range(s$year)), 10) - }) - }) -}) -``` - -- [ ] **Step 2: Run it gated off (default) — it must skip, not fail** - -Run: `Rscript -e 'devtools::test(filter = "live-corpus")'` -Expected: SKIP on both, with the message about `USCOGDATA_LIVE_TEST`. - -- [ ] **Step 3: Run it live** - -Run: `USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test(filter = "live-corpus")'` -Expected: PASS. This is the first time the package has ever read a remote corpus successfully. If it fails, Task 1 is incomplete — do not proceed. - -- [ ] **Step 4: Record the measured verb latency** - -Run and note the wall time, which the README needs (raw scans were 1.5 s / 2.8 s; real verbs do more work): - -```bash -USCOGDATA_LIVE_TEST=true Rscript -e ' - Sys.unsetenv("USCOGDATA_URL") - library(uscogdata) - g <- cog_gov_search(name = "Madison", state = "WI", type = 2) - print(system.time(cog_spending(g$canonical_govid[1], years = 2022))) - print(system.time(cog_spending(g$canonical_govid[1]))) -' -``` - -Carry these two numbers into Task 8. Do not reuse the raw-scan figures. - -- [ ] **Step 5: Commit** - -```bash -git add tests/testthat/test-live-corpus.R -git commit -m "test: exercise the public corpus end to end, unconfigured - -The remote-read defect survived because every test path used a local -corpus. This is the only test that runs the package the way a new user -does." -``` - ---- - -## Task 4: DESCRIPTION metadata - -**Files:** -- Modify: `DESCRIPTION` - -**Interfaces:** -- Produces: `URL`/`BugReports` that Task 7 (`_pkgdown.yml`) and Task 8 (README) both reference; the `Authors@R` that `citation("uscogdata")` renders. - -- [ ] **Step 1: Write the failing test** - -Append to `tests/testthat/test-config.R`: - -```r -test_that("DESCRIPTION carries release metadata", { - skip_if_no_source_tree("DESCRIPTION") - d <- read.dcf(source_tree_path("DESCRIPTION")) - fields <- colnames(d) - - expect_true(all(c("URL", "BugReports") %in% fields)) - expect_match(d[1, "Authors@R"], "Knowles", fixed = TRUE) - expect_match(d[1, "Authors@R"], "0000-0003-0005-9478", fixed = TRUE) - expect_match(d[1, "Authors@R"], "Civilytics Consulting LLC", fixed = TRUE) - - # The gate in .validate_schema() accepts up to 7 and the published corpus - # IS 7; DESCRIPTION must not claim otherwise. - expect_equal(as.integer(d[1, "MaxCorpusSchema"]), 7L) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: FAIL — `URL`/`BugReports` absent, `MaxCorpusSchema` is 5. - -- [ ] **Step 3: Edit DESCRIPTION** - -Replace the `Authors@R` block: - -``` -Authors@R: c( - person(c("Jared", "E."), "Knowles", - email = "jared@civilytics.com", - role = c("aut", "cre"), - comment = c(ORCID = "0000-0003-0005-9478")), - person("Civilytics Consulting LLC", role = c("cph", "fnd"))) -``` - -The given-name vector `c("Jared", "E.")` with family name `"Knowles"` matches -`merTools` exactly. That structural match is what lets ORCID and r-universe -collate both packages as one person's work — `person("Jared", "E. Knowles")` -would render identically but put the middle initial in the family-name slot. - -Add after `Description:`: - -``` -URL: https://github.com/civilytics/uscogdata, https://civilytics.r-universe.dev/uscogdata -BugReports: https://github.com/civilytics/uscogdata/issues -``` - -Change: - -``` -MaxCorpusSchema: 7 -``` - -Bump the version — this release changes user-visible behaviour (remote reads -go from broken to working; the default URL from placeholder to live corpus), -which is a minor bump, not a patch: - -``` -Version: 0.3.0 -``` - -- [ ] **Step 4: Verify the person object parses** - -Run: `Rscript -e 'print(eval(parse(text = read.dcf("DESCRIPTION")[1, "Authors@R"])))'` -Expected: prints two entries — `Jared E. Knowles [aut, cre] (ORCID: ...)` and `Civilytics Consulting LLC [cph, fnd]`. A parse error here means a malformed `person()` call. - -- [ ] **Step 5: Run the tests** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: PASS. - -- [ ] **Step 6: Commit** - -```bash -git add DESCRIPTION tests/testthat/test-config.R -git commit -m "chore: release metadata -- author of record, URLs, schema ceiling - -Authors@R was an org with no human, so citation() and the r-universe -maintainer page had nothing to render. MaxCorpusSchema claimed 5 while the -code accepts 7 and the published corpus is 7." -``` - ---- - -## Task 5: License files - -**Files:** -- Modify: `LICENSE` -- Create: `LICENSE.md` - -- [ ] **Step 1: Generate both files** - -Run: `Rscript -e 'usethis::use_mit_license("Civilytics Consulting LLC")'` - -This rewrites `LICENSE` to the two-line stub with the corrected holder and creates `LICENSE.md` with the full MIT text. It may also add `^LICENSE\.md$` to `.Rbuildignore` — that is correct and already present. - -- [ ] **Step 2: Verify** - -Run: `Rscript -e 'cat(readLines("LICENSE"), sep = "\n")'` -Expected: -``` -YEAR: 2026 -COPYRIGHT HOLDER: Civilytics Consulting LLC -``` - -Run: `Rscript -e 'cat(length(readLines("LICENSE.md")), "lines\n")'` -Expected: a non-zero count (the full MIT text, ~21 lines). - -- [ ] **Step 3: Confirm DESCRIPTION still declares the license correctly** - -Run: `Rscript -e 'cat(read.dcf("DESCRIPTION")[1, "License"], "\n")'` -Expected: `MIT + file LICENSE`. If `usethis` changed it, that is fine — leave whatever it wrote. - -- [ ] **Step 4: Commit** - -```bash -git add LICENSE LICENSE.md DESCRIPTION .Rbuildignore -git commit -m "chore: add full MIT text, name the copyright holder properly - -LICENSE held only the two-line stub and no LICENSE.md existed, so the -repo carried no license text for a human or for GitHub's detector." -``` - ---- - -## Task 6: Ship the vignettes - -`.Rbuildignore` excludes `^vignettes$`, so an installed package has no vignettes at all — while README tells users to run `vignette("total-spending", package = "uscogdata")`. Both vignettes build offline: `total-spending.Rmd` points `USCOGDATA_URL` at the bundled fixture, `population-denominators.Rmd` is `eval = FALSE`. - -**Files:** -- Modify: `.Rbuildignore` -- Test: `tests/testthat/test-config.R` (add) - -- [ ] **Step 1: Write the failing test** - -Append to `tests/testthat/test-config.R`: - -```r -test_that("vignettes are not excluded from the build", { - skip_if_no_source_tree(".Rbuildignore") - ignore <- readLines(source_tree_path(".Rbuildignore"), warn = FALSE) - expect_false(any(grepl("^\\^vignettes\\$$", ignore))) - # The fixture is what lets R CMD check run offline with no credentials on - # r-universe and GitHub Actions. It must never be excluded. - expect_false(any(grepl("fixture_corpus", ignore, fixed = TRUE))) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: FAIL on the first expectation. - -- [ ] **Step 3: Remove the three lines** - -Delete these lines from `.Rbuildignore`: - -``` -^vignettes$ -^doc$ -^Meta$ -``` - -Leave every other line untouched — in particular `^_pkgdown\.yml$`, `^docs$`, `^data-raw$`, `^specs$`, `^plans$`, `^\.gitea$`, `^CLAUDE\.md$` are all correct exclusions. - -- [ ] **Step 4: Build the tarball and confirm the vignettes are in it** - -```bash -Rscript -e 'devtools::build(path = tempdir())' -``` -Then list the tarball contents: -```bash -tar -tzf "$(ls -t $(Rscript -e 'cat(tempdir())')/uscogdata_*.tar.gz | head -1)" | grep -E 'vignettes|inst/doc' -``` -Expected: both `.Rmd` files appear under `uscogdata/vignettes/`. - -- [ ] **Step 5: Run the tests** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: PASS. - -- [ ] **Step 6: Commit** - -```bash -git add .Rbuildignore tests/testthat/test-config.R -git commit -m "fix: ship the vignettes - -.Rbuildignore excluded ^vignettes$, so vignette(\"total-spending\") failed -for every user -- while the README instructed them to run it. Both build -offline against the bundled fixture." -``` - ---- - -## Task 7: pkgdown reference index - -`_pkgdown.yml` lists 6 of 14 exports. pkgdown errors on topics missing from the index, so the docs site does not build. - -**Files:** -- Modify: `_pkgdown.yml` - -- [ ] **Step 1: Write the failing test** - -Append to `tests/testthat/test-config.R`: - -```r -test_that("_pkgdown.yml indexes every exported topic", { - skip_if_no_source_tree("_pkgdown.yml", "NAMESPACE") - exports <- grep("^export\\(", readLines(source_tree_path("NAMESPACE"), warn = FALSE), value = TRUE) - exports <- sub("^export\\((.*)\\)$", "\\1", exports) - yml <- paste(readLines(source_tree_path("_pkgdown.yml"), warn = FALSE), collapse = "\n") - missing <- exports[!vapply(exports, function(e) grepl(paste0("\\b", e, "\\b"), yml), logical(1))] - expect_equal(missing, character(0)) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: FAIL listing the eight missing: `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`. - -- [ ] **Step 3: Rewrite `_pkgdown.yml`** - -```yaml -url: https://civilytics.r-universe.dev/uscogdata - -template: - bootstrap: 5 - -reference: - - title: Financial data - desc: Spending, revenue and balance-sheet holdings for one or more governments. - contents: - - cog_spending - - cog_revenue - - cog_balances - - title: Search & basket - desc: Resolve place names into canonical govids. - contents: - - cog_gov_search - - cog_basket_resolution - - cog_basket_unresolved - - title: Comparison & aggregation - desc: Peer cohorts and geographic rollups. - contents: - - cog_find_peers - - cog_peer_compare - - cog_geographic_rollup - - title: Corpus metadata - desc: What the corpus contains, where it came from, and how to hold a local copy. - contents: - - cog_categories - - cog_recipes - - cog_manifest - - cog_explain - - cog_mirror - -articles: - - title: Concepts - navbar: ~ - contents: - - total-spending - - population-denominators -``` - -- [ ] **Step 4: Build the site** - -Run: `Rscript -e 'pkgdown::build_site(preview = FALSE)'` -Expected: completes without error. Any "Topics missing from index" warning means an export was missed. - -- [ ] **Step 5: Run the tests** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: PASS. - -- [ ] **Step 6: Commit** - -```bash -git add _pkgdown.yml tests/testthat/test-config.R -git commit -m "docs: index all 14 exports in pkgdown, set the site url - -The reference index covered 6 of 14, so pkgdown errored on the missing -topics and the docs site did not build." -``` - -`docs/` is gitignored via `.Rbuildignore`/`.gitignore`; do not commit built output. - ---- - -## Task 8: README rewrite - -The current README addresses someone inside the repo tree: status reads "Under active development (Phase 2 of the cog_pipeline project)", it points at `../cog_pipeline/docs/reader-specification.md`, the install line is commented out, and developer/testing/release sections sit above anything a user needs. - -**Files:** -- Rewrite: `README.md` - -**Interfaces:** -- Consumes: `URL`/`BugReports` from Task 4, the default URL from Task 2, the measured latencies from Task 3 Step 4. - -- [ ] **Step 1: Write the failing test** - -Append to `tests/testthat/test-config.R`: - -```r -test_that("README is written for a stranger, not a repo insider", { - skip_if_no_source_tree("README.md") - r <- paste(readLines(source_tree_path("README.md"), warn = FALSE), collapse = "\n") - - # No paths that only resolve inside Jared's checkout. - expect_false(grepl("../cog_pipeline", r, fixed = TRUE)) - # A real, uncommented install line. - expect_match(r, "install.packages", fixed = TRUE) - expect_false(grepl("# pak::pkg_install", r, fixed = TRUE)) - # The errata most likely to produce a plausible-looking wrong answer. - expect_match(r, "full US dollars", fixed = TRUE) - # The release-instructions section that conflicts with public CI is gone. - expect_false(grepl("Rbuildignore", r, fixed = TRUE)) - # Both read paths documented. - expect_match(r, "cog_mirror", fixed = TRUE) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: FAIL. - -- [ ] **Step 3: Rewrite README.md in this order** - -Write these sections, in this sequence. Content requirements are exact; prose is yours. - -1. **Title + one-paragraph what-it-is.** Curated R reader over the Civilytics US Census of Governments finance corpus. State the coverage: government types 0–3 (state, county, municipality, township), **FY1967–FY2024** (no source data FY1968, FY1969), **46,148,034 rows**, **190.6 MB**. Add the r-universe version badge. - -2. **Install.** - ````markdown - ```r - install.packages("uscogdata", - repos = c("https://civilytics.r-universe.dev", - "https://cloud.r-project.org")) - ``` - Or from source: - ```r - pak::pkg_install("git::https://gitea.civilytics.org/Civilytics/uscogdata.git") - ``` - ```` - -3. **Quickstart — no configuration step.** Verbatim: - ````markdown - ```r - library(uscogdata) - - # Resolve a place name to a canonical government id - madison <- cog_gov_search(name = "Madison", state = "WI", type = 2) - - # Full spending history, real dollars, per capita - spend <- cog_spending(madison$canonical_govid[1]) - - # Every result carries its own provenance - cog_explain(spend) - ``` - ```` - -4. **Two ways to read the corpus.** Remote (default, zero setup) vs `cog_mirror()` (190 MB once). Include the measured table — **use the verb latencies recorded in Task 3 Step 4, not the raw-scan figures**: - - | | Remote (default) | Mirrored | - |---|---|---| - | Setup | none | `cog_mirror()`, 190.6 MB once | - | Disk | 0 MB — HTTP range requests | 190.6 MB | - | Per-query | network round trip | local | - | Right for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility | - - State plainly: the default reads from a public HuggingFace mirror, and **the escape hatch is one function call** — after `cog_mirror()`, no analysis touches an external service. - -5. **Amounts are in full US dollars.** Keep the existing text nearly verbatim — it is correct and carefully argued. Keep the `attr(r, "provenance")$transformations$units_conversion` example and the "do not multiply again" warning. - -6. **Concepts.** Condense the existing primary/direct/total and general/total revenue sections to ~1/3 their length, each ending with a pointer to `vignette("total-spending")`. Add a short **Reporting coverage** paragraph: the Census is a complete census only in years ending in 2 and 7; every other year is a sample; `provenance$coverage` reports `n_units_reporting` per year, and `coverage = "census"` / `"consistent"` control the mode. - -7. **How to cite.** `citation("uscogdata")` for the package; corpus is CC-BY-4.0, cite *Civilytics Consulting, US Census of Governments finance corpus*. - -8. **Contributing.** Two sentences plus a link to `CONTRIBUTING.md` (Task 10). - -**Delete outright:** the "Status" block, the `../cog_pipeline/...` reference, the entire "Developer notes / Testing / Releasing against the live corpus" section (it moves to `CONTRIBUTING.md`, minus the fixture-stripping advice, which is wrong and must not be carried over). - -- [ ] **Step 4: Run the quickstart verbatim in a clean session** - -```bash -Rscript -e ' - Sys.unsetenv("USCOGDATA_URL") - devtools::load_all(".", quiet = TRUE) - madison <- cog_gov_search(name = "Madison", state = "WI", type = 2) - spend <- cog_spending(madison$canonical_govid[1]) - cat("rows:", nrow(spend), "years:", paste(range(spend$year), collapse = "-"), "\n") - cog_explain(spend) -' -``` -Expected: runs clean with no configuration. If it errors, the README is wrong — fix the README, not the test. - -- [ ] **Step 5: Run the tests** - -Run: `Rscript -e 'devtools::test(filter = "config")'` -Expected: PASS. - -- [ ] **Step 6: Commit** - -```bash -git add README.md tests/testthat/test-config.R -git commit -m "docs: rewrite README for a stranger - -Reordered around a new user -- what it is, install, a quickstart that runs -with no configuration, then the dollars warning and the concepts. Drops -the sibling-repo path, the commented-out install line, and the -fixture-stripping release advice that conflicts with public CI." -``` - ---- - -## Task 9: NEWS.md rewrite - -The current NEWS is a pre-release churn log spanning the package's entire development (2026-04-23 to 2026-08-04, 140 commits), describing changes relative to states no user has seen. To a newcomer it reads as instability. - -**Files:** -- Rewrite: `NEWS.md` -- Modify: `README.md` and roxygen blocks receiving migrated content - -- [ ] **Step 1: Migrate the load-bearing content first** - -Before deleting anything, move each of these to its documentation home. Verify each lands before proceeding: - -| Content in current NEWS | Destination | -|---|---| -| Coverage disclosure — census vs sample years, the 597-vs-112 Wisconsin example, `coverage` modes | README §6 (Task 8) **and** `@details` in `R/rollup.R`'s roxygen for `cog_geographic_rollup()` | -| `complete = TRUE` — the `reported` / `census_zero` / `not_reported` table | `@details` in `R/spending.R` and `R/revenue.R` roxygen | -| Series breaks and `corpus_break_refs` | `@details` in `R/provenance.R`'s `cog_explain()` roxygen | -| Per-year F-33 population denominators | already in `vignette("population-denominators")` — verify, do not duplicate | - -Run `Rscript -e 'devtools::document()'` after editing roxygen. - -- [ ] **Step 2: Restructure NEWS.md** - -Three edits, in this order: - -1. **Prepend** the `0.3.0` section below. -2. **Keep** the existing `# uscogdata 0.2.0` section verbatim — it is a real - changelog (`"All Categories"`, the coverage-signposting fix, the - `n_units_reporting` documentation) and users deserve it. -3. **Delete** the entire `# uscogdata 0.1.0 (development)` section and - everything under it. That is pre-release churn; it stays in git. - -The new top section: - -```markdown -# uscogdata 0.3.0 - -First public release. - -`uscogdata` provides curated R verbs over the Civilytics US Census of -Governments finance corpus: unit-level financial profiles, geographic -rollups, and peer comparisons, with auditable provenance on every result. - -## What it covers - -Government types 0–3 (state, county, municipality, township), FY1967–FY2024 -(no source data for FY1968 or FY1969) — 46,148,034 rows across 56 fiscal -years. The corpus is published under CC-BY-4.0 and reads directly over -HTTPS, or locally after `cog_mirror()`. - -## The verbs - -`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and -holdings; `cog_gov_search()` to resolve place names (including basket mode -for many at once); `cog_find_peers()` and `cog_peer_compare()` for cohorts; -`cog_geographic_rollup()` for aggregates; `cog_categories()`, -`cog_recipes()`, `cog_manifest()` and `cog_explain()` for metadata and -provenance; `cog_mirror()` for a local copy. - -## Four things to know before your first query - -* **Amounts are in full US dollars.** The raw Census files report thousands; - the verbs multiply by 1000 on the way out. Do not multiply again. -* **Multi-government aggregates disclose their coverage.** The Census is a - complete census only in years ending in 2 and 7. Every result carries - `provenance$coverage` with per-year `n_units_reporting`. -* **Absence means two different things.** Before FY2012 an absent cell means - Census published $0; from FY2012 it means not reported. - `complete = TRUE` labels which. -* **Series breaks reach you unasked.** Catalogued breaks intersecting your - query appear in provenance and in `cog_explain()`. - -## Known limits - -* Special districts (type 4) and school districts (type 5) are out of scope. -* Per-capita rollups exclude governments with no F-33 population. -* Employee-retirement (`X`) codes stop at FY2016, when those systems moved - to the Annual Survey of Public Pensions. -``` - -- [ ] **Step 3: Verify no orphaned content** - -Run: `git show HEAD:NEWS.md > /tmp/news-old.md && wc -l /tmp/news-old.md NEWS.md` - -Read `/tmp/news-old.md` once more and confirm every substantive claim from the **deleted `0.1.0 (development)` section** either appears in the new `0.3.0` section, landed somewhere in Step 1, or is genuinely pre-release churn (version bumps, fixture regenerations, internal refactors). - -Then confirm the `0.2.0` section survived intact: - -```bash -diff <(git show HEAD:NEWS.md | sed -n '/^# uscogdata 0.2.0/,/^# uscogdata 0.1.0/p' | head -n -1) \ - <(sed -n '/^# uscogdata 0.2.0/,$p' NEWS.md) -``` -Expected: no output. Any diff means the `0.2.0` changelog was damaged — restore it. - -- [ ] **Step 4: Run the full suite** - -Run: `Rscript -e 'devtools::test()'` -Expected: PASS — `devtools::document()` in Step 1 regenerated `man/`, so this catches a malformed roxygen block. - -- [ ] **Step 5: Commit** - -```bash -git add NEWS.md README.md R/ man/ -git commit -m "docs: recast NEWS around the first public release - -The changelog described changes relative to states no user ever saw, which -reads as instability to someone deciding whether to depend on this. The -load-bearing caveats move into README and roxygen, where they belong; the -pre-release history stays in git." -``` - ---- - -## Task 10: CONTRIBUTING.md - -**Files:** -- Create: `CONTRIBUTING.md` -- Modify: `.Rbuildignore` - -- [ ] **Step 1: Write CONTRIBUTING.md** - -It must contain, in this order: - -1. **Canonical source note.** Development happens on `gitea.civilytics.org/Civilytics/uscogdata`; `github.com/civilytics/uscogdata` is a mirror that accepts issues and pull requests. - -2. **What happens to a GitHub PR.** Verbatim explanation: it is fetched and landed on the canonical repo, then closes itself as merged when the mirror syncs — because the maintainer merges with `--no-ff`, preserving the contributor's commits and SHAs. Say plainly that a PR closing without a "Merged by" click is normal and not a rejection. - -3. **Running the tests.** - ````markdown - ```r - devtools::test() # uses the bundled fixture; no network, no credentials - ``` - ```` - Note that `tests/testthat/setup.R` points `USCOGDATA_URL` at - `inst/extdata/fixture_corpus/` automatically. - -4. **Testing against the live corpus.** - ````markdown - ```r - USCOGDATA_LIVE_TEST=true devtools::test(filter = "live-corpus") - ``` - ```` - Explain why it exists: every other test path uses a local corpus, which is - how the remote-read defect fixed for 0.3.0 went unnoticed. - -5. **Do not exclude the fixture from the build.** State the reason — it is what lets `R CMD check` pass on r-universe and GitHub Actions with no credentials. - -6. **Release checklist**, moved from README: run the suite against both the fixture and the live corpus, `pkgdown::build_site()`, `R CMD check --as-cran`, tag, then update the r-universe registry pin. - -**Do not carry over** the README's instruction to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`. It is wrong. - -- [ ] **Step 2: Exclude it from the build** - -Add to `.Rbuildignore`: - -``` -^CONTRIBUTING\.md$ -``` - -- [ ] **Step 3: Verify the build is clean** - -Run: `Rscript -e 'devtools::check(document = FALSE, args = "--no-manual")'` -Expected: 0 errors, 0 warnings. Notes about package size (the 15 MB fixture) are expected and acceptable — this package is not going to CRAN. - -- [ ] **Step 4: Commit** - -```bash -git add CONTRIBUTING.md .Rbuildignore -git commit -m "docs: add CONTRIBUTING with the canonical-on-Gitea PR flow - -Moves developer and release instructions out of the README, minus the -fixture-stripping advice, which would break the vignette and leave public -CI unable to check without credentials." -``` - ---- - -## Final verification - -Run before declaring the release ready. Every one of these must pass. - -- [ ] **1. Full suite, offline, no credentials** - `Rscript -e 'devtools::test()'` — the property public CI depends on. - -- [ ] **2. Live corpus** - `USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test()'` - -- [ ] **3. `R CMD check --as-cran`** - `Rscript -e 'devtools::check(args = "--as-cran")'` — 0 errors, 0 warnings. - -- [ ] **4. pkgdown** - `Rscript -e 'pkgdown::build_site(preview = FALSE)'` - -- [ ] **5. Vignettes resolve from an installed copy** - ```bash - Rscript -e 'devtools::install(build_vignettes = TRUE, quiet = TRUE)' - Rscript -e 'v <- vignette("total-spending", package = "uscogdata"); stopifnot(nzchar(v$File)); cat("OK\n")' - ``` - -- [ ] **6. Cold-start check.** On a machine (or container) that has never had this package: install it, then run the README quickstart **verbatim with no environment variables set**. - ```bash - docker run --rm -v "$PWD":/pkg rocker/r-ver:4.4 bash -c ' - apt-get update -qq && apt-get install -y -qq libcurl4-openssl-dev libssl-dev >/dev/null - Rscript -e "install.packages(c(\"pak\"), repos=\"https://cloud.r-project.org\")" \ - -e "pak::pkg_install(\"local::/pkg\")" \ - -e "library(uscogdata); m <- cog_gov_search(name=\"Madison\", state=\"WI\", type=2); s <- cog_spending(m\$canonical_govid[1]); cat(\"rows:\", nrow(s), \"\n\")" - ' - ``` - **This is the only check that catches the P0 class of fault**, and its absence is why the fault survived. Do not skip it. - -- [ ] **7. Verb latency re-measured** against the live corpus, and the README table updated if the numbers moved from what Task 3 Step 4 recorded. - -## Out of scope for this plan - -Flipping the Gitea repo public, `gitleaks`, the GitHub mirror and its Actions matrix, the Gitea push workflow, the r-universe registry, and tagging `v0.3.0`. Those follow after this plan's final verification is green — r-universe publishes check results on registration, so registering before checks pass means a red badge on day one. Corrections intake and announcement posts are deferred by decision (see the spec).