Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest enumeration, working default URL, and the live-corpus test whose absence let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite README, NEWS and CONTRIBUTING. Distribution mechanics stay out of scope -- r-universe publishes check results on registration, so it comes after final verification is green.
1065 lines
41 KiB
Markdown
1065 lines
41 KiB
Markdown
# uscogdata 0.1.0 Public Release Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
|
||
**Goal:** Make `uscogdata` installable and usable by a stranger — fix the corpus-unreachable defect, correct release metadata, and rewrite README and NEWS for a first public release.
|
||
|
||
**Architecture:** Two code changes fix the P0 (manifest-driven file enumeration replacing an HTTP-incompatible glob; a working public default URL). Everything else is metadata, packaging config, and documentation. Distribution mechanics (Gitea public, GitHub mirror, r-universe) are **out of scope for this plan** — they follow after checks are green.
|
||
|
||
**Tech Stack:** R (>= 4.1), DuckDB 1.5.5 via `duckdb`/`DBI`, `httr2`, `jsonlite`, `cli`, testthat 3e, pkgdown, roxygen2 7.3.3.
|
||
|
||
**Design spec:** `specs/2026-08-08-public-release-design.md`
|
||
|
||
## Global Constraints
|
||
|
||
- Package license is **MIT**. Copyright holder is **Civilytics Consulting LLC**.
|
||
- Author of record: **Jared E. Knowles**, `jared@civilytics.com`, ORCID **0000-0003-0005-9478**, roles `aut`/`cre`. Civilytics Consulting LLC is `cph`/`fnd`.
|
||
- Public corpus base URL (note the trailing slash, which `.resolve_url()` requires):
|
||
`https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/`
|
||
- Published corpus is **schema_version 7**; bundled fixture is **schema_version 6**. `.validate_schema()` accepts `c(4L, 5L, 6L, 7L)` and that list does not change in this plan.
|
||
- Corpus facts for documentation, measured 2026-08-08: **46,148,034 rows**, **56 partitions**, **190.6 MB**, government types **0–3**, **FY1967–FY2024** (no source data FY1968, FY1969).
|
||
- The bundled fixture at `inst/extdata/fixture_corpus/` **ships in the release**. Never add it to `.Rbuildignore`.
|
||
- Every amount column a verb returns is in **full US dollars**, already multiplied by 1000. Documentation must never tell a user to apply the $1,000s rule to verb output.
|
||
- Run the full suite with `devtools::test()` from the package root. It uses the bundled fixture and requires no network or credentials.
|
||
|
||
---
|
||
|
||
## File Structure
|
||
|
||
**Modified — code**
|
||
- `R/views.R` — gains `.long_files_sql()`; `.register_views()` substitutes a second token. This is the only file that knows how the `long` table's paths are built.
|
||
- `inst/sql/10-long.sql` — the one file containing a glob. Becomes token-driven.
|
||
- `R/config.R` — default corpus URL.
|
||
- `R/manifest.R` — `.check_url_configured()` guidance text only; the sentinel check itself is unchanged.
|
||
|
||
**Modified — metadata/packaging**
|
||
- `DESCRIPTION`, `LICENSE`, `.Rbuildignore`, `_pkgdown.yml`
|
||
|
||
**Created**
|
||
- `LICENSE.md`, `CONTRIBUTING.md`
|
||
- `tests/testthat/test-long-files.R` — unit tests for enumeration
|
||
- `tests/testthat/test-live-corpus.R` — network-gated integration test
|
||
|
||
**Rewritten**
|
||
- `README.md`, `NEWS.md`
|
||
|
||
**Deliberately untouched:** every verb file (`R/spending.R`, `R/revenue.R`, `R/balances.R`, `R/rollup.R`, `R/peers.R`, `R/search.R`), `R/mirror.R`, and all other `inst/sql/*.sql`. The enumeration fix is confined to view registration by design.
|
||
|
||
---
|
||
|
||
## Task 1: Manifest-driven partition enumeration
|
||
|
||
The P0 defect, part one. `inst/sql/10-long.sql` globs `{url}data/long/**/*.parquet`. DuckDB 1.5.5 refuses globs over generic HTTP, and `allow_asterisks_in_http_paths` does not help — it forwards the literal `**/*` as a filename and 404s, because HTTP exposes no directory listing. The manifest already enumerates every partition under `files$long_partitions[]`.
|
||
|
||
**Files:**
|
||
- Modify: `R/views.R` (add helper; `.register_views()` at the `gsub` line)
|
||
- Modify: `inst/sql/10-long.sql:3`
|
||
- Test: `tests/testthat/test-long-files.R` (create)
|
||
|
||
**Interfaces:**
|
||
- Consumes: `.sql_lit_chr(x)` from `R/spending.R:553` — quotes each element, escapes `'` by doubling, joins with `,` and **no space**. `%||%` from `R/manifest.R:158`.
|
||
- Produces: `.long_files_sql(url, manifest)` returning a single SQL string — either a bracketed list literal `['a','b']` or, on fallback, a single quoted glob `'…/**/*.parquet'`. Task 3 relies on this being the only place partition paths are constructed.
|
||
|
||
**Critical constraint:** `tests/testthat/test-views.R:322` calls `.register_views(con, url, manifest = list(schema_version = 4L))` — a manifest with **no `files` element at all**. The helper must not error on it. The glob fallback exists for exactly this case and for local paths, where globbing works fine.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Create `tests/testthat/test-long-files.R`:
|
||
|
||
```r
|
||
test_that(".long_files_sql enumerates every partition the manifest lists", {
|
||
manifest <- list(files = list(long_partitions = list(
|
||
list(year = 2011L, path = "data/long/year=2011/part-0.parquet"),
|
||
list(year = 2012L, path = "data/long/year=2012/part-0.parquet")
|
||
)))
|
||
expect_equal(
|
||
uscogdata:::.long_files_sql("https://example.org/corpus/", manifest),
|
||
paste0(
|
||
"['https://example.org/corpus/data/long/year=2011/part-0.parquet',",
|
||
"'https://example.org/corpus/data/long/year=2012/part-0.parquet']"
|
||
)
|
||
)
|
||
})
|
||
|
||
test_that(".long_files_sql falls back to the glob when no partition list is present", {
|
||
# test-views.R registers views with a hand-built manifest that has no
|
||
# `files` element. That must keep working: the glob is valid for the
|
||
# local paths such a manifest is used with.
|
||
expect_equal(
|
||
uscogdata:::.long_files_sql("/tmp/corpus/", list(schema_version = 4L)),
|
||
"'/tmp/corpus/data/long/**/*.parquet'"
|
||
)
|
||
expect_equal(
|
||
uscogdata:::.long_files_sql("/tmp/corpus/", list(files = list(long_partitions = list()))),
|
||
"'/tmp/corpus/data/long/**/*.parquet'"
|
||
)
|
||
})
|
||
|
||
test_that("the enumerated list matches the bundled fixture's partition count", {
|
||
skip_if_no_corpus()
|
||
m <- jsonlite::fromJSON(
|
||
file.path(fixture_corpus_path(), "manifest.json"), simplifyVector = FALSE
|
||
)
|
||
out <- uscogdata:::.long_files_sql(fixture_corpus_path(), m)
|
||
expect_equal(
|
||
lengths(regmatches(out, gregexpr("part-0\\.parquet", out))),
|
||
length(m$files$long_partitions)
|
||
)
|
||
})
|
||
|
||
test_that("registered `long` view reads through the enumerated list", {
|
||
skip_if_no_corpus()
|
||
with_fixture_corpus({
|
||
con <- uscogdata:::.ensure_session()
|
||
n <- DBI::dbGetQuery(con, "SELECT count(*) AS n FROM long")$n
|
||
expect_gt(n, 0)
|
||
yrs <- DBI::dbGetQuery(con, "SELECT DISTINCT year FROM long ORDER BY year")$year
|
||
expect_true(all(c(2011, 2012, 2019, 2020) %in% yrs))
|
||
})
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run the tests to verify they fail**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "long-files")'`
|
||
Expected: FAIL — `could not find function ".long_files_sql"` on the first three; the fourth may pass already (it exercises the glob against a local fixture, which works).
|
||
|
||
- [ ] **Step 3: Add the helper to `R/views.R`**
|
||
|
||
Insert immediately above `#' Register DuckDB views from inst/sql/ SQL files`:
|
||
|
||
```r
|
||
#' Build the SQL path expression for the partitioned `long` table.
|
||
#'
|
||
#' DuckDB cannot expand a glob over generic HTTP: there is no directory
|
||
#' listing to expand against, and `allow_asterisks_in_http_paths` only
|
||
#' forwards the literal `**/*` as a filename, which 404s. Measured against
|
||
#' the published corpus on 2026-08-08, an explicit file list returns the
|
||
#' same 46,148,034 rows the (working) `hf://` glob does, and
|
||
#' `hive_partitioning = true` still recovers `year` from the paths.
|
||
#'
|
||
#' The manifest already enumerates every partition, so we build the list
|
||
#' from it. This is host-agnostic -- Nextcloud, HuggingFace and a local
|
||
#' fixture take the same path -- where an `hf://` URL would tie the reader
|
||
#' to one vendor's protocol and still need special-casing, since manifest
|
||
#' fetching goes through httr2, which cannot speak `hf://`.
|
||
#'
|
||
#' Falls back to the glob when the manifest carries no partition list: a
|
||
#' hand-built manifest in a test (see test-views.R) or a corpus predating
|
||
#' the field. Both are local, where globbing works.
|
||
#' @noRd
|
||
.long_files_sql <- function(url, manifest) {
|
||
parts <- manifest$files$long_partitions %||% list()
|
||
if (length(parts) == 0L) {
|
||
return(.sql_lit_chr(paste0(url, "data/long/**/*.parquet")))
|
||
}
|
||
paths <- vapply(parts, function(p) as.character(p$path), character(1))
|
||
paste0("[", .sql_lit_chr(paste0(url, paths)), "]")
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 4: Substitute the new token in `.register_views()`**
|
||
|
||
In `R/views.R`, replace this line:
|
||
|
||
```r
|
||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||
```
|
||
|
||
with:
|
||
|
||
```r
|
||
sql <- gsub("\\{long_files\\}", .long_files_sql(url, manifest), sql, fixed = FALSE)
|
||
sql <- gsub("\\{url\\}", url, sql, fixed = FALSE)
|
||
```
|
||
|
||
Order matters: `{long_files}` expands to a string containing the url, so it must be substituted first or the `{url}` pass would have nothing to do and the token would survive.
|
||
|
||
- [ ] **Step 5: Change the SQL to use the token**
|
||
|
||
`inst/sql/10-long.sql` line 3 becomes:
|
||
|
||
```sql
|
||
FROM read_parquet({long_files}, hive_partitioning = true);
|
||
```
|
||
|
||
Note there are **no surrounding quotes** — `.long_files_sql()` returns its own quoting, whether a bracketed list or a single quoted glob.
|
||
|
||
- [ ] **Step 6: Run the new tests**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "long-files")'`
|
||
Expected: PASS, all four.
|
||
|
||
- [ ] **Step 7: Run the full suite for regressions**
|
||
|
||
Run: `Rscript -e 'devtools::test()'`
|
||
Expected: PASS. Pay particular attention to `test-views.R` — it is the file that exercises `.register_views()` with a `files`-less manifest.
|
||
|
||
- [ ] **Step 8: Commit**
|
||
|
||
```bash
|
||
git add R/views.R inst/sql/10-long.sql tests/testthat/test-long-files.R
|
||
git commit -m "fix: enumerate long partitions from the manifest instead of globbing
|
||
|
||
DuckDB cannot expand a glob over generic HTTP -- no directory listing --
|
||
so every remote corpus read failed. Only local paths worked, which is how
|
||
the API and the test fixture run, so nothing caught it.
|
||
|
||
Measured against the published corpus: the explicit list returns the same
|
||
46,148,034 rows, with hive_partitioning still recovering year."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 2: Ship a working default corpus URL
|
||
|
||
The P0 defect, part two. The default is the literal `REPLACE_WITH_SHARE_TOKEN` sentinel and no file in the repo supplies a real URL, so a new user has no path to a working session.
|
||
|
||
**Files:**
|
||
- Modify: `R/config.R:7`
|
||
- Modify: `R/manifest.R` (the `i =` guidance line in `.check_url_configured()`)
|
||
- Modify: `tests/testthat/test-manifest.R:11`, `:38`
|
||
- Test: `tests/testthat/test-config.R` (add cases)
|
||
|
||
**Interfaces:**
|
||
- Consumes: `.cfg("url")`, `.resolve_url()` from `R/config.R`.
|
||
- Produces: a `.uscogdata_defaults$url` that is a real, reachable, credential-free URL. Task 3's integration test depends on this being the default.
|
||
|
||
**Critical constraint:** `tests/testthat/test-manifest.R` hardcodes the placeholder at lines 11 and 38 and asserts `cog_open()` aborts with `uscogdata_url_not_configured`. Changing the default **breaks those two tests** and they must be updated in this task. The test at line 26 passes an explicit sentinel-bearing URL via env var — it keeps passing untouched, and it is what proves the sentinel guard still works.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Append to `tests/testthat/test-config.R`:
|
||
|
||
```r
|
||
test_that("the default corpus URL is real, not a placeholder", {
|
||
withr::with_envvar(c(USCOGDATA_URL = NA), {
|
||
withr::with_options(list(uscogdata.url = NULL), {
|
||
url <- uscogdata:::.resolve_url()
|
||
expect_false(grepl("REPLACE_WITH", url, fixed = TRUE))
|
||
expect_match(url, "^https://", perl = TRUE)
|
||
expect_match(url, "/$", perl = TRUE)
|
||
})
|
||
})
|
||
})
|
||
|
||
test_that("an explicitly-set sentinel URL still aborts", {
|
||
# The guard must survive the default change: a user who half-edited a
|
||
# copied config still gets the actionable error.
|
||
withr::with_envvar(
|
||
c(USCOGDATA_URL = "https://other.example/s/REPLACE_WITH_SHARE_TOKEN/x/"), {
|
||
expect_error(
|
||
uscogdata:::.check_url_configured(uscogdata:::.resolve_url()),
|
||
class = "uscogdata_url_not_configured"
|
||
)
|
||
})
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run to verify failure**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: FAIL on the first test — the resolved default still contains `REPLACE_WITH`.
|
||
|
||
- [ ] **Step 3: Change the default**
|
||
|
||
`R/config.R`, in `.uscogdata_defaults`:
|
||
|
||
```r
|
||
.uscogdata_defaults <- list(
|
||
# Public HuggingFace mirror of the published corpus: CC-BY-4.0, no
|
||
# credential, CDN-backed. Chosen as the default so `library(uscogdata)`
|
||
# followed by a verb works with zero configuration. Trailing slash is
|
||
# required -- every consumer concatenates onto this (see .resolve_url()).
|
||
url = "https://huggingface.co/datasets/civilytics/us-cog-finance/resolve/main/",
|
||
cache_dir = NULL,
|
||
manifest_ttl_secs = 3600L
|
||
)
|
||
```
|
||
|
||
- [ ] **Step 4: Update the stale guidance line**
|
||
|
||
In `R/manifest.R`, inside `.check_url_configured()`, replace:
|
||
|
||
```r
|
||
i = "For the live Civilytics corpus, request the Nextcloud share URL from the package maintainer."
|
||
```
|
||
|
||
with:
|
||
|
||
```r
|
||
i = "The public corpus is the default; unset USCOGDATA_URL to use it, or point it at a local copy from {.code cog_mirror()}."
|
||
```
|
||
|
||
- [ ] **Step 5: Correct the two now-inaccurate test names**
|
||
|
||
Both affected tests in `tests/testthat/test-manifest.R` set `USCOGDATA_URL` **explicitly** via `withr::with_envvar` before calling `cog_open()`, so they keep passing unchanged. Only their wording becomes wrong — the sentinel URL is no longer the default.
|
||
|
||
Line 7, rename the test:
|
||
|
||
```r
|
||
test_that("cog_open aborts with actionable error when URL contains the sentinel", {
|
||
```
|
||
|
||
Lines 11 and 38, rename the variable and say why it is still here:
|
||
|
||
```r
|
||
# No longer the package default (that is the public HF corpus). This is a
|
||
# user who copied a config template and did not finish editing it.
|
||
sentinel_url <- "https://cloud.civilytics.org/s/REPLACE_WITH_SHARE_TOKEN/download/"
|
||
```
|
||
|
||
Update the two `withr::with_envvar(c(USCOGDATA_URL = placeholder), ...)` call sites in those tests to use `sentinel_url`. No functional change — do not alter the assertions.
|
||
|
||
- [ ] **Step 6: Run the affected files**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config|manifest")'`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 7: Run the full suite**
|
||
|
||
Run: `Rscript -e 'devtools::test()'`
|
||
Expected: PASS. `setup.R` points `USCOGDATA_URL` at the bundled fixture for the whole suite, so the default change should not affect any other file.
|
||
|
||
- [ ] **Step 8: Commit**
|
||
|
||
```bash
|
||
git add R/config.R R/manifest.R tests/testthat/test-config.R tests/testthat/test-manifest.R
|
||
git commit -m "feat: default to the public corpus so the package works unconfigured
|
||
|
||
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no file in the
|
||
repo supplied a working URL, so a new user had no path to a session. The
|
||
sentinel guard stays for half-edited configs."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 3: Live-corpus integration test
|
||
|
||
Nothing in the suite exercises a remote corpus — that is why the P0 survived. This test is network-gated so it skips in offline CI but runs on demand and before release.
|
||
|
||
**Files:**
|
||
- Test: `tests/testthat/test-live-corpus.R` (create)
|
||
|
||
**Interfaces:**
|
||
- Consumes: `.long_files_sql()` (Task 1), the default URL (Task 2), and the public verbs `cog_gov_search()`, `cog_spending()`, `cog_explain()`.
|
||
- Produces: nothing consumed downstream.
|
||
|
||
- [ ] **Step 1: Write the test**
|
||
|
||
Create `tests/testthat/test-live-corpus.R`:
|
||
|
||
```r
|
||
# Network-gated. Set USCOGDATA_LIVE_TEST=true to run.
|
||
#
|
||
# This file exists because the P0 fixed in this release -- no remote corpus
|
||
# was readable at all -- survived precisely because every other test path
|
||
# used a LOCAL corpus (the bundled fixture) and so did the API in
|
||
# production. Nothing ever exercised the code the way a new user does.
|
||
skip_live <- function() {
|
||
testthat::skip_if_not(
|
||
identical(tolower(Sys.getenv("USCOGDATA_LIVE_TEST", "")), "true"),
|
||
"live-corpus test: set USCOGDATA_LIVE_TEST=true to run"
|
||
)
|
||
}
|
||
|
||
test_that("the package reads the public corpus with no configuration at all", {
|
||
skip_live()
|
||
withr::with_envvar(c(USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA), {
|
||
withr::with_options(list(uscogdata.url = NULL), {
|
||
uscogdata:::cog_close()
|
||
on.exit(uscogdata:::cog_close(), add = TRUE)
|
||
|
||
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||
expect_gt(nrow(g), 0)
|
||
|
||
s <- cog_spending(g$canonical_govid[1], years = 2022)
|
||
expect_gt(nrow(s), 0)
|
||
expect_true(all(c("amt_nominal", "year", "category") %in% names(s)))
|
||
|
||
# Amounts are full dollars, already x1000. A city's total annual
|
||
# spending is millions, not thousands -- this catches a regression
|
||
# that reintroduced the double conversion.
|
||
expect_gt(sum(s$amt_nominal, na.rm = TRUE), 1e6)
|
||
|
||
p <- attr(s, "provenance")
|
||
expect_true(isTRUE(p$transformations$units_conversion$applied))
|
||
expect_equal(p$transformations$units_conversion$multiplier, 1000)
|
||
})
|
||
})
|
||
})
|
||
|
||
test_that("a full-history query spans the published range", {
|
||
skip_live()
|
||
withr::with_envvar(c(USCOGDATA_URL = NA, USCOGDATA_FIXTURE_URL = NA), {
|
||
withr::with_options(list(uscogdata.url = NULL), {
|
||
uscogdata:::cog_close()
|
||
on.exit(uscogdata:::cog_close(), add = TRUE)
|
||
|
||
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||
s <- cog_spending(g$canonical_govid[1])
|
||
# The corpus publishes FY1967-FY2024. Any single government's span is
|
||
# narrower, but a full-history query must cross more than one decade
|
||
# -- if enumeration silently returned one partition, this fails.
|
||
expect_gt(diff(range(s$year)), 10)
|
||
})
|
||
})
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run it gated off (default) — it must skip, not fail**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "live-corpus")'`
|
||
Expected: SKIP on both, with the message about `USCOGDATA_LIVE_TEST`.
|
||
|
||
- [ ] **Step 3: Run it live**
|
||
|
||
Run: `USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test(filter = "live-corpus")'`
|
||
Expected: PASS. This is the first time the package has ever read a remote corpus successfully. If it fails, Task 1 is incomplete — do not proceed.
|
||
|
||
- [ ] **Step 4: Record the measured verb latency**
|
||
|
||
Run and note the wall time, which the README needs (raw scans were 1.5 s / 2.8 s; real verbs do more work):
|
||
|
||
```bash
|
||
USCOGDATA_LIVE_TEST=true Rscript -e '
|
||
Sys.unsetenv("USCOGDATA_URL")
|
||
library(uscogdata)
|
||
g <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||
print(system.time(cog_spending(g$canonical_govid[1], years = 2022)))
|
||
print(system.time(cog_spending(g$canonical_govid[1])))
|
||
'
|
||
```
|
||
|
||
Carry these two numbers into Task 8. Do not reuse the raw-scan figures.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git add tests/testthat/test-live-corpus.R
|
||
git commit -m "test: exercise the public corpus end to end, unconfigured
|
||
|
||
The remote-read defect survived because every test path used a local
|
||
corpus. This is the only test that runs the package the way a new user
|
||
does."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 4: DESCRIPTION metadata
|
||
|
||
**Files:**
|
||
- Modify: `DESCRIPTION`
|
||
|
||
**Interfaces:**
|
||
- Produces: `URL`/`BugReports` that Task 7 (`_pkgdown.yml`) and Task 8 (README) both reference; the `Authors@R` that `citation("uscogdata")` renders.
|
||
|
||
- [ ] **Step 1: Write the failing test**
|
||
|
||
Append to `tests/testthat/test-config.R`:
|
||
|
||
```r
|
||
test_that("DESCRIPTION carries release metadata", {
|
||
skip_if_no_source_tree("DESCRIPTION")
|
||
d <- read.dcf(source_tree_path("DESCRIPTION"))
|
||
fields <- colnames(d)
|
||
|
||
expect_true(all(c("URL", "BugReports") %in% fields))
|
||
expect_match(d[1, "Authors@R"], "Knowles", fixed = TRUE)
|
||
expect_match(d[1, "Authors@R"], "0000-0003-0005-9478", fixed = TRUE)
|
||
expect_match(d[1, "Authors@R"], "Civilytics Consulting LLC", fixed = TRUE)
|
||
|
||
# The gate in .validate_schema() accepts up to 7 and the published corpus
|
||
# IS 7; DESCRIPTION must not claim otherwise.
|
||
expect_equal(as.integer(d[1, "MaxCorpusSchema"]), 7L)
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run to verify failure**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: FAIL — `URL`/`BugReports` absent, `MaxCorpusSchema` is 5.
|
||
|
||
- [ ] **Step 3: Edit DESCRIPTION**
|
||
|
||
Replace the `Authors@R` block:
|
||
|
||
```
|
||
Authors@R: c(
|
||
person(c("Jared", "E."), "Knowles",
|
||
email = "jared@civilytics.com",
|
||
role = c("aut", "cre"),
|
||
comment = c(ORCID = "0000-0003-0005-9478")),
|
||
person("Civilytics Consulting LLC", role = c("cph", "fnd")))
|
||
```
|
||
|
||
The given-name vector `c("Jared", "E.")` with family name `"Knowles"` matches
|
||
`merTools` exactly. That structural match is what lets ORCID and r-universe
|
||
collate both packages as one person's work — `person("Jared", "E. Knowles")`
|
||
would render identically but put the middle initial in the family-name slot.
|
||
|
||
Add after `Description:`:
|
||
|
||
```
|
||
URL: https://github.com/civilytics/uscogdata, https://civilytics.r-universe.dev/uscogdata
|
||
BugReports: https://github.com/civilytics/uscogdata/issues
|
||
```
|
||
|
||
Change:
|
||
|
||
```
|
||
MaxCorpusSchema: 7
|
||
```
|
||
|
||
- [ ] **Step 4: Verify the person object parses**
|
||
|
||
Run: `Rscript -e 'print(eval(parse(text = read.dcf("DESCRIPTION")[1, "Authors@R"])))'`
|
||
Expected: prints two entries — `Jared E. Knowles <jared@civilytics.com> [aut, cre] (ORCID: ...)` and `Civilytics Consulting LLC [cph, fnd]`. A parse error here means a malformed `person()` call.
|
||
|
||
- [ ] **Step 5: Run the tests**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add DESCRIPTION tests/testthat/test-config.R
|
||
git commit -m "chore: release metadata -- author of record, URLs, schema ceiling
|
||
|
||
Authors@R was an org with no human, so citation() and the r-universe
|
||
maintainer page had nothing to render. MaxCorpusSchema claimed 5 while the
|
||
code accepts 7 and the published corpus is 7."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 5: License files
|
||
|
||
**Files:**
|
||
- Modify: `LICENSE`
|
||
- Create: `LICENSE.md`
|
||
|
||
- [ ] **Step 1: Generate both files**
|
||
|
||
Run: `Rscript -e 'usethis::use_mit_license("Civilytics Consulting LLC")'`
|
||
|
||
This rewrites `LICENSE` to the two-line stub with the corrected holder and creates `LICENSE.md` with the full MIT text. It may also add `^LICENSE\.md$` to `.Rbuildignore` — that is correct and already present.
|
||
|
||
- [ ] **Step 2: Verify**
|
||
|
||
Run: `Rscript -e 'cat(readLines("LICENSE"), sep = "\n")'`
|
||
Expected:
|
||
```
|
||
YEAR: 2026
|
||
COPYRIGHT HOLDER: Civilytics Consulting LLC
|
||
```
|
||
|
||
Run: `Rscript -e 'cat(length(readLines("LICENSE.md")), "lines\n")'`
|
||
Expected: a non-zero count (the full MIT text, ~21 lines).
|
||
|
||
- [ ] **Step 3: Confirm DESCRIPTION still declares the license correctly**
|
||
|
||
Run: `Rscript -e 'cat(read.dcf("DESCRIPTION")[1, "License"], "\n")'`
|
||
Expected: `MIT + file LICENSE`. If `usethis` changed it, that is fine — leave whatever it wrote.
|
||
|
||
- [ ] **Step 4: Commit**
|
||
|
||
```bash
|
||
git add LICENSE LICENSE.md DESCRIPTION .Rbuildignore
|
||
git commit -m "chore: add full MIT text, name the copyright holder properly
|
||
|
||
LICENSE held only the two-line stub and no LICENSE.md existed, so the
|
||
repo carried no license text for a human or for GitHub's detector."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 6: Ship the vignettes
|
||
|
||
`.Rbuildignore` excludes `^vignettes$`, so an installed package has no vignettes at all — while README tells users to run `vignette("total-spending", package = "uscogdata")`. Both vignettes build offline: `total-spending.Rmd` points `USCOGDATA_URL` at the bundled fixture, `population-denominators.Rmd` is `eval = FALSE`.
|
||
|
||
**Files:**
|
||
- Modify: `.Rbuildignore`
|
||
- Test: `tests/testthat/test-config.R` (add)
|
||
|
||
- [ ] **Step 1: Write the failing test**
|
||
|
||
Append to `tests/testthat/test-config.R`:
|
||
|
||
```r
|
||
test_that("vignettes are not excluded from the build", {
|
||
skip_if_no_source_tree(".Rbuildignore")
|
||
ignore <- readLines(source_tree_path(".Rbuildignore"), warn = FALSE)
|
||
expect_false(any(grepl("^\\^vignettes\\$$", ignore)))
|
||
# The fixture is what lets R CMD check run offline with no credentials on
|
||
# r-universe and GitHub Actions. It must never be excluded.
|
||
expect_false(any(grepl("fixture_corpus", ignore, fixed = TRUE)))
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run to verify failure**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: FAIL on the first expectation.
|
||
|
||
- [ ] **Step 3: Remove the three lines**
|
||
|
||
Delete these lines from `.Rbuildignore`:
|
||
|
||
```
|
||
^vignettes$
|
||
^doc$
|
||
^Meta$
|
||
```
|
||
|
||
Leave every other line untouched — in particular `^_pkgdown\.yml$`, `^docs$`, `^data-raw$`, `^specs$`, `^plans$`, `^\.gitea$`, `^CLAUDE\.md$` are all correct exclusions.
|
||
|
||
- [ ] **Step 4: Build the tarball and confirm the vignettes are in it**
|
||
|
||
```bash
|
||
Rscript -e 'devtools::build(path = tempdir())'
|
||
```
|
||
Then list the tarball contents:
|
||
```bash
|
||
tar -tzf "$(ls -t $(Rscript -e 'cat(tempdir())')/uscogdata_*.tar.gz | head -1)" | grep -E 'vignettes|inst/doc'
|
||
```
|
||
Expected: both `.Rmd` files appear under `uscogdata/vignettes/`.
|
||
|
||
- [ ] **Step 5: Run the tests**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add .Rbuildignore tests/testthat/test-config.R
|
||
git commit -m "fix: ship the vignettes
|
||
|
||
.Rbuildignore excluded ^vignettes$, so vignette(\"total-spending\") failed
|
||
for every user -- while the README instructed them to run it. Both build
|
||
offline against the bundled fixture."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 7: pkgdown reference index
|
||
|
||
`_pkgdown.yml` lists 6 of 14 exports. pkgdown errors on topics missing from the index, so the docs site does not build.
|
||
|
||
**Files:**
|
||
- Modify: `_pkgdown.yml`
|
||
|
||
- [ ] **Step 1: Write the failing test**
|
||
|
||
Append to `tests/testthat/test-config.R`:
|
||
|
||
```r
|
||
test_that("_pkgdown.yml indexes every exported topic", {
|
||
skip_if_no_source_tree("_pkgdown.yml", "NAMESPACE")
|
||
exports <- grep("^export\\(", readLines(source_tree_path("NAMESPACE"), warn = FALSE), value = TRUE)
|
||
exports <- sub("^export\\((.*)\\)$", "\\1", exports)
|
||
yml <- paste(readLines(source_tree_path("_pkgdown.yml"), warn = FALSE), collapse = "\n")
|
||
missing <- exports[!vapply(exports, function(e) grepl(paste0("\\b", e, "\\b"), yml), logical(1))]
|
||
expect_equal(missing, character(0))
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run to verify failure**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: FAIL listing the eight missing: `cog_categories`, `cog_explain`, `cog_find_peers`, `cog_geographic_rollup`, `cog_manifest`, `cog_mirror`, `cog_peer_compare`, `cog_recipes`.
|
||
|
||
- [ ] **Step 3: Rewrite `_pkgdown.yml`**
|
||
|
||
```yaml
|
||
url: https://civilytics.r-universe.dev/uscogdata
|
||
|
||
template:
|
||
bootstrap: 5
|
||
|
||
reference:
|
||
- title: Financial data
|
||
desc: Spending, revenue and balance-sheet holdings for one or more governments.
|
||
contents:
|
||
- cog_spending
|
||
- cog_revenue
|
||
- cog_balances
|
||
- title: Search & basket
|
||
desc: Resolve place names into canonical govids.
|
||
contents:
|
||
- cog_gov_search
|
||
- cog_basket_resolution
|
||
- cog_basket_unresolved
|
||
- title: Comparison & aggregation
|
||
desc: Peer cohorts and geographic rollups.
|
||
contents:
|
||
- cog_find_peers
|
||
- cog_peer_compare
|
||
- cog_geographic_rollup
|
||
- title: Corpus metadata
|
||
desc: What the corpus contains, where it came from, and how to hold a local copy.
|
||
contents:
|
||
- cog_categories
|
||
- cog_recipes
|
||
- cog_manifest
|
||
- cog_explain
|
||
- cog_mirror
|
||
|
||
articles:
|
||
- title: Concepts
|
||
navbar: ~
|
||
contents:
|
||
- total-spending
|
||
- population-denominators
|
||
```
|
||
|
||
- [ ] **Step 4: Build the site**
|
||
|
||
Run: `Rscript -e 'pkgdown::build_site(preview = FALSE)'`
|
||
Expected: completes without error. Any "Topics missing from index" warning means an export was missed.
|
||
|
||
- [ ] **Step 5: Run the tests**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add _pkgdown.yml tests/testthat/test-config.R
|
||
git commit -m "docs: index all 14 exports in pkgdown, set the site url
|
||
|
||
The reference index covered 6 of 14, so pkgdown errored on the missing
|
||
topics and the docs site did not build."
|
||
```
|
||
|
||
`docs/` is gitignored via `.Rbuildignore`/`.gitignore`; do not commit built output.
|
||
|
||
---
|
||
|
||
## Task 8: README rewrite
|
||
|
||
The current README addresses someone inside the repo tree: status reads "Under active development (Phase 2 of the cog_pipeline project)", it points at `../cog_pipeline/docs/reader-specification.md`, the install line is commented out, and developer/testing/release sections sit above anything a user needs.
|
||
|
||
**Files:**
|
||
- Rewrite: `README.md`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `URL`/`BugReports` from Task 4, the default URL from Task 2, the measured latencies from Task 3 Step 4.
|
||
|
||
- [ ] **Step 1: Write the failing test**
|
||
|
||
Append to `tests/testthat/test-config.R`:
|
||
|
||
```r
|
||
test_that("README is written for a stranger, not a repo insider", {
|
||
skip_if_no_source_tree("README.md")
|
||
r <- paste(readLines(source_tree_path("README.md"), warn = FALSE), collapse = "\n")
|
||
|
||
# No paths that only resolve inside Jared's checkout.
|
||
expect_false(grepl("../cog_pipeline", r, fixed = TRUE))
|
||
# A real, uncommented install line.
|
||
expect_match(r, "install.packages", fixed = TRUE)
|
||
expect_false(grepl("# pak::pkg_install", r, fixed = TRUE))
|
||
# The errata most likely to produce a plausible-looking wrong answer.
|
||
expect_match(r, "full US dollars", fixed = TRUE)
|
||
# The release-instructions section that conflicts with public CI is gone.
|
||
expect_false(grepl("Rbuildignore", r, fixed = TRUE))
|
||
# Both read paths documented.
|
||
expect_match(r, "cog_mirror", fixed = TRUE)
|
||
})
|
||
```
|
||
|
||
- [ ] **Step 2: Run to verify failure**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: FAIL.
|
||
|
||
- [ ] **Step 3: Rewrite README.md in this order**
|
||
|
||
Write these sections, in this sequence. Content requirements are exact; prose is yours.
|
||
|
||
1. **Title + one-paragraph what-it-is.** Curated R reader over the Civilytics US Census of Governments finance corpus. State the coverage: government types 0–3 (state, county, municipality, township), **FY1967–FY2024** (no source data FY1968, FY1969), **46,148,034 rows**, **190.6 MB**. Add the r-universe version badge.
|
||
|
||
2. **Install.**
|
||
````markdown
|
||
```r
|
||
install.packages("uscogdata",
|
||
repos = c("https://civilytics.r-universe.dev",
|
||
"https://cloud.r-project.org"))
|
||
```
|
||
Or from source:
|
||
```r
|
||
pak::pkg_install("git::https://gitea.civilytics.org/Civilytics/uscogdata.git")
|
||
```
|
||
````
|
||
|
||
3. **Quickstart — no configuration step.** Verbatim:
|
||
````markdown
|
||
```r
|
||
library(uscogdata)
|
||
|
||
# Resolve a place name to a canonical government id
|
||
madison <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||
|
||
# Full spending history, real dollars, per capita
|
||
spend <- cog_spending(madison$canonical_govid[1])
|
||
|
||
# Every result carries its own provenance
|
||
cog_explain(spend)
|
||
```
|
||
````
|
||
|
||
4. **Two ways to read the corpus.** Remote (default, zero setup) vs `cog_mirror()` (190 MB once). Include the measured table — **use the verb latencies recorded in Task 3 Step 4, not the raw-scan figures**:
|
||
|
||
| | Remote (default) | Mirrored |
|
||
|---|---|---|
|
||
| Setup | none | `cog_mirror()`, 190.6 MB once |
|
||
| Disk | 0 MB — HTTP range requests | 190.6 MB |
|
||
| Per-query | network round trip | local |
|
||
| Right for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility |
|
||
|
||
State plainly: the default reads from a public HuggingFace mirror, and **the escape hatch is one function call** — after `cog_mirror()`, no analysis touches an external service.
|
||
|
||
5. **Amounts are in full US dollars.** Keep the existing text nearly verbatim — it is correct and carefully argued. Keep the `attr(r, "provenance")$transformations$units_conversion` example and the "do not multiply again" warning.
|
||
|
||
6. **Concepts.** Condense the existing primary/direct/total and general/total revenue sections to ~1/3 their length, each ending with a pointer to `vignette("total-spending")`. Add a short **Reporting coverage** paragraph: the Census is a complete census only in years ending in 2 and 7; every other year is a sample; `provenance$coverage` reports `n_units_reporting` per year, and `coverage = "census"` / `"consistent"` control the mode.
|
||
|
||
7. **How to cite.** `citation("uscogdata")` for the package; corpus is CC-BY-4.0, cite *Civilytics Consulting, US Census of Governments finance corpus*.
|
||
|
||
8. **Contributing.** Two sentences plus a link to `CONTRIBUTING.md` (Task 10).
|
||
|
||
**Delete outright:** the "Status" block, the `../cog_pipeline/...` reference, the entire "Developer notes / Testing / Releasing against the live corpus" section (it moves to `CONTRIBUTING.md`, minus the fixture-stripping advice, which is wrong and must not be carried over).
|
||
|
||
- [ ] **Step 4: Run the quickstart verbatim in a clean session**
|
||
|
||
```bash
|
||
Rscript -e '
|
||
Sys.unsetenv("USCOGDATA_URL")
|
||
devtools::load_all(".", quiet = TRUE)
|
||
madison <- cog_gov_search(name = "Madison", state = "WI", type = 2)
|
||
spend <- cog_spending(madison$canonical_govid[1])
|
||
cat("rows:", nrow(spend), "years:", paste(range(spend$year), collapse = "-"), "\n")
|
||
cog_explain(spend)
|
||
'
|
||
```
|
||
Expected: runs clean with no configuration. If it errors, the README is wrong — fix the README, not the test.
|
||
|
||
- [ ] **Step 5: Run the tests**
|
||
|
||
Run: `Rscript -e 'devtools::test(filter = "config")'`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add README.md tests/testthat/test-config.R
|
||
git commit -m "docs: rewrite README for a stranger
|
||
|
||
Reordered around a new user -- what it is, install, a quickstart that runs
|
||
with no configuration, then the dollars warning and the concepts. Drops
|
||
the sibling-repo path, the commented-out install line, and the
|
||
fixture-stripping release advice that conflicts with public CI."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 9: NEWS.md rewrite
|
||
|
||
The current NEWS is a pre-release churn log spanning the package's entire development (2026-04-23 to 2026-08-04, 140 commits), describing changes relative to states no user has seen. To a newcomer it reads as instability.
|
||
|
||
**Files:**
|
||
- Rewrite: `NEWS.md`
|
||
- Modify: `README.md` and roxygen blocks receiving migrated content
|
||
|
||
- [ ] **Step 1: Migrate the load-bearing content first**
|
||
|
||
Before deleting anything, move each of these to its documentation home. Verify each lands before proceeding:
|
||
|
||
| Content in current NEWS | Destination |
|
||
|---|---|
|
||
| Coverage disclosure — census vs sample years, the 597-vs-112 Wisconsin example, `coverage` modes | README §6 (Task 8) **and** `@details` in `R/rollup.R`'s roxygen for `cog_geographic_rollup()` |
|
||
| `complete = TRUE` — the `reported` / `census_zero` / `not_reported` table | `@details` in `R/spending.R` and `R/revenue.R` roxygen |
|
||
| Series breaks and `corpus_break_refs` | `@details` in `R/provenance.R`'s `cog_explain()` roxygen |
|
||
| Per-year F-33 population denominators | already in `vignette("population-denominators")` — verify, do not duplicate |
|
||
|
||
Run `Rscript -e 'devtools::document()'` after editing roxygen.
|
||
|
||
- [ ] **Step 2: Replace NEWS.md entirely**
|
||
|
||
```markdown
|
||
# uscogdata 0.1.0
|
||
|
||
First public release.
|
||
|
||
`uscogdata` provides curated R verbs over the Civilytics US Census of
|
||
Governments finance corpus: unit-level financial profiles, geographic
|
||
rollups, and peer comparisons, with auditable provenance on every result.
|
||
|
||
## What it covers
|
||
|
||
Government types 0–3 (state, county, municipality, township), FY1967–FY2024
|
||
(no source data for FY1968 or FY1969) — 46,148,034 rows across 56 fiscal
|
||
years. The corpus is published under CC-BY-4.0 and reads directly over
|
||
HTTPS, or locally after `cog_mirror()`.
|
||
|
||
## The verbs
|
||
|
||
`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and
|
||
holdings; `cog_gov_search()` to resolve place names (including basket mode
|
||
for many at once); `cog_find_peers()` and `cog_peer_compare()` for cohorts;
|
||
`cog_geographic_rollup()` for aggregates; `cog_categories()`,
|
||
`cog_recipes()`, `cog_manifest()` and `cog_explain()` for metadata and
|
||
provenance; `cog_mirror()` for a local copy.
|
||
|
||
## Four things to know before your first query
|
||
|
||
* **Amounts are in full US dollars.** The raw Census files report thousands;
|
||
the verbs multiply by 1000 on the way out. Do not multiply again.
|
||
* **Multi-government aggregates disclose their coverage.** The Census is a
|
||
complete census only in years ending in 2 and 7. Every result carries
|
||
`provenance$coverage` with per-year `n_units_reporting`.
|
||
* **Absence means two different things.** Before FY2012 an absent cell means
|
||
Census published $0; from FY2012 it means not reported.
|
||
`complete = TRUE` labels which.
|
||
* **Series breaks reach you unasked.** Catalogued breaks intersecting your
|
||
query appear in provenance and in `cog_explain()`.
|
||
|
||
## Known limits
|
||
|
||
* Special districts (type 4) and school districts (type 5) are out of scope.
|
||
* Per-capita rollups exclude governments with no F-33 population.
|
||
* Employee-retirement (`X`) codes stop at FY2016, when those systems moved
|
||
to the Annual Survey of Public Pensions.
|
||
```
|
||
|
||
- [ ] **Step 3: Verify no orphaned content**
|
||
|
||
Run: `git show HEAD:NEWS.md > /tmp/news-old.md && wc -l /tmp/news-old.md NEWS.md`
|
||
|
||
Read `/tmp/news-old.md` once more and confirm every substantive claim either appears in the new NEWS, landed somewhere in Step 1, or is genuinely pre-release churn (version bumps, fixture regenerations, internal refactors). The pre-release history stays in git; it does not need preserving in NEWS.
|
||
|
||
- [ ] **Step 4: Run the full suite**
|
||
|
||
Run: `Rscript -e 'devtools::test()'`
|
||
Expected: PASS — `devtools::document()` in Step 1 regenerated `man/`, so this catches a malformed roxygen block.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git add NEWS.md README.md R/ man/
|
||
git commit -m "docs: rewrite NEWS as an initial release
|
||
|
||
The changelog described changes relative to states no user ever saw, which
|
||
reads as instability to someone deciding whether to depend on this. The
|
||
load-bearing caveats move into README and roxygen, where they belong; the
|
||
pre-release history stays in git."
|
||
```
|
||
|
||
---
|
||
|
||
## Task 10: CONTRIBUTING.md
|
||
|
||
**Files:**
|
||
- Create: `CONTRIBUTING.md`
|
||
- Modify: `.Rbuildignore`
|
||
|
||
- [ ] **Step 1: Write CONTRIBUTING.md**
|
||
|
||
It must contain, in this order:
|
||
|
||
1. **Canonical source note.** Development happens on `gitea.civilytics.org/Civilytics/uscogdata`; `github.com/civilytics/uscogdata` is a mirror that accepts issues and pull requests.
|
||
|
||
2. **What happens to a GitHub PR.** Verbatim explanation: it is fetched and landed on the canonical repo, then closes itself as merged when the mirror syncs — because the maintainer merges with `--no-ff`, preserving the contributor's commits and SHAs. Say plainly that a PR closing without a "Merged by" click is normal and not a rejection.
|
||
|
||
3. **Running the tests.**
|
||
````markdown
|
||
```r
|
||
devtools::test() # uses the bundled fixture; no network, no credentials
|
||
```
|
||
````
|
||
Note that `tests/testthat/setup.R` points `USCOGDATA_URL` at
|
||
`inst/extdata/fixture_corpus/` automatically.
|
||
|
||
4. **Testing against the live corpus.**
|
||
````markdown
|
||
```r
|
||
USCOGDATA_LIVE_TEST=true devtools::test(filter = "live-corpus")
|
||
```
|
||
````
|
||
Explain why it exists: every other test path uses a local corpus, which is
|
||
how the remote-read defect in 0.1.0 went unnoticed.
|
||
|
||
5. **Do not exclude the fixture from the build.** State the reason — it is what lets `R CMD check` pass on r-universe and GitHub Actions with no credentials.
|
||
|
||
6. **Release checklist**, moved from README: run the suite against both the fixture and the live corpus, `pkgdown::build_site()`, `R CMD check --as-cran`, tag, then update the r-universe registry pin.
|
||
|
||
**Do not carry over** the README's instruction to add `^inst/extdata/fixture_corpus$` to `.Rbuildignore`. It is wrong.
|
||
|
||
- [ ] **Step 2: Exclude it from the build**
|
||
|
||
Add to `.Rbuildignore`:
|
||
|
||
```
|
||
^CONTRIBUTING\.md$
|
||
```
|
||
|
||
- [ ] **Step 3: Verify the build is clean**
|
||
|
||
Run: `Rscript -e 'devtools::check(document = FALSE, args = "--no-manual")'`
|
||
Expected: 0 errors, 0 warnings. Notes about package size (the 15 MB fixture) are expected and acceptable — this package is not going to CRAN.
|
||
|
||
- [ ] **Step 4: Commit**
|
||
|
||
```bash
|
||
git add CONTRIBUTING.md .Rbuildignore
|
||
git commit -m "docs: add CONTRIBUTING with the canonical-on-Gitea PR flow
|
||
|
||
Moves developer and release instructions out of the README, minus the
|
||
fixture-stripping advice, which would break the vignette and leave public
|
||
CI unable to check without credentials."
|
||
```
|
||
|
||
---
|
||
|
||
## Final verification
|
||
|
||
Run before declaring the release ready. Every one of these must pass.
|
||
|
||
- [ ] **1. Full suite, offline, no credentials**
|
||
`Rscript -e 'devtools::test()'` — the property public CI depends on.
|
||
|
||
- [ ] **2. Live corpus**
|
||
`USCOGDATA_LIVE_TEST=true Rscript -e 'devtools::test()'`
|
||
|
||
- [ ] **3. `R CMD check --as-cran`**
|
||
`Rscript -e 'devtools::check(args = "--as-cran")'` — 0 errors, 0 warnings.
|
||
|
||
- [ ] **4. pkgdown**
|
||
`Rscript -e 'pkgdown::build_site(preview = FALSE)'`
|
||
|
||
- [ ] **5. Vignettes resolve from an installed copy**
|
||
```bash
|
||
Rscript -e 'devtools::install(build_vignettes = TRUE, quiet = TRUE)'
|
||
Rscript -e 'v <- vignette("total-spending", package = "uscogdata"); stopifnot(nzchar(v$File)); cat("OK\n")'
|
||
```
|
||
|
||
- [ ] **6. Cold-start check.** On a machine (or container) that has never had this package: install it, then run the README quickstart **verbatim with no environment variables set**.
|
||
```bash
|
||
docker run --rm -v "$PWD":/pkg rocker/r-ver:4.4 bash -c '
|
||
apt-get update -qq && apt-get install -y -qq libcurl4-openssl-dev libssl-dev >/dev/null
|
||
Rscript -e "install.packages(c(\"pak\"), repos=\"https://cloud.r-project.org\")" \
|
||
-e "pak::pkg_install(\"local::/pkg\")" \
|
||
-e "library(uscogdata); m <- cog_gov_search(name=\"Madison\", state=\"WI\", type=2); s <- cog_spending(m\$canonical_govid[1]); cat(\"rows:\", nrow(s), \"\n\")"
|
||
'
|
||
```
|
||
**This is the only check that catches the P0 class of fault**, and its absence is why the fault survived. Do not skip it.
|
||
|
||
- [ ] **7. Verb latency re-measured** against the live corpus, and the README table updated if the numbers moved from what Task 3 Step 4 recorded.
|
||
|
||
## Out of scope for this plan
|
||
|
||
Flipping the Gitea repo public, `gitleaks`, the GitHub mirror and its Actions matrix, the Gitea push workflow, the r-universe registry, and tagging `v0.1.0`. Those follow after this plan's final verification is green — r-universe publishes check results on registration, so registering before checks pass means a red badge on day one. Corrections intake and announcement posts are deferred by decision (see the spec).
|