Swaps every hardcoded 9-char canonical_govid literal (Broward County,
Fort Lauderdale City, Florida/Alabama state govts, Bexar/Tarrant/Wayne
counties, San Diego/Oakland/Miami/Austin cities) for its 12-char Phase P
equivalent, resolved by name+type+state against the regenerated fixture
xwalk. Also updates two gov_name search patterns that no longer match
under Phase P canonical naming ("FLORIDA STATE GOVT" -> "FLORIDA"; the
"Miami" substring test now pins type = "city" since MIAMI-DADE COUNTY's
canonical name now also contains "Miami", which would otherwise make the
match ambiguous across govs_types instead of resolving via largest-pop).
Underlying per-year population figures for Broward County and Alabama
are unchanged, so no expected data-value literals needed recomputation.
Suite: 126 test blocks / 336 expectations, 0 FAIL / 0 WARN / 0 SKIP.
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
`cog_gov_search()` (and every other verb) used to fail with a cryptic
`jsonlite` lexical error when the package's placeholder default URL was
hit and the server returned an HTML welcome page that got cached as
`manifest.json`. Three guards added:
1. `.check_url_configured()` aborts with class `uscogdata_url_not_configured`
when the resolved URL is empty or still contains the
`REPLACE_WITH_SHARE_TOKEN` sentinel. Message names both
`Sys.setenv(USCOGDATA_URL = ...)` and `options(uscogdata.url = ...)`
remediations and points at the bundled fixture.
2. `.fetch_or_cache_manifest()` parses the response body before persisting
it. Non-JSON payloads raise class `uscogdata_invalid_manifest` (URL,
Content-Type, parse error) and never touch the on-disk cache.
3. Cache writes are atomic via a sibling tempfile + `file.rename`, and
existing caches with non-JSON content are silently refetched instead
of returning a parse error to the caller.
Local-path manifests that aren't valid JSON now surface the same
`uscogdata_invalid_manifest` class with file context.
R CMD build (run by rcmdcheck before R CMD check) rebuilds vignettes
from source regardless of --no-vignettes. Our vignette declared
%\VignetteEngine{knitr::knitr} which requires the 'markdown' package
that isn't in CI's dependency tree. Switch to %\VignetteEngine{knitr::rmarkdown},
which uses the already-Suggests-listed 'rmarkdown' package and matches
the output: rmarkdown::html_vignette directive in the YAML header.
Also add ^\.gitea$ and ^CLAUDE\.md$ to .Rbuildignore so R CMD check
stops emitting the "hidden file" / "non-standard top-level file" notes.
Local rcmdcheck (mirroring CI's exact args) now reports
0 errors / 0 warnings / 0 notes.
The pdflatex notice in the build log is unrelated — R CMD build prints
"Not building PDF manual" and continues; with --no-manual it's silenced
entirely.
Final-review followups:
1. cog_explain rendered the popyear range as raw 2-digit values
("popyear range: 19-20"), which a user could read as the years 19-20.
Added .expand_popyear() helper to format as 4-digit calendar years
(2019-2020). Pivot at 70 to handle pre-2000 vintages if the corpus
ever extends backward.
2. cog_peer_compare provenance was missing pop_range and is_ratio,
omitted from the spec-required reproducibility metadata.
cog_find_peers now stamps both as tibble attributes; cog_peer_compare
reads them through to provenance$pop_range and provenance$is_ratio.
Test coverage extended: explain test asserts the 4-digit format and
rejects the old 2-digit form; peer-compare test asserts pop_range +
is_ratio propagate end-to-end.
326 PASS / 0 FAIL.
Explains the four population sources, why F-33 is the default, type-4/5
coverage gap, the popyear quirk, and how to build moving-window peer
cohorts manually.
Updates transformations\$per_capita with the new denominator_source string,
popyear_range, and pop_source_counts. .attach_per_capita stashes
popyear_range on the result; .verb_spendrev strips the helper attr after
provenance is built.
cog_geographic_rollup(per_capita = TRUE) now drops rows whose government
has no observed population for that year (pop_source == 'unavailable'),
matching the spec's exclusion rule. Records included/excluded govids in
provenance$rollup.
Reads attr(peers, 'cohort_year') when the caller passed a cog_find_peers()
tibble; NA when the caller passed a bare character vector. Stamped as a
constant column on the result and recorded in provenance alongside the
cohort govids.
Adds optional 'year' argument (defaults to most recent observed year for
the target). Filters and ranks candidates by gov_population_yearly.population
at that year. Returned column renamed population_acs -> population.
Cohort year attached as attr(x, 'cohort_year').
Adds .resolve_cohort_year() helper. Updates test assertions to use
'population' column name. Regenerates man/cog_find_peers.Rd.
Three RED tests that drive Task 6's cog_find_peers() rewrite:
- defaults year to most recent observed (expects cohort_year attr + population column)
- honors explicit year= argument (expects cohort_year attr)
- errors with "no observed population" for unobserved year
The plan-supplied predicate combined three redundant checks
(is.null + any + %in% TRUE). Element-wise behavior was correct via
scalar recycling, but the form was confusing — a code-quality reviewer
misread it as a multi-row false-positive bug. Simplify to mirror the
parts[[2]] structure: gate on column presence, then element-wise
%in% TRUE check. Equivalent semantics, fewer ways to misread.
.notes_column now joins multiple per-row notes with '; '. Adds the
'No population denominator available for this gov type' note when
pop_source is 'unavailable'. Gracefully handles absent pop_source
(per_capita = FALSE). Two new tests: one corpus-level (census_f33
branch) and one synthetic unit test covering multi-note concatenation.
cog_spending(per_capita = TRUE) and cog_revenue(per_capita = TRUE) now
divide each year's amount by that gov-year's population from
gov_population_yearly (drawn from long.population) instead of a single
static ACS 2018-2022 value. Adds pop_source column with values
'census_f33' or 'unavailable'.
Adds a RED test asserting that cog_spending(per_capita = TRUE) divides
by the per-year F-33 population (Broward 2019: 1,935,878; 2020: 1,952,778)
rather than the static ACS value (1,940,907). Uses absolute-tolerance
expect_true(abs(...) < 1) instead of expect_equal(tolerance=1) because
testthat 3 treats the tolerance argument as relative.
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:
- Salt Lake County now resolves to 20 distinct cities/townships instead
of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
a separate gap, documented in cog_pipeline).
manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.
All 283 uscogdata tests pass against the new fixture.
15-task TDD plan covering: gov_population_yearly view, .attach_per_capita
per-year join, pop_source column + multi-note concatenation, cog_find_peers
year arg, cog_peer_compare cohort_year, rollup unavailable-pop exclusion,
provenance updates, cog_explain rendering, vignette, cog_pipeline data
dictionary, and NEWS entry.
Also corrects spec to match existing rollup semantics (side-by-side, not
summed) and adds /plans to .Rbuildignore.
Existing function takes a peers argument (caller supplies cohort) rather
than building one internally. Cohort year flows through via an attr on
the peers tibble produced by cog_find_peers.
Design doc for switching cog_spending / cog_revenue / cog_geographic_rollup
per-capita calculations from a static ACS 2018-2022 population to per-year
F-33 population already present in long.population. Covers shifting peer
matching to a user-selectable cohort year (defaulting to most recent
observed year), type-4/5 NA policy, provenance updates, and a new vignette
enumerating denominator sources for future extensibility.
Specs live in /specs (added to .Rbuildignore) since docs/ is reserved for
pkgdown output.
Cross-task review found two edge cases that violated the basket-mode
soft-fail contract:
- Per-row excluded type (e.g. type = c(NA, "special_district")) hit
.coerce_type()'s abort inside the per-row resolver, killing the
whole basket call. Now treated as no_match in the sidecar.
- Malformed regex in the substring fallback (e.g. name = "San(Diego")
propagated DuckDB engine errors. .escape_regex() now backslash-
escapes meta characters before the regexp_matches call. Utility-
mode regex behavior is unchanged.
Plus a new public-surface test for the all-no-match case.
Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.
Sidecar accessors for basket-mode results. cog_basket_resolution()
returns the full resolution tibble (drops candidates list-col by
default for readable printing). cog_basket_unresolved() filters to
ambiguous/no_match rows for iterative refinement.
Moves .build_sidecar(), .type_to_label(), .basket_summary_message()
out of R/search.R and into a new R/basket.R. No behavior change;
brings R/search.R back under its 350-line budget. R/basket.R will
also host the public sidecar accessors added in the next commit.
Vector name + state + type arguments dispatch to a per-row resolver
that produces a basket tibble with a 'resolution' sidecar attribute.
Utility mode (length-1 name) is unchanged.
Multi-row matches within a single govs_type pick the largest-population
row (status=largest_pop). Multi-row matches spanning >=2 types return no
basket row (status=ambiguous) with all candidates preserved for the
sidecar.
.resolve_basket_row() now falls back to case-insensitive substring
match when no exact match is found, and short-circuits empty/whitespace
input to no_match. Disambiguation stub raises pending Task 5.
Pre-rename in preparation for basket mode. All existing callers in this
package and cog_explorer/ pass the first argument positionally, so this
rename is non-breaking. No deprecation alias added per design spec
(no external consumers; package is pre-release v0.1.0).
@param roxygen also updated to match the new formal; man/cog_gov_search.Rd
regenerated via devtools::document().
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.
setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).
helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.
test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).
Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
Discovery verb over the summary_categories view, grouped one row per
(category, subtype). Parallels cog_gov_search: analysts use it to
find the valid `category` values to pass into cog_spending(),
cog_revenue(), cog_geographic_rollup().
Columns: category, category_type, subtype, n_codes, item_codes
(comma-separated, alphabetical). Optional filters:
type = NULL | "spending" | "revenue"
pattern = regex matched case-insensitively on category
The user-facing "spending" alias is translated internally to the
corpus-native "expenditure" so callers don't have to learn Census
vocabulary, while the returned category_type column preserves the
native value for auditability.
Also: fix @noRd placement in session.R so devtools::document() stops
warning.
Tests: +16 new / 181 total pass. check 0E/0W/0N.
Two UX fixes surfaced by first real-user use:
1. cog_spending / cog_revenue / cog_geographic_rollup now accept either
a character vector OR a data.frame with a canonical_govid column
(e.g. output of cog_gov_search() or cog_find_peers()). Shared
.coerce_govid_input() helper in session.R. This lets the natural
pipe work:
cog_gov_search('MIAMI', state='FL', type='city') |>
cog_spending(years=2022, category='Police')
cog_peer_compare already accepted a data.frame for the peer arg;
behavior there is unchanged.
2. .check_govids_in_scope() message reworded. The old text led with
'v0.1 covers gov_types 0-3' which falsely implied the missing govids
were scope-excluded types when the more common real cause is a typo
or a guessed value. New message leads with typo + pre-2017 PID,
mentions scope exclusion as one possibility, and points at
cog_gov_search() as the recovery path.
Tests: 165 pass / 0 fail. check 0E/0W/0N.