docs: roxygen, NEWS, and pkgdown for basket mode

Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.
This commit is contained in:
2026-04-28 12:51:52 -04:00
parent 7475696853
commit 24e67791a0
7 changed files with 211 additions and 32 deletions
+10
View File
@@ -23,3 +23,13 @@ Returns the resolution tibble attached to a basket-mode result of
the `candidates` list-column is dropped for readable printing; pass
`expand_candidates = TRUE` to keep it.
}
\examples{
\dontrun{
basket <- cog_gov_search(
name = c("Broward", "San Diego", "Notarealplace"),
state = c("FL", "CA", "NY")
)
cog_basket_resolution(basket)
cog_basket_resolution(basket, expand_candidates = TRUE)
}
}
+9
View File
@@ -18,3 +18,12 @@ Convenience wrapper that returns just the rows where `status` is
refine before piping into a query verb. The `candidates` list-column
is preserved so the user can drill into ambiguous match sets.
}
\examples{
\dontrun{
basket <- cog_gov_search(
name = c("Broward", "San Diego", "Notarealplace"),
state = c("FL", "CA", "NY")
)
cog_basket_unresolved(basket)
}
}
+78 -14
View File
@@ -7,24 +7,88 @@
cog_gov_search(name = NULL, state = NULL, type = NULL)
}
\arguments{
\item{name}{Character regex matched case-insensitively against
`gov_name`. `NULL` (default) means no name filter.}
\item{name}{Character vector of place name(s). Length 1 = utility mode;
length >1 = basket mode.}
\item{state}{Either a 2-letter USPS abbreviation (e.g. `"FL"`), a FIPS
integer (e.g. `12`), or `NULL`.}
\item{state}{2-letter USPS abbreviation (e.g. `"FL"`), FIPS integer
(e.g. `12`), or `NULL`. In basket mode, length 1 recycles across
all entries; otherwise must match `length(name)`.}
\item{type}{Government type: an integer in `0:3` or one of `"state"`,
`"county"`, `"city"`, `"township"`. Passing `4`, `5`,
`"special_district"`, or `"school_district"` emits an explanatory
message and returns an empty tibble (v0.1 corpus excludes those types).}
\item{type}{Government type: integer in `0:3` or one of `"state"`,
`"county"`, `"city"`, `"township"`, or `NA`/`NULL`. Per-row optional
in basket mode (recycles from length 1). Excluded types `4`/`5` (or
`"special_district"` / `"school_district"`) trigger an explanatory
message and an empty result.}
}
\value{
Tibble from `canonical_fips_xwalk` sorted by `population_acs`
descending (`NULL`s last).
A tibble of `canonical_fips_xwalk` rows. In utility mode, all
matches sorted by `population_acs` desc. In basket mode, resolved
rows in input order, with `attr(., "resolution")` set to the
sidecar tibble.
}
\description{
Returns rows from `canonical_fips_xwalk` matching the supplied filters.
Intended as the entry point users call to resolve a human-readable place
name into one or more `canonical_govid` values before calling
[cog_spending()] / [cog_revenue()] / etc.
Resolves human-readable place names into rows of `canonical_fips_xwalk`,
the cross-vintage canonical-government registry. Operates in two modes:
}
\details{
* **Utility mode** (single `name`, the original behavior): returns all
rows whose `gov_name` matches the regex case-insensitively, sorted by
`population_acs` descending. Useful for exploratory lookups.
* **Basket mode** (`length(name) > 1`): resolves each input row to a
single canonical govid and returns a tibble in input order, suitable
for piping straight into [cog_spending()] / [cog_revenue()] /
[cog_geographic_rollup()]. Carries an audit sidecar accessible via
[cog_basket_resolution()] / [cog_basket_unresolved()].
**Basket-mode resolution algorithm** (per input row):
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
2. **Exact pass:** case-insensitive equality against `gov_name`.
Single hit -> resolved. Multiple -> step 4.
3. **Substring fallback:** case-insensitive regex against `gov_name`.
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
`status = "no_match"`. Multiple hits -> step 4.
4. **Disambiguation:** if matches share one `govs_type`, pick the
largest-population row (`status = "largest_pop"`). If they span >=2
types, no row is added (`status = "ambiguous"`); the user should
re-run with `type` specified.
Resolved rows form the returned tibble in input order. Unresolved
inputs (`ambiguous` / `no_match`) appear only in the sidecar.
}
\examples{
\dontrun{
# Utility mode — exploratory regex lookup
cog_gov_search("broward", state = "FL")
# Basket mode — resolve a known cohort
basket <- cog_gov_search(
name = c("BROWARD COUNTY", "SAN DIEGO CITY", "AUSTIN CITY"),
state = c("FL", "CA", "TX")
)
basket
# Inspect resolution audit
cog_basket_resolution(basket)
# Pipe into a spending query
library(dplyr)
basket |> cog_spending(years = 2019:2020, category = "Police")
# Iteratively refine ambiguous matches
partial <- cog_gov_search(
name = c("Broward", "San Diego"), # San Diego is ambiguous
state = c("FL", "CA")
)
cog_basket_unresolved(partial)
refined <- cog_gov_search(
name = c("Broward", "San Diego"),
state = c("FL", "CA"),
type = c(NA, "city") # disambiguate
)
}
}
\seealso{
[cog_basket_resolution()], [cog_basket_unresolved()],
[cog_spending()], [cog_revenue()].
}