Files
uscogdata/man/cog_gov_search.Rd
T
jared d006dea6e4
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00

99 lines
3.6 KiB
R

% Generated by roxygen2: do not edit by hand
% Please edit documentation in R/search.R
\name{cog_gov_search}
\alias{cog_gov_search}
\title{Search for governments by name, state, and/or type}
\usage{
cog_gov_search(name = NULL, state = NULL, type = NULL)
}
\arguments{
\item{name}{Character vector of place name(s). Length 1 = utility mode;
length >1 = basket mode.}
\item{state}{2-letter USPS abbreviation (e.g. `"FL"`), FIPS integer
(e.g. `12`), or `NULL`. In basket mode, length 1 recycles across
all entries; otherwise must match `length(name)`.}
\item{type}{Government type: integer in `0:3` or one of `"state"`,
`"county"`, `"city"`, `"township"`, or `NA`/`NULL`. Per-row optional
in basket mode (recycles from length 1). Excluded types `4`/`5` (or
`"special_district"` / `"school_district"`) trigger an explanatory
message and an empty result.}
}
\value{
A tibble of `canonical_fips_xwalk` rows. In utility mode, all
matches sorted by `population_acs` desc. In basket mode, resolved
rows in input order, with `attr(., "resolution")` set to the
sidecar tibble.
}
\description{
Resolves human-readable place names into rows of `canonical_fips_xwalk`,
the cross-vintage canonical-government registry. Operates in two modes:
}
\details{
* **Utility mode** (single `name`, the original behavior): returns all
rows whose `gov_name` contains `name` as a **literal, case-insensitive
substring**, sorted by `population_acs` descending. Useful for
exploratory lookups. Regex metacharacters in `name` are escaped, so a
government is findable by its own complete name even when that name
contains parentheses or a period.
* **Basket mode** (`length(name) > 1`): resolves each input row to a
single canonical govid and returns a tibble in input order, suitable
for piping straight into [cog_spending()] / [cog_revenue()] /
[cog_geographic_rollup()]. Carries an audit sidecar accessible via
[cog_basket_resolution()] / [cog_basket_unresolved()].
**Basket-mode resolution algorithm** (per input row):
1. Filter `canonical_fips_xwalk` by `state` and (if non-NA) `type`.
2. **Exact pass:** case-insensitive equality against `gov_name`.
Single hit -> resolved. Multiple -> step 4.
3. **Substring fallback:** case-insensitive literal substring against
`gov_name` (metacharacters escaped).
Single hit -> resolved (`match_method = "substring"`). Zero hits ->
`status = "no_match"`. Multiple hits -> step 4.
4. **Disambiguation:** if matches share one `govs_type`, pick the
largest-population row (`status = "largest_pop"`). If they span >=2
types, no row is added (`status = "ambiguous"`); the user should
re-run with `type` specified.
Resolved rows form the returned tibble in input order. Unresolved
inputs (`ambiguous` / `no_match`) appear only in the sidecar.
}
\examples{
\dontrun{
# Utility mode — exploratory substring lookup
cog_gov_search("broward", state = "FL")
# Basket mode — resolve a known cohort
basket <- cog_gov_search(
name = c("BROWARD COUNTY", "SAN DIEGO CITY", "AUSTIN CITY"),
state = c("FL", "CA", "TX")
)
basket
# Inspect resolution audit
cog_basket_resolution(basket)
# Pipe into a spending query
library(dplyr)
basket |> cog_spending(years = 2019:2020, category = "Police")
# Iteratively refine ambiguous matches
partial <- cog_gov_search(
name = c("Broward", "San Diego"), # San Diego is ambiguous
state = c("FL", "CA")
)
cog_basket_unresolved(partial)
refined <- cog_gov_search(
name = c("Broward", "San Diego"),
state = c("FL", "CA"),
type = c(NA, "city") # disambiguate
)
}
}
\seealso{
[cog_basket_resolution()], [cog_basket_unresolved()],
[cog_spending()], [cog_revenue()].
}