Commit Graph
14 Commits
Author SHA1 Message Date
jared 0fbae00e27 feat: name a cohort by state/type predicate instead of a 40k-id IN list
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00
jared d006dea6e4 fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
R-CMD-check / check (push) Failing after 3m4s
R-CMD-check / check (pull_request) Failing after 3m4s
The three kodor/fix issues, taken over after a day with no branch, PR or
comment on any of them. Batched because each is single-file with a committed
acceptance test, and two share documentation surfaces.

#16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight
into regexp_matches() unescaped, while basket mode in the same file already
routed it through .escape_regex() with the comment "so `name` is treated as
a literal substring". Two failure modes, both HTTP 200 through the API:
a government could not be found by its own complete name when that name
contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a
bare "." matched all 608 Wisconsin cities. Malformed pattern text reached
the engine as an error, which cog-api surfaced as a 500 -- reachable by
typing a real name one character at a time ("Athens-Clarke County (bal").

Utility mode now calls the escaper that already existed. Roxygen updated:
utility mode is documented as a literal case-insensitive substring match,
and the basket-mode "substring fallback" step no longer describes itself as
a regex either.

  BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop
  working, because there is no regex left to anchor. Two existing tests used
  "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now
  search for those characters literally. Updated to the bare names, which
  still resolve to exactly one row each once scoped by state/type (verified,
  not assumed). There is no exact-match option in utility mode any more --
  noted on the issue, since that is a real if small capability loss.

#15 (F-004) -- the raw Census files report thousands of dollars; this
package multiplies by 1000 and returns full US dollars. Correct, and already
stated in ?cog_spending / ?cog_revenue @return, in provenance, and in
cog-api's data-dictionary. Absent from every surface a reader meets FIRST.
Added to README.md as its own section and to both vignettes' openings.

The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule
("All raw `amt` values are in $1,000s") without scoping it to the raw column
-- a reader applying that to amt_nominal overstates by 1000x and gets a
plausible-looking number rather than an obvious error. Fixed there too; that
directory has no git remote, so it rides in no PR and is left uncommitted
for the owner.

#14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately
inside each (year, spend_subtype, category) cell, so a summary_p50 row is
"the median peer's value in that one category", never "the value of the
median peer's total" -- the median peer for Police and for Fire are usually
different governments. Summing them across categories misstated a
total-spending band by -32.7% to +251.0% across 24 years, with a sign flip
at FY2012. The verb is right and its documented use (facet by role AND
category) is unaffected, so the fix is @return prose plus a worked snippet
showing the correct computation: sum each peer's own categories first, then
take the quantile of those per-government totals.

This is the R-side counterpart of cog-api#9, fixed on the API surface
earlier today; the wording is deliberately consistent across the two.

Note the phrase "not additive" has to stay on one roxygen source line --
the test greps the generated Rd, where a line wrap turns it into
"not   additive" and stops matching. Cost one red run to find.

man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so
cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was
version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with
no content change. The two Rd files kept carry only the edits above.

Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips
are #11, #12 and #13.
2026-07-30 11:23:53 -04:00
jared e635a1fc9e feat!: require corpus schema_version 4 (Phase P canonical ids)
BREAKING CHANGE: canonical_govid is now uniformly 12 characters across
every vintage; corpora built against schema_version 3 are rejected.
Bumps MinCorpusSchema/MaxCorpusSchema to 4 and expected_version in
cog_open(). canonical_fips_xwalk grows to the 14-column Phase P master
schema (adds legacy_govs_id, census_geoid, id_source; confidence is
renamed to pop_confidence); .empty_xwalk_tibble() is rewritten to match.
2026-07-11 09:31:33 -04:00
jared efc0bd16b1 fix(search): soft-fail on per-row excluded type and malformed regex name
R-CMD-check / check (push) Successful in 1m37s
R-CMD-check / check (pull_request) Successful in 1m37s
Cross-task review found two edge cases that violated the basket-mode
soft-fail contract:

- Per-row excluded type (e.g. type = c(NA, "special_district")) hit
  .coerce_type()'s abort inside the per-row resolver, killing the
  whole basket call. Now treated as no_match in the sidecar.
- Malformed regex in the substring fallback (e.g. name = "San(Diego")
  propagated DuckDB engine errors. .escape_regex() now backslash-
  escapes meta characters before the regexp_matches call. Utility-
  mode regex behavior is unchanged.

Plus a new public-surface test for the all-no-match case.
2026-04-28 14:16:37 -04:00
jared 24e67791a0 docs: roxygen, NEWS, and pkgdown for basket mode
Adds full @description, @details (algorithm), @examples on
cog_gov_search(); examples on cog_basket_resolution() /
cog_basket_unresolved(); NEWS.md entry covering the new mode and
the pattern->name rename; pkgdown reference entries for the two
new exports.
2026-04-28 12:51:52 -04:00
jared ae04d4f48f refactor(search): split sidecar helpers into R/basket.R
Moves .build_sidecar(), .type_to_label(), .basket_summary_message()
out of R/search.R and into a new R/basket.R. No behavior change;
brings R/search.R back under its 350-line budget. R/basket.R will
also host the public sidecar accessors added in the next commit.
2026-04-28 11:53:40 -04:00
jared d9bf0552b7 feat(search): post-resolution summary message in basket mode
Emits a single cli::cli_inform when any input was ambiguous, missed,
or fell back to largest-population. Silent on clean baskets.
2026-04-28 11:41:28 -04:00
jared 98ed6318a6 feat(search): basket mode for cog_gov_search()
Vector name + state + type arguments dispatch to a per-row resolver
that produces a basket tibble with a 'resolution' sidecar attribute.
Utility mode (length-1 name) is unchanged.
2026-04-28 11:39:53 -04:00
jared 2f47f7ae67 feat(search): add disambiguation — largest_pop and ambiguous branches
Multi-row matches within a single govs_type pick the largest-population
row (status=largest_pop). Multi-row matches spanning >=2 types return no
basket row (status=ambiguous) with all candidates preserved for the
sidecar.
2026-04-28 11:08:26 -04:00
jared 392e818f52 feat(search): add substring fallback and no_match handling
.resolve_basket_row() now falls back to case-insensitive substring
match when no exact match is found, and short-circuits empty/whitespace
input to no_match. Disambiguation stub raises pending Task 5.
2026-04-28 11:06:30 -04:00
jared 4dcf1c72b2 feat(search): add per-row resolver — exact match branch
.resolve_basket_row() handles the exact-match case. Substring fallback
and disambiguation branches follow in subsequent commits.
2026-04-28 11:05:14 -04:00
jared 7fd29c919d feat(search): add basket-mode argument validator
Internal .validate_basket_args() handles length validation and
recycling of state/type from length 1. Foundation for basket mode.
2026-04-28 11:02:03 -04:00
jared 68b72faeae refactor(search): rename cog_gov_search() first argument to name
Pre-rename in preparation for basket mode. All existing callers in this
package and cog_explorer/ pass the first argument positionally, so this
rename is non-breaking. No deprecation alias added per design spec
(no external consumers; package is pre-release v0.1.0).

@param roxygen also updated to match the new formal; man/cog_gov_search.Rd
regenerated via devtools::document().
2026-04-28 10:47:56 -04:00
jared 377eed1240 feat: cog_gov_search + cog_mirror + scope-aware verb behavior
Three pieces:

1. cog_gov_search: name/state/type search over canonical_fips_xwalk
   for resolving human-readable place names into canonical_govids.
   Accepts USPS abbrev ('FL') or FIPS int (12) for state; integer
   0-3 or name ('state','county','city','township') for type. Types
   4/5 emit an explanatory cli message and return an empty tibble
   (v0.1 corpus excludes them). USPS<->FIPS table hardcoded with
   50 states + DC + territories; FIPS 66 = GU (not GA).

2. cog_mirror: downloads manifest-listed files to a local directory
   with SHA-256 idempotency (files with matching hash return status
   'cached'). Supports HTTP and local-path fixture URLs. Round-trip
   test: mirror + re-open against the mirror + query Broward 2020
   returns identical results.

3. Scope-aware verbs: .check_govids_in_scope() helper in session.R
   queries canonical_fips_xwalk for the requested govids, emits a
   cli_inform listing any missing ones, and records the found/missing
   sets under provenance$scope. Wired into cog_spending (and
   transitively into cog_revenue, cog_geographic_rollup,
   cog_peer_compare via their cog_spending calls).

Also: dropped dbplyr from Imports (unused).

Tests: +29 (22 search + 12 mirror - 5 refactored) / 159 total pass.
devtools::check() now clean: 0E / 0W / 0N.
2026-04-24 11:38:32 -04:00