Commit Graph
212 Commits
Author SHA1 Message Date
jaredandClaude Sonnet 5 9adb921170 docs: record session (#36)
Mirror to GitHub / mirror (push) Successful in 7s
R-CMD-check / check (push) Successful in 3m37s
Journal entries for the #33/#34 split and the #36 coverage counter
(two bugs caught before merge in the latter), plus the vignette
obligation filed as #72. Status board regenerated and republished.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 13:38:25 -04:00
jaredandClaude Sonnet 5 9f5cd988b2 docs(suggestions): restore dropped M/L-counterpart rationale (roborev job 224)
Mirror to GitHub / mirror (push) Successful in 7s
R-CMD-check / check (push) Successful in 3m46s
The #33 decomposition (0c7c7eb) silently dropped four roxygen lines
explaining why .attach_ig_counterparts()'s second flow-family check is
needed: condition 1 alone doesn't block ig_federal_b47_wide under
cog_revenue(), since its own "B" IS inside revenue's own flow_prefixes.
Restored them, plus the backticks around "B" that were also dropped as
an unstated formatting change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 13:35:18 -04:00
jared 9eaa759ccb Merge pull request 'feat(coverage): n_units_collected separates sampling from real zeros (#36)' (#71) from fix/coverage-collected-36 into main
Mirror to GitHub / mirror (push) Successful in 8s
R-CMD-check / check (push) Successful in 3m57s
2026-09-09 11:51:06 -04:00
jaredandClaude Sonnet 5 b41d5ee2aa feat(coverage): n_units_collected separates sampling from real zeros (#36)
R-CMD-check / check (push) Successful in 3m39s
R-CMD-check / check (pull_request) Successful in 3m42s
provenance$coverage's n_units_reporting is category-conditional: it
counts governments with rows for the SPECIFIC requested category, which
conflates two different things -- a government never collected that
year (sampling), and one collected but genuinely spending nothing in
that category (a real zero). FY2012 Georgia Police is the motivating
case from the issue: a complete census year reads as a 69% "response
rate" because most of the gap is cities that contract policing to the
county sheriff, not non-response.

Adds a second counter, n_units_collected: how many of the caller's
expected cohort appear in the corpus that year for ANY category.
n_units_collected / n_units_expected is the true collection rate;
n_units_reporting / n_units_collected is category participation among
collected units. cog_geographic_rollup() and cog_peer_compare() both
carry it; cog_explain() prints it alongside n_units_reporting.

Two real bugs caught and fixed while finishing this (both against the
already-written, previously-uncommitted draft):

- .coverage_table()'s candidate list for the collection query was
  derived from the category-filtered result rows, not the caller's
  full expected cohort. A government with zero rows in the requested
  category across every requested year never appears in that result,
  so it was silently excluded from n_units_collected too -- collapsing
  the new counter back to the old, broken one for exactly the
  governments it exists to count. Fixed by threading an explicit
  `expected_ids` (all_govids / peer_govids) through instead.
- The collection query hardcoded long_view = "spending_long_harmonized",
  which does not exist on a corpus with schema_version < 5 (R/basis.R
  resolves basis = "raw" there; R/views.R only registers the
  harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare()
  would hard-error on a corpus vintage the package otherwise explicitly
  supports. Fixed by deriving long_view from the basis cog_spending()
  actually resolved (prov$basis) via the existing .select_long_view()
  helper, matching how every other basis-aware query in the package
  already does this.

Also: cog_explain()'s general "complete census only in years ending in
2 or 7" footnote was gated on the OLD counter's absence, making it
permanently unreachable now that both callers always supply the new
one -- ungated it, since the explanation is orthogonal to which
counter set is present. Dropped a dead conditional branch, fixed two
stale roxygen blocks in R/peers.R/R/rollup.R still describing the old
two-counter model, fixed the same staleness in README.md, and switched
two `uscogdata:::` self-references to the package's own convention of
calling internal helpers unqualified.

1101 tests pass (2 skipped live-corpus), including new direct
regression tests for both bugs above (one exercising a government
collected-but-absent from a category result, one running the full
rollup/peer-compare path against a doctored schema_version 4 corpus).

Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW
-- all addressed above).

Closes #36.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:46:17 -04:00
jaredandClaude Sonnet 5 d0d724c4c8 fix(suggestions): guard empty candidates and restore year typing in .query_covered_years()
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m15s
Addresses roborev jobs 205/208 (reviewing pre-merge draft/original commits
of #33/#34, now split and merged as #69/#70):

- Restore `res$year <- as.integer(res$year)` in .query_covered_years(),
  dropped during the #33 rebuild as apparently-dead code. Its own roxygen
  promises an integer `year` column; without the coercion the
  implementation no longer matches that documented contract.
- Guard .query_covered_years() against empty `candidates`, mirroring the
  existing `gap_years` guard. Not reachable via .build_suggestions() today
  (candidates is checked non-empty before this is called), but the
  extracted helper is independently callable and previously built a
  malformed `WHERE r.recipe_id IN ()` clause for a hypothetical direct
  caller with no candidates.
- Add a direct fixture-only unit test of .query_candidate_recipes()'s
  category_type filtering (#34) in both flow directions -- queries only
  summary_categories/harmonization_recipes metadata, so it runs without
  skip_if_no_corpus(), unlike the two existing end-to-end tests.

1080 tests pass (2 skipped live-corpus).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:28:01 -04:00
jared 56f610ea3f Merge pull request 'fix(suggestions): scope candidate recipes by category_type (#34)' (#70) from fix/category-type-34 into main
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m36s
2026-09-09 11:18:52 -04:00
jared 62741343ee Merge pull request 'refactor(suggestions): decompose .build_suggestions() into named helpers (#33)' (#69) from issue-33 into main
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m33s
2026-09-09 11:18:01 -04:00
jaredandClaude Sonnet 5 392643bd74 fix(suggestions): scope candidate recipes by category_type (#34)
R-CMD-check / check (push) Successful in 4m15s
R-CMD-check / check (pull_request) Successful in 4m25s
.query_candidate_recipes() (extracted in #33) now filters candidates by
category_type ('expenditure' vs 'revenue'), derived from the calling
verb's own flow_prefixes (E/F/G -> 'expenditure', else 'revenue').

Without this, a category shared across both flow families in
summary_categories leaked cross-family recipes: cog_revenue(category =
"Corrections") surfaced the expenditure-only corrections_combined recipe
(E04/E05) merely because "Corrections" is also a spending category name,
and cog_spending(category = "IG Federal") surfaced the revenue-only
ig_federal_b47_wide recipe. Both are wrong: following either hint would
attribute dollars to the wrong flow, or (IG Federal) fire the
coverage-gap machinery for a category the calling verb structurally
cannot report on at all.

Updates the two tests this changes the expected behavior of:
- "a mis-scoped cog_spending() call never attaches an M/L counterpart to
  a revenue-flavored recipe" (test-expenditure-concept.R): IG Federal is
  revenue-only, so a spending call now finds zero candidates outright
  rather than firing the suggestion and then blocking its M/L
  counterpart as a second-order check.
- "cog_revenue never suggests expenditure-only recipes"
  (test-recipes.R, was "I1: ... never fabricates suppressed dollars"):
  corrections_combined is expenditure-only, so a revenue call now never
  considers it as a candidate, rather than considering it and reporting
  zero suppressed dollars.

All 1078 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree stacked on the #33 refactor.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:11:16 -04:00
jaredandClaude Sonnet 5 0c7c7eb299 refactor(suggestions): decompose .build_suggestions() into named helpers (#33)
R-CMD-check / check (pull_request) Successful in 4m12s
R-CMD-check / check (push) Successful in 4m20s
Extract three functions from the ~140-line .build_suggestions()
orchestrator to comply with the 'functions under 50 lines' convention:

- .query_candidate_recipes(): candidate recipe lookup by category/subtype
  scope, plus M/L self-exclusion
- .query_recipe_meta(): metadata lookup for labels and year spans
- .query_covered_years(): Path 1 gap-year coverage query via the recipe's
  own generic join; returns empty data frame when gap_years is empty

Kept inline per design: the for-loop that merges covered-years +
suppressed-components into suggestion objects, the M/L-exclusion comment
block as call-site rationale, and .attach_ig_counterparts() at the end.

Pure extraction, no behavior change -- SQL text is unchanged apart from
whitespace. Restored real multi-line SQL string literals in the two new
helpers (the original candidate/covered-years queries were written that
way; keep it consistent with .query_recipe_meta()) and normal roxygen
'#'' comment-marker spacing throughout, both of which drifted during
extraction in an earlier pass.

All 1084 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:09:01 -04:00
jared 7274ce3bfe chore: correct the roborev exclusions and refresh the guidelines
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m46s
roborev matches excluded_commit_patterns as substrings, and compass's cadence rule
requires `type(ws): subject (#N)`, so `chore:` never matched `chore(engine): ...`.
Adds `chore(`.

Does not add `docs(`. Those commits carry the journal entry and the board
narrative, and the plain-language rule exists to check exactly that prose -- it has
no other commit to fire on.

The guidelines were also stale: they were composed before base.md gained the
plain-language rule, and nothing re-composes them on its own. Refreshed, which is
what put that rule in this repository for the first time.
2026-08-23 23:57:05 -04:00
jared 24e86ed598 chore: retire the three plans whose work has shipped
Mirror to GitHub / mirror (push) Successful in 6s
R-CMD-check / check (push) Successful in 3m21s
186 unticked checkboxes across three plans, none of them outstanding work.
Superpowers-style plans are execution transcripts: nobody ticks the boxes, and
the plan is abandoned at the point the work is done. Left in place they are
indistinguishable from a live backlog -- the old compass retrofit rule, "open
checkboxes become issues", would have filed 186 issues for finished work.

Evidence, from scripts/plans.py plus a check by hand:

  2026-08-04-partial-coverage-signposting  NEWS: "Coverage signposting ..."
  2026-08-08-public-release                NEWS: "First public release."
  2026-04-29-per-year-population-denominator
      no NEWS line matched and 19 of the 21 files it names exist, so the
      classifier called it ambiguous. Confirmed shipped by hand:
      33c0274 docs(news): per-year population denominators, plus the feat
      commits behind it.

specs/ is untouched. A spec explains why the design is what it is and stays
useful; a plan is scaffolding, and once the building stands it is in the way.
All three remain in git history.
2026-08-23 23:32:11 -04:00
jared 1ec20174b7 chore: move compass out of docs/, which pkgdown deletes
Mirror to GitHub / mirror (push) Successful in 7s
R-CMD-check / check (push) Successful in 3m28s
compass.toml, JOURNAL.md, STATUS.md and decisions/ were sitting inside pkgdown's
output directory. Asked directly, pkgdown listed docs/pm and docs/decisions among
the 28 top-level entries clean_site() would delete, and the guard that would have
refused -- check_dest_is_pkgdown() -- was satisfied by docs/pkgdown.yml. After the
move it lists 26 and none of them are compass's.

Nothing was lost. The journal had no entries and there were no decision records
yet, so this was the cheapest moment to move.

The .gitignore workaround goes with it. Re-including two children of an excluded
docs/ forced the rule to be written as /docs/* plus two negations, which changed
the anchoring and made the fixture corpus's own docs/ need a separate rule. A bare
docs/ matches at any depth again, so both are unnecessary.

.Rbuildignore gains ^pm$ -- R CMD check flags a non-standard top-level directory.

Compass reads both layouts, so this repository worked either way; the point is
that docs/ is a directory another tool empties.
2026-08-23 23:26:33 -04:00
Jared Knowles 2e317d3a0f chore: adopt compass for project tracking
Mirror to GitHub / mirror (push) Successful in 10s
R-CMD-check / check (push) Successful in 3m55s
Three workstreams -- the query verbs, the corpus and how it is mounted, and the
docs that explain both. Each can go stale independently, which is what the
workstream boundary is for: a change to a verb obligates the vignettes, a change
to the corpus obligates NEWS.

All six open issues now carry ws/ and type/ labels, applied additively so #36
kept its existing kodor/, severity/, south-guide and verdict/ labels. The
generated board is pinned as issue #68.

.roborev.toml carries review guidelines composed from a shared baseline, the R
package overlay, and this project's own conventions read out of CLAUDE.md: the
ensure_session-then-dbGetQuery order, the tbl_df-with-provenance return
contract, coerce_govid_input at the boundary, the two SQL layers, no arrow, and
withr as tests-only.

The .gitignore change is load-bearing. pkgdown output made docs/ ignored, which
would have left every compass file untracked and unable to travel to another
machine. Git cannot re-include anything beneath an excluded directory, so the
rule had to list children instead. That forced an anchoring change: a bare
"docs/" matches at any depth, "/docs/*" only at the root, so the fixture corpus
docs directory needed an explicit rule to stay excluded as before.
2026-08-23 16:17:20 -04:00
jared f7c606984f Merge pull request 'ci: mirror canonical tag objects, not the lightweight refs checkout builds' (#67) from ci/mirror-canonical-tags into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m14s
Reviewed-on: #67
2026-08-11 10:03:43 -04:00
jared 9617b86a26 ci: mirror canonical tag objects, not the lightweight refs checkout builds
R-CMD-check / check (push) Successful in 3m45s
R-CMD-check / check (pull_request) Successful in 3m29s
v0.4.0 failed to mirror, and it would have failed at every future release.

On a tag-triggered run, checkout materializes refs/tags/<tag> as a LIGHTWEIGHT
tag at the commit SHA; the annotated tag object Gitea holds is never fetched.
Run 2070 therefore pushed a lightweight v0.4.0 to GitHub and reported success.
Run 2076, on main with fetch-depth 0, did fetch the real annotated object and
was rejected with "already exists" trying to correct it -- git will not clobber
an existing tag. Gitea had e138eeb (annotated), GitHub had d2caa6d (the commit).

Re-fetch canonical tag objects from Gitea before pushing. --force rewrites LOCAL
tag refs only; it is not a force push and does not weaken the non-force
guarantee on main. It is required: without it the fetch is rejected with "would
clobber existing tag" and the lightweight ref survives to be mirrored again.

Verified against a scratch clone -- lightweight d2caa6d becomes annotated
e138eeb, peeling back to the same commit; without --force the tag is unchanged.

The GitHub tag was repaired by hand out of band, so the two remotes already
agree; this stops it recurring.
2026-08-11 09:57:14 -04:00
jared 224e5d0530 Merge pull request 'chore: CI badge, mirror PR explainer, and the tag-pin release step (#47)' (#66) from chore/release-47-badges-mirror-pr into main
Mirror to GitHub / mirror (push) Failing after 11s
R-CMD-check / check (push) Successful in 3m15s
Reviewed-on: #66
2026-08-11 09:44:33 -04:00
jared 5f81ae386b chore: CI badge, mirror PR explainer, and the tag-pin release step (#47)
R-CMD-check / check (pull_request) Successful in 3m49s
R-CMD-check / check (push) Successful in 3m52s
Three items #47 absorbed from #46, held back so a dead badge would not sit
beside an unresolved r-universe one. Both resolve now that v0.4.0 is tagged.

- R-CMD-check badge pointing at the GitHub mirror's workflow, where the
  4-platform matrix actually runs.

- A pull_request_target workflow explaining the mirror flow on every incoming
  PR. A PR here is landed on Gitea and syncs back, and because the merge
  preserves the contributor's commits at their original SHAs, GitHub marks the
  PR 'Merged' with nobody visibly clicking Merge. To a first-time contributor
  that reads as rejection. Say so before it happens.

  pull_request_target rather than pull_request because a fork PR's token is
  read-only under the latter -- it could not comment, which is the entire job.
  That is only safe because this never checks out or runs contributor code; the
  file says so and says not to add a checkout.

- CONTRIBUTING's release checklist now spells out that the tag goes on Gitea and
  the mirror carries it, and that r-universe does NOT pick up a release until
  packages.json's branch pin is edited. '*release' would automate it but needs a
  GitHub Release object, and the mirror pushes tags only -- so it would silently
  never update. Learned while doing this release.
2026-08-10 19:47:26 -04:00
jared d2caa6de97 Merge pull request 'docs: re-measure the corpus-access table against the published corpus (#56)' (#65) from docs/readme-perf-remeasure-56 into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m25s
v0.4.0
2026-08-10 19:36:42 -04:00
jared 303aa59b07 docs: order the 0.4.0 NEWS sections by user impact, not merge order
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m38s
The three 0.4.0 features landed in the order their PRs merged, which buried the
headline change (cohort predicates, 4.8x) below an operator config knob and a
documentation note. Reordered to: cohorts, pagination, DuckDB budget, docs,
Fixes -- with Fixes last, where it was already.

The stacked PRs each appended their own '## Fixes' heading, so resolving the
conflicts also folded two of them into the single section that belongs there.

Content is byte-identical to what merged; only section order changed. Verified
by diffing the sorted non-blank lines against the previous commit.
2026-08-10 19:29:43 -04:00
jared 81f72321ee docs: re-measure the corpus-access table against the published corpus (#56)
figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09)
and reported the mirrored column as 'local speed' with no number -- hiding the
largest difference available to a user.

Measured 2026-08-10, fresh R session per arm, against the live corpus at
pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation.

Three findings the old table could not express:

- A local mirror is 60-80x faster. A one-off question is ~12 s end to end
  remotely against ~0.15 s mirrored. Stated outright now, because it is a
  bigger and cheaper win for users than anything in the R code.

- Opening the session is the LARGEST remote cost (~7.5 s), bigger than any
  individual query, and it lands on the user's first query rather than on
  library(). The old table accounted for it nowhere, so every per-query figure
  was quietly missing it.

- The remote cost is round-trips, not scanning: a repeat query over
  already-touched partitions is ~1.5 s against ~4 s cold, and a full-history
  query costs ~7 s whether it runs first or last (verified by running the arms
  in both orders). This is why #93's 1.4-1.7x, measured through cog-api against
  a local mount, does not show up on the remote path -- there, network latency
  swamps scan time.

Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was
ambiguous between MB and MiB besides. Measured from the manifest and on disk.
The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release.

Also documented HTTP 429: a burst of remote queries gets rate-limited by the
host. Hit while taking these measurements.
2026-08-10 19:29:05 -04:00
jared 5cb83d8f1d Merge pull request 'feat: limit/offset on cog_gov_search() and cog_balances() (#57)' (#63) from feat/pagination-search-balances-57 into main
Mirror to GitHub / mirror (push) Successful in 12s
R-CMD-check / check (push) Successful in 4m26s
2026-08-10 19:28:09 -04:00
jared 2bb9d41d72 feat: limit/offset on cog_gov_search() and cog_balances() (#57)
R-CMD-check / check (pull_request) Successful in 4m22s
R-CMD-check / check (push) Successful in 4m28s
verbs were left materializing everything and slicing in R -- the pattern behind
the 2026-08-06 production incident. cog_gov_search() had no LIMIT at all, so an
unfiltered call returns the entire 40,336-row crosswalk.

Extracted the #39 machinery into R/pagination.R first (.validate_pagination(),
.paginate_sql(), .take_pagination_total()) rather than growing a third inline
copy: three definitions of what total_rows means is three places for it to
drift. Conflict refusals stay at the call sites because each verb's conflict
set differs. .verb_spendrev() now uses the shared helpers and is unchanged in
behaviour.

The empty-page fallback query is now passed as a thunk, so the unpaginated SQL
is only BUILT when an offset actually lands past the end instead of on every
paged call.

Two things #57 did not anticipate:

- cog_gov_search()'s ORDER BY was not a total order. population_acs DESC NULLS
  LAST leaves ties -- and the whole NULL block -- in scan order, so two requests
  can order them differently and a paged sweep duplicates one row while dropping
  another. Added canonical_govid as tiebreaker. Unpaginated output changes only
  in the relative order of already-tied rows.

- Basket mode returns one resolved row per requested name plus a sidecar
  covering all of them, so a page of it is not a page of anything the caller
  asked for. Refused with uscogdata_basket_pagination_conflict rather than
  silently ignoring the arguments.

Both default to NULL, so cog-api adopts them behind its existing formals()
probe with no lockstep deploy.

Suite: 1067 passed, 0 failed, 0 warnings (2 pre-existing live-corpus skips).
2026-08-10 19:27:11 -04:00
jared 700ae93c9c Merge pull request 'feat: cog_open() honours a DuckDB thread and memory budget (#60)' (#62) from feat/duckdb-threads-60 into main
Mirror to GitHub / mirror (push) Successful in 10s
R-CMD-check / check (push) Successful in 4m6s
Reviewed-on: #62
2026-08-10 19:23:14 -04:00
jared f4ab9b6d90 feat: cog_open() honours a DuckDB thread and memory budget (#60)
R-CMD-check / check (push) Successful in 4m11s
R-CMD-check / check (pull_request) Successful in 4m11s
cog_open() connected with a bare dbConnect() and set no resource pragmas, so
DuckDB claimed every visible core. Right for one interactive session on a
dedicated machine; wrong for a server, where cog-api runs two replicas on an
8-core host budgeted 4 and each replica independently claims all 8.

USCOGDATA_DUCKDB_THREADS and USCOGDATA_DUCKDB_MEMORY_LIMIT now resolve through
.cfg() -- inheriting the env var > option > default precedence USCOGDATA_URL
already had -- and are applied as pragmas when the connection is created.

Unset issues NO pragma, so an unconfigured session is byte-identical to before.
That negative property is asserted directly against a connection opened the
pre-change way rather than against a hardcoded core count.

.cfg() returns an env var as character, so both resolvers coerce and validate
rather than trusting the type: sprintf("SET threads TO %d", "4") would
otherwise abort inside the connection path with an error naming the pragma
instead of the setting the operator got wrong.

Replaces cog-api's getFromNamespace(".ensure_session", "uscogdata") workaround,
which depended on a private name and on the session already being open.
2026-08-10 18:51:56 -04:00
jared 9508b98676 Merge pull request 'feat: cohort predicates instead of a 40k-id IN list (#58)' (#61) from feat/cohort-predicates-58 into main
Mirror to GitHub / mirror (push) Successful in 9s
R-CMD-check / check (push) Successful in 3m18s
Reviewed-on: #61
2026-08-09 14:31:53 -04:00
jared 0fbae00e27 feat: name a cohort by state/type predicate instead of a 40k-id IN list
R-CMD-check / check (push) Successful in 4m29s
R-CMD-check / check (pull_request) Successful in 4m19s
cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.

The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.

Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.

Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:

  IN (20,106 literals) -- 0.3.0            432 ms
  join against a temp cohort table         132 ms
  predicate on canonical_fips_xwalk         102 ms
  no cohort filter at all (the floor)      105 ms

The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.

Design decisions, both made explicitly rather than left implicit:

  - govid AND state/type INTERSECT. "These ids, narrowed to that
    state/type" is a real query, and an error here could never be
    relaxed later without breaking callers.
  - A predicate cohort has no id list to report, so
    provenance$scope$govids_found/govids_missing stay empty and a new
    scope$cohort block carries state, type and n_governments. Resolving
    the ids just to report them would put 20,000 govids in every
    fleet-scale response body -- the cost this change removes. A
    govid-named cohort's provenance is untouched.

state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.

.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.

Fixes uscogdata#58.
2026-08-09 14:15:21 -04:00
jared 698812a25c Merge pull request 'fix: local corpus paths were unreadable on Windows (backslashes eaten)' (#55) from fix/windows-backslash-paths into main
Mirror to GitHub / mirror (push) Failing after 8s
R-CMD-check / check (push) Successful in 3m34s
2026-08-09 09:08:41 -04:00
jared d74ecdd4a5 Merge pull request 'ci: mirror main and tags to the GitHub mirror' (#54) from ci/mirror-to-github into main
R-CMD-check / check (push) Successful in 3m29s
Mirror to GitHub / mirror (push) Failing after 11s
2026-08-09 09:04:28 -04:00
jared 72b2cc3a26 fix: local corpus paths were unreadable on Windows (backslashes eaten)
R-CMD-check / check (push) Successful in 3m35s
R-CMD-check / check (pull_request) Successful in 3m49s
gsub() in regex mode treats backslashes in the REPLACEMENT string as
escape sequences and silently drops them. A Windows corpus path is full of
them, so C:\Users\RUNNER\AppData\... was substituted into the view SQL
as C:UsersRUNNERAppData... and every DuckDB read failed with 'No files
found that match the pattern'.

Effect: uscogdata could not read a LOCAL corpus on Windows at all -- the
bundled fixture included, so the entire test suite failed there, and any
cog_mirror() copy was unusable. Remote https URLs were unaffected, having
no backslashes, which is part of why it stayed hidden.

The bug predates the {long_files} token; it lived in the original {url}
substitution since that code was written. Nothing ever ran on Windows
until the GitHub mirror's check matrix existed, which found it on its
first run: 14 of 15 Windows failures were this, the 15th a downstream
consequence of view registration failing.

fixed = TRUE treats pattern and replacement as literal text. The
regression test reproduces on any platform -- it is string handling, not
filesystem behaviour, so it needs no Windows runner.
2026-08-09 09:03:44 -04:00
jared 28c500f47a ci: mirror main and tags to the GitHub mirror
R-CMD-check / check (pull_request) Successful in 3m23s
R-CMD-check / check (push) Successful in 4m2s
A plain non-force git push rather than Gitea's built-in push mirror. A
push mirror force-updates the refs it owns, so a Merge clicked on a GitHub
PR would be silently overwritten on the next sync -- the PR still reading
'Merged' while its commit became unreachable. A non-force push is rejected
instead, which turns that into a red CI run.

Secret is PAT_GH, not GITHUB_MIRROR_PAT: Gitea reserves the GITHUB_ and
GITEA_ prefixes for its own injected variables and refuses secrets using
them.
2026-08-08 20:20:56 -04:00
jared fe9238a6ef Merge pull request 'ci: multi-platform R CMD check for the GitHub mirror' (#53) from ci/github-actions-matrix into main
R-CMD-check / check (push) Successful in 3m5s
2026-08-08 19:54:01 -04:00
jared 6392a74013 ci: multi-platform R CMD check for the GitHub mirror
R-CMD-check / check (push) Successful in 3m27s
R-CMD-check / check (pull_request) Successful in 3m19s
The canonical Gitea runner is Linux-only, and this package hard-depends on
duckdb and httr2 -- both compiled, both with real platform variance --
while having never been checked on Windows or macOS. A large share of the
audience is on Windows.

Inert here: Gitea reads .gitea/workflows, GitHub reads .github/workflows,
so both configs coexist and the Gitea CI stays authoritative for deploys.
This goes live the moment the mirror repo exists (#44).

The matrix needs no credentials -- setup.R points at the bundled fixture --
which is why that fixture must never be .Rbuildignore'd.
2026-08-08 19:49:52 -04:00
jared de2ba0cfb9 Merge pull request 'feat: uscogdata 0.3.0 — first public release' (#41) from feat/public-release-0.3.0 into main
R-CMD-check / check (push) Successful in 3m27s
2026-08-08 18:56:14 -04:00
jared e912a926c2 Merge pull request 'chore: regenerate fixture against pipeline e7394a4 (SB203-SB209)' (#40) from chore/fixture-sb203 into main
R-CMD-check / check (push) Successful in 3m29s
2026-08-08 18:55:59 -04:00
jared 44e4953f94 fix: keep doc/ and Meta/ out of the build
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m52s
Removing ^vignettes$ was right; removing ^doc$ and ^Meta$ with it was
not. Those are devtools::build_vignettes() artefacts, not sources -- R CMD
build regenerates inst/doc/ from vignettes/ by itself, and shipping the
local copies earned a 'non-standard file/directory found at top level'
NOTE.

R CMD check --as-cran is now 0 errors, 0 warnings, 0 notes.
2026-08-08 18:25:00 -04:00
jared a5500f0b6a docs: add CONTRIBUTING with the canonical-on-Gitea PR flow
Moves developer, testing and release instructions out of the README,
minus the fixture-stripping advice, which was wrong.

Explains that a GitHub PR closes itself as merged once the mirror syncs,
because the merge preserves the contributor's SHAs -- so a PR closing
without a visible Merge click reads as success rather than rejection.

Documents the cog-api dependency: its CI clones this package at
USCOGDATA_REF, defaulting to main with no pin, so anything merged here
reaches the API's next build. Includes the commands to run its suite
against a branch first.
2026-08-08 18:20:42 -04:00
jared da2839f885 docs: recast NEWS around the first public release
NEWS described changes relative to states no user had ever seen --
'Breaking: corpus schema_version 4', 'the package now requires...' --
across the whole pre-release development. To someone deciding whether to
depend on this, that reads as instability.

0.3.0 is written as an announcement: what it covers, the verbs, that
reading the corpus now works out of the box, four things to know before a
first query, and the known limits. The 0.2.0 changelog is kept verbatim.
The 0.1.0 development log is dropped; that history is in git.

cog_explain() now documents what provenance actually holds, since the
README points readers at it -- in particular why series_break_refs and
corpus_break_refs are separate fields rather than one list.
2026-08-08 18:20:33 -04:00
jared 331399ab86 docs: rewrite README for a stranger
Reordered around a new user: what the data is, where it comes from,
install, a quickstart that runs with no configuration, then the
full-dollars warning and the concepts that decide whether a published
number is right.

Adds a 'Where the data comes from' section linking the API documentation
site, the live API, the Hugging Face corpus and the Census source, so
attribution and provenance are reachable from the top rather than implied.

Drops the sibling-repo path, the commented-out install line, the Status
block, and the release advice telling you to strip the fixture -- which
would break the vignette and leave public CI unable to check without
credentials.

The quickstart passes years=; cog_spending() has no full-history default,
so the obvious one-liner errors on a reader's first call.
2026-08-08 18:20:24 -04:00
jared 5582c6cb57 docs: index all 14 exports in pkgdown, set the site url
The reference index covered 6 of 14 exports, so pkgdown errored on the
eight missing topics and the docs site did not build at all. Adds a
Comparison & aggregation section and a Corpus metadata section, lists both
vignettes as articles, and sets url so canonical links and search resolve.

build_site() now completes clean: reference metadata ok, no problems.
2026-08-08 17:49:10 -04:00
jared 013af5b0d2 fix: ship the vignettes
.Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed
uscogdata had no vignettes at all -- while the README instructed users to
run vignette("total-spending"), which failed for every one of them.

Both build offline: total-spending points USCOGDATA_URL at the bundled
fixture, population-denominators is eval = FALSE. Confirmed present in the
built tarball as both source and rendered inst/doc/.

The test also pins the fixture as never-excluded -- it is what lets
R CMD check run with no credentials on r-universe and GitHub Actions.
2026-08-08 17:48:07 -04:00
jared 4a92f36d44 chore: add full MIT text, name the copyright holder properly
LICENSE held only the two-line stub and no LICENSE.md existed, so the
repo carried no license text for a human browsing it or for GitHub's
license detector.

usethis::use_mit_license() writes LICENSE.md but leaves an existing
LICENSE alone, so the stub kept saying 'Civilytics' while the full text
said 'Civilytics Consulting LLC'. Corrected by hand, with a test pinning
the two together.
2026-08-08 17:47:18 -04:00
jared a67735f121 chore: release metadata -- author of record, URLs, schema ceiling, 0.3.0
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.

MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.

0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
2026-08-08 17:46:39 -04:00
jared b9f7f7d8d3 test: exercise the public corpus end to end, unconfigured
The remote-read defect survived because every test path used a local
corpus, and so did the API in production. This is the only test that runs
the package the way a new user does: no USCOGDATA_URL, no option, no
fixture -- just install and call a verb.

Gated on USCOGDATA_LIVE_TEST so offline CI skips rather than fails.

Measured against the live corpus while writing this: cog_spending for one
government is 3.9s for a single year and 5.9s across 2000-2022. Well above
the 1.5-2.8s raw parquet scan, because the verbs also join crosswalks,
resolve categories and assemble provenance.
2026-08-08 17:21:53 -04:00
jared 4300b636b1 feat: default to the public corpus so the package works unconfigured
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in
the package supplied a working URL, so a new user installing uscogdata
had no path to a session at all -- just an actionable-looking error with
nothing actionable behind it.

The default is now the public HuggingFace mirror: CC-BY-4.0, no
credential, CDN-backed, and it keeps the origin's uplink out of the read
path. USCOGDATA_URL and options(uscogdata.url=) still override, so
Nextcloud and cog_mirror() copies are unaffected.

The sentinel guard stays for half-edited configs; the two tests covering
it set the URL explicitly, so they only needed renaming to stop calling
it 'the default'.
2026-08-08 17:18:56 -04:00
jared 99e1e86e37 fix: enumerate long partitions from the manifest instead of globbing
DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.

Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.

Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
2026-08-08 17:11:44 -04:00
jared c042ee0b90 docs: retarget the release at 0.3.0 and preserve the 0.2.0 changelog
Spec and plan were written against a branch 25 commits behind main, where
the package still read 0.1.0. It is 0.2.0, with a real 0.2.0 changelog in
NEWS that the plan would have deleted.

0.3.0 rather than 0.2.0 because remote corpus reads go from broken to
working and the default URL from placeholder to live -- user-visible
behaviour, so a minor bump. Not 1.0.0: types 4 and 5 remain out of scope.

Task 9 now prepends a 0.3.0 section, keeps 0.2.0 verbatim with a diff
check to prove it, and drops only the 0.1.0 development churn. Task 4
gains the Version bump.
2026-08-08 17:07:19 -04:00
jared 29dc8199e0 docs: implementation plan for the uscogdata 0.1.0 public release
Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest
enumeration, working default URL, and the live-corpus test whose absence
let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite
README, NEWS and CONTRIBUTING.

Distribution mechanics stay out of scope -- r-universe publishes check
results on registration, so it comes after final verification is green.
2026-08-08 17:03:40 -04:00
jared 619b167ab1 docs: design spec for the uscogdata 0.1.0 public release
Covers the P0 finding that the package cannot read the corpus remotely at
all -- no working default URL, and Hive globs are unsupported over generic
HTTP by DuckDB 1.5.5. Fix is manifest-driven file enumeration (measured:
46,148,034 rows over plain https, identical to the hf:// glob) plus a
working public default.

Also: seven release-readiness fixes, a README restructured for a stranger,
NEWS rewritten as an initial release rather than a pre-release churn log,
and the Gitea-canonical/GitHub-mirror/r-universe distribution mechanics.
2026-08-08 17:03:40 -04:00
jared dfda39051e docs: point agents at the Civilytics values file before they start
Mirrors the same block in cog_explorer/CLAUDE.md. A project-level @ import
does not preload, so the instruction to read it is the mechanism.
2026-08-08 17:03:39 -04:00
jared 785f3af16d chore: regenerate fixture against pipeline e7394a4 (SB203-SB209)
R-CMD-check / check (pull_request) Successful in 3m50s
R-CMD-check / check (push) Successful in 4m9s
Picks up seven new catalogued series breaks, all break_year 2017,
recording that the employee-retirement X-codes (X21, X30, X42, X44, X47,
Z77, Z78) were last collected in the annual finance file at FY2016 before
those systems moved to the Annual Survey of Public Pensions.

series_breaks.parquet 202 -> 209 rows; nothing removed. manifest built_at
and pipeline_commit updated, and the series_breaks sha256 with them.
2026-08-08 17:02:14 -04:00