main
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b41d5ee2aa
|
feat(coverage): n_units_collected separates sampling from real zeros (#36)
provenance$coverage's n_units_reporting is category-conditional: it counts governments with rows for the SPECIFIC requested category, which conflates two different things -- a government never collected that year (sampling), and one collected but genuinely spending nothing in that category (a real zero). FY2012 Georgia Police is the motivating case from the issue: a complete census year reads as a 69% "response rate" because most of the gap is cities that contract policing to the county sheriff, not non-response. Adds a second counter, n_units_collected: how many of the caller's expected cohort appear in the corpus that year for ANY category. n_units_collected / n_units_expected is the true collection rate; n_units_reporting / n_units_collected is category participation among collected units. cog_geographic_rollup() and cog_peer_compare() both carry it; cog_explain() prints it alongside n_units_reporting. Two real bugs caught and fixed while finishing this (both against the already-written, previously-uncommitted draft): - .coverage_table()'s candidate list for the collection query was derived from the category-filtered result rows, not the caller's full expected cohort. A government with zero rows in the requested category across every requested year never appears in that result, so it was silently excluded from n_units_collected too -- collapsing the new counter back to the old, broken one for exactly the governments it exists to count. Fixed by threading an explicit `expected_ids` (all_govids / peer_govids) through instead. - The collection query hardcoded long_view = "spending_long_harmonized", which does not exist on a corpus with schema_version < 5 (R/basis.R resolves basis = "raw" there; R/views.R only registers the harmonized views on v5+). cog_geographic_rollup()/cog_peer_compare() would hard-error on a corpus vintage the package otherwise explicitly supports. Fixed by deriving long_view from the basis cog_spending() actually resolved (prov$basis) via the existing .select_long_view() helper, matching how every other basis-aware query in the package already does this. Also: cog_explain()'s general "complete census only in years ending in 2 or 7" footnote was gated on the OLD counter's absence, making it permanently unreachable now that both callers always supply the new one -- ungated it, since the explanation is orthogonal to which counter set is present. Dropped a dead conditional branch, fixed two stale roxygen blocks in R/peers.R/R/rollup.R still describing the old two-counter model, fixed the same staleness in README.md, and switched two `uscogdata:::` self-references to the package's own convention of calling internal helpers unqualified. 1101 tests pass (2 skipped live-corpus), including new direct regression tests for both bugs above (one exercising a government collected-but-absent from a category result, one running the full rollup/peer-compare path against a doctored schema_version 4 corpus). Reviewed by an independent code-reviewer pass (1 HIGH, 1 MEDIUM, 3 LOW -- all addressed above). Closes #36. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5f81ae386b
|
chore: CI badge, mirror PR explainer, and the tag-pin release step (#47)
Three items #47 absorbed from #46, held back so a dead badge would not sit beside an unresolved r-universe one. Both resolve now that v0.4.0 is tagged. - R-CMD-check badge pointing at the GitHub mirror's workflow, where the 4-platform matrix actually runs. - A pull_request_target workflow explaining the mirror flow on every incoming PR. A PR here is landed on Gitea and syncs back, and because the merge preserves the contributor's commits at their original SHAs, GitHub marks the PR 'Merged' with nobody visibly clicking Merge. To a first-time contributor that reads as rejection. Say so before it happens. pull_request_target rather than pull_request because a fork PR's token is read-only under the latter -- it could not comment, which is the entire job. That is only safe because this never checks out or runs contributor code; the file says so and says not to add a checkout. - CONTRIBUTING's release checklist now spells out that the tag goes on Gitea and the mirror carries it, and that r-universe does NOT pick up a release until packages.json's branch pin is edited. '*release' would automate it but needs a GitHub Release object, and the mirror pushes tags only -- so it would silently never update. Learned while doing this release. |
||
|
|
81f72321ee
|
docs: re-measure the corpus-access table against the published corpus (#56)
figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09) and reported the mirrored column as 'local speed' with no number -- hiding the largest difference available to a user. Measured 2026-08-10, fresh R session per arm, against the live corpus at pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation. Three findings the old table could not express: - A local mirror is 60-80x faster. A one-off question is ~12 s end to end remotely against ~0.15 s mirrored. Stated outright now, because it is a bigger and cheaper win for users than anything in the R code. - Opening the session is the LARGEST remote cost (~7.5 s), bigger than any individual query, and it lands on the user's first query rather than on library(). The old table accounted for it nowhere, so every per-query figure was quietly missing it. - The remote cost is round-trips, not scanning: a repeat query over already-touched partitions is ~1.5 s against ~4 s cold, and a full-history query costs ~7 s whether it runs first or last (verified by running the arms in both orders). This is why #93's 1.4-1.7x, measured through cog-api against a local mount, does not show up on the remote path -- there, network latency swamps scan time. Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was ambiguous between MB and MiB besides. Measured from the manifest and on disk. The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release. Also documented HTTP 429: a burst of remote queries gets rate-limited by the host. Hit while taking these measurements. |
||
|
|
f4ab9b6d90
|
feat: cog_open() honours a DuckDB thread and memory budget (#60)
cog_open() connected with a bare dbConnect() and set no resource pragmas, so
DuckDB claimed every visible core. Right for one interactive session on a
dedicated machine; wrong for a server, where cog-api runs two replicas on an
8-core host budgeted 4 and each replica independently claims all 8.
USCOGDATA_DUCKDB_THREADS and USCOGDATA_DUCKDB_MEMORY_LIMIT now resolve through
.cfg() -- inheriting the env var > option > default precedence USCOGDATA_URL
already had -- and are applied as pragmas when the connection is created.
Unset issues NO pragma, so an unconfigured session is byte-identical to before.
That negative property is asserted directly against a connection opened the
pre-change way rather than against a hardcoded core count.
.cfg() returns an env var as character, so both resolvers coerce and validate
rather than trusting the type: sprintf("SET threads TO %d", "4") would
otherwise abort inside the connection path with an error naming the pragma
instead of the setting the operator got wrong.
Replaces cog-api's getFromNamespace(".ensure_session", "uscogdata") workaround,
which depended on a private name and on the session already being open.
|
||
|
|
331399ab86
|
docs: rewrite README for a stranger
Reordered around a new user: what the data is, where it comes from, install, a quickstart that runs with no configuration, then the full-dollars warning and the concepts that decide whether a published number is right. Adds a 'Where the data comes from' section linking the API documentation site, the live API, the Hugging Face corpus and the Census source, so attribution and provenance are reachable from the top rather than implied. Drops the sibling-repo path, the commented-out install line, the Status block, and the release advice telling you to strip the fixture -- which would break the vignette and leave public CI unable to check without credentials. The quickstart passes years=; cog_spending() has no full-history default, so the obvious one-liner errors on a reader's first call. |
||
|
|
4b23dbd9f4
|
feat: revenue_concept = c("general", "total") off the crosswalk (#12)
Closes the last blocked test in the suite. Owner ruled both halves of the open question yes on 2026-07-30. `cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`, with Census's two published concepts defined as crosswalk `revenue_subtype` sets rather than item-code prefixes: general = own_source + federal + state + local_aid (the default) total = general + utility + liquor_store + insurance_trust The manual defines the first by subtracting the other three from the second (4.3), so both are computable only once all four families are named -- which cog_pipeline#79 does. Insurance trust now includes the employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes. - inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue subtype; the concept narrows in R via the existing subtype_scope machinery, exactly as expenditure_concept narrows spending_long. - cog_explain() now prints each verb's OWN concept. It previously printed `expenditure_concept` unconditionally, so a cog_revenue() caller was told "Concept: primary" -- a spending concept their result has nothing to do with. - Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows). Corrected two stale expectations in the blocked test while un-skipping it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state governments and is nonzero for Wisconsin; X04 is an exhibit code for an INTRAgovernmental transfer that Census's own "Total Emp Ret Rev" excludes. Verified against that Census field: the right set is X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of $33,377,093k predated the Y codes being classified -- complete Total Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785 + X 2,283,883). Behaviour change worth knowing: `general` is now STRICT Census General Revenue, so utility and liquor store revenue leave the default. Measured on the fixture that is 15.9% of what cog_revenue() returned for cities, vs 1.2% for states and 1.7% for counties. Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had no skipped tests. Closes #12 |
||
|
|
93300ae0c1
|
feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:
primary = operations + capital + assistance
direct = primary + interest + insurance_benefits (Census Direct)
total = direct + intergovernmental (M/L/Q via ig views)
- inst/sql: flow views (20-25) select by crosswalk membership;
summary_categories moves to 11- so it registers before them (DuckDB
binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
harmonization exclusion count, and the complete = TRUE grid;
flow_prefixes survives only to scope recipe suggestions.
cog_geographic_rollup/cog_peer_compare accept primary|direct, still
refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
(uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
proof is asserted against the crosswalk, not the default
cog_revenue() call (which stays General Revenue pending #12).
Suite: 696 pass / 0 fail / 1 skip (#12, expected).
Closes #11
|
||
|
|
d006dea6e4
|
fix: literal name search, units docs, peer-summary semantics (#16, #15, #14)
The three kodor/fix issues, taken over after a day with no branch, PR or comment on any of them. Batched because each is single-file with a committed acceptance test, and two share documentation surfaces. #16 (F-025) -- cog_gov_search() utility mode interpolated `name` straight into regexp_matches() unescaped, while basket mode in the same file already routed it through .escape_regex() with the comment "so `name` is treated as a literal substring". Two failure modes, both HTTP 200 through the API: a government could not be found by its own complete name when that name contains a metacharacter (FREDONIA (BRISCOE) CITY returned nothing), and a bare "." matched all 608 Wisconsin cities. Malformed pattern text reached the engine as an error, which cog-api surfaced as a 500 -- reachable by typing a real name one character at a time ("Athens-Clarke County (bal"). Utility mode now calls the escaper that already existed. Roxygen updated: utility mode is documented as a literal case-insensitive substring match, and the basket-mode "substring fallback" step no longer describes itself as a regex either. BEHAVIOUR CHANGE worth flagging: anchored exact-match searches stop working, because there is no regex left to anchor. Two existing tests used "^BROWARD COUNTY$" and "^FLORIDA$" as their exact-match idiom; both now search for those characters literally. Updated to the bare names, which still resolve to exactly one row each once scoped by state/type (verified, not assumed). There is no exact-match option in utility mode any more -- noted on the issue, since that is a real if small capability loss. #15 (F-004) -- the raw Census files report thousands of dollars; this package multiplies by 1000 and returns full US dollars. Correct, and already stated in ?cog_spending / ?cog_revenue @return, in provenance, and in cog-api's data-dictionary. Absent from every surface a reader meets FIRST. Added to README.md as its own section and to both vignettes' openings. The dangerous one is cog_explorer/CLAUDE.md, which states the opposite rule ("All raw `amt` values are in $1,000s") without scoping it to the raw column -- a reader applying that to amt_nominal overstates by 1000x and gets a plausible-looking number rather than an obvious error. Fixed there too; that directory has no git remote, so it rides in no PR and is left uncommitted for the owner. #14 (F-021) -- .peer_summary_rows() computes stats::quantile() separately inside each (year, spend_subtype, category) cell, so a summary_p50 row is "the median peer's value in that one category", never "the value of the median peer's total" -- the median peer for Police and for Fire are usually different governments. Summing them across categories misstated a total-spending band by -32.7% to +251.0% across 24 years, with a sign flip at FY2012. The verb is right and its documented use (facet by role AND category) is unaffected, so the fix is @return prose plus a worked snippet showing the correct computation: sum each peer's own categories first, then take the quantile of those per-government totals. This is the R-side counterpart of cog-api#9, fixed on the API surface earlier today; the wording is deliberately consistent across the two. Note the phrase "not additive" has to stay on one roxygen source line -- the test greps the generated Rd, where a line wrap turns it into "not additive" and stops matching. Cost one red run to find. man/ regenerated with roxygen 8.0.0 against a repo built with 7.3.3, so cog_spending.Rd and DESCRIPTION were reverted -- their entire diff was version churn (reindentation, RoxygenNote -> Config/roxygen2/version) with no content change. The two Rd files kept carry only the edits above. Suite: 629 pass / 0 fail / 3 skip (was 606/0/6). The three remaining skips are #11, #12 and #13. |
||
|
|
e7d3a7a310 |
docs: fix stale fixture description and Direct/Total vignette figures (M1, M2, M5)
M5: README.md described the bundled fixture as a "3.6 MB two-year slice (2019 + 2020)"; it's now a 15 MB four-year slice (2011, 2012, 2019, 2020), matching the regenerated fixture and the vignette's own description. M2: total-spending.Rmd cited County 1.8% / City 0.8% intergovernmental- to-Direct and County 91.6% L/M, all roughly 2x off against the bundled fixture. Measured directly against the fixture (all 50 states, each of its four years): County IG/Direct 3.4%-5.1%, City IG/Direct 2.6%-3.1%, County L/M 43%-51% (all varying by year). State 17.2%, AL 7.6%, national 11.6%, and City L/M 188.3% were re-checked and left as-is. M1: the "Why Total = Direct + M + L" paragraph described money a local government *receives* and the state "redistributing as M" -- backwards. M and L are both the *queried* government's own payments *out*: M to other local governments, L up to its state. Rewrote the explanation; the conclusion and non-M2-flagged figures are unchanged. |
||
|
|
3bd9b1f011 |
docs: total-spending vignette + README section on Direct vs Total
Task 7 (final) of the expenditure_concept plan. The vignette leads with the two archetype questions -- a single government's own trend (either concept works, held fixed across years) vs a cross-government rollup (direct only, with the refusal error from cog_geographic_rollup() shown and explained) -- walked through with code that runs against the bundled fixture corpus (years 2011/2012/2019/2020, substituting for "2017 vs today"). Explains the double-counting mechanism (a state's M44 payment to a county is the same dollar as the county's own E44/F44), why Total = Direct + M + L rather than Direct + M, and the composition rules (expenditure_concept is orthogonal to basis, mutually exclusive with recipe). README gets a short pointer section with the one-line rule. |
||
|
|
af6f18f9d7
|
docs: add Developer notes section to README (testing + live-corpus release) | ||
|
|
f7035049e7
|
feat: package skeleton — DESCRIPTION, NAMESPACE, session/manifest/cache/views
Minimal skeleton for uscogdata v0.1. Internal session layer with lazy cog_open(), manifest fetch+validate+cache, view registration placeholder. Depends on DuckDB >=1.0, httr2, jsonlite. inst/schemas/provenance-v1.json ships the structured provenance JSON Schema. |