From 81f72321ee6fa9e99a2e1a8d21f988b20a538f0f Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 10 Aug 2026 19:16:03 -0400 Subject: [PATCH 1/2] docs: re-measure the corpus-access table against the published corpus (#56) figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09) and reported the mirrored column as 'local speed' with no number -- hiding the largest difference available to a user. Measured 2026-08-10, fresh R session per arm, against the live corpus at pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation. Three findings the old table could not express: - A local mirror is 60-80x faster. A one-off question is ~12 s end to end remotely against ~0.15 s mirrored. Stated outright now, because it is a bigger and cheaper win for users than anything in the R code. - Opening the session is the LARGEST remote cost (~7.5 s), bigger than any individual query, and it lands on the user's first query rather than on library(). The old table accounted for it nowhere, so every per-query figure was quietly missing it. - The remote cost is round-trips, not scanning: a repeat query over already-touched partitions is ~1.5 s against ~4 s cold, and a full-history query costs ~7 s whether it runs first or last (verified by running the arms in both orders). This is why #93's 1.4-1.7x, measured through cog-api against a local mount, does not show up on the remote path -- there, network latency swamps scan time. Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was ambiguous between MB and MiB besides. Measured from the manifest and on disk. The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release. Also documented HTTP 429: a burst of remote queries gets rate-limited by the host. Hit while taking these measurements. --- NEWS.md | 24 ++++++++++++++++++++++++ README.md | 34 +++++++++++++++++++++++++++++----- 2 files changed, 53 insertions(+), 5 deletions(-) diff --git a/NEWS.md b/NEWS.md index 1091972..9fe944b 100644 --- a/NEWS.md +++ b/NEWS.md @@ -41,6 +41,30 @@ Two refusals rather than silent surprises: per requested name with a sidecar covering all of them; a page of that is not a page of anything the caller asked for. +## Documentation: the corpus-access table is re-measured and honest + +The README's "two ways to read the corpus" table carried figures taken before +the corpus was re-chunked into row groups (cog_pipeline#93, published +2026-08-09) and reported the mirrored column as "local speed" with no number at +all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit +3d28ddd`), fresh R session per arm: + +* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s + end to end remotely against ~0.15 s mirrored. That is the largest single + difference available to a user and it is now stated outright rather than left + as "local speed". +* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch + plus 23 view registrations over HTTPS), larger than any individual query, and + it lands on the first query rather than on `library(uscogdata)`. The old table + did not account for it anywhere. +* **The remote cost is round-trips, not scanning.** A repeat query over + already-touched partitions is ~1.5 s against ~4 s cold, and a full-history + query costs ~7 s whether it runs first or last. +* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4% + and the old figure was ambiguous between MB and MiB besides. +* Documented that a burst of remote queries can be rate-limited by the host + (`HTTP 429`), which is another reason to mirror for real work. + ## Cohorts can be named by predicate, not just by id `cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state` diff --git a/README.md b/README.md index 1578400..a6ac34b 100644 --- a/README.md +++ b/README.md @@ -18,7 +18,7 @@ carries provenance describing what was converted, what was aggregated, and which known series breaks intersect your query. **Scope:** government types 0–3 (state, county, municipality, township). -56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968 +56 fiscal years, 46,148,034 rows, ~201 MB. There is no source data for FY1968 or FY1969. Special districts (type 4) and school districts (type 5) are excluded pending validation. @@ -85,14 +85,38 @@ cog_explain(spend) | | Remote (default) | Mirrored | |---|---|---| -| Setup | none | `cog_mirror(dest)`, 190.6 MB once | -| Disk used | **0 MB** — HTTP range requests only | 190.6 MB | -| Per query | ~4 s (one government, one year)
~6 s (one government, 23 years) | local speed | +| Setup | none | `cog_mirror(dest)`, ~201 MB once | +| Disk used | **0 MB** — HTTP range requests only | ~201 MB | +| Opening a session | ~7.5 s | ~0.1 s | +| One government, one year | ~4 s | ~0.05 s | +| One government, full history | ~7 s | ~0.1 s | +| Later queries, same session | ~1.5 s | ~0.05 s | | Good for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility | +**A local mirror is roughly 60–80x faster, and it is one function call.** That is +by far the largest difference any of these settings makes. If you are going to +ask more than a handful of questions, mirror first. + +Measured 2026-08-10 on a 16-core Linux workstation against the published corpus +(schema v7, `pipeline_commit 3d28ddd`), fresh R session per arm. A one-off +question costs about **12 seconds end to end remotely and 0.15 seconds +mirrored**, session setup included. + +Two things the per-query rows hide: + +- **Opening the session is the single largest remote cost** — larger than any + one query. It fetches the manifest and registers 23 SQL views over HTTPS, and + it lands on your first query, not on `library(uscogdata)`. +- **The cost is network round-trips, not scanning.** A repeat query against + partitions this session has already touched is ~1.5 s rather than ~4 s, and a + full-history query costs ~7 s whether it runs first or last. What you are + paying for is reaching each of the 56 yearly files over HTTPS the first time. + Nothing is written to disk in remote mode: DuckDB fetches the parquet footer, works out which row groups it needs, and reads only those. Nothing is cached -between sessions either, so every query goes back to the network. +between sessions either, so every query goes back to the network — and a session +that issues many remote queries in quick succession can be rate-limited by the +host (`HTTP Error: ... 429`). Both are further reasons to mirror for real work. The default points at a public HuggingFace mirror of the corpus. If you would rather not depend on a third party — for reproducibility, for an air-gapped -- 2.54.0 From 303aa59b074400726260fc560a56331e4bfe03ad Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 10 Aug 2026 19:29:43 -0400 Subject: [PATCH 2/2] docs: order the 0.4.0 NEWS sections by user impact, not merge order The three 0.4.0 features landed in the order their PRs merged, which buried the headline change (cohort predicates, 4.8x) below an operator config knob and a documentation note. Reordered to: cohorts, pagination, DuckDB budget, docs, Fixes -- with Fixes last, where it was already. The stacked PRs each appended their own '## Fixes' heading, so resolving the conflicts also folded two of them into the single section that belongs there. Content is byte-identical to what merged; only section order changed. Verified by diffing the sorted non-blank lines against the previous commit. --- NEWS.md | 130 ++++++++++++++++++++++++++++---------------------------- 1 file changed, 65 insertions(+), 65 deletions(-) diff --git a/NEWS.md b/NEWS.md index 9fe944b..ea1282e 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,70 +1,5 @@ # uscogdata 0.4.0 -## DuckDB's resource budget is configurable - -`USCOGDATA_DUCKDB_THREADS` and `USCOGDATA_DUCKDB_MEMORY_LIMIT` (with matching -`options(uscogdata.duckdb_threads = )` / `options(uscogdata.duckdb_memory_limit = )` -spellings) cap the DuckDB connection the package opens. Both follow the same -env-var > option > default precedence as `USCOGDATA_URL`. - -Unset, **no pragma is issued at all** and DuckDB's own defaults apply exactly as -before -- every visible core. That is right for one interactive session on a -dedicated machine and wrong for a server: where several readers share a host, each -otherwise claims the whole machine and they contend. Capping measured ~5% on a -single-government all-years query (502 ms at 2 threads vs 475 ms uncapped on 16 -cores), which is cheap enough that a server should always cap. - -This replaces a workaround in which a consumer reached into the package namespace -at boot -- `getFromNamespace(".ensure_session", "uscogdata")()` followed by a manual -`SET threads` -- depending both on a private name and on the session already being -open. - -## `cog_gov_search()` and `cog_balances()` gain `limit`/`offset` - -Pagination arrived on `cog_spending()`/`cog_revenue()` in 0.3.0; the other two -verbs were left materializing everything and slicing in R. Both now take -`limit`/`offset` with the same semantics: `NULL` default, the page applied in -SQL behind a deterministic `ORDER BY`, and the unpaginated count returned as a -`total_rows` attribute computed by `COUNT(*) OVER()` in the same scan rather -than a second query. - -`cog_gov_search()` had no `LIMIT` at all, which made it the one verb that -returns the entire 40,336-row crosswalk when called with no filter. - -Two refusals rather than silent surprises: - -* `cog_balances(recipe = , limit = )` aborts with class - `uscogdata_recipe_pagination_conflict` -- a recipe's result comes from a - separate query that pagination is not wired into. -* `cog_gov_search()` in basket mode (`length(name) > 1`) aborts with class - `uscogdata_basket_pagination_conflict`. Basket mode returns one resolved row - per requested name with a sidecar covering all of them; a page of that is not - a page of anything the caller asked for. - -## Documentation: the corpus-access table is re-measured and honest - -The README's "two ways to read the corpus" table carried figures taken before -the corpus was re-chunked into row groups (cog_pipeline#93, published -2026-08-09) and reported the mirrored column as "local speed" with no number at -all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit -3d28ddd`), fresh R session per arm: - -* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s - end to end remotely against ~0.15 s mirrored. That is the largest single - difference available to a user and it is now stated outright rather than left - as "local speed". -* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch - plus 23 view registrations over HTTPS), larger than any individual query, and - it lands on the first query rather than on `library(uscogdata)`. The old table - did not account for it anywhere. -* **The remote cost is round-trips, not scanning.** A repeat query over - already-touched partitions is ~1.5 s against ~4 s cold, and a full-history - query costs ~7 s whether it runs first or last. -* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4% - and the old figure was ambiguous between MB and MiB besides. -* Documented that a burst of remote queries can be rate-limited by the host - (`HTTP 429`), which is another reason to mirror for real work. - ## Cohorts can be named by predicate, not just by id `cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state` @@ -111,6 +46,71 @@ When the cohort is named by predicate there is no id list to report, so `provenance$scope$cohort` carries `state`, `type` and `n_governments` instead. A `govid`-named cohort's provenance is unchanged. +## `cog_gov_search()` and `cog_balances()` gain `limit`/`offset` + +Pagination arrived on `cog_spending()`/`cog_revenue()` in 0.3.0; the other two +verbs were left materializing everything and slicing in R. Both now take +`limit`/`offset` with the same semantics: `NULL` default, the page applied in +SQL behind a deterministic `ORDER BY`, and the unpaginated count returned as a +`total_rows` attribute computed by `COUNT(*) OVER()` in the same scan rather +than a second query. + +`cog_gov_search()` had no `LIMIT` at all, which made it the one verb that +returns the entire 40,336-row crosswalk when called with no filter. + +Two refusals rather than silent surprises: + +* `cog_balances(recipe = , limit = )` aborts with class + `uscogdata_recipe_pagination_conflict` -- a recipe's result comes from a + separate query that pagination is not wired into. +* `cog_gov_search()` in basket mode (`length(name) > 1`) aborts with class + `uscogdata_basket_pagination_conflict`. Basket mode returns one resolved row + per requested name with a sidecar covering all of them; a page of that is not + a page of anything the caller asked for. + +## DuckDB's resource budget is configurable + +`USCOGDATA_DUCKDB_THREADS` and `USCOGDATA_DUCKDB_MEMORY_LIMIT` (with matching +`options(uscogdata.duckdb_threads = )` / `options(uscogdata.duckdb_memory_limit = )` +spellings) cap the DuckDB connection the package opens. Both follow the same +env-var > option > default precedence as `USCOGDATA_URL`. + +Unset, **no pragma is issued at all** and DuckDB's own defaults apply exactly as +before -- every visible core. That is right for one interactive session on a +dedicated machine and wrong for a server: where several readers share a host, each +otherwise claims the whole machine and they contend. Capping measured ~5% on a +single-government all-years query (502 ms at 2 threads vs 475 ms uncapped on 16 +cores), which is cheap enough that a server should always cap. + +This replaces a workaround in which a consumer reached into the package namespace +at boot -- `getFromNamespace(".ensure_session", "uscogdata")()` followed by a manual +`SET threads` -- depending both on a private name and on the session already being +open. + +## Documentation: the corpus-access table is re-measured and honest + +The README's "two ways to read the corpus" table carried figures taken before +the corpus was re-chunked into row groups (cog_pipeline#93, published +2026-08-09) and reported the mirrored column as "local speed" with no number at +all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit +3d28ddd`), fresh R session per arm: + +* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s + end to end remotely against ~0.15 s mirrored. That is the largest single + difference available to a user and it is now stated outright rather than left + as "local speed". +* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch + plus 23 view registrations over HTTPS), larger than any individual query, and + it lands on the first query rather than on `library(uscogdata)`. The old table + did not account for it anywhere. +* **The remote cost is round-trips, not scanning.** A repeat query over + already-touched partitions is ~1.5 s against ~4 s cold, and a full-history + query costs ~7 s whether it runs first or last. +* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4% + and the old figure was ambiguous between MB and MiB besides. +* Documented that a burst of remote queries can be rate-limited by the host + (`HTTP 429`), which is another reason to mirror for real work. + ## Fixes * `cog_gov_search()` now orders by `population_acs DESC NULLS LAST, -- 2.54.0