From 8bd16bb0859f88b829164a7401a39fd28880945d Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Mon, 10 Aug 2026 19:16:03 -0400 Subject: [PATCH] docs: re-measure the corpus-access table against the published corpus (#56) #56 step 4 asked for the README table to be re-measured after the pass. The old figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09) and reported the mirrored column as 'local speed' with no number -- hiding the largest difference available to a user. Measured 2026-08-10, fresh R session per arm, against the live corpus at pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation. Three findings the old table could not express: - A local mirror is 60-80x faster. A one-off question is ~12 s end to end remotely against ~0.15 s mirrored. Stated outright now, because it is a bigger and cheaper win for users than anything in the R code. - Opening the session is the LARGEST remote cost (~7.5 s), bigger than any individual query, and it lands on the user's first query rather than on library(). The old table accounted for it nowhere, so every per-query figure was quietly missing it. - The remote cost is round-trips, not scanning: a repeat query over already-touched partitions is ~1.5 s against ~4 s cold, and a full-history query costs ~7 s whether it runs first or last (verified by running the arms in both orders). This is why #93's 1.4-1.7x, measured through cog-api against a local mount, does not show up on the remote path -- there, network latency swamps scan time. Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was ambiguous between MB and MiB besides. Measured from the manifest and on disk. The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release. Also documented HTTP 429: a burst of remote queries gets rate-limited by the host. Hit while taking these measurements. --- NEWS.md | 24 ++++++++++++++++++++++++ README.md | 34 +++++++++++++++++++++++++++++----- 2 files changed, 53 insertions(+), 5 deletions(-) diff --git a/NEWS.md b/NEWS.md index 6c3b108..f619841 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,5 +1,29 @@ # uscogdata 0.4.0 +## Documentation: the corpus-access table is re-measured and honest + +The README's "two ways to read the corpus" table carried figures taken before +the corpus was re-chunked into row groups (cog_pipeline#93, published +2026-08-09) and reported the mirrored column as "local speed" with no number at +all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit +3d28ddd`), fresh R session per arm: + +* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s + end to end remotely against ~0.15 s mirrored. That is the largest single + difference available to a user and it is now stated outright rather than left + as "local speed". +* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch + plus 23 view registrations over HTTPS), larger than any individual query, and + it lands on the first query rather than on `library(uscogdata)`. The old table + did not account for it anywhere. +* **The remote cost is round-trips, not scanning.** A repeat query over + already-touched partitions is ~1.5 s against ~4 s cold, and a full-history + query costs ~7 s whether it runs first or last. +* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4% + and the old figure was ambiguous between MB and MiB besides. +* Documented that a burst of remote queries can be rate-limited by the host + (`HTTP 429`), which is another reason to mirror for real work. + ## Cohorts can be named by predicate, not just by id `cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state` diff --git a/README.md b/README.md index 7590d76..4fa7b73 100644 --- a/README.md +++ b/README.md @@ -18,7 +18,7 @@ carries provenance describing what was converted, what was aggregated, and which known series breaks intersect your query. **Scope:** government types 0–3 (state, county, municipality, township). -56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968 +56 fiscal years, 46,148,034 rows, ~201 MB. There is no source data for FY1968 or FY1969. Special districts (type 4) and school districts (type 5) are excluded pending validation. @@ -85,14 +85,38 @@ cog_explain(spend) | | Remote (default) | Mirrored | |---|---|---| -| Setup | none | `cog_mirror(dest)`, 190.6 MB once | -| Disk used | **0 MB** — HTTP range requests only | 190.6 MB | -| Per query | ~4 s (one government, one year)
~6 s (one government, 23 years) | local speed | +| Setup | none | `cog_mirror(dest)`, ~201 MB once | +| Disk used | **0 MB** — HTTP range requests only | ~201 MB | +| Opening a session | ~7.5 s | ~0.1 s | +| One government, one year | ~4 s | ~0.05 s | +| One government, full history | ~7 s | ~0.1 s | +| Later queries, same session | ~1.5 s | ~0.05 s | | Good for | trying it out, teaching, one-off questions | repeated analysis, offline work, reproducibility | +**A local mirror is roughly 60–80x faster, and it is one function call.** That is +by far the largest difference any of these settings makes. If you are going to +ask more than a handful of questions, mirror first. + +Measured 2026-08-10 on a 16-core Linux workstation against the published corpus +(schema v7, `pipeline_commit 3d28ddd`), fresh R session per arm. A one-off +question costs about **12 seconds end to end remotely and 0.15 seconds +mirrored**, session setup included. + +Two things the per-query rows hide: + +- **Opening the session is the single largest remote cost** — larger than any + one query. It fetches the manifest and registers 23 SQL views over HTTPS, and + it lands on your first query, not on `library(uscogdata)`. +- **The cost is network round-trips, not scanning.** A repeat query against + partitions this session has already touched is ~1.5 s rather than ~4 s, and a + full-history query costs ~7 s whether it runs first or last. What you are + paying for is reaching each of the 56 yearly files over HTTPS the first time. + Nothing is written to disk in remote mode: DuckDB fetches the parquet footer, works out which row groups it needs, and reads only those. Nothing is cached -between sessions either, so every query goes back to the network. +between sessions either, so every query goes back to the network — and a session +that issues many remote queries in quick succession can be rate-limited by the +host (`HTTP Error: ... 429`). Both are further reasons to mirror for real work. The default points at a public HuggingFace mirror of the corpus. If you would rather not depend on a third party — for reproducibility, for an air-gapped