#56 step 4: "Re-measure and update the README table."
The old figures were taken 2026-08-08, before the corpus was re-chunked into row groups
(cog_pipeline#93, published 2026-08-09), and the mirrored column read local speed with
no number at all — hiding the largest difference available to a user.
Method
Measured 2026-08-10 on a 16-core Linux workstation, fresh R session per arm, against
the live published corpus (schema v7, pipeline_commit 3d28ddd) and a byte-identical
local mirror of the same build. Madison WI (552025209777), the README quickstart's own
example. Session open is timed separately so it is never folded into a per-query number.
Findings
A local mirror is 60–80x faster. A one-off question costs ~12 s end to end remotely
against ~0.15 s mirrored. The #56 comment argued the README "arguably should say so
outright"; it now does, in bold, above the fold of that section.
Opening the session is the largest remote cost — ~7.5 s. Larger than any individual
query. It fetches the manifest and registers 23 SQL views over HTTPS, and it lands on the
user's first query, not on library(uscogdata). The old table accounted for it nowhere,
so every per-query figure it printed was quietly missing it.
The remote cost is round-trips, not scanning. Verified by running the arms in both
orders: a full-history query costs ~7 s whether it runs first or last, while a one-year
query drops from ~4 s to ~1.5 s once the session has touched those partitions. So the
1.4–1.7x from cog_pipeline#93 — measured through cog-api against a local mount, where
scan time dominates — does not show up on the remote path, where network latency swamps
it. Worth saying plainly rather than implying a speedup that did not reach this path.
Corpus size corrected to ~201 MB. Row-group chunking added ~3.4%, and 190.6 was
ambiguous between MB and MiB besides. Taken from the manifest (56 partitions, 197.8 MB)
and from disk (whole mirrored tree, 200.9 MB). The 0.3.0 NEWS section keeps 190.6 MB —
that was correct for that release, and a changelog should not be rewritten.
HTTP 429. A burst of remote queries gets rate-limited by the host. Hit while taking
these measurements, so it is now documented as another reason to mirror.
Numbers
Remote (default)
Mirrored
Opening a session
~7.5 s
~0.1 s
One government, one year
~4 s
~0.05 s
One government, full history
~7 s
~0.1 s
Later queries, same session
~1.5 s
~0.05 s
One-off question, end to end
~12 s
~0.15 s
Run-to-run spread across three sessions per arm was under 10% on the remote side and
under 0.02 s on the local side.
Verification
test-config.R asserts on README content; suite subset covering it and cog_mirror()
passes 54/54, 0 failures, 0 warnings.
#56 step 4: "Re-measure and update the README table."
The old figures were taken 2026-08-08, before the corpus was re-chunked into row groups
(`cog_pipeline#93`, published 2026-08-09), and the mirrored column read `local speed` with
no number at all — hiding the largest difference available to a user.
## Method
Measured 2026-08-10 on a 16-core Linux workstation, **fresh R session per arm**, against
the live published corpus (schema v7, `pipeline_commit 3d28ddd`) and a byte-identical
local mirror of the same build. Madison WI (`552025209777`), the README quickstart's own
example. Session open is timed separately so it is never folded into a per-query number.
## Findings
**A local mirror is 60–80x faster.** A one-off question costs ~12 s end to end remotely
against ~0.15 s mirrored. The #56 comment argued the README "arguably should say so
outright"; it now does, in bold, above the fold of that section.
**Opening the session is the largest remote cost — ~7.5 s.** Larger than any individual
query. It fetches the manifest and registers 23 SQL views over HTTPS, and it lands on the
user's *first query*, not on `library(uscogdata)`. The old table accounted for it nowhere,
so every per-query figure it printed was quietly missing it.
**The remote cost is round-trips, not scanning.** Verified by running the arms in both
orders: a full-history query costs ~7 s whether it runs first or last, while a one-year
query drops from ~4 s to ~1.5 s once the session has touched those partitions. So the
1.4–1.7x from `cog_pipeline#93` — measured through cog-api against a *local* mount, where
scan time dominates — does not show up on the remote path, where network latency swamps
it. Worth saying plainly rather than implying a speedup that did not reach this path.
**Corpus size corrected to ~201 MB.** Row-group chunking added ~3.4%, and `190.6` was
ambiguous between MB and MiB besides. Taken from the manifest (56 partitions, 197.8 MB)
and from disk (whole mirrored tree, 200.9 MB). The 0.3.0 NEWS section keeps `190.6 MB` —
that was correct for that release, and a changelog should not be rewritten.
**HTTP 429.** A burst of remote queries gets rate-limited by the host. Hit while taking
these measurements, so it is now documented as another reason to mirror.
## Numbers
| | Remote (default) | Mirrored |
|---|---:|---:|
| Opening a session | ~7.5 s | ~0.1 s |
| One government, one year | ~4 s | ~0.05 s |
| One government, full history | ~7 s | ~0.1 s |
| Later queries, same session | ~1.5 s | ~0.05 s |
| **One-off question, end to end** | **~12 s** | **~0.15 s** |
Run-to-run spread across three sessions per arm was under 10% on the remote side and
under 0.02 s on the local side.
## Verification
`test-config.R` asserts on README content; suite subset covering it and `cog_mirror()`
passes 54/54, 0 failures, 0 warnings.
figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09)
and reported the mirrored column as 'local speed' with no number -- hiding the
largest difference available to a user.
Measured 2026-08-10, fresh R session per arm, against the live corpus at
pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation.
Three findings the old table could not express:
- A local mirror is 60-80x faster. A one-off question is ~12 s end to end
remotely against ~0.15 s mirrored. Stated outright now, because it is a
bigger and cheaper win for users than anything in the R code.
- Opening the session is the LARGEST remote cost (~7.5 s), bigger than any
individual query, and it lands on the user's first query rather than on
library(). The old table accounted for it nowhere, so every per-query figure
was quietly missing it.
- The remote cost is round-trips, not scanning: a repeat query over
already-touched partitions is ~1.5 s against ~4 s cold, and a full-history
query costs ~7 s whether it runs first or last (verified by running the arms
in both orders). This is why #93's 1.4-1.7x, measured through cog-api against
a local mount, does not show up on the remote path -- there, network latency
swamps scan time.
Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was
ambiguous between MB and MiB besides. Measured from the manifest and on disk.
The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release.
Also documented HTTP 429: a burst of remote queries gets rate-limited by the
host. Hit while taking these measurements.
The three 0.4.0 features landed in the order their PRs merged, which buried the
headline change (cohort predicates, 4.8x) below an operator config knob and a
documentation note. Reordered to: cohorts, pagination, DuckDB budget, docs,
Fixes -- with Fixes last, where it was already.
The stacked PRs each appended their own '## Fixes' heading, so resolving the
conflicts also folded two of them into the single section that belongs there.
Content is byte-identical to what merged; only section order changed. Verified
by diffing the sorted non-blank lines against the previous commit.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
#56 step 4: "Re-measure and update the README table."
The old figures were taken 2026-08-08, before the corpus was re-chunked into row groups
(
cog_pipeline#93, published 2026-08-09), and the mirrored column readlocal speedwithno number at all — hiding the largest difference available to a user.
Method
Measured 2026-08-10 on a 16-core Linux workstation, fresh R session per arm, against
the live published corpus (schema v7,
pipeline_commit 3d28ddd) and a byte-identicallocal mirror of the same build. Madison WI (
552025209777), the README quickstart's ownexample. Session open is timed separately so it is never folded into a per-query number.
Findings
A local mirror is 60–80x faster. A one-off question costs ~12 s end to end remotely
against ~0.15 s mirrored. The #56 comment argued the README "arguably should say so
outright"; it now does, in bold, above the fold of that section.
Opening the session is the largest remote cost — ~7.5 s. Larger than any individual
query. It fetches the manifest and registers 23 SQL views over HTTPS, and it lands on the
user's first query, not on
library(uscogdata). The old table accounted for it nowhere,so every per-query figure it printed was quietly missing it.
The remote cost is round-trips, not scanning. Verified by running the arms in both
orders: a full-history query costs ~7 s whether it runs first or last, while a one-year
query drops from ~4 s to ~1.5 s once the session has touched those partitions. So the
1.4–1.7x from
cog_pipeline#93— measured through cog-api against a local mount, wherescan time dominates — does not show up on the remote path, where network latency swamps
it. Worth saying plainly rather than implying a speedup that did not reach this path.
Corpus size corrected to ~201 MB. Row-group chunking added ~3.4%, and
190.6wasambiguous between MB and MiB besides. Taken from the manifest (56 partitions, 197.8 MB)
and from disk (whole mirrored tree, 200.9 MB). The 0.3.0 NEWS section keeps
190.6 MB—that was correct for that release, and a changelog should not be rewritten.
HTTP 429. A burst of remote queries gets rate-limited by the host. Hit while taking
these measurements, so it is now documented as another reason to mirror.
Numbers
Run-to-run spread across three sessions per arm was under 10% on the remote side and
under 0.02 s on the local side.
Verification
test-config.Rasserts on README content; suite subset covering it andcog_mirror()passes 54/54, 0 failures, 0 warnings.
8bd16bb085to303aa59b07