docs: re-measure the corpus-access table against the published corpus (#56)
figures predate the row-group rechunk (cog_pipeline#93, published 2026-08-09) and reported the mirrored column as 'local speed' with no number -- hiding the largest difference available to a user. Measured 2026-08-10, fresh R session per arm, against the live corpus at pipeline_commit 3d28ddd. Madison WI, 16-core Linux workstation. Three findings the old table could not express: - A local mirror is 60-80x faster. A one-off question is ~12 s end to end remotely against ~0.15 s mirrored. Stated outright now, because it is a bigger and cheaper win for users than anything in the R code. - Opening the session is the LARGEST remote cost (~7.5 s), bigger than any individual query, and it lands on the user's first query rather than on library(). The old table accounted for it nowhere, so every per-query figure was quietly missing it. - The remote cost is round-trips, not scanning: a repeat query over already-touched partitions is ~1.5 s against ~4 s cold, and a full-history query costs ~7 s whether it runs first or last (verified by running the arms in both orders). This is why #93's 1.4-1.7x, measured through cog-api against a local mount, does not show up on the remote path -- there, network latency swamps scan time. Corpus size corrected to ~201 MB: row-group chunking added ~3.4%, and 190.6 was ambiguous between MB and MiB besides. Measured from the manifest and on disk. The 0.3.0 NEWS section keeps 190.6 -- it was correct for that release. Also documented HTTP 429: a burst of remote queries gets rate-limited by the host. Hit while taking these measurements.
This commit is contained in:
@@ -41,6 +41,30 @@ Two refusals rather than silent surprises:
|
||||
per requested name with a sidecar covering all of them; a page of that is not
|
||||
a page of anything the caller asked for.
|
||||
|
||||
## Documentation: the corpus-access table is re-measured and honest
|
||||
|
||||
The README's "two ways to read the corpus" table carried figures taken before
|
||||
the corpus was re-chunked into row groups (cog_pipeline#93, published
|
||||
2026-08-09) and reported the mirrored column as "local speed" with no number at
|
||||
all. Re-measured 2026-08-10 against the published corpus (`pipeline_commit
|
||||
3d28ddd`), fresh R session per arm:
|
||||
|
||||
* **A local mirror is roughly 60-80x faster.** A one-off question costs ~12 s
|
||||
end to end remotely against ~0.15 s mirrored. That is the largest single
|
||||
difference available to a user and it is now stated outright rather than left
|
||||
as "local speed".
|
||||
* **Opening the session is the largest remote cost** (~7.5 s -- manifest fetch
|
||||
plus 23 view registrations over HTTPS), larger than any individual query, and
|
||||
it lands on the first query rather than on `library(uscogdata)`. The old table
|
||||
did not account for it anywhere.
|
||||
* **The remote cost is round-trips, not scanning.** A repeat query over
|
||||
already-touched partitions is ~1.5 s against ~4 s cold, and a full-history
|
||||
query costs ~7 s whether it runs first or last.
|
||||
* The corpus size is **~201 MB**, not 190.6 MB -- row-group chunking added ~3.4%
|
||||
and the old figure was ambiguous between MB and MiB besides.
|
||||
* Documented that a burst of remote queries can be rate-limited by the host
|
||||
(`HTTP 429`), which is another reason to mirror for real work.
|
||||
|
||||
## Cohorts can be named by predicate, not just by id
|
||||
|
||||
`cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state`
|
||||
|
||||
Reference in New Issue
Block a user