Partition-level caching: R/cache.R is still a stub, and the remote path pays for it every session #64

Open
opened 2026-08-10 19:10:59 -04:00 by jared · 0 comments
Owner

Carried out of the #56 performance pass, which named this "likely the single
highest-leverage change"
on the remote path and then deferred it. Filing so it survives
that issue closing.

The finding

R/cache.R is a six-line stub:

# Local partition cache; SHA-based invalidation.
# Phase N v0.1 implementation: DuckDB httpfs handles actual reads directly
# from Nextcloud; cache_dir holds only manifest.json. Richer partition
# caching (pre-fetch hot partitions) is a v0.2 feature.
# Stub here for cog_mirror to compose against.

USCOGDATA_CACHE_DIR therefore holds exactly one file — manifest.json. Nothing else is
cached between sessions, so every query goes back to the network, and every session
re-pays the cost of touching partitions it touched last time.

Why it is the leverage

The #56 pass established that the reader's cost on the default (remote) path is dominated
by HTTP, not by the R pipeline: the same cog_spending() call is 66 ms against a local
mirror and 3.9 s remote
, ~50–60x. Everything else the pass looked at — provenance
assembly, crosswalk joins, the suggestion and suppression machinery — sums to tens of
milliseconds and cannot move that number. Caching bytes locally is the only thing in this
package that addresses the actual cost.

Measured on efron against the live corpus (schema v7, pipeline_commit 3d28ddd), fresh
R session per arm, Madison WI:

remote (default) local mirror
session open ~7.5 s 0.10 s
first query, one year ~3.9 s 0.05 s
same query again, same session ~1.4 s 0.05 s
full history (56 partitions) ~5.9 s 0.08 s

Two things that table shows which the issue text should not lose:

  1. Session open is ~7.5 s — larger than any single query, and paid before the user's
    first result. That is the manifest fetch plus registering 23 SQL views over HTTP.
  2. A repeat query is ~3x cheaper than the first within a session, so DuckDB is already
    getting value from holding parquet footers in memory. A persistent cache would extend
    that across sessions instead of discarding it at every exit.

Shape

Whatever is built should decide explicitly:

  • What is cached — footers only (small, cheap, most of the session-open and
    first-touch win) versus whole partitions (190 MB, at which point it converges on
    cog_mirror()).
  • Invalidation — the manifest already carries a per-partition sha256, so correctness
    is available without guessing. Use it; do not invent a TTL.
  • Whether this is just cog_mirror() with better ergonomics. A serious answer might
    be "no new cache; make mirroring the documented default for repeat work" — the README
    now says outright that a mirror is dramatically faster. That would be a legitimate
    resolution of this issue rather than a dodge, and it is cheaper than building a cache.

Not a blocker

Filed as follow-up, not as a release gate. cog_mirror() already gives a user the whole
win in one function call, and the README documents it. This is about the default path
being better for people who never mirror.

Related: #56 (the pass this came from), cog_pipeline#93 (row-group layout, the other half
of the remote-read cost, shipped and published).

Carried out of the #56 performance pass, which named this **"likely the single highest-leverage change"** on the remote path and then deferred it. Filing so it survives that issue closing. ## The finding `R/cache.R` is a six-line stub: ```r # Local partition cache; SHA-based invalidation. # Phase N v0.1 implementation: DuckDB httpfs handles actual reads directly # from Nextcloud; cache_dir holds only manifest.json. Richer partition # caching (pre-fetch hot partitions) is a v0.2 feature. # Stub here for cog_mirror to compose against. ``` `USCOGDATA_CACHE_DIR` therefore holds exactly one file — `manifest.json`. Nothing else is cached between sessions, so **every query goes back to the network**, and every session re-pays the cost of touching partitions it touched last time. ## Why it is the leverage The #56 pass established that the reader's cost on the default (remote) path is dominated by HTTP, not by the R pipeline: the same `cog_spending()` call is **66 ms against a local mirror and 3.9 s remote**, ~50–60x. Everything else the pass looked at — provenance assembly, crosswalk joins, the suggestion and suppression machinery — sums to tens of milliseconds and cannot move that number. Caching bytes locally is the only thing in this package that addresses the actual cost. Measured on `efron` against the live corpus (schema v7, `pipeline_commit 3d28ddd`), fresh R session per arm, Madison WI: | | remote (default) | local mirror | |---|---:|---:| | session open | ~7.5 s | 0.10 s | | first query, one year | ~3.9 s | 0.05 s | | same query again, same session | ~1.4 s | 0.05 s | | full history (56 partitions) | ~5.9 s | 0.08 s | Two things that table shows which the issue text should not lose: 1. **Session open is ~7.5 s** — larger than any single query, and paid before the user's first result. That is the manifest fetch plus registering 23 SQL views over HTTP. 2. **A repeat query is ~3x cheaper than the first** within a session, so DuckDB is already getting value from holding parquet footers in memory. A persistent cache would extend that across sessions instead of discarding it at every exit. ## Shape Whatever is built should decide explicitly: - **What is cached** — footers only (small, cheap, most of the session-open and first-touch win) versus whole partitions (190 MB, at which point it converges on `cog_mirror()`). - **Invalidation** — the manifest already carries a per-partition `sha256`, so correctness is available without guessing. Use it; do not invent a TTL. - **Whether this is just `cog_mirror()` with better ergonomics.** A serious answer might be "no new cache; make mirroring the documented default for repeat work" — the README now says outright that a mirror is dramatically faster. That would be a legitimate resolution of this issue rather than a dodge, and it is cheaper than building a cache. ## Not a blocker Filed as follow-up, not as a release gate. `cog_mirror()` already gives a user the whole win in one function call, and the README documents it. This is about the default path being better for people who never mirror. Related: #56 (the pass this came from), cog_pipeline#93 (row-group layout, the other half of the remote-read cost, shipped and published).
jared added the
origin
review
type
feature
ws
corpus
labels 2026-08-23 16:07:40 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#64