Performance pass over the reader before final release #56

Closed
opened 2026-08-09 09:18:47 -04:00 by jared · 3 comments
Owner

Gate on the final release. Tagging and r-universe registration (#47) wait for this.

Why now

The package is correct — 972 tests, R CMD check clean, 4/4 platforms green — but nothing has ever been profiled. The first public release sets expectations, and a reader that feels slow on a first query is the impression people keep.

What is actually measured

Against the live corpus from a workstation on a good connection, 2026-08-08:

Operation Time
Raw parquet scan, 1 partition, filtered by govid 1.5 s
Raw parquet scan, all 56 partitions, filtered by govid 2.8 s
cog_spending(), 1 government × 1 year 3.9 s
cog_spending(), 1 government × 23 years 5.9 s
cog_revenue(), 1 government × 23 years 2.0 s
cog_spending(), per-capita + real dollars, 23 years 3.1 s

The gap between a raw scan and a verb is the target: roughly 2–3 s of overhead above the data access itself, on every call. The variance is also unexplained — revenue over 23 years is faster than spending over 1 year, which suggests the cost is not dominated by I/O.

These numbers are in the README's "two ways to read the corpus" table, so anything that moves materially means updating that table.

Candidate hot spots, unverified

Guesses to be confirmed or killed by profiling, not treated as a work list:

  • Session setup — cog_open() fetches the manifest and registers 23 SQL views on first use. Paid once per session, but it lands on the user's first query.
  • Partition enumeration — the fix in #41 embeds all 56 fully-qualified URLs into the long view's SQL. Correct, but it is a large string and DuckDB re-plans it per query. Worth checking whether a year predicate prunes before the file list is opened, or after.
  • Suggestion machinery — #35 already records that the suppression query runs on every category-scoped call. #34 records that the candidate query is scoped too broadly. Both are performance findings filed before anyone was measuring.
  • Provenance assembly — every result builds coverage, series-break and corpus-break references. Several separate queries per call.
  • Crosswalk joins — canonical_fips_xwalk and summary_categories join on every verb.

Approach

  1. Profile a representative call with profvis against both a local mirror and the remote corpus — the split matters, because remote latency and local compute need different fixes.
  2. Separate per-session cost from per-query cost. If most of the 3.9 s is session warm-up, the fix is documentation and a warm-up helper, not query tuning.
  3. Fix what profiling actually implicates. Absorb #34 and #35 if they are confirmed; close them as speculative if they are not.
  4. Re-measure and update the README table.

Related

  • #33, #34, #35 — pre-existing issues in this area
  • Downstream: the API has its own pass (cog-api), since it adds an HTTP layer and its own caching over these same verbs
**Gate on the final release.** Tagging and r-universe registration (#47) wait for this. ## Why now The package is correct — 972 tests, `R CMD check` clean, 4/4 platforms green — but nothing has ever been profiled. The first public release sets expectations, and a reader that feels slow on a first query is the impression people keep. ## What is actually measured Against the live corpus from a workstation on a good connection, 2026-08-08: | Operation | Time | |---|---| | Raw parquet scan, 1 partition, filtered by govid | 1.5 s | | Raw parquet scan, all 56 partitions, filtered by govid | 2.8 s | | `cog_spending()`, 1 government × 1 year | **3.9 s** | | `cog_spending()`, 1 government × 23 years | **5.9 s** | | `cog_revenue()`, 1 government × 23 years | 2.0 s | | `cog_spending()`, per-capita + real dollars, 23 years | 3.1 s | The gap between a raw scan and a verb is the target: roughly **2–3 s of overhead** above the data access itself, on every call. The variance is also unexplained — revenue over 23 years is faster than spending over 1 year, which suggests the cost is not dominated by I/O. These numbers are in the README's "two ways to read the corpus" table, so anything that moves materially means updating that table. ## Candidate hot spots, unverified Guesses to be confirmed or killed by profiling, not treated as a work list: - **Session setup** — `cog_open()` fetches the manifest and registers 23 SQL views on first use. Paid once per session, but it lands on the user's first query. - **Partition enumeration** — the fix in #41 embeds all 56 fully-qualified URLs into the `long` view's SQL. Correct, but it is a large string and DuckDB re-plans it per query. Worth checking whether a year predicate prunes before the file list is opened, or after. - **Suggestion machinery** — #35 already records that the suppression query runs on *every* category-scoped call. #34 records that the candidate query is scoped too broadly. Both are performance findings filed before anyone was measuring. - **Provenance assembly** — every result builds coverage, series-break and corpus-break references. Several separate queries per call. - **Crosswalk joins** — `canonical_fips_xwalk` and `summary_categories` join on every verb. ## Approach 1. Profile a representative call with `profvis` against **both** a local mirror and the remote corpus — the split matters, because remote latency and local compute need different fixes. 2. Separate **per-session** cost from **per-query** cost. If most of the 3.9 s is session warm-up, the fix is documentation and a warm-up helper, not query tuning. 3. Fix what profiling actually implicates. Absorb #34 and #35 if they are confirmed; close them as speculative if they are not. 4. Re-measure and update the README table. ## Related - #33, #34, #35 — pre-existing issues in this area - Downstream: the API has its own pass (`cog-api`), since it adds an HTTP layer and its own caching over these same verbs
Author
Owner

Evidence from the cog-api side: the overhead is remote I/O, not compute

While doing a performance pass on cog-api (which is a thin wrapper over these verbs), I
measured the same call shapes against a local corpus mirror. The comparison answers
the "profile against both a local mirror and the remote corpus — the split matters"
question in the Approach section above, and I think it reframes this issue.

Measured 2026-08-09 on efron (16 cores), full production corpus (schema v7, 56
partitions, 46.1M rows) bind-mounted locally. These are end-to-end HTTP requests
through cog-api — so each number includes plumber dispatch, parameter validation, the
verb call, and JSON envelope construction. The verb itself is therefore faster than
what is shown:

Call This issue's table (remote) Measured via cog-api (local mirror)
cog_spending(), 1 government × 1 year 3.9 s 66 ms
cog_spending(), 1 government × 5 years — 96 ms
cog_spending(), 1 government × all 56 years 5.9 s (23 yr) 337 ms

Roughly 50–60x faster on a local mirror, and the shape of the remaining cost is
clean: about 50 ms fixed plus ~5 ms per year partition scanned.

Conclusion: the 2–3 s of "overhead above data access" is dominated by remote HTTP
round-trips, not by the verb pipeline.
Session setup, provenance assembly, crosswalk
joins and the suggestion machinery are all real costs, but on a local corpus they sum to
tens of milliseconds, not seconds. I would not spend the release gate on micro-optimising
the R pipeline; the leverage is in how many HTTP range requests a query makes against the
remote corpus.

That suggests reframing the candidate list toward:

  • Round-trip count per query. With 56 fully-qualified URLs in the long view, how
    many separate HTTP range requests does a single-year query actually issue, and does the
    year predicate prune before the file list is opened? On a local FS this is invisible;
    over HTTPS it is the whole cost.
  • A documented warm-up / mirror path. cog_mirror() already exists. If a local mirror
    is 50x faster, the README's "two ways to read the corpus" table arguably should say so
    outright — that is a bigger, cheaper win for users than anything in the R code.
  • Partition-level caching. R/cache.R is still a stub deferring this. On the remote
    path it is likely the single highest-leverage change.

A separate finding: sort canonical_govid within each year partition

Not an uscogdata change — this is a corpus-layout change for whichever pipeline writes
the parquet — but it showed up clearly in the same measurements and is worth recording
where the performance discussion lives.

A single-government query has to scan every year partition, because nothing tells DuckDB
where that govid lives. If rows were sorted by canonical_govid within each year
partition
, parquet row-group statistics (min/max per row group) would let DuckDB skip
almost every row group for a single-govid predicate. The per-partition cost would fall
from "decompress and filter the partition" to "read the footer, open one or two row
groups."

That is the largest remaining structural win I can identify for /profile and
per-government /spending, which are the slowest routes in the API and the ones that do
not benefit from adding server replicas (measured: 1.27x from a second replica, versus
1.90x for bounded-work endpoints — they are resource-bound on the scan, not
concurrency-bound).

Context

Full write-up, method and raw numbers:
docs/benchmarks/ in Civilytics/cog-api (2026-08-baseline.md,
2026-08-phase3-replicas.md, 2026-08-final.md), reproducible with scripts/bench.sh
in that repo.

## Evidence from the cog-api side: the overhead is remote I/O, not compute While doing a performance pass on `cog-api` (which is a thin wrapper over these verbs), I measured the same call shapes against a **local corpus mirror**. The comparison answers the "profile against both a local mirror and the remote corpus — the split matters" question in the Approach section above, and I think it reframes this issue. Measured 2026-08-09 on `efron` (16 cores), full production corpus (schema v7, 56 partitions, 46.1M rows) bind-mounted **locally**. These are **end-to-end HTTP requests** through cog-api — so each number includes plumber dispatch, parameter validation, the verb call, and JSON envelope construction. The verb itself is therefore *faster* than what is shown: | Call | This issue's table (remote) | Measured via cog-api (local mirror) | |---|---:|---:| | `cog_spending()`, 1 government × 1 year | 3.9 s | **66 ms** | | `cog_spending()`, 1 government × 5 years | — | 96 ms | | `cog_spending()`, 1 government × all 56 years | 5.9 s (23 yr) | **337 ms** | Roughly **50–60x faster on a local mirror**, and the shape of the remaining cost is clean: about 50 ms fixed plus ~5 ms per year partition scanned. **Conclusion: the 2–3 s of "overhead above data access" is dominated by remote HTTP round-trips, not by the verb pipeline.** Session setup, provenance assembly, crosswalk joins and the suggestion machinery are all real costs, but on a local corpus they sum to tens of milliseconds, not seconds. I would not spend the release gate on micro-optimising the R pipeline; the leverage is in how many HTTP range requests a query makes against the remote corpus. That suggests reframing the candidate list toward: - **Round-trip count per query.** With 56 fully-qualified URLs in the `long` view, how many separate HTTP range requests does a single-year query actually issue, and does the year predicate prune *before* the file list is opened? On a local FS this is invisible; over HTTPS it is the whole cost. - **A documented warm-up / mirror path.** `cog_mirror()` already exists. If a local mirror is 50x faster, the README's "two ways to read the corpus" table arguably should say so outright — that is a bigger, cheaper win for users than anything in the R code. - **Partition-level caching.** `R/cache.R` is still a stub deferring this. On the remote path it is likely the single highest-leverage change. ## A separate finding: sort `canonical_govid` within each year partition Not an `uscogdata` change — this is a corpus-layout change for whichever pipeline writes the parquet — but it showed up clearly in the same measurements and is worth recording where the performance discussion lives. A single-government query has to scan every year partition, because nothing tells DuckDB where that govid lives. If rows were **sorted by `canonical_govid` within each year partition**, parquet row-group statistics (min/max per row group) would let DuckDB skip almost every row group for a single-govid predicate. The per-partition cost would fall from "decompress and filter the partition" to "read the footer, open one or two row groups." That is the largest remaining structural win I can identify for `/profile` and per-government `/spending`, which are the slowest routes in the API and the ones that do not benefit from adding server replicas (measured: 1.27x from a second replica, versus 1.90x for bounded-work endpoints — they are resource-bound on the scan, not concurrency-bound). ## Context Full write-up, method and raw numbers: `docs/benchmarks/` in `Civilytics/cog-api` (`2026-08-baseline.md`, `2026-08-phase3-replicas.md`, `2026-08-final.md`), reproducible with `scripts/bench.sh` in that repo.
Author
Owner

Closing. Against this issue's own four-step Approach.

1. Profile against both a local mirror and the remote corpus

Done, and it overturned this issue's framing. The "2–3 s of overhead above the data
access itself" is not the verb pipeline — it is remote HTTP round-trips. The same
cog_spending() call measures 66 ms against a local mirror and 3.9 s remote, ~50–60x.
Session setup, provenance assembly, crosswalk joins and the suggestion machinery are all
real costs, and on a local corpus they sum to tens of milliseconds.

The candidate list in this issue was therefore mostly a list of things not worth fixing.
That is a good outcome for a profiling pass — it is what stopped the release gate being
spent on micro-optimising R.

Re-measured again 2026-08-10 against the current corpus (pipeline_commit 3d28ddd), fresh
session per arm:

remote (default) mirrored
opening a session ~7.5 s ~0.1 s
one government, one year ~4 s ~0.05 s
one government, full history ~7 s ~0.1 s
later queries, same session ~1.5 s ~0.05 s
one-off question, end to end ~12 s ~0.15 s

2. Separate per-session from per-query cost

Done, and it produced the finding this issue guessed at but understated. Session open is
the single largest remote cost (~7.5 s)
— larger than any individual query, and it lands
on the user's first query rather than on library(). The old README table accounted for it
nowhere, so every per-query number it printed was quietly missing it.

This issue anticipated that "if most of the 3.9 s is session warm-up, the fix is
documentation and a warm-up helper." It is a large share, but no warm-up helper was
built, deliberately
: cog_open() stays unexported, because a helper would only let a
user move the 7.5 s earlier, not remove it. Mirroring removes it. Documentation was the
right half of that prediction; the helper was not. Recording the reasoning so it reads as
decided rather than forgotten.

3. Fix what profiling actually implicates

outcome
#58 cohort predicates Shipped (PR #61). 4.8x, and at the no-filter floor — 102 ms against 105 ms for no cohort restriction at all. End to end through the verb, 3.99x.
#59 rollup pushdown Closed as measured-not-there. Targets ~1% of runtime: the many-to-many join is 6 ms and .coverage_table() 7 ms against 1442 ms inside cog_spending(). Pagination cannot help either — GROUP BY completes before LIMIT, and both coverage computations need the full result.
#35 suppression skip Closed as speculative, per this issue's instruction. ~12.45 ms against seconds of remote I/O, and a first attempt had already shipped and been reverted for being net-negative.
#34 suggestion scoping Kept open, re-scoped. This issue files it as a performance finding; it is not one. Its actual content is a correctness defect — cog_revenue(category = "Corrections") offers expenditure recipes — which profiling leaves untouched. Retitled so it is not re-triaged as perf and dropped.
#33 decompose .build_suggestions() Untouched. A refactor, listed here only under "Related", and out of scope either way.
#57 pagination gap Shipped (PR #63). Also fixed cog_gov_search()'s ORDER BY, which was not a total order and made a paged sweep unsound.
#60 DuckDB thread cap Shipped (PR #62). Removes cog-api's getFromNamespace(".ensure_session", …) reach into package internals — a consumer depending on a private name, which should not be frozen at a public release.
corpus row-group layout Shipped upstream as cog_pipeline#93 and already published (manifest pipeline_commit 3d28ddd, built 2026-08-09T20:48Z). Worth noting the finding evolved: the corpus turned out to be already sorted by canonical_govid, so the sort was a no-op and single-row-group partitions were the whole problem.

4. Re-measure and update the README table

Done — PR #65. Three things the old table could not express: the mirror is 60–80x faster
(previously written as local speed, no number); session open is the largest remote cost;
and the remote cost is round-trips rather than scanning. Verified by running the arms in
both orders — a full-history query costs ~7 s whether it runs first or last, while a
one-year query drops from ~4 s to ~1.5 s once its partitions have been touched.

That last point is worth stating plainly: cog_pipeline#93's 1.4–1.7x does not appear
on the remote path.
It was measured through cog-api against a local mount, where scan
time dominates; over HTTPS, network latency swamps it. The win is real and it is why the
mirrored column is now so fast — it just is not a remote win.

Corpus size corrected to ~201 MB (row-group chunking added ~3.4%, and 190.6 was
ambiguous between MB and MiB). Also documented HTTP 429: a burst of remote queries gets
rate-limited by the host, which I hit while taking these measurements.

Carried forward, not lost

#64 — partition-level caching. My comment above called R/cache.R "likely the single
highest-leverage change" on the remote path and then deferred it, and it existed nowhere
else. Now filed with the measurements attached, including the honest option that the right
answer may be "no new cache; make mirroring the documented default," which the README
change already half-does.

maxwell capacity. All benchmarks ran on efron (16 cores). maxwell is 8 with 4
budgeted for the API, so production numbers are lower — scripts/bench.sh should be run
there before capacity is quoted publicly. Tracked on the cog-api side, not here.

The gate

This was the blocker on #47. It is cleared — with the note that #47's title said v0.3.0
while main is already 0.4.0
; corrected there.

Method and raw numbers: docs/benchmarks/ in Civilytics/cog-api
(2026-08-baseline.md, 2026-08-phase3-replicas.md, 2026-08-phase4-bulk.md,
2026-08-final.md), reproducible via scripts/bench.sh; plus the 2026-08-10 re-measure
recorded in PR #65.

## Closing. Against this issue's own four-step Approach. ### 1. Profile against both a local mirror and the remote corpus Done, and it **overturned this issue's framing**. The "2–3 s of overhead above the data access itself" is not the verb pipeline — it is remote HTTP round-trips. The same `cog_spending()` call measures **66 ms against a local mirror and 3.9 s remote**, ~50–60x. Session setup, provenance assembly, crosswalk joins and the suggestion machinery are all real costs, and on a local corpus they sum to tens of milliseconds. The candidate list in this issue was therefore mostly a list of things not worth fixing. That is a good outcome for a profiling pass — it is what stopped the release gate being spent on micro-optimising R. Re-measured again 2026-08-10 against the current corpus (`pipeline_commit 3d28ddd`), fresh session per arm: | | remote (default) | mirrored | |---|---:|---:| | opening a session | ~7.5 s | ~0.1 s | | one government, one year | ~4 s | ~0.05 s | | one government, full history | ~7 s | ~0.1 s | | later queries, same session | ~1.5 s | ~0.05 s | | **one-off question, end to end** | **~12 s** | **~0.15 s** | ### 2. Separate per-session from per-query cost Done, and it produced the finding this issue guessed at but understated. **Session open is the single largest remote cost (~7.5 s)** — larger than any individual query, and it lands on the user's first query rather than on `library()`. The old README table accounted for it nowhere, so every per-query number it printed was quietly missing it. This issue anticipated that "if most of the 3.9 s is session warm-up, the fix is documentation and a warm-up helper." It is a large share, but **no warm-up helper was built, deliberately**: `cog_open()` stays unexported, because a helper would only let a user move the 7.5 s earlier, not remove it. Mirroring removes it. Documentation was the right half of that prediction; the helper was not. Recording the reasoning so it reads as decided rather than forgotten. ### 3. Fix what profiling actually implicates | | outcome | |---|---| | **#58** cohort predicates | **Shipped** (PR #61). 4.8x, and at the no-filter floor — 102 ms against 105 ms for no cohort restriction at all. End to end through the verb, 3.99x. | | **#59** rollup pushdown | **Closed as measured-not-there.** Targets ~1% of runtime: the many-to-many join is 6 ms and `.coverage_table()` 7 ms against 1442 ms inside `cog_spending()`. Pagination cannot help either — `GROUP BY` completes before `LIMIT`, and both coverage computations need the full result. | | **#35** suppression skip | **Closed as speculative**, per this issue's instruction. ~12.45 ms against seconds of remote I/O, and a first attempt had already shipped and been reverted for being net-negative. | | **#34** suggestion scoping | **Kept open, re-scoped.** This issue files it as a performance finding; it is not one. Its actual content is a correctness defect — `cog_revenue(category = "Corrections")` offers *expenditure* recipes — which profiling leaves untouched. Retitled so it is not re-triaged as perf and dropped. | | **#33** decompose `.build_suggestions()` | Untouched. A refactor, listed here only under "Related", and out of scope either way. | | **#57** pagination gap | **Shipped** (PR #63). Also fixed `cog_gov_search()`'s `ORDER BY`, which was not a total order and made a paged sweep unsound. | | **#60** DuckDB thread cap | **Shipped** (PR #62). Removes cog-api's `getFromNamespace(".ensure_session", …)` reach into package internals — a consumer depending on a private name, which should not be frozen at a public release. | | corpus row-group layout | **Shipped upstream** as `cog_pipeline#93` and **already published** (manifest `pipeline_commit 3d28ddd`, built 2026-08-09T20:48Z). Worth noting the finding evolved: the corpus turned out to be *already sorted* by `canonical_govid`, so the sort was a no-op and single-row-group partitions were the whole problem. | ### 4. Re-measure and update the README table Done — PR #65. Three things the old table could not express: the mirror is 60–80x faster (previously written as `local speed`, no number); session open is the largest remote cost; and the remote cost is round-trips rather than scanning. Verified by running the arms in both orders — a full-history query costs ~7 s whether it runs first or last, while a one-year query drops from ~4 s to ~1.5 s once its partitions have been touched. That last point is worth stating plainly: **`cog_pipeline#93`'s 1.4–1.7x does not appear on the remote path.** It was measured through cog-api against a local mount, where scan time dominates; over HTTPS, network latency swamps it. The win is real and it is why the mirrored column is now so fast — it just is not a remote win. Corpus size corrected to **~201 MB** (row-group chunking added ~3.4%, and `190.6` was ambiguous between MB and MiB). Also documented `HTTP 429`: a burst of remote queries gets rate-limited by the host, which I hit while taking these measurements. ## Carried forward, not lost **#64 — partition-level caching.** My comment above called `R/cache.R` "likely the single highest-leverage change" on the remote path and then deferred it, and it existed nowhere else. Now filed with the measurements attached, including the honest option that the right answer may be "no new cache; make mirroring the documented default," which the README change already half-does. **maxwell capacity.** All benchmarks ran on `efron` (16 cores). maxwell is 8 with 4 budgeted for the API, so production numbers are lower — `scripts/bench.sh` should be run there before capacity is quoted publicly. Tracked on the cog-api side, not here. ## The gate This was the blocker on #47. It is cleared — with the note that **#47's title said v0.3.0 while `main` is already 0.4.0**; corrected there. Method and raw numbers: `docs/benchmarks/` in `Civilytics/cog-api` (`2026-08-baseline.md`, `2026-08-phase3-replicas.md`, `2026-08-phase4-bulk.md`, `2026-08-final.md`), reproducible via `scripts/bench.sh`; plus the 2026-08-10 re-measure recorded in PR #65.
Author
Owner

Closed. All four Approach steps done, everything merged.

main is at d2caa6d with all three PRs in:

PR
#62 #60 cog_open() honours a DuckDB thread and memory budget
#63 #57 limit/offset on cog_gov_search() and cog_balances()
#65 #56 step 4 corpus-access table re-measured

They stacked and collided on the NEWS.md 0.4.0 anchor. Resolved by rebasing each onto
the previous merge; the collision was more than textual — each branch had appended its own
## Fixes heading, so a naive merge would have shipped 0.4.0 with two of them. Folded into
one, and reordered the release's sections by user impact rather than merge order (cohorts,
pagination, DuckDB budget, docs, Fixes) — the headline 4.8x change had ended up fourth,
below a config knob. Content byte-identical, verified by diffing sorted non-blank lines.

Verified on merged main: suite 1084 passed, 0 failed, 0 warnings (2 pre-existing
test-live-corpus.R skips); R CMD check --as-cran 0 errors, 0 warnings, 1 NOTE — the
pre-existing one, whose r-universe 404 is #47 itself and clears on registration.

Final state of everything this issue touched

  • #58 shipped (4.8x, at the no-filter floor) · #59 closed, measured at ~1% of runtime
  • #57, #60 shipped — both were spawned by this pass and both gated the tag
  • #35 closed as speculative, per this issue's own instruction
  • #34 kept open but re-scoped: this issue filed it as performance, and it is not —
    the defect is cog_revenue() offering expenditure recipes, which profiling leaves
    untouched. Retitled so it is not re-triaged as perf and dropped
  • #64 filed for the deferred partition-cache work, which this issue called the highest
    remaining leverage and which existed nowhere else
  • cog_pipeline#93 shipped and published (pipeline_commit 3d28ddd)
  • #47 corrected from v0.3.0 to v0.4.0 — main had moved past 0.3.0

The one thing worth carrying forward

Every benchmark ran on efron (16 cores). maxwell is 8 with 4 budgeted for the API, so
production numbers are lower — run scripts/bench.sh there before quoting capacity
publicly. Not a blocker for the tag.

#47 is unblocked.

## Closed. All four Approach steps done, everything merged. `main` is at `d2caa6d` with all three PRs in: | PR | | | |---|---|---| | #62 | #60 | `cog_open()` honours a DuckDB thread and memory budget | | #63 | #57 | `limit`/`offset` on `cog_gov_search()` and `cog_balances()` | | #65 | #56 step 4 | corpus-access table re-measured | They stacked and collided on the `NEWS.md` 0.4.0 anchor. Resolved by rebasing each onto the previous merge; the collision was more than textual — each branch had appended its own `## Fixes` heading, so a naive merge would have shipped 0.4.0 with two of them. Folded into one, and reordered the release's sections by user impact rather than merge order (cohorts, pagination, DuckDB budget, docs, Fixes) — the headline 4.8x change had ended up fourth, below a config knob. Content byte-identical, verified by diffing sorted non-blank lines. **Verified on merged `main`:** suite **1084 passed, 0 failed, 0 warnings** (2 pre-existing `test-live-corpus.R` skips); `R CMD check --as-cran` **0 errors, 0 warnings, 1 NOTE** — the pre-existing one, whose r-universe 404 is #47 itself and clears on registration. ### Final state of everything this issue touched - **#58** shipped (4.8x, at the no-filter floor) · **#59** closed, measured at ~1% of runtime - **#57**, **#60** shipped — both were spawned by this pass and both gated the tag - **#35** closed as speculative, per this issue's own instruction - **#34** kept open but **re-scoped**: this issue filed it as performance, and it is not — the defect is `cog_revenue()` offering expenditure recipes, which profiling leaves untouched. Retitled so it is not re-triaged as perf and dropped - **#64** filed for the deferred partition-cache work, which this issue called the highest remaining leverage and which existed nowhere else - **cog_pipeline#93** shipped and published (`pipeline_commit 3d28ddd`) - **#47** corrected from v0.3.0 to **v0.4.0** — `main` had moved past 0.3.0 ### The one thing worth carrying forward Every benchmark ran on `efron` (16 cores). maxwell is 8 with 4 budgeted for the API, so production numbers are lower — run `scripts/bench.sh` there before quoting capacity publicly. Not a blocker for the tag. **#47 is unblocked.**
jared closed this issue 2026-08-10 19:39:13 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#56