docs: rewrite AGENTS.md and the README chart list for the rebuilt app
Deploy to git-pages / deploy (push) Successful in 23s

AGENTS.md described 6 charts, D3 selections, and DistrictVsNational /
ModelDrawsComparison / ExceedanceProbability — none of which exist. An agent
reading it as authoritative would have been actively misled, so it now opens by
saying src/ wins any disagreement.

Rewritten: the real component tree and data flow, ChartPanel as the owner of all
cross-chart state, the pooled/unpooled key namespace trap, the four properties
of the draws pipeline that are easy to break, the multi-part shard layout and
why discovery uses the tree listing, the "posterior predictive draws" wording
rule and the reason for it, the 95% convention, and a table of which tuning
decisions carry stated rationale and should not be re-derived (the CVD-validated
palette, the KDE bandwidth clamp, the pooling threshold, the mass/KDE cutoff,
the axis cap).

Adds a testing section — there are automated tests now — and notes that the
deployment check should use a multi-part state like California, since a
single-part state cannot catch a regression in part discovery. Records that
App.jsx's "Search another district" is a full page reload that discards the
shard cache.

README: the chart list becomes the summary table plus the two ported figures,
the file tree matches src/, the d3 role is stated precisely (scales and paths,
not selections), and the endpoint table warns that /estimates?state= ranks by
LEAID on a short read.

Also commits the plan this work followed.
This commit is contained in:
2026-08-12 08:41:01 -04:00
parent c1920c11fb
commit 4396a56a4a
3 changed files with 481 additions and 123 deletions
@@ -0,0 +1,219 @@
# CRDC demo — rebuild the visuals around the white-paper figures
## Context
The demo app at `pages.civilytics.org/crdc-demo/` works, but the charts don't show off the
modelling. The navigation (state → district) is good and stays. The three existing charts get
cut to one, replaced by a summary table and two charts ported from the white paper:
`wp_fig_group_density` (Fig 6) and `wp_fig_group_difference` (Fig 7) in
`crdc-arrests/R/paper_figures.R:472-568`.
Two real defects surfaced while scoping:
1. **The "suggested districts" list is ranked over a truncated set.** `DistrictSearch.jsx:22`
calls `/estimates?state=XX&year=21-22&limit=500`, but that endpoint returns rows
`ORDER BY LEAID` (`crdc-arrests/api/R/handlers_estimates.R:46`) at 8 rows per district. So
the app ranks the ~62 lowest-LEAID districts in the state. California has 11,488 rows.
2. **The analytic fallback assumes 90% bounds but the API returns 95%.**
`distributionApprox.js` defaults `intervalMass = 0.90`; the API default is `interval=95`
(`crdc-arrests/api/R/handlers_estimates.R`). The new charts compute intervals from real
draws, so this only affects the fallback path — fix the default while we're in there.
### What the data supports (verified, not assumed)
- The HF parquet carries `LEAID, RACE, SEX, pred, draw_id, subgroup_id, batch_num`. `draw_id`
is dense 1–500 for every group. The current query at `useDrawDistribution.js:93` just
doesn't select it — **that one column is what unlocks both pooling and differences.**
- **Sex pooling is the project's own house method.** `build_state_summary()` in
`crdc-arrests/R/summarize_draws.R:152-182` pools across LEAs by summing `pred` and
`stu_enroll` within each draw, then summarizing across draws. Pooling M+F within a district
is the identical operation on a different axis. Honest effort estimate: **~3–4 hours**, most
of it UI and labelling, not statistics.
- **Caveat to word carefully:** `draw_id` is renumbered 1–500 per write batch
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in different batches,
so cross-group draw pairing is effectively independent. Measured correlation between
Black-male and Hispanic-male `pred` in Clark County was 0.019 even *within* a batch —
observation noise from `posterior_predict` dominates. Published Fig 7 has the same property,
so the app matches the paper. Captions should say "posterior predictive draws", never
"paired parameter draws".
- Enrollment covers only AM/BL/HI/WH (verified against the API). Clark County sums to 148,928
against a district enrollment near 304,000. The table must label this.
## Decisions taken
| Question | Decision |
|---|---|
| Model specs | One selected spec by default; opt-in "compare all four" expands to 4 ridge rows |
| Existing charts | Keep `ArrestsOverTime`; delete `RateByGroupBar` and `RateDensityRidgeline` |
| Pooling trigger | Whole-district: pool when total observed arrests across the 8 cells < 20 |
| Rendering | React owns the DOM; add d3 submodules for scales/paths/interpolation |
---
## Work
### 1. Draws pipeline — expose `draw_id`, return counts not rates
**`src/hooks/useDrawDistribution.js`** — the one structural change everything else rests on.
- Query becomes `SELECT RACE, SEX, draw_id, pred FROM read_parquet(...) WHERE LEAID = ?`.
- Return **raw counts indexed by draw**, not rates: `countsByGroup[key][draw_id - 1] = pred`.
Indexing by `draw_id` rather than push-order means row ordering from DuckDB is irrelevant.
- Accept `models: string[]` instead of a single `model`, returning `{status, byModel, nDraws}`.
A fixed-length array avoids conditional hooks when "compare all four" is on. The existing
module-level `shardCache` already keys on `(model, year, state)`, so four models is four
cache entries with no other change.
- Reject a group whose count array has holes (fewer entries than `nDraws`) — a partial group
must fall back, not silently render a short draw set.
- Keep the reset-before-guard ordering at `useDrawDistribution.js:79-85`; it exists to stop one
model's draws being shown under another model's label.
**`src/utils/pooling.js`** (new, + test) — pure functions, no React:
- `poolBySex(countsByGroup, enrollByGroup)` → sums counts within each draw index across
`SEX ∈ {F,M}` and sums enrollment, keyed by race alone.
- `toRates(counts, enroll)` → per-1,000 array.
- `shouldPoolBySex(rows)` → total `observed_arrests` across rows < `POOL_BY_SEX_ARREST_THRESHOLD`
(20, a named constant with the rationale in a comment).
**`src/utils/drawGroups.js`** — extend `groupKey` to handle a pooled key (race only) and update
`hasDrawsForAll` for the new count-array shape.
### 2. Frequentist interval
**`src/utils/agrestiCoull.js`** (new, + test) — direct port of
`crdc-arrests/R/paper_figures.R:219-237`, including the zero-numerator branch
(`ci_upper = -log(1 - level)`, the rule of three; `ci_lower = 0`). Note the R function returns
`c(upper, lower, sd, se, phat)` — **upper first**. Return a named object here instead. Needs a
`qnorm`/probit; `distributionApprox.js` already has one — reuse it rather than adding a second.
Tests should pin at least one case against R output (e.g. `agresti_coull(15, 499, 0.95)`).
### 3. Density profile — handle discrete posteriors honestly
**`src/utils/densityProfile.js`** (new, + test).
In sparse districts the posterior predictive is a discrete count distribution. Carson City NV
(`3200390`) has a 53-student AI/AN female cell where one arrest is 18.9 per 1,000 — the draws
take four distinct values and a Gaussian KDE renders them as a lumpy smear that reads as a
rendering bug.
- `densityProfile(counts, enroll, domain)` returns `{kind: 'kde'|'mass', points}`.
- `kind: 'mass'` when the draws take ≤ 12 distinct values: probability mass at each achievable
rate, drawn as a filled staircase so it visually rhymes with the smooth areas beside it.
- Otherwise delegate to the existing `kdeCurve` in `src/utils/kde.js` — its bandwidth clamp
(`BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`) was tuned for exactly these
zero-inflated posteriors and should not be touched.
### 4. Summary table (top of results)
**`src/components/DistrictSummaryTable.jsx`** (new). One row per student group plus a total:
| Student group | Students | Observed arrests | Rate per 1,000 |
- Sorted by observed arrests descending. Zero-arrest rows de-emphasized, not hidden.
- Each row carries the checkbox that drives chart A — the table *is* the legend and the control.
Default checked = `observed_arrests > 0`; if no group has any, check the two largest by
enrollment and say so.
- Footnote: students counted are those in the four modeled race groups (AI/AN, Black, Hispanic,
White), not total district enrollment.
- When pooling is active, rows collapse to four races and a banner states the rule in one
sentence, with a switch to force it off.
### 5. Chart A — "Arrest rate probability density"
**`src/charts/RateDensityPanel.jsx`** (new). Replaces `RateDensityRidgeline.jsx`.
- Two stacked sub-panels, Female over Male, sharing one x-axis (per 1,000). Collapses to a
single panel when pooled. This preserves the palette contract documented at
`src/utils/colors.js:9-13`: race is hue, sex is position — never a second hue.
- Within a sub-panel, selected groups overlap as filled areas (fill ~0.4 opacity, 2px stroke in
the race color), direct-labelled at each peak so there's no legend hunting.
- Below each sub-panel's baseline, a thin rail stacks one Agresti–Coull point-range per selected
group in the matching color — the R figure's `position_nudge` idea, but un-overplotted.
- Segmented control for the four unified quadrant specs. A "Compare all four specifications"
switch expands to four ridge rows (matching Fig 1's structure) and triggers four shard
fetches — cheap for NV (~100KB each), ~6.3MB each for CA, so it stays opt-in with a spinner.
- x-domain: max of the density supports and the frequentist upper bounds, with the existing cap
logic from `RateDensityRidgeline.jsx:101` and `niceTicks`.
- Caption states 500 posterior predictive draws and a 95% Agresti–Coull observed interval.
### 6. Chart B — "Model Estimated Differences"
**`src/charts/GroupDifference.jsx`** (new).
- Two group pickers; defaults are the two groups with the most observed arrests (pooled groups
when pooling is on). Δ = rate(A) − rate(B) per 1,000, computed per draw index.
- Single density, filled with an SVG `linearGradient` mapped across x. Use a **diverging ramp
centered at zero** — navy for Δ<0, paper at 0, ember for Δ>0 — rather than the paper's YlOrRd:
the quantity is signed, and diverging-at-zero is the honest encoding. On-brand via
`tokens.css`.
- Dashed vertical rule at 0 in `--cv-danger`, matching the paper's red line.
- Large readout `Pr(Δ > 0)` with a plain-language sentence beneath ("In 94.4% of posterior
draws, the Black male arrest rate exceeds the White male rate"), plus median Δ and an 80%/95%
interval as a point-range.
- Degrade explicitly when fewer than two groups have usable draws — say why, don't render empty.
### 7. Fix the district suggestions
**`scripts/build-top-districts.mjs`** (new) — pages `/estimates?state=XX&year=21-22&limit=1000`
using `meta.total` (confirmed present in the envelope) across all 51 states, aggregates observed
arrests per LEAID, and writes `public/data/top_districts.json` with the top 15 per state
(leaid, name, arrests, enrollment, rate). Roughly 140 requests as a one-off; the output is
~50KB and gets committed, following the `public/data/national_rates.json` precedent.
**`src/components/DistrictSearch.jsx`** — read the fixture instead of calling
`fetchStateDistricts` at runtime. The search screen loses a multi-second fetch and the ranking
becomes correct. Keep live name search on `/districts` unchanged. Drop the hardcoded
"Try Derby (KS), Paterson (NJ)…" hint at `DistrictSearch.jsx:164` — the real list supersedes it.
### 8. Wiring, deletions, docs
- **`src/components/ChartPanel.jsx`** — owns pooling state, selected groups, selected spec, and
the difference pair; passes them down. Keep `ArrestsOverTime`. Drop `QUADRANT_MODELS`
prefetch of all four models' *summaries* if only the selected one is needed.
- **Delete**: `src/charts/RateByGroupBar.jsx`, `src/charts/RateDensityRidgeline.jsx`.
- **Keep**: `distributionApprox.js` and `ApproxNote.jsx` — still the fallback when draws can't
be fetched (`AGENTS.md:91-101`). Fix its `intervalMass` default to 0.95 to match the API.
- **`package.json`** — add `d3-scale`, `d3-shape`, `d3-array`, `d3-interpolate` as real
`dependencies` (the existing deps are all miscategorised under `devDependencies`; leave that
alone unless it's breaking the build).
- **Docs**: `AGENTS.md` still describes 6 charts, D3 selections, and `DistrictVsNational` /
`ModelDrawsComparison` / `ExceedanceProbability` — none of which exist. Rewrite the
architecture, data-flow, and interval sections. Update `README.md`'s chart list.
---
## Verification
1. `npm run build` clean; `npm test` (`node --test 'src/**/*.test.js'`) passes, including new
tests for `agrestiCoull`, `pooling`, `densityProfile`, and the extended `drawGroups`.
2. `npm run dev`, then walk these districts:
- **Clark County NV (`3200060`, 100 arrests / 148,928 students)** — pooling stays off, all
four races render, differences chart defaults to the top two groups.
- **Carson City NV (`3200390`, 6 arrests / 4,073 students)** — pooling auto-engages (6 < 20),
banner appears, table collapses to four races. The AI/AN cell should render as a discrete
mass profile, not a smear. Verified against the draws: pooling narrows AI/AN's 90% interval
from 37.7 to 27.0 per 1,000, and Hispanic male's from 4.9 to 2.5.
- **Washoe County NV (`3200480`)** — cross-check the summary table's observed counts and
rates against `/api/v1/estimates/3200480?model=unified_m4_mod&year=21-22`.
- **A California district** — confirm the suggestion fixture ranks correctly (this is the
case the current code gets wrong), and that "compare all four" warns/spins before pulling
~25MB of shards.
3. Toggle every group off, then on; switch specs; flip pooling manually — no stale draws from a
previous model should ever appear under a new label.
4. Compare chart A against `wp_fig_group_density` output for Clark County: same curve shapes,
same point-range positions.
5. Browser console clean on hard refresh; check the Network tab shows one shard fetch per
(model, state) and no repeats when navigating between districts.
## Effort
Roughly **2–3 focused days** end to end: ~1 day for the draws/pooling/util layer with tests,
~1 day for the two charts, ~half a day for the table, the suggestion fixture, and docs. At ~10
hours a week that's about two calendar weeks.
The R Shiny alternative would be slower, not faster — it trades a zero-server static site for a
container, an R runtime, and server-side access to either the 91GB draws DuckDB or the 51-state
parquet tree, and turns every toggle into a round-trip re-render. The React app already fetches
real draws client-side and already carries the design tokens.