docs: rewrite AGENTS.md and the README chart list for the rebuilt app
Deploy to git-pages / deploy (push) Successful in 23s

AGENTS.md described 6 charts, D3 selections, and DistrictVsNational /
ModelDrawsComparison / ExceedanceProbability — none of which exist. An agent
reading it as authoritative would have been actively misled, so it now opens by
saying src/ wins any disagreement.

Rewritten: the real component tree and data flow, ChartPanel as the owner of all
cross-chart state, the pooled/unpooled key namespace trap, the four properties
of the draws pipeline that are easy to break, the multi-part shard layout and
why discovery uses the tree listing, the "posterior predictive draws" wording
rule and the reason for it, the 95% convention, and a table of which tuning
decisions carry stated rationale and should not be re-derived (the CVD-validated
palette, the KDE bandwidth clamp, the pooling threshold, the mass/KDE cutoff,
the axis cap).

Adds a testing section — there are automated tests now — and notes that the
deployment check should use a multi-part state like California, since a
single-part state cannot catch a regression in part discovery. Records that
App.jsx's "Search another district" is a full page reload that discards the
shard cache.

README: the chart list becomes the summary table plus the two ported figures,
the file tree matches src/, the d3 role is stated precisely (scales and paths,
not selections), and the endpoint table warns that /estimates?state= ranks by
LEAID on a short read.

Also commits the plan this work followed.
This commit is contained in:
2026-08-12 08:41:01 -04:00
parent c1920c11fb
commit 4396a56a4a
3 changed files with 481 additions and 123 deletions
+30 -14
View File
@@ -18,18 +18,21 @@ npm run preview # serve built files locally
## What It Does
Visitors select a U.S. state, search for a school district (with suggestions of districts that have the most arrests), and see 3 charts comparing observed data against Bayesian model estimates:
Visitors select a U.S. state, search for a school district (with suggestions of the districts reporting the most arrests), and see what was actually reported alongside what the Bayesian models estimate. The results page ports the white paper's Figs 6 and 7 (`wp_fig_group_density` / `wp_fig_group_difference`):
1. **Arrests over time** (Chart 1) — raw counts by CRDC wave with per-1k rate labels, built with inline SVG
2. **Rate by student group** (Chart 2) — box-and-whisker per race×sex group for the most recent year, split into Female/Male panels. The box is the 25th–75th percentile of that group's real 500-draw posterior (same duckdb-wasm fetch as Chart 3), the whisker is the model's reported 90% interval, and a diamond marks the observed rate; individual boxes fall back to an analytic approximation when a group's draws can't be fetched.
3. **Predicted rates by student group** (Chart 3) — density ridges built from each group's real 500-draw posterior (fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset), with a model-selector dropdown for the four Bayesian specifications and diamond markers for observed rates; falls back to an analytic approximation if the draws can't be fetched.
1. **Reported arrests and enrollment** — a summary table, one row per student group plus a district total: students, observed arrests, and rate per 1,000. Its checkboxes double as the legend and the group control for the density panel. In sparse districts (fewer than 20 arrests district-wide) Female and Male are pooled within each race, with a banner explaining the rule and a switch to override it.
2. **Arrest rate probability density** — each selected group's posterior predictive distribution as a filled area, Female over Male sharing one axis, direct-labelled at the peak. Beneath each panel, a rail of 95% Agresti–Coull point ranges for the observed rate. A segmented control switches between the four Bayesian specifications; an opt-in toggle compares all four at once. Draws that take only a handful of distinct values are drawn as discrete probability mass rather than smoothed — in a small district the posterior predictive genuinely *is* discrete.
3. **Model estimated differences** — the posterior of Δ = rate(A) − rate(B) per 1,000, computed at each draw, filled with a diverging ramp centred at zero, with a dashed rule at no-difference and a plain-language `Pr(Δ > 0)` readout.
4. **Arrests over time** — observed counts by CRDC wave against the three-year model's median and 95% interval, inline SVG.
Distributions come from the published 500-draw posteriors, fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset, and fall back per group to an analytic approximation (with a visible note) when those draws can't be fetched.
## Architecture
### Tech Stack
- **React 19** + **Vite** (static site generation, no backend required)
- Plain CSS custom properties for styling (matches Civilytics design tokens exactly)
- Inline SVG rendering with hand-rolled scales — no charting library, no D3 dependency
- Inline SVG rendering that **React owns** — no charting library. `d3-scale`, `d3-shape`, `d3-array` and `d3-interpolate` supply scales, path generators and colour interpolation only; no d3 selections and no `useEffect` DOM mutation
- Embeds **DuckDB-Wasm** (`@duckdb/duckdb-wasm`, ~39MB uncompressed / ~8.8MB gzipped, loaded on demand only after a district is selected) to query real posterior draws client-side from a public Hugging Face Parquet dataset
- Calls the public read-only API directly from the browser
@@ -40,27 +43,38 @@ crdc-demo/
├── vite.config.mjs # Vite build config
├── public/ # Static assets (wordmark, favicon, fixtures)
│ ├── civilytics-wordmark.svg # Civilytics wordmark from civilyticsR package
│ └── data/national_rates.json # Static national rates fixture for comparisons
│ └── data/
│ ├── national_rates.json # National rates fixture for comparisons
│ └── top_districts.json # Top 15 districts per state by observed arrests
├── scripts/
│ └── build-top-districts.mjs # One-off generator for top_districts.json
├── src/
│ ├── main.jsx # React entry
│ ├── App.jsx # Main router (state → search → loading → charts)
│ ├── hooks/
│ │ ├── useApi.js # API client with retry/backoff + endpoint wrappers
│ │ └── useDrawDistribution.js # Fetches real posterior draws (duckdb-wasm + HF parquet)
│ │ └── useDrawDistribution.js # Posterior draw counts by draw_id (duckdb-wasm + HF parquet)
│ ├── components/
│ │ ├── StateSelector.jsx # Landing screen — state dropdown/grid
│ │ ├── DistrictSearch.jsx # Search + "interesting" district suggestions
│ │ ├── DistrictSearch.jsx # Search + suggestions from the committed fixture
│ │ ├── LoadingAnimation.jsx # Animated histogram grid during data fetch
│ │ ├── ChartLegend.jsx # Shared legend row
│ │ ├── ApproxNote.jsx # "shape estimated from interval bounds" caption
│ │ └── ChartPanel.jsx # Orchestrates all 3 charts + data fetching
│ │ ├── DistrictSummaryTable.jsx # Observed arrests table — also the chart's legend/control
│ │ └── ChartPanel.jsx # Owns cross-chart state + data fetching
│ ├── charts/
│ │ ├── ArrestsOverTime.jsx # Chart 1 — counts by wave (SVG)
│ │ ├── RateByGroupBar.jsx # Chart 2 — box-and-whisker by group (SVG)
│ │ └── RateDensityRidgeline.jsx # Chart 3 — density ridges per student group (SVG)
│ │ ├── ArrestsOverTime.jsx # Observed vs. modelled counts by wave (SVG)
│ │ ├── RateDensityPanel.jsx # Chart A — posterior density per group + AC rail
│ │ └── GroupDifference.jsx # Chart B — posterior of Δ between two groups
│ ├── utils/
│ │ ├── duckdbClient.js # Lazy duckdb-wasm bundle loader (dynamic import)
│ │ ├── kde.js # Empirical density from real draws (+ kde.test.js)
│ │ ├── kde.js # Empirical density from real draws (+ .test.js)
│ │ ├── densityProfile.js # Discrete-mass vs. KDE profile choice (+ .test.js)
│ │ ├── agrestiCoull.js # Frequentist interval, ported from R (+ .test.js)
│ │ ├── pooling.js # Sex pooling for sparse districts (+ .test.js)
│ │ ├── districtGroups.js # Display-row derivation and defaults (+ .test.js)
│ │ ├── groupDifference.js # Per-draw Δ and its summary (+ .test.js)
│ │ ├── rateDomain.js # Shared x-axis domain and clip flag (+ .test.js)
│ │ ├── drawGroups.js # Draw-map key format + coverage check (+ .test.js)
│ │ └── distributionApprox.js # Analytic fallback when draws are unavailable
│ └── styles/tokens.css # Civilytics design tokens (colors, fonts, spacing)
@@ -74,8 +88,10 @@ crdc-demo/
|---|---|---|
| `/api/v1/models` | List available Bayesian model specs | Once (cached) |
| `/api/v1/districts?q=&state=` | District name/geo lookup → LEAID | On keystroke |
| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | ~40 calls per district |
| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | 3 waves + 1 per selected spec |
| `/api/v1/estimates?state=XX&year=Y` | Not called at runtime — rows come back `ORDER BY LEAID` at 8 per district, so any short read ranks the lowest-LEAID districts. Paged with `meta.total` by `scripts/build-top-districts.mjs` | Build-time only |
| `/api/v1/draws?...` | Locate raw-posterior Parquet shard | Not called from app — the app fetches shards directly from Hugging Face via duckdb-wasm; see `src/hooks/useDrawDistribution.js` |
| `/data/top_districts.json` | Suggested districts per state (committed fixture) | Once per session |
| `/data/national_rates.json` | Static national rates fixture (committed) | Once per session |
## Deployment