From 4396a56a4a4512a83facceeefad2a3e1fc3f6ac9 Mon Sep 17 00:00:00 2001 From: Jared Knowles Date: Wed, 12 Aug 2026 08:41:01 -0400 Subject: [PATCH] docs: rewrite AGENTS.md and the README chart list for the rebuilt app MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit AGENTS.md described 6 charts, D3 selections, and DistrictVsNational / ModelDrawsComparison / ExceedanceProbability — none of which exist. An agent reading it as authoritative would have been actively misled, so it now opens by saying src/ wins any disagreement. Rewritten: the real component tree and data flow, ChartPanel as the owner of all cross-chart state, the pooled/unpooled key namespace trap, the four properties of the draws pipeline that are easy to break, the multi-part shard layout and why discovery uses the tree listing, the "posterior predictive draws" wording rule and the reason for it, the 95% convention, and a table of which tuning decisions carry stated rationale and should not be re-derived (the CVD-validated palette, the KDE bandwidth clamp, the pooling threshold, the mass/KDE cutoff, the axis cap). Adds a testing section — there are automated tests now — and notes that the deployment check should use a multi-part state like California, since a single-part state cannot catch a regression in part discovery. Records that App.jsx's "Search another district" is a full page reload that discards the shard cache. README: the chart list becomes the summary table plus the two ported figures, the file tree matches src/, the d3 role is stated precisely (scales and paths, not selections), and the endpoint table warns that /estimates?state= ranks by LEAID on a short read. Also commits the plan this work followed. --- AGENTS.md | 341 ++++++++++++------ README.md | 44 ++- .../plans/2026-08-12-visual-rebuild.md | 219 +++++++++++ 3 files changed, 481 insertions(+), 123 deletions(-) create mode 100644 docs/superpowers/plans/2026-08-12-visual-rebuild.md diff --git a/AGENTS.md b/AGENTS.md index 8575683..9fbb3b5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,154 +1,277 @@ # Agent Guide — CRDC Demo App -This document captures context, decisions, and guidance for agents working on this codebase. It is the primary source of truth for how to make changes safely. +Context, decisions, and guidance for agents working on this codebase. + +**`src/` is the source of truth.** If this file and the code disagree, the code +wins and this file is the bug — fix it in the same change. ## Project Overview -A React + Vite static web app demonstrating the [CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors select a state, search for a school district, and see 6 charts comparing observed arrest data against Bayesian model estimates. Deployed via Gitea Actions to `pages.civilytics.org/crdc-demo/`. +A React + Vite static web app demonstrating the +[CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors +pick a state, search for a school district, and see what was actually reported +alongside what the Bayesian models estimate. Deployed via Gitea Actions to +`pages.civilytics.org/crdc-demo/`. + +The results page is a port of the white paper's Figs 6 and 7 +(`wp_fig_group_density` / `wp_fig_group_difference` in +`crdc-arrests/R/paper_figures.R:472-568`). Keeping it recognisably the same +figure is the point — someone who has read the paper should see the paper. ## Architecture Summary -- **Frontend**: React 19 + Vite (static site generation) -- **Styling**: Plain CSS custom properties matching Civilytics design tokens (`src/styles/tokens.css`) -- **Charts**: Mixed approach: - - Charts 1–3 use inline SVG with manual scales (no D3 dependency for these) - - Chart 4 (ModelDrawsComparison) uses D3.js v7 for data-driven rendering of quadrant comparisons - - Chart 5 (RateDensityRidgeline) uses D3.js v7 for density ridge visualizations -- **API**: Calls public read-only API directly from browser; no backend required -- **Deployment**: Static site deployed via Gitea Actions (`.gitea/workflows/pages.yml`) +- **Frontend**: React 19 + Vite (static site) +- **Styling**: CSS custom properties mirroring the Civilytics design tokens + (`src/styles/tokens.css`) +- **Charts**: inline SVG that **React owns**. `d3-scale`, `d3-shape`, + `d3-array` and `d3-interpolate` supply scales, path generators and colour + interpolation only. **No d3 selections, no `useEffect` DOM mutation** — + if you find yourself reaching for `d3.select`, the answer is a render. +- **API**: public read-only API called directly from the browser; no backend +- **Deployment**: Gitea Actions (`.gitea/workflows/pages.yml`) ## Key Components and Data Flow ``` App.jsx (router) - → StateSelector (landing screen: state dropdown/grid) - → DistrictSearch (search + "interesting" suggestions from /estimates?state=&year=) - → LoadingAnimation (fetches all data in parallel, shows animated histogram grid) - → ChartPanel (receives district object, fetches structured estimates for 6 charts) - ├── ArrestsOverTime — SVG line chart by wave (3 years) - ├── RateByGroupBar — SVG bar chart: observed vs modeled per group - ├── DistrictVsNational — SVG comparison to national average - ├── ModelDrawsComparison — D3 quadrant charts (4 model types × 1 year) - ├── RateDensityRidgeline — D3 density ridges per race×sex group - └── ExceedanceProbability — P(district > national) per student group + → StateSelector landing screen: state dropdown/grid + → DistrictSearch live /districts search + suggestions from the + │ committed public/data/top_districts.json fixture + → LoadingAnimation warms the API, animated histogram grid + → ChartPanel OWNS all cross-chart state (see below) + ├── DistrictSummaryTable observed arrests + enrollment; the checkboxes + │ here are the density panel's group control + ├── RateDensityPanel Chart A — posterior density per group, + │ Female over Male, Agresti–Coull rail beneath + ├── GroupDifference Chart B — posterior of Δ between two groups + └── ArrestsOverTime observed vs. modelled totals across 3 waves ``` -### Data Fetching Strategy (`ChartPanel.jsx`) +### `ChartPanel` owns the state -`ChartPanel` fetches all chart data on mount (after `LoadingAnimation` pre-fetched via batch calls): +Selected model specification, sex pooling, which groups are checked, and the +difference pair all live in `ChartPanel` and are passed down. Charts hold none +of it. That is what keeps the table's checkboxes and the density panel from +drifting apart. -1. **Wave data** for Charts 1–3: Fetches `unified_m3_mod` model estimates for years `['21-22', '17-18', '15-16']`. - - Uses three-year models because they return observed arrest counts across all waves (one-year models only have data for the most recent wave). +One trap worth knowing: pooled group keys (`'BL'`) and unpooled ones (`'BL_F'`) +are different namespaces. Selections are therefore stored **with the pooling +mode they were made in** and fall back to defaults when the mode changes. -2. **Quad data** for Charts 4–6: Fetches estimates from all four quadrant models (`unified_m1_mod`, `unified_m2_mod`, `unified_m3_mod`, `unified_m4_mod`) for year `21-22` only (one year, as the most recent wave). +### Data fetching -3. **National rates**: Loaded once from a static JSON fixture or cached by `LoadingAnimation`. +1. **Wave data** (`ArrestsOverTime`): `unified_m3_mod` for `['15-16', '17-18', + '21-22']`. Three-year models are used because they return observed counts + across all waves; one-year models only cover the most recent one. +2. **Current-wave summary**: fetched for the **selected specification only**. + Enrollment and observed arrests are district facts, not model outputs, and + the modelled shapes now come from real draws — prefetching all four specs' + summaries would be four requests for data three of which are never read. +3. **Posterior draws**: `useDrawDistribution` (see below). -### Error Bar Convention: 90% Intervals +## The draws pipeline — read this before touching `useDrawDistribution.js` -The API returns 95% HPD intervals (`count_lower`, `count_upper`). However, all chart labels and calculations in this app use **90% intervals**. +`useDrawDistribution({leaid, state, models, year})` fetches the published +posterior draws from the Hugging Face parquet dataset +(`civilytics/crdc-school-arrest-rates`) and queries them client-side with +`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js`). There is no server-side +draws endpoint. -The 90% convention is baked into the analytic fallback in `src/utils/distributionApprox.js` (the path used when real posterior draws can't be fetched — see the duckdb-wasm section below). That fallback generates **no draws**: `fitSkewedInterval` fits a two-piece normal directly from `{median, lower, upper}`, converting each half-interval to its own sigma with `z = probit((1 + intervalMass) / 2)` and `intervalMass` defaulting to `0.90` (so z ≈ 1.645). For a symmetric interval that is equivalent to the old fixed "full width / 3.29" divisor, but it is computed from `intervalMass` rather than hardcoded, and each side gets its own sigma so the fitted shape stays skewed. +**A state's draws are split across multiple parquet parts.** `data_0.parquet`, +`data_1.parquet`, … and the count varies by state: Nevada is one file, +California is **eight** (6.2MB total, of which `data_0` is 37KB and holds 11 of +California's 1,715 districts). The app fetched only `data_0` until 2026-08-12, +which made every CA district except those 11 look absent from the published +data and silently fall back to the approximation — invisible in testing because +Nevada, the district everyone tests with, has exactly one part. -If you change this convention, update: -- `distributionApprox.js` — the `intervalMass = 0.90` default in `fitSkewedInterval`, and its JSDoc claim that the API's bounds are a 90% interval -- `RateByGroupBar.jsx` — the caption and component JSDoc, both of which say "the model's reported 90% interval" -- `ArrestsOverTime.jsx` — the legend label "Modeled (median + 90% interval)" -- Any documentation referencing confidence/credible intervals +Parts are discovered from the Hugging Face **tree listing API** +(`/api/datasets/{id}/tree/main/parquet/...`), not by probing `data_N` until a +404. A 404 is logged as a console error by the browser's network layer however +cleanly the fetch handles it, and a red error on every page load is +indistinguishable from a real one. Probing (via HEAD) remains the fallback if +the listing API is unavailable. All parts are fetched in parallel, registered +individually, and queried as `read_parquet([...])`. + +Four more properties it is easy to break: + +- **It returns counts, not rates**, indexed by `draw_id - 1`. Counts are what + make sex pooling and between-group differences possible: both have to sum or + subtract numerators and denominators separately. Callers divide. +- **Indexing by `draw_id`, not push order**, makes DuckDB's row ordering + irrelevant and turns a missing draw into a hole rather than a short array. +- **Incomplete groups are dropped** (`isCompleteDrawSet`). A group present for + 300 of 500 draws would otherwise get an interval computed off a biased + subsample that looks identical on screen to a complete one. +- **The reset-before-guard ordering in the effect is load-bearing.** State is + cleared on *every* input change, including ones with nothing to fetch, so a + failed fetch can never leave the previous model's draws on screen under the + new model's label. + +`status` is an ANY-model, ANY-group signal. To claim "these are all real draws" +for a specific set of rendered groups, use `hasDrawsForAll`. + +### Caption wording: "posterior predictive draws" + +`draw_id` is renumbered 1–500 per write batch upstream +(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in +different batches, so draw *k* of one group is **not** the same parameter draw +as draw *k* of another. Measured correlation between Black-male and +Hispanic-male `pred` in Clark County was 0.019 even within a batch — +`posterior_predict` observation noise dominates. + +The published Fig 7 has the same property, so the app matches the paper. What +neither can claim is a paired-parameter contrast. **Captions must say "posterior +predictive draws" and must never say "paired parameter draws".** + +### Fallback path — do not delete + +If the draws fetch fails (network, unsupported browser, HF outage), +`RateDensityPanel` falls back **per group** to the analytic approximation in +`src/utils/distributionApprox.js` and shows ``. Neither +`distributionApprox.js` nor `ApproxNote.jsx` is dead code. + +A pooled group has no fallback shape: adding two groups' interval *bounds* +together is not a pooled interval, so `buildDisplayGroups` sets `modeled: null` +when pooling and the chart omits that group rather than inventing a curve. + +The duckdb-wasm engine is ~39MB uncompressed / ~8.86MB gzipped (measured +against the shipped package). It loads via dynamic `import()` only once a +district is selected — never on initial page load — and is browser-cached +thereafter, but it is a real one-time cost. + +### Error bar convention: 95% + +The API returns **95%** intervals. `validate_interval()` in +`crdc-arrests/api/R/validate.R` defaults to `95L` and this app never passes +`interval=`. `fitSkewedInterval`'s `intervalMass` therefore defaults to `0.95`, +and `ArrestsOverTime`'s legend says "95% interval". (Both said 90% before +2026-08-12; that was a bug, not a convention.) + +The observed-data point ranges are a different thing again: a 95% +Agresti–Coull interval computed from observed counts +(`src/utils/agrestiCoull.js`), a direct port of `agresti_coull()` in +`crdc-arrests/R/paper_figures.R:219-237`. Two faithfulness quirks are pinned by +tests and must not be "fixed": the bounds are on the **count** scale, and +`lower` can be **negative** for very small numerators (charts clamp at draw +time, the port does not). + +## Tuning decisions with stated rationale — don't re-derive + +- **`src/utils/colors.js:9-13`** — the race palette passed the dataviz skill's + CVD validator. Re-run `validate_palette.js` before changing any hex value. + Race is hue; **sex is position, never a second hue**; observed-vs-modelled is + mark type, never a second hue. +- **`src/utils/kde.js`** — `BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR` + were tuned for zero-inflated sparse-district posteriors. Both ends matter. +- **`src/utils/pooling.js`** — `POOL_BY_SEX_ARREST_THRESHOLD = 20`, applied to + the district total, not per cell. +- **`src/utils/densityProfile.js`** — `MASS_MAX_DISTINCT = 12`. Below it the + posterior predictive is drawn as discrete mass, because it *is* discrete; a + Gaussian KDE over four achievable values renders as a lumpy smear that reads + as a rendering bug. +- **`src/utils/rateDomain.js`** — `MAX_RATE_DOMAIN = 30` caps the axis so a + four-student cell can't squash every other curve. It reports `clipped` so the + chart says so instead of silently cropping. ## Common Pitfalls & Gotchas -### 1. Null Safety in Chart Components +### 1. Null safety in chart components + +API responses can have empty arrays or missing fields for districts with no +arrests. Use optional chaining and explicit defaults, and never divide by a +denominator you haven't checked: -Several API responses may return empty arrays or missing fields for districts with no arrests: ```javascript -// Always use optional chaining and defaults: -const yearRow = (quadModels[q.key] || []).find(r => r.year === year) -const predMedian = yearRow?.count_median || 0 -const enroll = row.stu_enroll || 1 // Prevent division by zero +const enroll = row.stu_enroll || 0 +const rate = enroll > 0 ? (row.observed_arrests || 0) / enroll * 1000 : 0 ``` -### 2. D3 useEffect Dependency Arrays +A rate with no denominator is **not zero and not Infinity — it's undefined**. +`toRates` returns `[]`, `buildDisplayGroups` reports `0` and lets the +enrollment column explain why. -When using `useEffect` for D3 rendering, always include all data dependencies to prevent stale renders: -```javascript -// Correct — includes all props used inside the effect -}, [quadData, selectedModel, rateByGroup]) +### 2. SVG dimensions and responsiveness + +Charts set `width="100%"` with a fixed `viewBox`, wrapped in +`overflowX: 'auto'`. Chart cards are `max-width: 70rem` in `ChartPanel.jsx` +(vs. the ~60rem default text width). + +### 3. API endpoint availability + +- `/api/v1/estimates/{leaid}` — ✅ summary rows (median, bounds, enrollment, + observed arrests) +- `/api/v1/estimates?state=XX&...` — ✅ but returns rows `ORDER BY LEAID, RACE, + SEX` at 8 per district, capped at `limit=1000`. **Any short read ranks the + lowest-LEAID districts, not the busiest.** Page it with `meta.total`. +- `/api/v1/draws?...` — returns a shard URL + SQL, not draw data. The app goes + to the parquet directly. + +### 4. CORS + +The API sends no CORS headers. The app auto-detects a proxy via `VITE_PROXY_URL`; +unset, it fetches directly (works same-origin or behind the Docker/nginx proxy). + +## The suggestion fixture + +`public/data/top_districts.json` holds the top 15 districts per state by +observed arrests, and is generated by a one-off, read-only script: + +```bash +node scripts/build-top-districts.mjs # all 51, ~150 requests +node scripts/build-top-districts.mjs --states NV,CA # spot-check ``` -### 3. SVG Dimensions and Responsiveness +It is committed (the `national_rates.json` precedent). Re-run it only when a +new CRDC wave lands. `DistrictSearch` degrades to search-only if it's missing. -Charts use fixed dimensions with responsive containers (`overflowX: 'auto'` for wide content): -- Chart cards have `max-width: 70rem` in `ChartPanel.jsx` (vs the default text width of ~60rem) -- SVG elements should set both `width="100%"` and a fixed `viewBox` for proper scaling +## Testing -### 4. API Endpoint Availability +`npm test` runs `node --test 'src/**/*.test.js'`. The pure utilities are all +covered and **should be written test-first**: -Not all endpoints are available to browser-based clients: -- `/api/v1/estimates/{leaid}` — ✅ Returns estimates summary (median, lower, upper bounds) -- `/api/v1/draws?...` — Returns a Parquet shard URL + DuckDB SQL, not draw data itself. The app **does** use the real draws in the shard it points to — see below. +| Module | What its tests pin | +|---|---| +| `agrestiCoull.js` | five cases against real R output, incl. the negative lower bound | +| `pooling.js` | numerator and denominator always drawn from the same groups | +| `densityProfile.js` | the mass/KDE switch, and that KDE delegates to `kde.js` | +| `districtGroups.js` | display-row derivation, defaults, pooled vs unpooled keys | +| `groupDifference.js` | refusal to pair mismatched draw sets | +| `rateDomain.js` | the axis cap, and that it reports clipping | +| `drawGroups.js` | key shapes, and hole detection in draw arrays | +| `kde.js` | bandwidth clamp behaviour | -### 5. Real posterior draws via duckdb-wasm - -Charts 2 (`RateByGroupBar`) and 3 (`RateDensityRidgeline`) fetch the actual -500-draw-per-group posterior from the public Hugging Face parquet dataset -(`civilytics/crdc-school-arrest-rates`), queried client-side with -`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js` + -`src/hooks/useDrawDistribution.js`). No server-side draws endpoint is -involved. If that fetch fails (network, unsupported browser, HF outage), -both charts fall back to the `distributionApprox.js` analytic approximation -and show the "estimated shape" note — **do not delete `distributionApprox.js` -or `ApproxNote.jsx`**, they're the fallback path, not dead code. - -See `docs/superpowers/specs/2026-08-11-empirical-draws-wasm-design.md` for -the full design. - -The duckdb-wasm engine itself is ~39MB uncompressed / ~8.86MB gzipped (confirmed against the shipped `@duckdb/duckdb-wasm` package, not the design spec's original ~3-5MB estimate, which was wrong). It's loaded via dynamic `import()` only once a district is selected — never on initial page load — and cached by the browser thereafter, but it's a real one-time cost worth knowing about before touching this code path. - -### 6. CORS Configuration - -The CRDC API does not send CORS headers. When deployed to git-pages (static hosting), requests are blocked by same-origin policy unless a proxy is configured: -- The app auto-detects proxy availability via `VITE_PROXY_URL` environment variable -- If unset, the app attempts direct fetch — works when served from Docker/nginx or same-origin +Components are verified manually — there is no DOM test harness. Useful +districts: **Clark County NV `3200060`** (100 arrests / 148,928 students, +pooling off), **Carson City NV `3200390`** (6 arrests, pooling auto-engages, +discrete mass profile), **Washoe County NV `3200480`** (cross-check the table +against the API), and any California district (large shards, fixture ranking). ## Deployment Checklist -Before pushing to production: +1. `npm run build` succeeds +2. `npm test` passes +3. No console errors after a hard refresh +4. Both charts render for a sample district; Network tab shows **one fetch per + (model, state) shard part** and no repeats when switching specs +5. Test with a **multi-part state** (California), not just Nevada — a + single-part state cannot catch a regression in part discovery -1. **Build succeeds**: `npm run build` (check for new errors)2. **No console errors in browser** after hard refresh3. **All 6 charts render** with sample districts (test "Denver", "Mobile County") -4. **Loading animation** appears briefly, then transitions to ChartPanel - -### Commit Message Convention - -Use descriptive commit messages that explain the *why*, not just the what: -```bash -# Good -git commit -m "Fix: use three-year model for wave data (returns observed counts across all years)" -git commit -m "Add D3 density ridges to Chart 5 with diamond markers for observed rates" - -# Avoid vague messages -git commit -m "Fix charts" # Too generic -git commit -m "Update code" # No context -``` - -## Testing Strategy - -There are no automated tests in this project. Manual verification is required:1. **Visual check**: Load a district and verify all 6 charts render correctly2. **Error console**: Check browser DevTools for JavaScript errors3. **Data accuracy**: Compare observed values against API response (check Network tab) -4. **Responsiveness**: Resize window to ensure layout adapts - -## Style Guide References - -- R code style: Follows tidyverse principles (`r-style-guide` skill in agent knowledge base)- Chart aesthetic decisions should match patterns from `social_media_posts.md` and `white_paper.qmd` -- Colors, typography, spacing are defined as CSS custom properties in `src/styles/tokens.css` - -## Related Repositories - -- **crdc-arrests** — The API server (Plumber/R) at `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/` -- **civilyticsR** — R package with wordmark and visualization functions at `/home/jared/Nextcloud/Civilytics/Code/Civilytics/civilyticsR/` +> Caveat on the cache: `App.jsx:112`'s "Search another district" button does +> `window.location.href = '/crdc-demo/'`, a full page reload, which discards the +> module-level `shardCache`. Within one district view — switching specs, +> toggling compare-all — the cache works as intended. ## Git Conventions - Remote: `https://gitea.civilytics.org/Civilytics/crdc-demo.git` -- Default branch: `main` (not `master`) -- Gitea Actions workflow auto-deploys on push to `main` via `.gitea/workflows/pages.yml` -- Always pull before making changes: `git pull origin main` \ No newline at end of file +- Default branch: `main` +- Gitea Actions auto-deploys on push to `main` +- Commit messages explain the *why*: `fix: rank suggestions over the full state + (limit=500 was selecting the lowest 62 LEAIDs)` + +## Related Repositories + +- **crdc-arrests** — API server (Plumber/R) and the white paper, at + `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/` +- **civilyticsR** — wordmark and visualization functions diff --git a/README.md b/README.md index cd5931f..181b34b 100644 --- a/README.md +++ b/README.md @@ -18,18 +18,21 @@ npm run preview # serve built files locally ## What It Does -Visitors select a U.S. state, search for a school district (with suggestions of districts that have the most arrests), and see 3 charts comparing observed data against Bayesian model estimates: +Visitors select a U.S. state, search for a school district (with suggestions of the districts reporting the most arrests), and see what was actually reported alongside what the Bayesian models estimate. The results page ports the white paper's Figs 6 and 7 (`wp_fig_group_density` / `wp_fig_group_difference`): -1. **Arrests over time** (Chart 1) — raw counts by CRDC wave with per-1k rate labels, built with inline SVG -2. **Rate by student group** (Chart 2) — box-and-whisker per race×sex group for the most recent year, split into Female/Male panels. The box is the 25th–75th percentile of that group's real 500-draw posterior (same duckdb-wasm fetch as Chart 3), the whisker is the model's reported 90% interval, and a diamond marks the observed rate; individual boxes fall back to an analytic approximation when a group's draws can't be fetched. -3. **Predicted rates by student group** (Chart 3) — density ridges built from each group's real 500-draw posterior (fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset), with a model-selector dropdown for the four Bayesian specifications and diamond markers for observed rates; falls back to an analytic approximation if the draws can't be fetched. +1. **Reported arrests and enrollment** — a summary table, one row per student group plus a district total: students, observed arrests, and rate per 1,000. Its checkboxes double as the legend and the group control for the density panel. In sparse districts (fewer than 20 arrests district-wide) Female and Male are pooled within each race, with a banner explaining the rule and a switch to override it. +2. **Arrest rate probability density** — each selected group's posterior predictive distribution as a filled area, Female over Male sharing one axis, direct-labelled at the peak. Beneath each panel, a rail of 95% Agresti–Coull point ranges for the observed rate. A segmented control switches between the four Bayesian specifications; an opt-in toggle compares all four at once. Draws that take only a handful of distinct values are drawn as discrete probability mass rather than smoothed — in a small district the posterior predictive genuinely *is* discrete. +3. **Model estimated differences** — the posterior of Δ = rate(A) − rate(B) per 1,000, computed at each draw, filled with a diverging ramp centred at zero, with a dashed rule at no-difference and a plain-language `Pr(Δ > 0)` readout. +4. **Arrests over time** — observed counts by CRDC wave against the three-year model's median and 95% interval, inline SVG. + +Distributions come from the published 500-draw posteriors, fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset, and fall back per group to an analytic approximation (with a visible note) when those draws can't be fetched. ## Architecture ### Tech Stack - **React 19** + **Vite** (static site generation, no backend required) - Plain CSS custom properties for styling (matches Civilytics design tokens exactly) -- Inline SVG rendering with hand-rolled scales — no charting library, no D3 dependency +- Inline SVG rendering that **React owns** — no charting library. `d3-scale`, `d3-shape`, `d3-array` and `d3-interpolate` supply scales, path generators and colour interpolation only; no d3 selections and no `useEffect` DOM mutation - Embeds **DuckDB-Wasm** (`@duckdb/duckdb-wasm`, ~39MB uncompressed / ~8.8MB gzipped, loaded on demand only after a district is selected) to query real posterior draws client-side from a public Hugging Face Parquet dataset - Calls the public read-only API directly from the browser @@ -40,27 +43,38 @@ crdc-demo/ ├── vite.config.mjs # Vite build config ├── public/ # Static assets (wordmark, favicon, fixtures) │ ├── civilytics-wordmark.svg # Civilytics wordmark from civilyticsR package -│ └── data/national_rates.json # Static national rates fixture for comparisons +│ └── data/ +│ ├── national_rates.json # National rates fixture for comparisons +│ └── top_districts.json # Top 15 districts per state by observed arrests +├── scripts/ +│ └── build-top-districts.mjs # One-off generator for top_districts.json ├── src/ │ ├── main.jsx # React entry │ ├── App.jsx # Main router (state → search → loading → charts) │ ├── hooks/ │ │ ├── useApi.js # API client with retry/backoff + endpoint wrappers -│ │ └── useDrawDistribution.js # Fetches real posterior draws (duckdb-wasm + HF parquet) +│ │ └── useDrawDistribution.js # Posterior draw counts by draw_id (duckdb-wasm + HF parquet) │ ├── components/ │ │ ├── StateSelector.jsx # Landing screen — state dropdown/grid -│ │ ├── DistrictSearch.jsx # Search + "interesting" district suggestions +│ │ ├── DistrictSearch.jsx # Search + suggestions from the committed fixture │ │ ├── LoadingAnimation.jsx # Animated histogram grid during data fetch │ │ ├── ChartLegend.jsx # Shared legend row │ │ ├── ApproxNote.jsx # "shape estimated from interval bounds" caption -│ │ └── ChartPanel.jsx # Orchestrates all 3 charts + data fetching +│ │ ├── DistrictSummaryTable.jsx # Observed arrests table — also the chart's legend/control +│ │ └── ChartPanel.jsx # Owns cross-chart state + data fetching │ ├── charts/ -│ │ ├── ArrestsOverTime.jsx # Chart 1 — counts by wave (SVG) -│ │ ├── RateByGroupBar.jsx # Chart 2 — box-and-whisker by group (SVG) -│ │ └── RateDensityRidgeline.jsx # Chart 3 — density ridges per student group (SVG) +│ │ ├── ArrestsOverTime.jsx # Observed vs. modelled counts by wave (SVG) +│ │ ├── RateDensityPanel.jsx # Chart A — posterior density per group + AC rail +│ │ └── GroupDifference.jsx # Chart B — posterior of Δ between two groups │ ├── utils/ │ │ ├── duckdbClient.js # Lazy duckdb-wasm bundle loader (dynamic import) -│ │ ├── kde.js # Empirical density from real draws (+ kde.test.js) +│ │ ├── kde.js # Empirical density from real draws (+ .test.js) +│ │ ├── densityProfile.js # Discrete-mass vs. KDE profile choice (+ .test.js) +│ │ ├── agrestiCoull.js # Frequentist interval, ported from R (+ .test.js) +│ │ ├── pooling.js # Sex pooling for sparse districts (+ .test.js) +│ │ ├── districtGroups.js # Display-row derivation and defaults (+ .test.js) +│ │ ├── groupDifference.js # Per-draw Δ and its summary (+ .test.js) +│ │ ├── rateDomain.js # Shared x-axis domain and clip flag (+ .test.js) │ │ ├── drawGroups.js # Draw-map key format + coverage check (+ .test.js) │ │ └── distributionApprox.js # Analytic fallback when draws are unavailable │ └── styles/tokens.css # Civilytics design tokens (colors, fonts, spacing) @@ -74,8 +88,10 @@ crdc-demo/ |---|---|---| | `/api/v1/models` | List available Bayesian model specs | Once (cached) | | `/api/v1/districts?q=&state=` | District name/geo lookup → LEAID | On keystroke | -| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | ~40 calls per district | +| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | 3 waves + 1 per selected spec | +| `/api/v1/estimates?state=XX&year=Y` | Not called at runtime — rows come back `ORDER BY LEAID` at 8 per district, so any short read ranks the lowest-LEAID districts. Paged with `meta.total` by `scripts/build-top-districts.mjs` | Build-time only | | `/api/v1/draws?...` | Locate raw-posterior Parquet shard | Not called from app — the app fetches shards directly from Hugging Face via duckdb-wasm; see `src/hooks/useDrawDistribution.js` | +| `/data/top_districts.json` | Suggested districts per state (committed fixture) | Once per session | | `/data/national_rates.json` | Static national rates fixture (committed) | Once per session | ## Deployment diff --git a/docs/superpowers/plans/2026-08-12-visual-rebuild.md b/docs/superpowers/plans/2026-08-12-visual-rebuild.md new file mode 100644 index 0000000..2901e0f --- /dev/null +++ b/docs/superpowers/plans/2026-08-12-visual-rebuild.md @@ -0,0 +1,219 @@ +# CRDC demo — rebuild the visuals around the white-paper figures + +## Context + +The demo app at `pages.civilytics.org/crdc-demo/` works, but the charts don't show off the +modelling. The navigation (state → district) is good and stays. The three existing charts get +cut to one, replaced by a summary table and two charts ported from the white paper: +`wp_fig_group_density` (Fig 6) and `wp_fig_group_difference` (Fig 7) in +`crdc-arrests/R/paper_figures.R:472-568`. + +Two real defects surfaced while scoping: + +1. **The "suggested districts" list is ranked over a truncated set.** `DistrictSearch.jsx:22` + calls `/estimates?state=XX&year=21-22&limit=500`, but that endpoint returns rows + `ORDER BY LEAID` (`crdc-arrests/api/R/handlers_estimates.R:46`) at 8 rows per district. So + the app ranks the ~62 lowest-LEAID districts in the state. California has 11,488 rows. +2. **The analytic fallback assumes 90% bounds but the API returns 95%.** + `distributionApprox.js` defaults `intervalMass = 0.90`; the API default is `interval=95` + (`crdc-arrests/api/R/handlers_estimates.R`). The new charts compute intervals from real + draws, so this only affects the fallback path — fix the default while we're in there. + +### What the data supports (verified, not assumed) + +- The HF parquet carries `LEAID, RACE, SEX, pred, draw_id, subgroup_id, batch_num`. `draw_id` + is dense 1–500 for every group. The current query at `useDrawDistribution.js:93` just + doesn't select it — **that one column is what unlocks both pooling and differences.** +- **Sex pooling is the project's own house method.** `build_state_summary()` in + `crdc-arrests/R/summarize_draws.R:152-182` pools across LEAs by summing `pred` and + `stu_enroll` within each draw, then summarizing across draws. Pooling M+F within a district + is the identical operation on a different axis. Honest effort estimate: **~3–4 hours**, most + of it UI and labelling, not statistics. +- **Caveat to word carefully:** `draw_id` is renumbered 1–500 per write batch + (`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in different batches, + so cross-group draw pairing is effectively independent. Measured correlation between + Black-male and Hispanic-male `pred` in Clark County was 0.019 even *within* a batch — + observation noise from `posterior_predict` dominates. Published Fig 7 has the same property, + so the app matches the paper. Captions should say "posterior predictive draws", never + "paired parameter draws". +- Enrollment covers only AM/BL/HI/WH (verified against the API). Clark County sums to 148,928 + against a district enrollment near 304,000. The table must label this. + +## Decisions taken + +| Question | Decision | +|---|---| +| Model specs | One selected spec by default; opt-in "compare all four" expands to 4 ridge rows | +| Existing charts | Keep `ArrestsOverTime`; delete `RateByGroupBar` and `RateDensityRidgeline` | +| Pooling trigger | Whole-district: pool when total observed arrests across the 8 cells < 20 | +| Rendering | React owns the DOM; add d3 submodules for scales/paths/interpolation | + +--- + +## Work + +### 1. Draws pipeline — expose `draw_id`, return counts not rates + +**`src/hooks/useDrawDistribution.js`** — the one structural change everything else rests on. + +- Query becomes `SELECT RACE, SEX, draw_id, pred FROM read_parquet(...) WHERE LEAID = ?`. +- Return **raw counts indexed by draw**, not rates: `countsByGroup[key][draw_id - 1] = pred`. + Indexing by `draw_id` rather than push-order means row ordering from DuckDB is irrelevant. +- Accept `models: string[]` instead of a single `model`, returning `{status, byModel, nDraws}`. + A fixed-length array avoids conditional hooks when "compare all four" is on. The existing + module-level `shardCache` already keys on `(model, year, state)`, so four models is four + cache entries with no other change. +- Reject a group whose count array has holes (fewer entries than `nDraws`) — a partial group + must fall back, not silently render a short draw set. +- Keep the reset-before-guard ordering at `useDrawDistribution.js:79-85`; it exists to stop one + model's draws being shown under another model's label. + +**`src/utils/pooling.js`** (new, + test) — pure functions, no React: + +- `poolBySex(countsByGroup, enrollByGroup)` → sums counts within each draw index across + `SEX ∈ {F,M}` and sums enrollment, keyed by race alone. +- `toRates(counts, enroll)` → per-1,000 array. +- `shouldPoolBySex(rows)` → total `observed_arrests` across rows < `POOL_BY_SEX_ARREST_THRESHOLD` + (20, a named constant with the rationale in a comment). + +**`src/utils/drawGroups.js`** — extend `groupKey` to handle a pooled key (race only) and update +`hasDrawsForAll` for the new count-array shape. + +### 2. Frequentist interval + +**`src/utils/agrestiCoull.js`** (new, + test) — direct port of +`crdc-arrests/R/paper_figures.R:219-237`, including the zero-numerator branch +(`ci_upper = -log(1 - level)`, the rule of three; `ci_lower = 0`). Note the R function returns +`c(upper, lower, sd, se, phat)` — **upper first**. Return a named object here instead. Needs a +`qnorm`/probit; `distributionApprox.js` already has one — reuse it rather than adding a second. + +Tests should pin at least one case against R output (e.g. `agresti_coull(15, 499, 0.95)`). + +### 3. Density profile — handle discrete posteriors honestly + +**`src/utils/densityProfile.js`** (new, + test). + +In sparse districts the posterior predictive is a discrete count distribution. Carson City NV +(`3200390`) has a 53-student AI/AN female cell where one arrest is 18.9 per 1,000 — the draws +take four distinct values and a Gaussian KDE renders them as a lumpy smear that reads as a +rendering bug. + +- `densityProfile(counts, enroll, domain)` returns `{kind: 'kde'|'mass', points}`. +- `kind: 'mass'` when the draws take ≤ 12 distinct values: probability mass at each achievable + rate, drawn as a filled staircase so it visually rhymes with the smooth areas beside it. +- Otherwise delegate to the existing `kdeCurve` in `src/utils/kde.js` — its bandwidth clamp + (`BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`) was tuned for exactly these + zero-inflated posteriors and should not be touched. + +### 4. Summary table (top of results) + +**`src/components/DistrictSummaryTable.jsx`** (new). One row per student group plus a total: + +| Student group | Students | Observed arrests | Rate per 1,000 | + +- Sorted by observed arrests descending. Zero-arrest rows de-emphasized, not hidden. +- Each row carries the checkbox that drives chart A — the table *is* the legend and the control. + Default checked = `observed_arrests > 0`; if no group has any, check the two largest by + enrollment and say so. +- Footnote: students counted are those in the four modeled race groups (AI/AN, Black, Hispanic, + White), not total district enrollment. +- When pooling is active, rows collapse to four races and a banner states the rule in one + sentence, with a switch to force it off. + +### 5. Chart A — "Arrest rate probability density" + +**`src/charts/RateDensityPanel.jsx`** (new). Replaces `RateDensityRidgeline.jsx`. + +- Two stacked sub-panels, Female over Male, sharing one x-axis (per 1,000). Collapses to a + single panel when pooled. This preserves the palette contract documented at + `src/utils/colors.js:9-13`: race is hue, sex is position — never a second hue. +- Within a sub-panel, selected groups overlap as filled areas (fill ~0.4 opacity, 2px stroke in + the race color), direct-labelled at each peak so there's no legend hunting. +- Below each sub-panel's baseline, a thin rail stacks one Agresti–Coull point-range per selected + group in the matching color — the R figure's `position_nudge` idea, but un-overplotted. +- Segmented control for the four unified quadrant specs. A "Compare all four specifications" + switch expands to four ridge rows (matching Fig 1's structure) and triggers four shard + fetches — cheap for NV (~100KB each), ~6.3MB each for CA, so it stays opt-in with a spinner. +- x-domain: max of the density supports and the frequentist upper bounds, with the existing cap + logic from `RateDensityRidgeline.jsx:101` and `niceTicks`. +- Caption states 500 posterior predictive draws and a 95% Agresti–Coull observed interval. + +### 6. Chart B — "Model Estimated Differences" + +**`src/charts/GroupDifference.jsx`** (new). + +- Two group pickers; defaults are the two groups with the most observed arrests (pooled groups + when pooling is on). Δ = rate(A) − rate(B) per 1,000, computed per draw index. +- Single density, filled with an SVG `linearGradient` mapped across x. Use a **diverging ramp + centered at zero** — navy for Δ<0, paper at 0, ember for Δ>0 — rather than the paper's YlOrRd: + the quantity is signed, and diverging-at-zero is the honest encoding. On-brand via + `tokens.css`. +- Dashed vertical rule at 0 in `--cv-danger`, matching the paper's red line. +- Large readout `Pr(Δ > 0)` with a plain-language sentence beneath ("In 94.4% of posterior + draws, the Black male arrest rate exceeds the White male rate"), plus median Δ and an 80%/95% + interval as a point-range. +- Degrade explicitly when fewer than two groups have usable draws — say why, don't render empty. + +### 7. Fix the district suggestions + +**`scripts/build-top-districts.mjs`** (new) — pages `/estimates?state=XX&year=21-22&limit=1000` +using `meta.total` (confirmed present in the envelope) across all 51 states, aggregates observed +arrests per LEAID, and writes `public/data/top_districts.json` with the top 15 per state +(leaid, name, arrests, enrollment, rate). Roughly 140 requests as a one-off; the output is +~50KB and gets committed, following the `public/data/national_rates.json` precedent. + +**`src/components/DistrictSearch.jsx`** — read the fixture instead of calling +`fetchStateDistricts` at runtime. The search screen loses a multi-second fetch and the ranking +becomes correct. Keep live name search on `/districts` unchanged. Drop the hardcoded +"Try Derby (KS), Paterson (NJ)…" hint at `DistrictSearch.jsx:164` — the real list supersedes it. + +### 8. Wiring, deletions, docs + +- **`src/components/ChartPanel.jsx`** — owns pooling state, selected groups, selected spec, and + the difference pair; passes them down. Keep `ArrestsOverTime`. Drop `QUADRANT_MODELS` + prefetch of all four models' *summaries* if only the selected one is needed. +- **Delete**: `src/charts/RateByGroupBar.jsx`, `src/charts/RateDensityRidgeline.jsx`. +- **Keep**: `distributionApprox.js` and `ApproxNote.jsx` — still the fallback when draws can't + be fetched (`AGENTS.md:91-101`). Fix its `intervalMass` default to 0.95 to match the API. +- **`package.json`** — add `d3-scale`, `d3-shape`, `d3-array`, `d3-interpolate` as real + `dependencies` (the existing deps are all miscategorised under `devDependencies`; leave that + alone unless it's breaking the build). +- **Docs**: `AGENTS.md` still describes 6 charts, D3 selections, and `DistrictVsNational` / + `ModelDrawsComparison` / `ExceedanceProbability` — none of which exist. Rewrite the + architecture, data-flow, and interval sections. Update `README.md`'s chart list. + +--- + +## Verification + +1. `npm run build` clean; `npm test` (`node --test 'src/**/*.test.js'`) passes, including new + tests for `agrestiCoull`, `pooling`, `densityProfile`, and the extended `drawGroups`. +2. `npm run dev`, then walk these districts: + - **Clark County NV (`3200060`, 100 arrests / 148,928 students)** — pooling stays off, all + four races render, differences chart defaults to the top two groups. + - **Carson City NV (`3200390`, 6 arrests / 4,073 students)** — pooling auto-engages (6 < 20), + banner appears, table collapses to four races. The AI/AN cell should render as a discrete + mass profile, not a smear. Verified against the draws: pooling narrows AI/AN's 90% interval + from 37.7 to 27.0 per 1,000, and Hispanic male's from 4.9 to 2.5. + - **Washoe County NV (`3200480`)** — cross-check the summary table's observed counts and + rates against `/api/v1/estimates/3200480?model=unified_m4_mod&year=21-22`. + - **A California district** — confirm the suggestion fixture ranks correctly (this is the + case the current code gets wrong), and that "compare all four" warns/spins before pulling + ~25MB of shards. +3. Toggle every group off, then on; switch specs; flip pooling manually — no stale draws from a + previous model should ever appear under a new label. +4. Compare chart A against `wp_fig_group_density` output for Clark County: same curve shapes, + same point-range positions. +5. Browser console clean on hard refresh; check the Network tab shows one shard fetch per + (model, state) and no repeats when navigating between districts. + +## Effort + +Roughly **2–3 focused days** end to end: ~1 day for the draws/pooling/util layer with tests, +~1 day for the two charts, ~half a day for the table, the suggestion fixture, and docs. At ~10 +hours a week that's about two calendar weeks. + +The R Shiny alternative would be slower, not faster — it trades a zero-server static site for a +container, an R runtime, and server-side access to either the 91GB draws DuckDB or the 51-state +parquet tree, and turns every toggle into a round-trip re-render. The React app already fetches +real draws client-side and already carries the design tokens.