# Agent Guide — CRDC Demo App Context, decisions, and guidance for agents working on this codebase. **`src/` is the source of truth.** If this file and the code disagree, the code wins and this file is the bug — fix it in the same change. ## Project Overview A React + Vite static web app demonstrating the [CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors pick a state, search for a school district, and see what was actually reported alongside what the Bayesian models estimate. Deployed via Gitea Actions to `pages.civilytics.org/crdc-demo/`. The results page is a port of the white paper's Figs 6 and 7 (`wp_fig_group_density` / `wp_fig_group_difference` in `crdc-arrests/R/paper_figures.R:472-568`). Keeping it recognisably the same figure is the point — someone who has read the paper should see the paper. ## Architecture Summary - **Frontend**: React 19 + Vite (static site) - **Styling**: CSS custom properties mirroring the Civilytics design tokens (`src/styles/tokens.css`) - **Charts**: inline SVG that **React owns**. `d3-scale`, `d3-shape`, `d3-array` and `d3-interpolate` supply scales, path generators and colour interpolation only. **No d3 selections, no `useEffect` DOM mutation** — if you find yourself reaching for `d3.select`, the answer is a render. - **API**: public read-only API called directly from the browser; no backend - **Deployment**: Gitea Actions (`.gitea/workflows/pages.yml`) ## Key Components and Data Flow ``` App.jsx (router) → StateSelector landing screen: state dropdown/grid → DistrictSearch live /districts search + suggestions from the │ committed public/data/top_districts.json fixture → LoadingAnimation warms the API, animated histogram grid → ChartPanel OWNS all cross-chart state (see below) ├── DistrictSummaryTable observed arrests + enrollment; the checkboxes │ here are the density panel's group control ├── RateDensityPanel Chart A — posterior density per group, │ Female over Male, Agresti–Coull rail beneath ├── GroupDifference Chart B — posterior of Δ between two groups └── ArrestsOverTime observed vs. modelled totals across 3 waves ``` ### `ChartPanel` owns the state Selected model specification, sex pooling, which groups are checked, and the difference pair all live in `ChartPanel` and are passed down. Charts hold none of it. That is what keeps the table's checkboxes and the density panel from drifting apart. One trap worth knowing: pooled group keys (`'BL'`) and unpooled ones (`'BL_F'`) are different namespaces. Selections are therefore stored **with the pooling mode they were made in** and fall back to defaults when the mode changes. ### Data fetching 1. **Wave data** (`ArrestsOverTime`): `unified_m3_mod` for `['15-16', '17-18', '21-22']`. Three-year models are used because they return observed counts across all waves; one-year models only cover the most recent one. 2. **Current-wave summary**: fetched for the **selected specification only**. Enrollment and observed arrests are district facts, not model outputs, and the modelled shapes now come from real draws — prefetching all four specs' summaries would be four requests for data three of which are never read. 3. **Posterior draws**: `useDrawDistribution` (see below). ## The draws pipeline — read this before touching `useDrawDistribution.js` `useDrawDistribution({leaid, state, models, year})` fetches the published posterior draws from the Hugging Face parquet dataset (`civilytics/crdc-school-arrest-rates`) and queries them client-side with `@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js`). There is no server-side draws endpoint. **A state's draws are split across multiple parquet parts.** `data_0.parquet`, `data_1.parquet`, … and the count varies by state: Nevada is one file, California is **eight** (6.2MB total, of which `data_0` is 37KB and holds 11 of California's 1,715 districts). The app fetched only `data_0` until 2026-08-12, which made every CA district except those 11 look absent from the published data and silently fall back to the approximation — invisible in testing because Nevada, the district everyone tests with, has exactly one part. Parts are discovered from the Hugging Face **tree listing API** (`/api/datasets/{id}/tree/main/parquet/...`), not by probing `data_N` until a 404. A 404 is logged as a console error by the browser's network layer however cleanly the fetch handles it, and a red error on every page load is indistinguishable from a real one. Probing (via HEAD) remains the fallback if the listing API is unavailable. All parts are fetched in parallel, registered individually, and queried as `read_parquet([...])`. Four more properties it is easy to break: - **It returns counts, not rates**, indexed by `draw_id - 1`. Counts are what make sex pooling and between-group differences possible: both have to sum or subtract numerators and denominators separately. Callers divide. - **Indexing by `draw_id`, not push order**, makes DuckDB's row ordering irrelevant and turns a missing draw into a hole rather than a short array. - **Incomplete groups are dropped** (`isCompleteDrawSet`). A group present for 300 of 500 draws would otherwise get an interval computed off a biased subsample that looks identical on screen to a complete one. - **The reset-before-guard ordering in the effect is load-bearing.** State is cleared on *every* input change, including ones with nothing to fetch, so a failed fetch can never leave the previous model's draws on screen under the new model's label. `status` is an ANY-model, ANY-group signal. To claim "these are all real draws" for a specific set of rendered groups, use `hasDrawsForAll`. ### Caption wording: "posterior predictive draws" `draw_id` is renumbered 1–500 per write batch upstream (`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in different batches, so draw *k* of one group is **not** the same parameter draw as draw *k* of another. Measured correlation between Black-male and Hispanic-male `pred` in Clark County was 0.019 even within a batch — `posterior_predict` observation noise dominates. The published Fig 7 has the same property, so the app matches the paper. What neither can claim is a paired-parameter contrast. **Captions must say "posterior predictive draws" and must never say "paired parameter draws".** ### Fallback path — do not delete If the draws fetch fails (network, unsupported browser, HF outage), `RateDensityPanel` falls back **per group** to the analytic approximation in `src/utils/distributionApprox.js` and shows ``. Neither `distributionApprox.js` nor `ApproxNote.jsx` is dead code. A pooled group has no fallback shape: adding two groups' interval *bounds* together is not a pooled interval, so `buildDisplayGroups` sets `modeled: null` when pooling and the chart omits that group rather than inventing a curve. The duckdb-wasm engine is ~39MB uncompressed / ~8.86MB gzipped (measured against the shipped package). It loads via dynamic `import()` only once a district is selected — never on initial page load — and is browser-cached thereafter, but it is a real one-time cost. ### Error bar convention: 95% The API returns **95%** intervals. `validate_interval()` in `crdc-arrests/api/R/validate.R` defaults to `95L` and this app never passes `interval=`. `fitSkewedInterval`'s `intervalMass` therefore defaults to `0.95`, and `ArrestsOverTime`'s legend says "95% interval". (Both said 90% before 2026-08-12; that was a bug, not a convention.) The observed-data point ranges are a different thing again: a 95% Agresti–Coull interval computed from observed counts (`src/utils/agrestiCoull.js`), a direct port of `agresti_coull()` in `crdc-arrests/R/paper_figures.R:219-237`. Two faithfulness quirks are pinned by tests and must not be "fixed": the bounds are on the **count** scale, and `lower` can be **negative** for very small numerators (charts clamp at draw time, the port does not). ## Tuning decisions with stated rationale — don't re-derive - **`src/utils/colors.js:9-13`** — the race palette passed the dataviz skill's CVD validator. Re-run `validate_palette.js` before changing any hex value. Race is hue; **sex is position, never a second hue**; observed-vs-modelled is mark type, never a second hue. - **`src/utils/kde.js`** — `BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR` were tuned for zero-inflated sparse-district posteriors. Both ends matter. - **`src/utils/pooling.js`** — `POOL_BY_SEX_ARREST_THRESHOLD = 20`, applied to the district total, not per cell. - **`src/utils/densityProfile.js`** — `MASS_MAX_DISTINCT = 12`. Below it the posterior predictive is drawn as discrete mass, because it *is* discrete; a Gaussian KDE over four achievable values renders as a lumpy smear that reads as a rendering bug. - **`src/utils/rateDomain.js`** — `MAX_RATE_DOMAIN = 100` caps the axis so a four-student cell can't squash every other curve (up from 30, which was exceeded by over 60% of districts). It reports `clipped` so the chart says so instead of silently cropping. ## Common Pitfalls & Gotchas ### 1. Null safety in chart components API responses can have empty arrays or missing fields for districts with no arrests. Use optional chaining and explicit defaults, and never divide by a denominator you haven't checked: ```javascript const enroll = row.stu_enroll || 0 const rate = enroll > 0 ? (row.observed_arrests || 0) / enroll * 1000 : 0 ``` A rate with no denominator is **not zero and not Infinity — it's undefined**. `toRates` returns `[]`, `buildDisplayGroups` reports `0` and lets the enrollment column explain why. ### 2. SVG dimensions and responsiveness Charts set `width="100%"` with a fixed `viewBox`, wrapped in `overflowX: 'auto'`. Chart cards are `max-width: 70rem` in `ChartPanel.jsx` (vs. the ~60rem default text width). ### 3. API endpoint availability - `/api/v1/estimates/{leaid}` — ✅ summary rows (median, bounds, enrollment, observed arrests) - `/api/v1/estimates?state=XX&...` — ✅ but returns rows `ORDER BY LEAID, RACE, SEX` at 8 per district, capped at `limit=1000`. **Any short read ranks the lowest-LEAID districts, not the busiest.** Page it with `meta.total`. - `/api/v1/draws?...` — returns a shard URL + SQL, not draw data. The app goes to the parquet directly. ### 4. CORS The API sends no CORS headers. The app auto-detects a proxy via `VITE_PROXY_URL`; unset, it fetches directly (works same-origin or behind the Docker/nginx proxy). ## The suggestion fixture `public/data/top_districts.json` holds the top 15 districts per state by observed arrests, and is generated by a one-off, read-only script: ```bash node scripts/build-top-districts.mjs # all 51, ~150 requests node scripts/build-top-districts.mjs --states NV,CA # spot-check ``` It is committed (the `national_rates.json` precedent). Re-run it only when a new CRDC wave lands. `DistrictSearch` degrades to search-only if it's missing. ## Testing `npm test` runs `node --test 'src/**/*.test.js'`. The pure utilities are all covered and **should be written test-first**: | Module | What its tests pin | |---|---| | `agrestiCoull.js` | five cases against real R output, incl. the negative lower bound | | `pooling.js` | numerator and denominator always drawn from the same groups | | `densityProfile.js` | the mass/KDE switch, and that KDE delegates to `kde.js` | | `districtGroups.js` | display-row derivation, defaults, pooled vs unpooled keys | | `groupDifference.js` | refusal to pair mismatched draw sets | | `rateDomain.js` | the axis cap, and that it reports clipping | | `drawGroups.js` | key shapes, and hole detection in draw arrays | | `kde.js` | bandwidth clamp behaviour | Components are verified manually — there is no DOM test harness. Useful districts: **Clark County NV `3200060`** (100 arrests / 148,928 students, pooling off), **Carson City NV `3200390`** (6 arrests, pooling auto-engages, discrete mass profile), **Washoe County NV `3200480`** (cross-check the table against the API), and any California district (large shards, fixture ranking). ## Deployment Checklist 1. `npm run build` succeeds 2. `npm test` passes 3. No console errors after a hard refresh 4. Both charts render for a sample district; Network tab shows **one fetch per (model, state) shard part** and no repeats when switching specs 5. Test with a **multi-part state** (California), not just Nevada — a single-part state cannot catch a regression in part discovery > Caveat on the cache: `App.jsx:112`'s "Search another district" button does > `window.location.href = '/crdc-demo/'`, a full page reload, which discards the > module-level `shardCache`. Within one district view — switching specs, > toggling compare-all — the cache works as intended. ## Git Conventions - Remote: `https://gitea.civilytics.org/Civilytics/crdc-demo.git` - Default branch: `main` - Gitea Actions auto-deploys on push to `main` - Commit messages explain the *why*: `fix: rank suggestions over the full state (limit=500 was selecting the lowest 62 LEAIDs)` ## Related Repositories - **crdc-arrests** — API server (Plumber/R) and the white paper, at `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/` - **civilyticsR** — wordmark and visualization functions