Files
crdc-demo/AGENTS.md
T
jared 654b42ca71
Deploy to git-pages / deploy (push) Successful in 19s
fix: raise MAX_RATE_DOMAIN from 30 to 100 per 1,000 students
The previous cap of 30 was exceeded by over 60% of districts — mostly sparse
groups with small enrollment cells whose Agresti-Coull upper bounds genuinely
extend that far. Analysis across all 51 states (16,279 districts) showed the
median max x-value is already ~56 per 1,000; only truly degenerate cases like a
single predicted arrest in a four-student cell (~250/1000) need capping.

A cap of 100 still guards against these outliers while letting realistic data
drive the axis for the vast majority of districts. The existing 'clipped' flag
and note mechanism remain unchanged — they activate only when extreme values are
encountered.
2026-08-22 20:29:32 -04:00

278 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Agent Guide — CRDC Demo App
Context, decisions, and guidance for agents working on this codebase.
**`src/` is the source of truth.** If this file and the code disagree, the code
wins and this file is the bug — fix it in the same change.
## Project Overview
A React + Vite static web app demonstrating the
[CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors
pick a state, search for a school district, and see what was actually reported
alongside what the Bayesian models estimate. Deployed via Gitea Actions to
`pages.civilytics.org/crdc-demo/`.
The results page is a port of the white paper's Figs 6 and 7
(`wp_fig_group_density` / `wp_fig_group_difference` in
`crdc-arrests/R/paper_figures.R:472-568`). Keeping it recognisably the same
figure is the point — someone who has read the paper should see the paper.
## Architecture Summary
- **Frontend**: React 19 + Vite (static site)
- **Styling**: CSS custom properties mirroring the Civilytics design tokens
(`src/styles/tokens.css`)
- **Charts**: inline SVG that **React owns**. `d3-scale`, `d3-shape`,
`d3-array` and `d3-interpolate` supply scales, path generators and colour
interpolation only. **No d3 selections, no `useEffect` DOM mutation** —
if you find yourself reaching for `d3.select`, the answer is a render.
- **API**: public read-only API called directly from the browser; no backend
- **Deployment**: Gitea Actions (`.gitea/workflows/pages.yml`)
## Key Components and Data Flow
```
App.jsx (router)
→ StateSelector landing screen: state dropdown/grid
→ DistrictSearch live /districts search + suggestions from the
│ committed public/data/top_districts.json fixture
→ LoadingAnimation warms the API, animated histogram grid
→ ChartPanel OWNS all cross-chart state (see below)
├── DistrictSummaryTable observed arrests + enrollment; the checkboxes
│ here are the density panel's group control
├── RateDensityPanel Chart A — posterior density per group,
│ Female over Male, Agresti–Coull rail beneath
├── GroupDifference Chart B — posterior of Δ between two groups
└── ArrestsOverTime observed vs. modelled totals across 3 waves
```
### `ChartPanel` owns the state
Selected model specification, sex pooling, which groups are checked, and the
difference pair all live in `ChartPanel` and are passed down. Charts hold none
of it. That is what keeps the table's checkboxes and the density panel from
drifting apart.
One trap worth knowing: pooled group keys (`'BL'`) and unpooled ones (`'BL_F'`)
are different namespaces. Selections are therefore stored **with the pooling
mode they were made in** and fall back to defaults when the mode changes.
### Data fetching
1. **Wave data** (`ArrestsOverTime`): `unified_m3_mod` for `['15-16', '17-18',
'21-22']`. Three-year models are used because they return observed counts
across all waves; one-year models only cover the most recent one.
2. **Current-wave summary**: fetched for the **selected specification only**.
Enrollment and observed arrests are district facts, not model outputs, and
the modelled shapes now come from real draws — prefetching all four specs'
summaries would be four requests for data three of which are never read.
3. **Posterior draws**: `useDrawDistribution` (see below).
## The draws pipeline — read this before touching `useDrawDistribution.js`
`useDrawDistribution({leaid, state, models, year})` fetches the published
posterior draws from the Hugging Face parquet dataset
(`civilytics/crdc-school-arrest-rates`) and queries them client-side with
`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js`). There is no server-side
draws endpoint.
**A state's draws are split across multiple parquet parts.** `data_0.parquet`,
`data_1.parquet`, … and the count varies by state: Nevada is one file,
California is **eight** (6.2MB total, of which `data_0` is 37KB and holds 11 of
California's 1,715 districts). The app fetched only `data_0` until 2026-08-12,
which made every CA district except those 11 look absent from the published
data and silently fall back to the approximation — invisible in testing because
Nevada, the district everyone tests with, has exactly one part.
Parts are discovered from the Hugging Face **tree listing API**
(`/api/datasets/{id}/tree/main/parquet/...`), not by probing `data_N` until a
404. A 404 is logged as a console error by the browser's network layer however
cleanly the fetch handles it, and a red error on every page load is
indistinguishable from a real one. Probing (via HEAD) remains the fallback if
the listing API is unavailable. All parts are fetched in parallel, registered
individually, and queried as `read_parquet([...])`.
Four more properties it is easy to break:
- **It returns counts, not rates**, indexed by `draw_id - 1`. Counts are what
make sex pooling and between-group differences possible: both have to sum or
subtract numerators and denominators separately. Callers divide.
- **Indexing by `draw_id`, not push order**, makes DuckDB's row ordering
irrelevant and turns a missing draw into a hole rather than a short array.
- **Incomplete groups are dropped** (`isCompleteDrawSet`). A group present for
300 of 500 draws would otherwise get an interval computed off a biased
subsample that looks identical on screen to a complete one.
- **The reset-before-guard ordering in the effect is load-bearing.** State is
cleared on *every* input change, including ones with nothing to fetch, so a
failed fetch can never leave the previous model's draws on screen under the
new model's label.
`status` is an ANY-model, ANY-group signal. To claim "these are all real draws"
for a specific set of rendered groups, use `hasDrawsForAll`.
### Caption wording: "posterior predictive draws"
`draw_id` is renumbered 1–500 per write batch upstream
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in
different batches, so draw *k* of one group is **not** the same parameter draw
as draw *k* of another. Measured correlation between Black-male and
Hispanic-male `pred` in Clark County was 0.019 even within a batch —
`posterior_predict` observation noise dominates.
The published Fig 7 has the same property, so the app matches the paper. What
neither can claim is a paired-parameter contrast. **Captions must say "posterior
predictive draws" and must never say "paired parameter draws".**
### Fallback path — do not delete
If the draws fetch fails (network, unsupported browser, HF outage),
`RateDensityPanel` falls back **per group** to the analytic approximation in
`src/utils/distributionApprox.js` and shows `<ApproxNote />`. Neither
`distributionApprox.js` nor `ApproxNote.jsx` is dead code.
A pooled group has no fallback shape: adding two groups' interval *bounds*
together is not a pooled interval, so `buildDisplayGroups` sets `modeled: null`
when pooling and the chart omits that group rather than inventing a curve.
The duckdb-wasm engine is ~39MB uncompressed / ~8.86MB gzipped (measured
against the shipped package). It loads via dynamic `import()` only once a
district is selected — never on initial page load — and is browser-cached
thereafter, but it is a real one-time cost.
### Error bar convention: 95%
The API returns **95%** intervals. `validate_interval()` in
`crdc-arrests/api/R/validate.R` defaults to `95L` and this app never passes
`interval=`. `fitSkewedInterval`'s `intervalMass` therefore defaults to `0.95`,
and `ArrestsOverTime`'s legend says "95% interval". (Both said 90% before
2026-08-12; that was a bug, not a convention.)
The observed-data point ranges are a different thing again: a 95%
Agresti–Coull interval computed from observed counts
(`src/utils/agrestiCoull.js`), a direct port of `agresti_coull()` in
`crdc-arrests/R/paper_figures.R:219-237`. Two faithfulness quirks are pinned by
tests and must not be "fixed": the bounds are on the **count** scale, and
`lower` can be **negative** for very small numerators (charts clamp at draw
time, the port does not).
## Tuning decisions with stated rationale — don't re-derive
- **`src/utils/colors.js:9-13`** — the race palette passed the dataviz skill's
CVD validator. Re-run `validate_palette.js` before changing any hex value.
Race is hue; **sex is position, never a second hue**; observed-vs-modelled is
mark type, never a second hue.
- **`src/utils/kde.js`** — `BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`
were tuned for zero-inflated sparse-district posteriors. Both ends matter.
- **`src/utils/pooling.js`** — `POOL_BY_SEX_ARREST_THRESHOLD = 20`, applied to
the district total, not per cell.
- **`src/utils/densityProfile.js`** — `MASS_MAX_DISTINCT = 12`. Below it the
posterior predictive is drawn as discrete mass, because it *is* discrete; a
Gaussian KDE over four achievable values renders as a lumpy smear that reads
as a rendering bug.
- **`src/utils/rateDomain.js`** — `MAX_RATE_DOMAIN = 100` caps the axis so a
four-student cell can't squash every other curve (up from 30, which was exceeded by over 60% of districts). It reports `clipped` so the
chart says so instead of silently cropping.
## Common Pitfalls & Gotchas
### 1. Null safety in chart components
API responses can have empty arrays or missing fields for districts with no
arrests. Use optional chaining and explicit defaults, and never divide by a
denominator you haven't checked:
```javascript
const enroll = row.stu_enroll || 0
const rate = enroll > 0 ? (row.observed_arrests || 0) / enroll * 1000 : 0
```
A rate with no denominator is **not zero and not Infinity — it's undefined**.
`toRates` returns `[]`, `buildDisplayGroups` reports `0` and lets the
enrollment column explain why.
### 2. SVG dimensions and responsiveness
Charts set `width="100%"` with a fixed `viewBox`, wrapped in
`overflowX: 'auto'`. Chart cards are `max-width: 70rem` in `ChartPanel.jsx`
(vs. the ~60rem default text width).
### 3. API endpoint availability
- `/api/v1/estimates/{leaid}` — ✅ summary rows (median, bounds, enrollment,
observed arrests)
- `/api/v1/estimates?state=XX&...` — ✅ but returns rows `ORDER BY LEAID, RACE,
SEX` at 8 per district, capped at `limit=1000`. **Any short read ranks the
lowest-LEAID districts, not the busiest.** Page it with `meta.total`.
- `/api/v1/draws?...` — returns a shard URL + SQL, not draw data. The app goes
to the parquet directly.
### 4. CORS
The API sends no CORS headers. The app auto-detects a proxy via `VITE_PROXY_URL`;
unset, it fetches directly (works same-origin or behind the Docker/nginx proxy).
## The suggestion fixture
`public/data/top_districts.json` holds the top 15 districts per state by
observed arrests, and is generated by a one-off, read-only script:
```bash
node scripts/build-top-districts.mjs # all 51, ~150 requests
node scripts/build-top-districts.mjs --states NV,CA # spot-check
```
It is committed (the `national_rates.json` precedent). Re-run it only when a
new CRDC wave lands. `DistrictSearch` degrades to search-only if it's missing.
## Testing
`npm test` runs `node --test 'src/**/*.test.js'`. The pure utilities are all
covered and **should be written test-first**:
| Module | What its tests pin |
|---|---|
| `agrestiCoull.js` | five cases against real R output, incl. the negative lower bound |
| `pooling.js` | numerator and denominator always drawn from the same groups |
| `densityProfile.js` | the mass/KDE switch, and that KDE delegates to `kde.js` |
| `districtGroups.js` | display-row derivation, defaults, pooled vs unpooled keys |
| `groupDifference.js` | refusal to pair mismatched draw sets |
| `rateDomain.js` | the axis cap, and that it reports clipping |
| `drawGroups.js` | key shapes, and hole detection in draw arrays |
| `kde.js` | bandwidth clamp behaviour |
Components are verified manually — there is no DOM test harness. Useful
districts: **Clark County NV `3200060`** (100 arrests / 148,928 students,
pooling off), **Carson City NV `3200390`** (6 arrests, pooling auto-engages,
discrete mass profile), **Washoe County NV `3200480`** (cross-check the table
against the API), and any California district (large shards, fixture ranking).
## Deployment Checklist
1. `npm run build` succeeds
2. `npm test` passes
3. No console errors after a hard refresh
4. Both charts render for a sample district; Network tab shows **one fetch per
(model, state) shard part** and no repeats when switching specs
5. Test with a **multi-part state** (California), not just Nevada — a
single-part state cannot catch a regression in part discovery
> Caveat on the cache: `App.jsx:112`'s "Search another district" button does
> `window.location.href = '/crdc-demo/'`, a full page reload, which discards the
> module-level `shardCache`. Within one district view — switching specs,
> toggling compare-all — the cache works as intended.
## Git Conventions
- Remote: `https://gitea.civilytics.org/Civilytics/crdc-demo.git`
- Default branch: `main`
- Gitea Actions auto-deploys on push to `main`
- Commit messages explain the *why*: `fix: rank suggestions over the full state
(limit=500 was selecting the lowest 62 LEAIDs)`
## Related Repositories
- **crdc-arrests** — API server (Plumber/R) and the white paper, at
`/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/`
- **civilyticsR** — wordmark and visualization functions