Compare commits

...
4 Commits
Author SHA1 Message Date
jared 4396a56a4a docs: rewrite AGENTS.md and the README chart list for the rebuilt app
Deploy to git-pages / deploy (push) Successful in 23s
AGENTS.md described 6 charts, D3 selections, and DistrictVsNational /
ModelDrawsComparison / ExceedanceProbability — none of which exist. An agent
reading it as authoritative would have been actively misled, so it now opens by
saying src/ wins any disagreement.

Rewritten: the real component tree and data flow, ChartPanel as the owner of all
cross-chart state, the pooled/unpooled key namespace trap, the four properties
of the draws pipeline that are easy to break, the multi-part shard layout and
why discovery uses the tree listing, the "posterior predictive draws" wording
rule and the reason for it, the 95% convention, and a table of which tuning
decisions carry stated rationale and should not be re-derived (the CVD-validated
palette, the KDE bandwidth clamp, the pooling threshold, the mass/KDE cutoff,
the axis cap).

Adds a testing section — there are automated tests now — and notes that the
deployment check should use a multi-part state like California, since a
single-part state cannot catch a regression in part discovery. Records that
App.jsx's "Search another district" is a full page reload that discards the
shard cache.

README: the chart list becomes the summary table plus the two ported figures,
the file tree matches src/, the d3 role is stated precisely (scales and paths,
not selections), and the endpoint table warns that /estimates?state= ranks by
LEAID on a short read.

Also commits the plan this work followed.
2026-08-12 08:41:01 -04:00
jared c1920c11fb feat: rebuild the results page around the white-paper figures
Ports Figs 6 and 7 from crdc-arrests/R/paper_figures.R (wp_fig_group_density and
wp_fig_group_difference) so someone who has read the paper sees the paper. The
three old charts become a summary table and two charts; ArrestsOverTime stays.

Draws pipeline. useDrawDistribution now selects draw_id and returns predicted
counts indexed by draw rather than pre-divided rates. That one column is what
unlocks the rest: pooling has to sum numerators and denominators separately, and
a between-group difference has to be taken at a common draw index. Indexing by
draw_id rather than push order makes DuckDB's row ordering irrelevant and turns
a missing draw into a hole, which isCompleteDrawSet then rejects — a group
present for 300 of 500 draws would otherwise get an interval computed off a
biased subsample that looks identical on screen to a complete one. It also takes
models[] instead of a single model, so "compare all four" needs no conditional
hooks. The reset-before-guard ordering is preserved.

Also fixes a pre-existing bug in that pipeline: a state's draws are split across
data_0.parquet, data_1.parquet, … and the part count varies by state. The app
only ever fetched data_0. Nevada has one part, so this was invisible in every
Nevada test; California has eight, totalling 6.2MB, of which data_0 is 37KB and
holds 11 of California's 1,715 districts. Every other CA district looked absent
from the published data and silently fell back to the approximation. Parts are
now discovered from the Hugging Face tree listing API and fetched in parallel —
listing rather than probing data_N until a 404, because the browser logs a 404
as a console error however cleanly the fetch handles it, and a red error on
every load is indistinguishable from a real one. HEAD probing is the fallback.

New pure utils, all written against tests first:
  agrestiCoull    faithful port incl. the zero-numerator rule of three and the
                  negative lower bound at (1, 53); pinned to five R outputs
  pooling         sex pooling for sparse districts; numerator and denominator
                  are always drawn from the same set of groups
  densityProfile  discrete probability mass below 12 distinct values, KDE above;
                  KDE delegates to kde.js, whose bandwidth clamp is untouched
  districtGroups  display-row derivation, defaults, pooled vs unpooled keys
  groupDifference per-draw delta; refuses to pair mismatched draw sets
  rateDomain      shared x-axis, with a clip flag so the cap is never silent

Chart A keeps the palette contract: race is hue, sex is position. Female and
Male are stacked panels sharing one axis, collapsing to one panel when pooled —
never a second hue. Chart B uses a diverging ramp centred at zero rather than
the paper's sequential YlOrRd, because delta is signed and a sequential ramp
encodes "more" where the data means "which direction".

Captions say "posterior predictive draws", never "paired parameter draws":
draw_id is renumbered per write batch upstream and a district's groups land in
different batches, so cross-group pairing is effectively independent (measured
cor ~= 0.02). The published figure has the same property; what neither can claim
is a paired-parameter contrast.

Sparse districts (under 20 arrests district-wide) pool Female and Male within
each race, with a banner stating the rule and a switch to override it. A pooled
group carries no modelled interval — summing two groups' interval bounds is not
a pooled interval, and there is no honest way to fake one without the draws.

React still owns the DOM. d3-scale/shape/array/interpolate supply scales, path
generators and colour interpolation; no selections, no useEffect DOM mutation.

Deletes RateByGroupBar and RateDensityRidgeline. Keeps distributionApprox.js and
ApproxNote.jsx — still the per-group fallback when draws can't be fetched.

112 tests pass; npm run build clean.
2026-08-12 08:40:46 -04:00
jared 63413f9eb7 fix: rank district suggestions over the full state, not the first 500 rows
DistrictSearch built its "most arrests" list from
/estimates?state=XX&year=21-22&limit=500. That endpoint returns rows
ORDER BY LEAID, RACE, SEX at eight rows per district, so a 500-row cap is not a
sample of the state — it is the ~62 lowest-LEAID districts in it.

Measured against California (11,488 rows, 1,715 districts): the old read covered
68 districts, and 6 of the true top 8 were invisible to it. It suggested
districts with 1 and 2 arrests as the state's most notable, while San Diego
Unified (178), Fresno Unified (77) and Kern High (69) never appeared.

Replaced with a committed fixture, public/data/top_districts.json, generated by
scripts/build-top-districts.mjs. The script pages each state to completion using
meta.total from the response envelope and fails loudly on a short read, since a
silent truncation there would reintroduce exactly this bug. 51 states, 135
requests, ~104KB, following the national_rates.json precedent. Re-run it only
when a new CRDC wave lands.

The search screen also loses a multi-second fetch on every visit, and the
hardcoded "Try Derby (KS), Paterson (NJ)" hint goes with it — the real list
supersedes it. Degrades to search-only if the fixture is missing.

fetchStateDistricts() is kept for scripts and ad-hoc use, with its JSDoc now
warning that any short read ranks by LEAID.

pages.yml gains public/** in its paths filter: the fixture ships with the build,
so regenerating it has to be able to trigger a deploy on its own.
2026-08-12 08:40:23 -04:00
jared c62d1e3068 fix: report the API's 95% intervals as 95%, not 90%
The API returns 95% intervals: validate_interval() in
crdc-arrests/api/R/validate.R defaults to 95L and this app never passes
`interval=`. Two places claimed 90% anyway.

fitSkewedInterval's `intervalMass` defaulted to 0.90, so the analytic fallback
fitted 95% bounds as if they covered 90% of the mass. That divides each
half-interval by 1.645 instead of 1.960 and understates sigma by ~16% — the
fallback drew a distribution visibly narrower than the model's own, in the one
code path where we have no draws to check it against.

ArrestsOverTime's legend read "Modeled (median + 90% interval)" while plotting
count_lower/count_upper, which are the same 95% bounds.

Also exports probit() from distributionApprox.js so the Agresti-Coull port can
reuse it rather than carrying a second qnorm implementation.
2026-08-12 08:40:02 -04:00
35 changed files with 8492 additions and 691 deletions
+4
View File
@@ -10,6 +10,10 @@ on:
branches: [main]
paths:
- 'src/**'
# Committed fixtures ship with the build: public/data/top_districts.json
# is what the search screen ranks its suggestions from, so regenerating it
# has to be able to trigger a deploy on its own.
- 'public/**'
- 'index.html'
- 'vite.config.mjs'
- 'package.json'
+232 -109
View File
@@ -1,154 +1,277 @@
# Agent Guide — CRDC Demo App
This document captures context, decisions, and guidance for agents working on this codebase. It is the primary source of truth for how to make changes safely.
Context, decisions, and guidance for agents working on this codebase.
**`src/` is the source of truth.** If this file and the code disagree, the code
wins and this file is the bug — fix it in the same change.
## Project Overview
A React + Vite static web app demonstrating the [CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors select a state, search for a school district, and see 6 charts comparing observed arrest data against Bayesian model estimates. Deployed via Gitea Actions to `pages.civilytics.org/crdc-demo/`.
A React + Vite static web app demonstrating the
[CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors
pick a state, search for a school district, and see what was actually reported
alongside what the Bayesian models estimate. Deployed via Gitea Actions to
`pages.civilytics.org/crdc-demo/`.
The results page is a port of the white paper's Figs 6 and 7
(`wp_fig_group_density` / `wp_fig_group_difference` in
`crdc-arrests/R/paper_figures.R:472-568`). Keeping it recognisably the same
figure is the point — someone who has read the paper should see the paper.
## Architecture Summary
- **Frontend**: React 19 + Vite (static site generation)
- **Styling**: Plain CSS custom properties matching Civilytics design tokens (`src/styles/tokens.css`)
- **Charts**: Mixed approach:
- Charts 1–3 use inline SVG with manual scales (no D3 dependency for these)
- Chart 4 (ModelDrawsComparison) uses D3.js v7 for data-driven rendering of quadrant comparisons
- Chart 5 (RateDensityRidgeline) uses D3.js v7 for density ridge visualizations
- **API**: Calls public read-only API directly from browser; no backend required
- **Deployment**: Static site deployed via Gitea Actions (`.gitea/workflows/pages.yml`)
- **Frontend**: React 19 + Vite (static site)
- **Styling**: CSS custom properties mirroring the Civilytics design tokens
(`src/styles/tokens.css`)
- **Charts**: inline SVG that **React owns**. `d3-scale`, `d3-shape`,
`d3-array` and `d3-interpolate` supply scales, path generators and colour
interpolation only. **No d3 selections, no `useEffect` DOM mutation** —
if you find yourself reaching for `d3.select`, the answer is a render.
- **API**: public read-only API called directly from the browser; no backend
- **Deployment**: Gitea Actions (`.gitea/workflows/pages.yml`)
## Key Components and Data Flow
```
App.jsx (router)
→ StateSelector (landing screen: state dropdown/grid)
→ DistrictSearch (search + "interesting" suggestions from /estimates?state=&year=)
→ LoadingAnimation (fetches all data in parallel, shows animated histogram grid)
→ ChartPanel (receives district object, fetches structured estimates for 6 charts)
├── ArrestsOverTime — SVG line chart by wave (3 years)
├── RateByGroupBar — SVG bar chart: observed vs modeled per group
├── DistrictVsNational — SVG comparison to national average
├── ModelDrawsComparison — D3 quadrant charts (4 model types × 1 year)
├── RateDensityRidgeline — D3 density ridges per race×sex group
└── ExceedanceProbability — P(district > national) per student group
→ StateSelector landing screen: state dropdown/grid
→ DistrictSearch live /districts search + suggestions from the
│ committed public/data/top_districts.json fixture
→ LoadingAnimation warms the API, animated histogram grid
→ ChartPanel OWNS all cross-chart state (see below)
├── DistrictSummaryTable observed arrests + enrollment; the checkboxes
│ here are the density panel's group control
├── RateDensityPanel Chart A — posterior density per group,
│ Female over Male, Agresti–Coull rail beneath
├── GroupDifference Chart B — posterior of Δ between two groups
└── ArrestsOverTime observed vs. modelled totals across 3 waves
```
### Data Fetching Strategy (`ChartPanel.jsx`)
### `ChartPanel` owns the state
`ChartPanel` fetches all chart data on mount (after `LoadingAnimation` pre-fetched via batch calls):
Selected model specification, sex pooling, which groups are checked, and the
difference pair all live in `ChartPanel` and are passed down. Charts hold none
of it. That is what keeps the table's checkboxes and the density panel from
drifting apart.
1. **Wave data** for Charts 1–3: Fetches `unified_m3_mod` model estimates for years `['21-22', '17-18', '15-16']`.
- Uses three-year models because they return observed arrest counts across all waves (one-year models only have data for the most recent wave).
One trap worth knowing: pooled group keys (`'BL'`) and unpooled ones (`'BL_F'`)
are different namespaces. Selections are therefore stored **with the pooling
mode they were made in** and fall back to defaults when the mode changes.
2. **Quad data** for Charts 4–6: Fetches estimates from all four quadrant models (`unified_m1_mod`, `unified_m2_mod`, `unified_m3_mod`, `unified_m4_mod`) for year `21-22` only (one year, as the most recent wave).
### Data fetching
3. **National rates**: Loaded once from a static JSON fixture or cached by `LoadingAnimation`.
1. **Wave data** (`ArrestsOverTime`): `unified_m3_mod` for `['15-16', '17-18',
'21-22']`. Three-year models are used because they return observed counts
across all waves; one-year models only cover the most recent one.
2. **Current-wave summary**: fetched for the **selected specification only**.
Enrollment and observed arrests are district facts, not model outputs, and
the modelled shapes now come from real draws — prefetching all four specs'
summaries would be four requests for data three of which are never read.
3. **Posterior draws**: `useDrawDistribution` (see below).
### Error Bar Convention: 90% Intervals
## The draws pipeline — read this before touching `useDrawDistribution.js`
The API returns 95% HPD intervals (`count_lower`, `count_upper`). However, all chart labels and calculations in this app use **90% intervals**.
`useDrawDistribution({leaid, state, models, year})` fetches the published
posterior draws from the Hugging Face parquet dataset
(`civilytics/crdc-school-arrest-rates`) and queries them client-side with
`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js`). There is no server-side
draws endpoint.
The 90% convention is baked into the analytic fallback in `src/utils/distributionApprox.js` (the path used when real posterior draws can't be fetched — see the duckdb-wasm section below). That fallback generates **no draws**: `fitSkewedInterval` fits a two-piece normal directly from `{median, lower, upper}`, converting each half-interval to its own sigma with `z = probit((1 + intervalMass) / 2)` and `intervalMass` defaulting to `0.90` (so z ≈ 1.645). For a symmetric interval that is equivalent to the old fixed "full width / 3.29" divisor, but it is computed from `intervalMass` rather than hardcoded, and each side gets its own sigma so the fitted shape stays skewed.
**A state's draws are split across multiple parquet parts.** `data_0.parquet`,
`data_1.parquet`, … and the count varies by state: Nevada is one file,
California is **eight** (6.2MB total, of which `data_0` is 37KB and holds 11 of
California's 1,715 districts). The app fetched only `data_0` until 2026-08-12,
which made every CA district except those 11 look absent from the published
data and silently fall back to the approximation — invisible in testing because
Nevada, the district everyone tests with, has exactly one part.
If you change this convention, update:
- `distributionApprox.js` — the `intervalMass = 0.90` default in `fitSkewedInterval`, and its JSDoc claim that the API's bounds are a 90% interval
- `RateByGroupBar.jsx` — the caption and component JSDoc, both of which say "the model's reported 90% interval"
- `ArrestsOverTime.jsx` — the legend label "Modeled (median + 90% interval)"
- Any documentation referencing confidence/credible intervals
Parts are discovered from the Hugging Face **tree listing API**
(`/api/datasets/{id}/tree/main/parquet/...`), not by probing `data_N` until a
404. A 404 is logged as a console error by the browser's network layer however
cleanly the fetch handles it, and a red error on every page load is
indistinguishable from a real one. Probing (via HEAD) remains the fallback if
the listing API is unavailable. All parts are fetched in parallel, registered
individually, and queried as `read_parquet([...])`.
Four more properties it is easy to break:
- **It returns counts, not rates**, indexed by `draw_id - 1`. Counts are what
make sex pooling and between-group differences possible: both have to sum or
subtract numerators and denominators separately. Callers divide.
- **Indexing by `draw_id`, not push order**, makes DuckDB's row ordering
irrelevant and turns a missing draw into a hole rather than a short array.
- **Incomplete groups are dropped** (`isCompleteDrawSet`). A group present for
300 of 500 draws would otherwise get an interval computed off a biased
subsample that looks identical on screen to a complete one.
- **The reset-before-guard ordering in the effect is load-bearing.** State is
cleared on *every* input change, including ones with nothing to fetch, so a
failed fetch can never leave the previous model's draws on screen under the
new model's label.
`status` is an ANY-model, ANY-group signal. To claim "these are all real draws"
for a specific set of rendered groups, use `hasDrawsForAll`.
### Caption wording: "posterior predictive draws"
`draw_id` is renumbered 1–500 per write batch upstream
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in
different batches, so draw *k* of one group is **not** the same parameter draw
as draw *k* of another. Measured correlation between Black-male and
Hispanic-male `pred` in Clark County was 0.019 even within a batch —
`posterior_predict` observation noise dominates.
The published Fig 7 has the same property, so the app matches the paper. What
neither can claim is a paired-parameter contrast. **Captions must say "posterior
predictive draws" and must never say "paired parameter draws".**
### Fallback path — do not delete
If the draws fetch fails (network, unsupported browser, HF outage),
`RateDensityPanel` falls back **per group** to the analytic approximation in
`src/utils/distributionApprox.js` and shows `<ApproxNote />`. Neither
`distributionApprox.js` nor `ApproxNote.jsx` is dead code.
A pooled group has no fallback shape: adding two groups' interval *bounds*
together is not a pooled interval, so `buildDisplayGroups` sets `modeled: null`
when pooling and the chart omits that group rather than inventing a curve.
The duckdb-wasm engine is ~39MB uncompressed / ~8.86MB gzipped (measured
against the shipped package). It loads via dynamic `import()` only once a
district is selected — never on initial page load — and is browser-cached
thereafter, but it is a real one-time cost.
### Error bar convention: 95%
The API returns **95%** intervals. `validate_interval()` in
`crdc-arrests/api/R/validate.R` defaults to `95L` and this app never passes
`interval=`. `fitSkewedInterval`'s `intervalMass` therefore defaults to `0.95`,
and `ArrestsOverTime`'s legend says "95% interval". (Both said 90% before
2026-08-12; that was a bug, not a convention.)
The observed-data point ranges are a different thing again: a 95%
Agresti–Coull interval computed from observed counts
(`src/utils/agrestiCoull.js`), a direct port of `agresti_coull()` in
`crdc-arrests/R/paper_figures.R:219-237`. Two faithfulness quirks are pinned by
tests and must not be "fixed": the bounds are on the **count** scale, and
`lower` can be **negative** for very small numerators (charts clamp at draw
time, the port does not).
## Tuning decisions with stated rationale — don't re-derive
- **`src/utils/colors.js:9-13`** — the race palette passed the dataviz skill's
CVD validator. Re-run `validate_palette.js` before changing any hex value.
Race is hue; **sex is position, never a second hue**; observed-vs-modelled is
mark type, never a second hue.
- **`src/utils/kde.js`** — `BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`
were tuned for zero-inflated sparse-district posteriors. Both ends matter.
- **`src/utils/pooling.js`** — `POOL_BY_SEX_ARREST_THRESHOLD = 20`, applied to
the district total, not per cell.
- **`src/utils/densityProfile.js`** — `MASS_MAX_DISTINCT = 12`. Below it the
posterior predictive is drawn as discrete mass, because it *is* discrete; a
Gaussian KDE over four achievable values renders as a lumpy smear that reads
as a rendering bug.
- **`src/utils/rateDomain.js`** — `MAX_RATE_DOMAIN = 30` caps the axis so a
four-student cell can't squash every other curve. It reports `clipped` so the
chart says so instead of silently cropping.
## Common Pitfalls & Gotchas
### 1. Null Safety in Chart Components
### 1. Null safety in chart components
API responses can have empty arrays or missing fields for districts with no
arrests. Use optional chaining and explicit defaults, and never divide by a
denominator you haven't checked:
Several API responses may return empty arrays or missing fields for districts with no arrests:
```javascript
// Always use optional chaining and defaults:
const yearRow = (quadModels[q.key] || []).find(r => r.year === year)
const predMedian = yearRow?.count_median || 0
const enroll = row.stu_enroll || 1 // Prevent division by zero
const enroll = row.stu_enroll || 0
const rate = enroll > 0 ? (row.observed_arrests || 0) / enroll * 1000 : 0
```
### 2. D3 useEffect Dependency Arrays
A rate with no denominator is **not zero and not Infinity — it's undefined**.
`toRates` returns `[]`, `buildDisplayGroups` reports `0` and lets the
enrollment column explain why.
When using `useEffect` for D3 rendering, always include all data dependencies to prevent stale renders:
```javascript
// Correct — includes all props used inside the effect
}, [quadData, selectedModel, rateByGroup])
### 2. SVG dimensions and responsiveness
Charts set `width="100%"` with a fixed `viewBox`, wrapped in
`overflowX: 'auto'`. Chart cards are `max-width: 70rem` in `ChartPanel.jsx`
(vs. the ~60rem default text width).
### 3. API endpoint availability
- `/api/v1/estimates/{leaid}` — ✅ summary rows (median, bounds, enrollment,
observed arrests)
- `/api/v1/estimates?state=XX&...` — ✅ but returns rows `ORDER BY LEAID, RACE,
SEX` at 8 per district, capped at `limit=1000`. **Any short read ranks the
lowest-LEAID districts, not the busiest.** Page it with `meta.total`.
- `/api/v1/draws?...` — returns a shard URL + SQL, not draw data. The app goes
to the parquet directly.
### 4. CORS
The API sends no CORS headers. The app auto-detects a proxy via `VITE_PROXY_URL`;
unset, it fetches directly (works same-origin or behind the Docker/nginx proxy).
## The suggestion fixture
`public/data/top_districts.json` holds the top 15 districts per state by
observed arrests, and is generated by a one-off, read-only script:
```bash
node scripts/build-top-districts.mjs # all 51, ~150 requests
node scripts/build-top-districts.mjs --states NV,CA # spot-check
```
### 3. SVG Dimensions and Responsiveness
It is committed (the `national_rates.json` precedent). Re-run it only when a
new CRDC wave lands. `DistrictSearch` degrades to search-only if it's missing.
Charts use fixed dimensions with responsive containers (`overflowX: 'auto'` for wide content):
- Chart cards have `max-width: 70rem` in `ChartPanel.jsx` (vs the default text width of ~60rem)
- SVG elements should set both `width="100%"` and a fixed `viewBox` for proper scaling
## Testing
### 4. API Endpoint Availability
`npm test` runs `node --test 'src/**/*.test.js'`. The pure utilities are all
covered and **should be written test-first**:
Not all endpoints are available to browser-based clients:
- `/api/v1/estimates/{leaid}` — ✅ Returns estimates summary (median, lower, upper bounds)
- `/api/v1/draws?...` — Returns a Parquet shard URL + DuckDB SQL, not draw data itself. The app **does** use the real draws in the shard it points to — see below.
| Module | What its tests pin |
|---|---|
| `agrestiCoull.js` | five cases against real R output, incl. the negative lower bound |
| `pooling.js` | numerator and denominator always drawn from the same groups |
| `densityProfile.js` | the mass/KDE switch, and that KDE delegates to `kde.js` |
| `districtGroups.js` | display-row derivation, defaults, pooled vs unpooled keys |
| `groupDifference.js` | refusal to pair mismatched draw sets |
| `rateDomain.js` | the axis cap, and that it reports clipping |
| `drawGroups.js` | key shapes, and hole detection in draw arrays |
| `kde.js` | bandwidth clamp behaviour |
### 5. Real posterior draws via duckdb-wasm
Charts 2 (`RateByGroupBar`) and 3 (`RateDensityRidgeline`) fetch the actual
500-draw-per-group posterior from the public Hugging Face parquet dataset
(`civilytics/crdc-school-arrest-rates`), queried client-side with
`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js` +
`src/hooks/useDrawDistribution.js`). No server-side draws endpoint is
involved. If that fetch fails (network, unsupported browser, HF outage),
both charts fall back to the `distributionApprox.js` analytic approximation
and show the "estimated shape" note — **do not delete `distributionApprox.js`
or `ApproxNote.jsx`**, they're the fallback path, not dead code.
See `docs/superpowers/specs/2026-08-11-empirical-draws-wasm-design.md` for
the full design.
The duckdb-wasm engine itself is ~39MB uncompressed / ~8.86MB gzipped (confirmed against the shipped `@duckdb/duckdb-wasm` package, not the design spec's original ~3-5MB estimate, which was wrong). It's loaded via dynamic `import()` only once a district is selected — never on initial page load — and cached by the browser thereafter, but it's a real one-time cost worth knowing about before touching this code path.
### 6. CORS Configuration
The CRDC API does not send CORS headers. When deployed to git-pages (static hosting), requests are blocked by same-origin policy unless a proxy is configured:
- The app auto-detects proxy availability via `VITE_PROXY_URL` environment variable
- If unset, the app attempts direct fetch — works when served from Docker/nginx or same-origin
Components are verified manually — there is no DOM test harness. Useful
districts: **Clark County NV `3200060`** (100 arrests / 148,928 students,
pooling off), **Carson City NV `3200390`** (6 arrests, pooling auto-engages,
discrete mass profile), **Washoe County NV `3200480`** (cross-check the table
against the API), and any California district (large shards, fixture ranking).
## Deployment Checklist
Before pushing to production:
1. `npm run build` succeeds
2. `npm test` passes
3. No console errors after a hard refresh
4. Both charts render for a sample district; Network tab shows **one fetch per
(model, state) shard part** and no repeats when switching specs
5. Test with a **multi-part state** (California), not just Nevada — a
single-part state cannot catch a regression in part discovery
1. **Build succeeds**: `npm run build` (check for new errors)2. **No console errors in browser** after hard refresh3. **All 6 charts render** with sample districts (test "Denver", "Mobile County")
4. **Loading animation** appears briefly, then transitions to ChartPanel
### Commit Message Convention
Use descriptive commit messages that explain the *why*, not just the what:
```bash
# Good
git commit -m "Fix: use three-year model for wave data (returns observed counts across all years)"
git commit -m "Add D3 density ridges to Chart 5 with diamond markers for observed rates"
# Avoid vague messages
git commit -m "Fix charts" # Too generic
git commit -m "Update code" # No context
```
## Testing Strategy
There are no automated tests in this project. Manual verification is required:1. **Visual check**: Load a district and verify all 6 charts render correctly2. **Error console**: Check browser DevTools for JavaScript errors3. **Data accuracy**: Compare observed values against API response (check Network tab)
4. **Responsiveness**: Resize window to ensure layout adapts
## Style Guide References
- R code style: Follows tidyverse principles (`r-style-guide` skill in agent knowledge base)- Chart aesthetic decisions should match patterns from `social_media_posts.md` and `white_paper.qmd`
- Colors, typography, spacing are defined as CSS custom properties in `src/styles/tokens.css`
## Related Repositories
- **crdc-arrests** — The API server (Plumber/R) at `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/`
- **civilyticsR** — R package with wordmark and visualization functions at `/home/jared/Nextcloud/Civilytics/Code/Civilytics/civilyticsR/`
> Caveat on the cache: `App.jsx:112`'s "Search another district" button does
> `window.location.href = '/crdc-demo/'`, a full page reload, which discards the
> module-level `shardCache`. Within one district view — switching specs,
> toggling compare-all — the cache works as intended.
## Git Conventions
- Remote: `https://gitea.civilytics.org/Civilytics/crdc-demo.git`
- Default branch: `main` (not `master`)
- Gitea Actions workflow auto-deploys on push to `main` via `.gitea/workflows/pages.yml`
- Always pull before making changes: `git pull origin main`
- Default branch: `main`
- Gitea Actions auto-deploys on push to `main`
- Commit messages explain the *why*: `fix: rank suggestions over the full state
(limit=500 was selecting the lowest 62 LEAIDs)`
## Related Repositories
- **crdc-arrests** — API server (Plumber/R) and the white paper, at
`/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/`
- **civilyticsR** — wordmark and visualization functions
+30 -14
View File
@@ -18,18 +18,21 @@ npm run preview # serve built files locally
## What It Does
Visitors select a U.S. state, search for a school district (with suggestions of districts that have the most arrests), and see 3 charts comparing observed data against Bayesian model estimates:
Visitors select a U.S. state, search for a school district (with suggestions of the districts reporting the most arrests), and see what was actually reported alongside what the Bayesian models estimate. The results page ports the white paper's Figs 6 and 7 (`wp_fig_group_density` / `wp_fig_group_difference`):
1. **Arrests over time** (Chart 1) — raw counts by CRDC wave with per-1k rate labels, built with inline SVG
2. **Rate by student group** (Chart 2) — box-and-whisker per race×sex group for the most recent year, split into Female/Male panels. The box is the 25th–75th percentile of that group's real 500-draw posterior (same duckdb-wasm fetch as Chart 3), the whisker is the model's reported 90% interval, and a diamond marks the observed rate; individual boxes fall back to an analytic approximation when a group's draws can't be fetched.
3. **Predicted rates by student group** (Chart 3) — density ridges built from each group's real 500-draw posterior (fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset), with a model-selector dropdown for the four Bayesian specifications and diamond markers for observed rates; falls back to an analytic approximation if the draws can't be fetched.
1. **Reported arrests and enrollment** — a summary table, one row per student group plus a district total: students, observed arrests, and rate per 1,000. Its checkboxes double as the legend and the group control for the density panel. In sparse districts (fewer than 20 arrests district-wide) Female and Male are pooled within each race, with a banner explaining the rule and a switch to override it.
2. **Arrest rate probability density** — each selected group's posterior predictive distribution as a filled area, Female over Male sharing one axis, direct-labelled at the peak. Beneath each panel, a rail of 95% Agresti–Coull point ranges for the observed rate. A segmented control switches between the four Bayesian specifications; an opt-in toggle compares all four at once. Draws that take only a handful of distinct values are drawn as discrete probability mass rather than smoothed — in a small district the posterior predictive genuinely *is* discrete.
3. **Model estimated differences** — the posterior of Δ = rate(A) − rate(B) per 1,000, computed at each draw, filled with a diverging ramp centred at zero, with a dashed rule at no-difference and a plain-language `Pr(Δ > 0)` readout.
4. **Arrests over time** — observed counts by CRDC wave against the three-year model's median and 95% interval, inline SVG.
Distributions come from the published 500-draw posteriors, fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset, and fall back per group to an analytic approximation (with a visible note) when those draws can't be fetched.
## Architecture
### Tech Stack
- **React 19** + **Vite** (static site generation, no backend required)
- Plain CSS custom properties for styling (matches Civilytics design tokens exactly)
- Inline SVG rendering with hand-rolled scales — no charting library, no D3 dependency
- Inline SVG rendering that **React owns** — no charting library. `d3-scale`, `d3-shape`, `d3-array` and `d3-interpolate` supply scales, path generators and colour interpolation only; no d3 selections and no `useEffect` DOM mutation
- Embeds **DuckDB-Wasm** (`@duckdb/duckdb-wasm`, ~39MB uncompressed / ~8.8MB gzipped, loaded on demand only after a district is selected) to query real posterior draws client-side from a public Hugging Face Parquet dataset
- Calls the public read-only API directly from the browser
@@ -40,27 +43,38 @@ crdc-demo/
├── vite.config.mjs # Vite build config
├── public/ # Static assets (wordmark, favicon, fixtures)
│ ├── civilytics-wordmark.svg # Civilytics wordmark from civilyticsR package
│ └── data/national_rates.json # Static national rates fixture for comparisons
│ └── data/
│ ├── national_rates.json # National rates fixture for comparisons
│ └── top_districts.json # Top 15 districts per state by observed arrests
├── scripts/
│ └── build-top-districts.mjs # One-off generator for top_districts.json
├── src/
│ ├── main.jsx # React entry
│ ├── App.jsx # Main router (state → search → loading → charts)
│ ├── hooks/
│ │ ├── useApi.js # API client with retry/backoff + endpoint wrappers
│ │ └── useDrawDistribution.js # Fetches real posterior draws (duckdb-wasm + HF parquet)
│ │ └── useDrawDistribution.js # Posterior draw counts by draw_id (duckdb-wasm + HF parquet)
│ ├── components/
│ │ ├── StateSelector.jsx # Landing screen — state dropdown/grid
│ │ ├── DistrictSearch.jsx # Search + "interesting" district suggestions
│ │ ├── DistrictSearch.jsx # Search + suggestions from the committed fixture
│ │ ├── LoadingAnimation.jsx # Animated histogram grid during data fetch
│ │ ├── ChartLegend.jsx # Shared legend row
│ │ ├── ApproxNote.jsx # "shape estimated from interval bounds" caption
│ │ └── ChartPanel.jsx # Orchestrates all 3 charts + data fetching
│ │ ├── DistrictSummaryTable.jsx # Observed arrests table — also the chart's legend/control
│ │ └── ChartPanel.jsx # Owns cross-chart state + data fetching
│ ├── charts/
│ │ ├── ArrestsOverTime.jsx # Chart 1 — counts by wave (SVG)
│ │ ├── RateByGroupBar.jsx # Chart 2 — box-and-whisker by group (SVG)
│ │ └── RateDensityRidgeline.jsx # Chart 3 — density ridges per student group (SVG)
│ │ ├── ArrestsOverTime.jsx # Observed vs. modelled counts by wave (SVG)
│ │ ├── RateDensityPanel.jsx # Chart A — posterior density per group + AC rail
│ │ └── GroupDifference.jsx # Chart B — posterior of Δ between two groups
│ ├── utils/
│ │ ├── duckdbClient.js # Lazy duckdb-wasm bundle loader (dynamic import)
│ │ ├── kde.js # Empirical density from real draws (+ kde.test.js)
│ │ ├── kde.js # Empirical density from real draws (+ .test.js)
│ │ ├── densityProfile.js # Discrete-mass vs. KDE profile choice (+ .test.js)
│ │ ├── agrestiCoull.js # Frequentist interval, ported from R (+ .test.js)
│ │ ├── pooling.js # Sex pooling for sparse districts (+ .test.js)
│ │ ├── districtGroups.js # Display-row derivation and defaults (+ .test.js)
│ │ ├── groupDifference.js # Per-draw Δ and its summary (+ .test.js)
│ │ ├── rateDomain.js # Shared x-axis domain and clip flag (+ .test.js)
│ │ ├── drawGroups.js # Draw-map key format + coverage check (+ .test.js)
│ │ └── distributionApprox.js # Analytic fallback when draws are unavailable
│ └── styles/tokens.css # Civilytics design tokens (colors, fonts, spacing)
@@ -74,8 +88,10 @@ crdc-demo/
|---|---|---|
| `/api/v1/models` | List available Bayesian model specs | Once (cached) |
| `/api/v1/districts?q=&state=` | District name/geo lookup → LEAID | On keystroke |
| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | ~40 calls per district |
| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | 3 waves + 1 per selected spec |
| `/api/v1/estimates?state=XX&year=Y` | Not called at runtime — rows come back `ORDER BY LEAID` at 8 per district, so any short read ranks the lowest-LEAID districts. Paged with `meta.total` by `scripts/build-top-districts.mjs` | Build-time only |
| `/api/v1/draws?...` | Locate raw-posterior Parquet shard | Not called from app — the app fetches shards directly from Hugging Face via duckdb-wasm; see `src/hooks/useDrawDistribution.js` |
| `/data/top_districts.json` | Suggested districts per state (committed fixture) | Once per session |
| `/data/national_rates.json` | Static national rates fixture (committed) | Once per session |
## Deployment
@@ -0,0 +1,219 @@
# CRDC demo — rebuild the visuals around the white-paper figures
## Context
The demo app at `pages.civilytics.org/crdc-demo/` works, but the charts don't show off the
modelling. The navigation (state → district) is good and stays. The three existing charts get
cut to one, replaced by a summary table and two charts ported from the white paper:
`wp_fig_group_density` (Fig 6) and `wp_fig_group_difference` (Fig 7) in
`crdc-arrests/R/paper_figures.R:472-568`.
Two real defects surfaced while scoping:
1. **The "suggested districts" list is ranked over a truncated set.** `DistrictSearch.jsx:22`
calls `/estimates?state=XX&year=21-22&limit=500`, but that endpoint returns rows
`ORDER BY LEAID` (`crdc-arrests/api/R/handlers_estimates.R:46`) at 8 rows per district. So
the app ranks the ~62 lowest-LEAID districts in the state. California has 11,488 rows.
2. **The analytic fallback assumes 90% bounds but the API returns 95%.**
`distributionApprox.js` defaults `intervalMass = 0.90`; the API default is `interval=95`
(`crdc-arrests/api/R/handlers_estimates.R`). The new charts compute intervals from real
draws, so this only affects the fallback path — fix the default while we're in there.
### What the data supports (verified, not assumed)
- The HF parquet carries `LEAID, RACE, SEX, pred, draw_id, subgroup_id, batch_num`. `draw_id`
is dense 1–500 for every group. The current query at `useDrawDistribution.js:93` just
doesn't select it — **that one column is what unlocks both pooling and differences.**
- **Sex pooling is the project's own house method.** `build_state_summary()` in
`crdc-arrests/R/summarize_draws.R:152-182` pools across LEAs by summing `pred` and
`stu_enroll` within each draw, then summarizing across draws. Pooling M+F within a district
is the identical operation on a different axis. Honest effort estimate: **~3–4 hours**, most
of it UI and labelling, not statistics.
- **Caveat to word carefully:** `draw_id` is renumbered 1–500 per write batch
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in different batches,
so cross-group draw pairing is effectively independent. Measured correlation between
Black-male and Hispanic-male `pred` in Clark County was 0.019 even *within* a batch —
observation noise from `posterior_predict` dominates. Published Fig 7 has the same property,
so the app matches the paper. Captions should say "posterior predictive draws", never
"paired parameter draws".
- Enrollment covers only AM/BL/HI/WH (verified against the API). Clark County sums to 148,928
against a district enrollment near 304,000. The table must label this.
## Decisions taken
| Question | Decision |
|---|---|
| Model specs | One selected spec by default; opt-in "compare all four" expands to 4 ridge rows |
| Existing charts | Keep `ArrestsOverTime`; delete `RateByGroupBar` and `RateDensityRidgeline` |
| Pooling trigger | Whole-district: pool when total observed arrests across the 8 cells < 20 |
| Rendering | React owns the DOM; add d3 submodules for scales/paths/interpolation |
---
## Work
### 1. Draws pipeline — expose `draw_id`, return counts not rates
**`src/hooks/useDrawDistribution.js`** — the one structural change everything else rests on.
- Query becomes `SELECT RACE, SEX, draw_id, pred FROM read_parquet(...) WHERE LEAID = ?`.
- Return **raw counts indexed by draw**, not rates: `countsByGroup[key][draw_id - 1] = pred`.
Indexing by `draw_id` rather than push-order means row ordering from DuckDB is irrelevant.
- Accept `models: string[]` instead of a single `model`, returning `{status, byModel, nDraws}`.
A fixed-length array avoids conditional hooks when "compare all four" is on. The existing
module-level `shardCache` already keys on `(model, year, state)`, so four models is four
cache entries with no other change.
- Reject a group whose count array has holes (fewer entries than `nDraws`) — a partial group
must fall back, not silently render a short draw set.
- Keep the reset-before-guard ordering at `useDrawDistribution.js:79-85`; it exists to stop one
model's draws being shown under another model's label.
**`src/utils/pooling.js`** (new, + test) — pure functions, no React:
- `poolBySex(countsByGroup, enrollByGroup)` → sums counts within each draw index across
`SEX ∈ {F,M}` and sums enrollment, keyed by race alone.
- `toRates(counts, enroll)` → per-1,000 array.
- `shouldPoolBySex(rows)` → total `observed_arrests` across rows < `POOL_BY_SEX_ARREST_THRESHOLD`
(20, a named constant with the rationale in a comment).
**`src/utils/drawGroups.js`** — extend `groupKey` to handle a pooled key (race only) and update
`hasDrawsForAll` for the new count-array shape.
### 2. Frequentist interval
**`src/utils/agrestiCoull.js`** (new, + test) — direct port of
`crdc-arrests/R/paper_figures.R:219-237`, including the zero-numerator branch
(`ci_upper = -log(1 - level)`, the rule of three; `ci_lower = 0`). Note the R function returns
`c(upper, lower, sd, se, phat)` — **upper first**. Return a named object here instead. Needs a
`qnorm`/probit; `distributionApprox.js` already has one — reuse it rather than adding a second.
Tests should pin at least one case against R output (e.g. `agresti_coull(15, 499, 0.95)`).
### 3. Density profile — handle discrete posteriors honestly
**`src/utils/densityProfile.js`** (new, + test).
In sparse districts the posterior predictive is a discrete count distribution. Carson City NV
(`3200390`) has a 53-student AI/AN female cell where one arrest is 18.9 per 1,000 — the draws
take four distinct values and a Gaussian KDE renders them as a lumpy smear that reads as a
rendering bug.
- `densityProfile(counts, enroll, domain)` returns `{kind: 'kde'|'mass', points}`.
- `kind: 'mass'` when the draws take ≤ 12 distinct values: probability mass at each achievable
rate, drawn as a filled staircase so it visually rhymes with the smooth areas beside it.
- Otherwise delegate to the existing `kdeCurve` in `src/utils/kde.js` — its bandwidth clamp
(`BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`) was tuned for exactly these
zero-inflated posteriors and should not be touched.
### 4. Summary table (top of results)
**`src/components/DistrictSummaryTable.jsx`** (new). One row per student group plus a total:
| Student group | Students | Observed arrests | Rate per 1,000 |
- Sorted by observed arrests descending. Zero-arrest rows de-emphasized, not hidden.
- Each row carries the checkbox that drives chart A — the table *is* the legend and the control.
Default checked = `observed_arrests > 0`; if no group has any, check the two largest by
enrollment and say so.
- Footnote: students counted are those in the four modeled race groups (AI/AN, Black, Hispanic,
White), not total district enrollment.
- When pooling is active, rows collapse to four races and a banner states the rule in one
sentence, with a switch to force it off.
### 5. Chart A — "Arrest rate probability density"
**`src/charts/RateDensityPanel.jsx`** (new). Replaces `RateDensityRidgeline.jsx`.
- Two stacked sub-panels, Female over Male, sharing one x-axis (per 1,000). Collapses to a
single panel when pooled. This preserves the palette contract documented at
`src/utils/colors.js:9-13`: race is hue, sex is position — never a second hue.
- Within a sub-panel, selected groups overlap as filled areas (fill ~0.4 opacity, 2px stroke in
the race color), direct-labelled at each peak so there's no legend hunting.
- Below each sub-panel's baseline, a thin rail stacks one Agresti–Coull point-range per selected
group in the matching color — the R figure's `position_nudge` idea, but un-overplotted.
- Segmented control for the four unified quadrant specs. A "Compare all four specifications"
switch expands to four ridge rows (matching Fig 1's structure) and triggers four shard
fetches — cheap for NV (~100KB each), ~6.3MB each for CA, so it stays opt-in with a spinner.
- x-domain: max of the density supports and the frequentist upper bounds, with the existing cap
logic from `RateDensityRidgeline.jsx:101` and `niceTicks`.
- Caption states 500 posterior predictive draws and a 95% Agresti–Coull observed interval.
### 6. Chart B — "Model Estimated Differences"
**`src/charts/GroupDifference.jsx`** (new).
- Two group pickers; defaults are the two groups with the most observed arrests (pooled groups
when pooling is on). Δ = rate(A) − rate(B) per 1,000, computed per draw index.
- Single density, filled with an SVG `linearGradient` mapped across x. Use a **diverging ramp
centered at zero** — navy for Δ<0, paper at 0, ember for Δ>0 — rather than the paper's YlOrRd:
the quantity is signed, and diverging-at-zero is the honest encoding. On-brand via
`tokens.css`.
- Dashed vertical rule at 0 in `--cv-danger`, matching the paper's red line.
- Large readout `Pr(Δ > 0)` with a plain-language sentence beneath ("In 94.4% of posterior
draws, the Black male arrest rate exceeds the White male rate"), plus median Δ and an 80%/95%
interval as a point-range.
- Degrade explicitly when fewer than two groups have usable draws — say why, don't render empty.
### 7. Fix the district suggestions
**`scripts/build-top-districts.mjs`** (new) — pages `/estimates?state=XX&year=21-22&limit=1000`
using `meta.total` (confirmed present in the envelope) across all 51 states, aggregates observed
arrests per LEAID, and writes `public/data/top_districts.json` with the top 15 per state
(leaid, name, arrests, enrollment, rate). Roughly 140 requests as a one-off; the output is
~50KB and gets committed, following the `public/data/national_rates.json` precedent.
**`src/components/DistrictSearch.jsx`** — read the fixture instead of calling
`fetchStateDistricts` at runtime. The search screen loses a multi-second fetch and the ranking
becomes correct. Keep live name search on `/districts` unchanged. Drop the hardcoded
"Try Derby (KS), Paterson (NJ)…" hint at `DistrictSearch.jsx:164` — the real list supersedes it.
### 8. Wiring, deletions, docs
- **`src/components/ChartPanel.jsx`** — owns pooling state, selected groups, selected spec, and
the difference pair; passes them down. Keep `ArrestsOverTime`. Drop `QUADRANT_MODELS`
prefetch of all four models' *summaries* if only the selected one is needed.
- **Delete**: `src/charts/RateByGroupBar.jsx`, `src/charts/RateDensityRidgeline.jsx`.
- **Keep**: `distributionApprox.js` and `ApproxNote.jsx` — still the fallback when draws can't
be fetched (`AGENTS.md:91-101`). Fix its `intervalMass` default to 0.95 to match the API.
- **`package.json`** — add `d3-scale`, `d3-shape`, `d3-array`, `d3-interpolate` as real
`dependencies` (the existing deps are all miscategorised under `devDependencies`; leave that
alone unless it's breaking the build).
- **Docs**: `AGENTS.md` still describes 6 charts, D3 selections, and `DistrictVsNational` /
`ModelDrawsComparison` / `ExceedanceProbability` — none of which exist. Rewrite the
architecture, data-flow, and interval sections. Update `README.md`'s chart list.
---
## Verification
1. `npm run build` clean; `npm test` (`node --test 'src/**/*.test.js'`) passes, including new
tests for `agrestiCoull`, `pooling`, `densityProfile`, and the extended `drawGroups`.
2. `npm run dev`, then walk these districts:
- **Clark County NV (`3200060`, 100 arrests / 148,928 students)** — pooling stays off, all
four races render, differences chart defaults to the top two groups.
- **Carson City NV (`3200390`, 6 arrests / 4,073 students)** — pooling auto-engages (6 < 20),
banner appears, table collapses to four races. The AI/AN cell should render as a discrete
mass profile, not a smear. Verified against the draws: pooling narrows AI/AN's 90% interval
from 37.7 to 27.0 per 1,000, and Hispanic male's from 4.9 to 2.5.
- **Washoe County NV (`3200480`)** — cross-check the summary table's observed counts and
rates against `/api/v1/estimates/3200480?model=unified_m4_mod&year=21-22`.
- **A California district** — confirm the suggestion fixture ranks correctly (this is the
case the current code gets wrong), and that "compare all four" warns/spins before pulling
~25MB of shards.
3. Toggle every group off, then on; switch specs; flip pooling manually — no stale draws from a
previous model should ever appear under a new label.
4. Compare chart A against `wp_fig_group_density` output for Clark County: same curve shapes,
same point-range positions.
5. Browser console clean on hard refresh; check the Network tab shows one shard fetch per
(model, state) and no repeats when navigating between districts.
## Effort
Roughly **2–3 focused days** end to end: ~1 day for the draws/pooling/util layer with tests,
~1 day for the two charts, ~half a day for the table, the suggestion fixture, and docs. At ~10
hours a week that's about two calendar weeks.
The R Shiny alternative would be slower, not faster — it trades a zero-server static site for a
container, an R runtime, and server-side access to either the 91GB draws DuckDB or the 51-state
parquet tree, and turns every toggle into a round-trip re-render. The React app already fetches
real draws client-side and already carries the design tokens.
+118
View File
@@ -8,6 +8,12 @@
"name": "crdc-arrests-demo",
"version": "0.1.0",
"license": "MIT",
"dependencies": {
"d3-array": "^3.2.4",
"d3-interpolate": "^3.0.1",
"d3-scale": "^4.0.2",
"d3-shape": "^3.2.0"
},
"devDependencies": {
"@duckdb/duckdb-wasm": "^1.32.0",
"@vitejs/plugin-react-swc": "^4.3.3",
@@ -1001,6 +1007,109 @@
"node": ">= 8"
}
},
"node_modules/d3-array": {
"version": "3.2.4",
"resolved": "https://registry.npmjs.org/d3-array/-/d3-array-3.2.4.tgz",
"integrity": "sha512-tdQAmyA18i4J7wprpYq8ClcxZy3SC31QMeByyCFyRt7BVHdREQZ5lpzoe5mFEYZUWe+oq8HBvk9JjpibyEV4Jg==",
"license": "ISC",
"dependencies": {
"internmap": "1 - 2"
},
"engines": {
"node": ">=12"
}
},
"node_modules/d3-color": {
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/d3-color/-/d3-color-3.1.0.tgz",
"integrity": "sha512-zg/chbXyeBtMQ1LbD/WSoW2DpC3I0mpmPdW+ynRTj/x2DAWYrIY7qeZIHidozwV24m4iavr15lNwIwLxRmOxhA==",
"license": "ISC",
"engines": {
"node": ">=12"
}
},
"node_modules/d3-format": {
"version": "3.1.2",
"resolved": "https://registry.npmjs.org/d3-format/-/d3-format-3.1.2.tgz",
"integrity": "sha512-AJDdYOdnyRDV5b6ArilzCPPwc1ejkHcoyFarqlPqT7zRYjhavcT3uSrqcMvsgh2CgoPbK3RCwyHaVyxYcP2Arg==",
"license": "ISC",
"engines": {
"node": ">=12"
}
},
"node_modules/d3-interpolate": {
"version": "3.0.1",
"resolved": "https://registry.npmjs.org/d3-interpolate/-/d3-interpolate-3.0.1.tgz",
"integrity": "sha512-3bYs1rOD33uo8aqJfKP3JWPAibgw8Zm2+L9vBKEHJ2Rg+viTR7o5Mmv5mZcieN+FRYaAOWX5SJATX6k1PWz72g==",
"license": "ISC",
"dependencies": {
"d3-color": "1 - 3"
},
"engines": {
"node": ">=12"
}
},
"node_modules/d3-path": {
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/d3-path/-/d3-path-3.1.0.tgz",
"integrity": "sha512-p3KP5HCf/bvjBSSKuXid6Zqijx7wIfNW+J/maPs+iwR35at5JCbLUT0LzF1cnjbCHWhqzQTIN2Jpe8pRebIEFQ==",
"license": "ISC",
"engines": {
"node": ">=12"
}
},
"node_modules/d3-scale": {
"version": "4.0.2",
"resolved": "https://registry.npmjs.org/d3-scale/-/d3-scale-4.0.2.tgz",
"integrity": "sha512-GZW464g1SH7ag3Y7hXjf8RoUuAFIqklOAq3MRl4OaWabTFJY9PN/E1YklhXLh+OQ3fM9yS2nOkCoS+WLZ6kvxQ==",
"license": "ISC",
"dependencies": {
"d3-array": "2.10.0 - 3",
"d3-format": "1 - 3",
"d3-interpolate": "1.2.0 - 3",
"d3-time": "2.1.1 - 3",
"d3-time-format": "2 - 4"
},
"engines": {
"node": ">=12"
}
},
"node_modules/d3-shape": {
"version": "3.2.0",
"resolved": "https://registry.npmjs.org/d3-shape/-/d3-shape-3.2.0.tgz",
"integrity": "sha512-SaLBuwGm3MOViRq2ABk3eLoxwZELpH6zhl3FbAoJ7Vm1gofKx6El1Ib5z23NUEhF9AsGl7y+dzLe5Cw2AArGTA==",
"license": "ISC",
"dependencies": {
"d3-path": "^3.1.0"
},
"engines": {
"node": ">=12"
}
},
"node_modules/d3-time": {
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/d3-time/-/d3-time-3.1.0.tgz",
"integrity": "sha512-VqKjzBLejbSMT4IgbmVgDjpkYrNWUYJnbCGo874u7MMKIWsILRX+OpX/gTk8MqjpT1A/c6HY2dCA77ZN0lkQ2Q==",
"license": "ISC",
"dependencies": {
"d3-array": "2 - 3"
},
"engines": {
"node": ">=12"
}
},
"node_modules/d3-time-format": {
"version": "4.1.0",
"resolved": "https://registry.npmjs.org/d3-time-format/-/d3-time-format-4.1.0.tgz",
"integrity": "sha512-dJxPBlzC7NugB2PDLwo9Q8JiTR3M3e4/XANkreKSUxF8vvXKqm1Yfq4Q5dl8budlunRVlUUaDUgFt7eA8D6NLg==",
"license": "ISC",
"dependencies": {
"d3-time": "1 - 3"
},
"engines": {
"node": ">=12"
}
},
"node_modules/debug": {
"version": "4.4.3",
"resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz",
@@ -1480,6 +1589,15 @@
"dev": true,
"license": "ISC"
},
"node_modules/internmap": {
"version": "2.0.3",
"resolved": "https://registry.npmjs.org/internmap/-/internmap-2.0.3.tgz",
"integrity": "sha512-5Hh7Y1wQbvY5ooGgPbDaL5iYLAPzMTUrjMulskHLH6wnv/A+1q5rgEaiuqEjB+oxGXIVZs1FF+R/KPN3ZSQYYg==",
"license": "ISC",
"engines": {
"node": ">=12"
}
},
"node_modules/is-extglob": {
"version": "2.1.1",
"resolved": "https://registry.npmjs.org/is-extglob/-/is-extglob-2.1.1.tgz",
+7 -1
View File
@@ -27,5 +27,11 @@
"react-dom": "^19.2.8",
"vite": "^8.2.1"
},
"type": "module"
"type": "module",
"dependencies": {
"d3-array": "^3.2.4",
"d3-interpolate": "^3.0.1",
"d3-scale": "^4.0.2",
"d3-shape": "^3.2.0"
}
}
File diff suppressed because it is too large Load Diff
+204
View File
@@ -0,0 +1,204 @@
#!/usr/bin/env node
/**
* Builds `public/data/top_districts.json` — the "suggested districts" list the
* search screen shows before you type anything.
*
* Why this exists: the app used to build that list at runtime from
* `/estimates?state=XX&year=21-22&limit=500`. That endpoint returns rows
* `ORDER BY LEAID, RACE, SEX` at eight rows per district, so a 500-row cap is
* the ~62 *lowest-LEAID* districts in the state, not the busiest ones —
* California alone has 11,488 rows. The list was therefore ranked over a
* truncated and essentially arbitrary slice of each state.
*
* This script pages the whole state using `meta.total` from the response
* envelope, aggregates observed arrests per district, and commits the answer as
* a fixture (the `public/data/national_rates.json` precedent). The search screen
* then loses a multi-second fetch and gets a correct ranking.
*
* Read-only against the public API. Roughly 150 requests as a one-off; re-run it
* only when a new CRDC wave lands.
*
* node scripts/build-top-districts.mjs
* node scripts/build-top-districts.mjs --states NV,CA # spot-check a few
*/
import { writeFile, mkdir } from 'node:fs/promises'
import { dirname, resolve } from 'node:path'
import { fileURLToPath } from 'node:url'
const BASE_URL = process.env.CRDC_API_BASE || 'https://crdc-api.civilytics.org/api/v1'
const YEAR = '21-22'
// Pinned rather than left to the API default so a change to that default can't
// silently alter the fixture. Enrollment and observed arrests are the same in
// every specification; only the modelled columns differ, and we read none.
const MODEL = 'unified_m2_mod'
const PAGE_SIZE = 1000 // the API's LIMIT_CAP
const TOP_N = 15
const CONCURRENCY = 3
const MAX_RETRIES = 4
const ALL_STATES = [
'AL', 'AK', 'AZ', 'AR', 'CA', 'CO', 'CT', 'DE', 'DC', 'FL', 'GA', 'HI',
'ID', 'IL', 'IN', 'IA', 'KS', 'KY', 'LA', 'ME', 'MD', 'MA', 'MI', 'MN',
'MS', 'MO', 'MT', 'NE', 'NV', 'NH', 'NJ', 'NM', 'NY', 'NC', 'ND', 'OH',
'OK', 'OR', 'PA', 'RI', 'SC', 'SD', 'TN', 'TX', 'UT', 'VT', 'VA', 'WA',
'WV', 'WI', 'WY',
]
const OUT_PATH = resolve(
dirname(fileURLToPath(import.meta.url)),
'..',
'public',
'data',
'top_districts.json',
)
function parseStates() {
const flag = process.argv.indexOf('--states')
if (flag === -1) return ALL_STATES
const requested = (process.argv[flag + 1] || '').split(',').map((s) => s.trim().toUpperCase())
const unknown = requested.filter((s) => !ALL_STATES.includes(s))
if (unknown.length) throw new Error(`Unknown state code(s): ${unknown.join(', ')}`)
return requested
}
const sleep = (ms) => new Promise((r) => setTimeout(r, ms))
/** GET one page, returning the full envelope (we need `meta.total`). */
async function fetchPage(state, page) {
const params = new URLSearchParams({
state,
year: YEAR,
model: MODEL,
limit: String(PAGE_SIZE),
page: String(page),
})
const url = `${BASE_URL}/estimates?${params}`
let lastError
for (let attempt = 0; attempt <= MAX_RETRIES; attempt++) {
try {
const res = await fetch(url, { signal: AbortSignal.timeout(60000) })
if (!res.ok) throw new Error(`HTTP ${res.status} ${res.statusText}`)
const envelope = await res.json()
if (envelope.status !== 'success') throw new Error(envelope.error || 'Unknown API error')
if (!envelope.meta || typeof envelope.meta.total !== 'number') {
throw new Error('Response envelope is missing meta.total — cannot page safely')
}
return envelope
} catch (err) {
lastError = err
if (attempt === MAX_RETRIES) break
await sleep(500 * 2 ** attempt)
}
}
throw new Error(`${state} page ${page}: ${lastError.message}`)
}
async function collectState(state) {
const first = await fetchPage(state, 0)
const total = first.meta.total
const rows = [...first.data]
const pages = Math.ceil(total / PAGE_SIZE)
for (let page = 1; page < pages; page++) {
const envelope = await fetchPage(state, page)
rows.push(...envelope.data)
}
if (rows.length !== total) {
// Loud rather than silent: a short read here would quietly produce a
// truncated ranking, which is the exact bug this script exists to fix.
throw new Error(`${state}: expected ${total} rows, collected ${rows.length}`)
}
const byLeaid = new Map()
for (const row of rows) {
const leaid = row.leaid
if (!leaid) continue
const prev = byLeaid.get(leaid) || { leaid, name: row.lea_name || leaid, arrests: 0, enrollment: 0 }
byLeaid.set(leaid, {
...prev,
name: prev.name || row.lea_name || leaid,
arrests: prev.arrests + (row.observed_arrests || 0),
enrollment: prev.enrollment + (row.stu_enroll || 0),
})
}
const ranked = [...byLeaid.values()]
.filter((d) => d.arrests > 0)
.sort((a, b) => b.arrests - a.arrests || a.leaid.localeCompare(b.leaid))
.slice(0, TOP_N)
.map((d) => ({
leaid: d.leaid,
name: d.name,
arrests: d.arrests,
enrollment: d.enrollment,
rate: d.enrollment > 0 ? Math.round((d.arrests / d.enrollment) * 1000 * 100) / 100 : 0,
}))
return { state, districts: ranked, districtsSeen: byLeaid.size, rows: total, pages }
}
/** Small fixed-size worker pool — polite to a single public API host. */
async function mapWithConcurrency(items, limit, worker) {
const results = new Array(items.length)
let next = 0
const runners = Array.from({ length: Math.min(limit, items.length) }, async () => {
while (next < items.length) {
const i = next++
results[i] = await worker(items[i], i)
}
})
await Promise.all(runners)
return results
}
async function main() {
const states = parseStates()
console.log(`Fetching ${states.length} state(s) from ${BASE_URL} (year ${YEAR}, model ${MODEL})…`)
let done = 0
let requests = 0
const collected = await mapWithConcurrency(states, CONCURRENCY, async (state) => {
const result = await collectState(state)
requests += result.pages
done += 1
console.log(
` [${String(done).padStart(2)}/${states.length}] ${state}: ` +
`${result.rows} rows / ${result.pages} page(s), ` +
`${result.districtsSeen} districts, top ${result.districts.length} kept`,
)
return result
})
const byState = {}
for (const { state, districts } of collected.sort((a, b) => a.state.localeCompare(b.state))) {
byState[state] = districts
}
const payload = {
metadata: {
source: 'CRDC School Arrest Rate API (Knowles & Miller 2025)',
endpoint: `${BASE_URL}/estimates`,
year: YEAR,
model: MODEL,
description:
`Top ${TOP_N} school districts per state by total observed arrests in ${YEAR}, ` +
'summed across the eight modelled race×sex groups. Generated by ' +
'scripts/build-top-districts.mjs over the complete paged result set for each state.',
generated_states: states.length,
generated_requests: requests,
},
states: byState,
}
await mkdir(dirname(OUT_PATH), { recursive: true })
await writeFile(OUT_PATH, `${JSON.stringify(payload, null, 2)}\n`, 'utf8')
console.log(`\nWrote ${OUT_PATH} (${requests} requests, ${states.length} states).`)
}
main().catch((err) => {
console.error('\nbuild-top-districts failed:', err.message)
process.exit(1)
})
+3 -1
View File
@@ -129,7 +129,9 @@ export default function ArrestsOverTime({ data, modelId }) {
<ChartLegend items={[
{ shape: 'diamond', color: OBSERVED_MARK_COLOR, label: 'Observed' },
{ shape: 'dot', color: MODELED_AGGREGATE_COLOR, label: 'Modeled (median + 90% interval)' },
// 95%, not 90%: these bars are the API's count_lower/count_upper, and
// validate_interval() defaults to 95 (the app never passes interval=).
{ shape: 'dot', color: MODELED_AGGREGATE_COLOR, label: 'Modeled (median + 95% interval)' },
]} />
<p style={{ fontSize: '0.75rem', color: 'var(--cv-ink-3)', marginTop: 'var(--space-1)' }}>
+394
View File
@@ -0,0 +1,394 @@
import { useMemo } from 'react'
import { area, curveMonotoneX, curveStep } from 'd3-shape'
import { scaleLinear } from 'd3-scale'
import { interpolateRgb } from 'd3-interpolate'
import { rateProfile } from '../utils/densityProfile.js'
import { differenceRates, differenceSummary } from '../utils/groupDifference.js'
import { displayDraws } from '../utils/districtGroups.js'
import { niceTicks } from '../utils/niceTicks.js'
/**
* Chart B — "Model estimated differences".
*
* Port of the white paper's Fig 7 (`wp_fig_group_difference`): the posterior
* distribution of Δ = rate(A) − rate(B), per 1,000 students, computed at each
* draw index.
*
* Two deliberate departures from the R figure:
*
* - The fill is a **diverging** ramp centred at zero (navy below, paper at
* zero, ember above) rather than the paper's sequential YlOrRd. Δ is a signed
* quantity; a sequential ramp encodes "more" where the data means "which
* direction", and would make a large negative difference read as a small one.
* - The readout is spelled out in a sentence. `Pr(Δ > 0) = 94.4%` is the number
* a reader is most likely to misread as "94.4% more arrests".
*/
const NEGATIVE_COLOR = '#22406A' // --cv-navy-600
const ZERO_COLOR = '#F2EDE4' // --cv-paper-2
const POSITIVE_COLOR = '#C25311' // --cv-accent
const GRADIENT_STOPS = 24
const WIDTH = 760
const HEIGHT = 250
const MARGIN = { top: 18, right: 22, bottom: 46, left: 22 }
export default function GroupDifference({
groups,
enrollByGroup,
byModel,
selectedModel,
status,
pooled,
pair,
onPairChange,
}) {
const { counts, enroll } = useMemo(
() => displayDraws(byModel?.[selectedModel]?.counts, enrollByGroup, pooled),
[byModel, selectedModel, enrollByGroup, pooled],
)
// Only groups with a usable draw set can be differenced at all — offering the
// others in the picker would produce an empty chart with no explanation.
const comparable = useMemo(
() => groups.filter((g) => counts?.[g.key]?.length > 0 && enroll?.[g.key] > 0),
[groups, counts, enroll],
)
const [keyA, keyB] = pair || []
const groupA = comparable.find((g) => g.key === keyA)
const groupB = comparable.find((g) => g.key === keyB)
const deltas = useMemo(
() =>
groupA && groupB
? differenceRates(counts[groupA.key], enroll[groupA.key], counts[groupB.key], enroll[groupB.key])
: [],
[groupA, groupB, counts, enroll],
)
const summary = useMemo(() => differenceSummary(deltas), [deltas])
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<div style={headerStyle}>
<div>
<h3 style={cardTitle}>Model estimated differences</h3>
<p style={subtitleStyle}>
How much higher is one group&rsquo;s modelled arrest rate than another&rsquo;s, and how
sure is the model of the direction?
</p>
</div>
{comparable.length >= 2 && (
<PairPickers
comparable={comparable}
keyA={keyA}
keyB={keyB}
onPairChange={onPairChange}
/>
)}
</div>
{status === 'loading' ? (
<p style={emptyStyle}>Loading posterior draws…</p>
) : comparable.length < 2 ? (
<Degraded
reason={
comparable.length === 1
? `Only ${comparable[0].label} has a usable set of posterior draws in this district, so there is no second group to compare it against.`
: 'No student group in this district has a usable set of posterior draws, so no difference can be computed. This is usually because the district is absent from the published draw shard for this model.'
}
/>
) : !groupA || !groupB || groupA.key === groupB.key ? (
<Degraded reason="Pick two different student groups to compare." />
) : !summary ? (
<Degraded
reason={`${groupA.label} and ${groupB.label} do not have matching draw sets in this model specification, so their difference cannot be computed draw by draw.`}
/>
) : (
<>
<Readout summary={summary} groupA={groupA} groupB={groupB} />
<DifferencePlot deltas={deltas} summary={summary} groupA={groupA} groupB={groupB} />
<p style={captionStyle}>
Δ is computed at each of {summary.n.toLocaleString()} posterior predictive draws as{' '}
{groupA.label} minus {groupB.label}, per 1,000 students. The dashed line marks no
difference. Bars beneath the curve are the 80% and 95% intervals around the median Δ.
</p>
</>
)}
</div>
)
}
function Readout({ summary, groupA, groupB }) {
const pct = summary.prGreater * 100
const higher = summary.median >= 0 ? groupA : groupB
const lower = summary.median >= 0 ? groupB : groupA
const share = summary.median >= 0 ? pct : 100 - pct
return (
<div style={readoutStyle}>
<div>
<span style={readoutNumber}>{formatPercent(pct)}</span>
<span style={readoutLabel}>Pr(Δ &gt; 0)</span>
</div>
<p style={{ margin: 0, fontSize: '0.88rem', maxWidth: '38rem' }}>
In {formatPercent(share)} of posterior predictive draws, the {higher.sentenceLabel} arrest
rate exceeds the {lower.sentenceLabel} rate. The median difference is{' '}
<strong>{formatDelta(summary.median)}</strong> per 1,000 students.
</p>
</div>
)
}
function DifferencePlot({ deltas, summary, groupA, groupB }) {
const innerWidth = WIDTH - MARGIN.left - MARGIN.right
const baselineY = HEIGHT - MARGIN.bottom
const { ticks, min, max } = useMemo(() => symmetricDomain(summary), [summary])
const x = scaleLinear().domain([min, max]).range([MARGIN.left, MARGIN.left + innerWidth])
const profile = useMemo(() => rateProfile(deltas, { min, max, n: 80 }), [deltas, min, max])
if (!profile) return null
const y = scaleLinear().domain([0, profile.maxY || 1]).range([baselineY, MARGIN.top])
const points =
profile.kind === 'mass' ? padMassPoints(profile.points, profile.step, min, max) : profile.points
const areaGen = area()
.x((p) => x(clamp(p.x, min, max)))
.y0(baselineY)
.y1((p) => y(p.y))
.curve(profile.kind === 'mass' ? curveStep : curveMonotoneX)
const gradientId = `delta-gradient-${groupA.key}-${groupB.key}`
return (
<div style={{ overflowX: 'auto' }}>
<svg
width="100%"
viewBox={`0 0 ${WIDTH} ${HEIGHT}`}
style={{ maxWidth: '100%', minWidth: '320px' }}
role="img"
aria-label={`Posterior distribution of the difference in arrest rate between ${groupA.label} and ${groupB.label}`}
>
<defs>
<linearGradient id={gradientId} x1="0" y1="0" x2="1" y2="0">
{divergingStops(min, max).map((s) => (
<stop key={s.offset} offset={`${s.offset * 100}%`} stopColor={s.color} />
))}
</linearGradient>
</defs>
{ticks.map((t) => (
<line key={t} x1={x(t)} y1={MARGIN.top} x2={x(t)} y2={baselineY} stroke="var(--cv-rule)" strokeWidth={1} />
))}
<path d={areaGen(points)} fill={`url(#${gradientId})`} opacity={0.85} />
<path
d={areaGen.lineY1()(points)}
fill="none"
stroke="var(--cv-ink-2)"
strokeWidth={1.5}
strokeLinejoin="round"
/>
{/* No difference — the paper's red vertical line. */}
<line
x1={x(0)}
y1={MARGIN.top - 4}
x2={x(0)}
y2={baselineY + 6}
stroke="var(--cv-danger)"
strokeWidth={2}
strokeDasharray="5 4"
/>
<text x={x(0)} y={MARGIN.top - 7} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-danger)">
no difference
</text>
<line x1={MARGIN.left} y1={baselineY} x2={MARGIN.left + innerWidth} y2={baselineY} stroke="var(--cv-rule-strong)" strokeWidth={1} />
{/* Median with 80% (thick) and 95% (thin) intervals. */}
<g>
<line x1={x(clamp(summary.lower95, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper95, min, max))} y2={baselineY + 13} stroke="var(--cv-ink-2)" strokeWidth={1.5} strokeLinecap="round" />
<line x1={x(clamp(summary.lower80, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper80, min, max))} y2={baselineY + 13} stroke="var(--cv-ink)" strokeWidth={4} strokeLinecap="round" />
<circle cx={x(clamp(summary.median, min, max))} cy={baselineY + 13} r={3.5} fill="var(--cv-paper)" stroke="var(--cv-ink)" strokeWidth={2} />
</g>
{ticks.map((t) => (
<text key={t} x={x(t)} y={baselineY + 32} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-ink-3)">
{formatTick(t)}
</text>
))}
<text x={MARGIN.left} y={HEIGHT - 4} textAnchor="start" fontSize="0.62rem" fill="var(--cv-ink-3)">
← {groupB.shortLabel} higher
</text>
<text x={MARGIN.left + innerWidth} y={HEIGHT - 4} textAnchor="end" fontSize="0.62rem" fill="var(--cv-ink-3)">
{groupA.shortLabel} higher →
</text>
</svg>
</div>
)
}
/**
* A domain centred on zero. A signed quantity drawn on an off-centre axis makes
* the eye read the *position* of the curve as the size of the difference, so
* zero sits in the middle even when every draw falls on one side of it.
*/
function symmetricDomain(summary) {
const extent = Math.max(
Math.abs(summary.lower95),
Math.abs(summary.upper95),
Math.abs(summary.median),
1e-6,
)
const { niceMax } = niceTicks(extent * 1.25, 4)
const step = niceMax / 4
const ticks = []
for (let i = -4; i <= 4; i++) ticks.push(Math.round(step * i * 1e6) / 1e6)
return { ticks, min: -niceMax, max: niceMax }
}
/** Diverging ramp, with the paper-coloured midpoint pinned to Δ = 0. */
function divergingStops(min, max) {
const zeroOffset = (0 - min) / (max - min)
const toNegative = interpolateRgb(NEGATIVE_COLOR, ZERO_COLOR)
const toPositive = interpolateRgb(ZERO_COLOR, POSITIVE_COLOR)
const stops = []
for (let i = 0; i <= GRADIENT_STOPS; i++) {
const offset = i / GRADIENT_STOPS
const color =
offset <= zeroOffset
? toNegative(zeroOffset > 0 ? offset / zeroOffset : 1)
: toPositive(zeroOffset < 1 ? (offset - zeroOffset) / (1 - zeroOffset) : 0)
stops.push({ offset, color })
}
return stops
}
function padMassPoints(points, step, min, max) {
const half = Math.max(step, 1e-6) / 2
return [
{ x: Math.max(points[0].x - half, min), y: 0 },
...points,
{ x: Math.min(points[points.length - 1].x + half, max), y: 0 },
]
}
function PairPickers({ comparable, keyA, keyB, onPairChange }) {
return (
<div style={{ display: 'flex', alignItems: 'center', gap: '0.4rem', flexWrap: 'wrap' }}>
<GroupSelect
label="Compare"
value={keyA}
options={comparable}
onChange={(v) => onPairChange([v, keyB])}
/>
<span style={{ fontSize: '0.8rem', color: 'var(--cv-ink-3)' }}>against</span>
<GroupSelect
label="against"
value={keyB}
options={comparable}
onChange={(v) => onPairChange([keyA, v])}
/>
</div>
)
}
function GroupSelect({ label, value, options, onChange }) {
return (
<select
aria-label={label}
value={value || ''}
onChange={(e) => onChange(e.target.value)}
style={{ padding: '0.25rem 0.4rem', fontFamily: 'var(--font-sans)', fontSize: '0.78rem' }}
>
{options.map((g) => (
<option key={g.key} value={g.key}>
{g.label}
</option>
))}
</select>
)
}
function Degraded({ reason }) {
return (
<div style={degradedStyle}>
<p style={{ margin: 0, fontSize: '0.85rem' }}>{reason}</p>
</div>
)
}
function clamp(v, min, max) {
return Math.min(Math.max(v, min), max)
}
function formatPercent(pct) {
return `${pct.toLocaleString(undefined, { minimumFractionDigits: 1, maximumFractionDigits: 1 })}%`
}
function formatDelta(v) {
const sign = v > 0 ? '+' : ''
return `${sign}${v.toLocaleString(undefined, { minimumFractionDigits: 2, maximumFractionDigits: 2 })}`
}
function formatTick(v) {
if (v === 0) return '0'
const abs = Math.abs(v)
return v.toLocaleString(undefined, {
minimumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
maximumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
})
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const headerStyle = {
display: 'flex',
justifyContent: 'space-between',
alignItems: 'flex-start',
flexWrap: 'wrap',
gap: 'var(--space-2)',
}
const subtitleStyle = { fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: 0, maxWidth: '32rem' }
const emptyStyle = { color: 'var(--cv-ink-3)', fontSize: '0.85rem', padding: 'var(--space-2) 0' }
const captionStyle = { fontSize: '0.74rem', color: 'var(--cv-ink-3)', margin: '0.4rem 0 0', lineHeight: 1.5 }
const readoutStyle = {
display: 'flex',
alignItems: 'center',
gap: 'var(--space-3)',
flexWrap: 'wrap',
padding: 'var(--space-2) 0',
borderTop: '3px double var(--cv-ink)',
borderBottom: '1px solid var(--cv-rule)',
margin: 'var(--space-2) 0',
}
const readoutNumber = {
fontFamily: "'Source Serif 4', Georgia, serif",
fontWeight: 700,
fontSize: '2.75rem',
lineHeight: 1,
letterSpacing: '-0.02em',
color: 'var(--cv-ink)',
fontVariantNumeric: 'tabular-nums',
display: 'block',
}
const readoutLabel = {
fontSize: '0.72rem',
fontWeight: 600,
textTransform: 'uppercase',
letterSpacing: '0.08em',
color: 'var(--cv-ink-3)',
display: 'block',
marginTop: '0.3rem',
}
const degradedStyle = {
background: 'var(--cv-paper-2)',
border: '1px solid var(--cv-rule)',
borderLeft: '3px solid var(--cv-ink-4)',
borderRadius: 'var(--radius-md)',
padding: 'var(--space-2)',
marginTop: 'var(--space-2)',
color: 'var(--cv-ink-2)',
}
-162
View File
@@ -1,162 +0,0 @@
import ChartLegend from '../components/ChartLegend.jsx'
import ApproxNote from '../components/ApproxNote.jsx'
import { raceColor, OBSERVED_MARK_COLOR, SHORT_RACE_LABEL } from '../utils/colors.js'
import { fitSkewedInterval } from '../utils/distributionApprox.js'
import { groupKey, hasDrawsForAll } from '../utils/drawGroups.js'
import { quantile } from '../utils/kde.js'
import { niceTicks } from '../utils/niceTicks.js'
import { useDrawDistribution } from '../hooks/useDrawDistribution.js'
/**
* Arrest rate by student group, most recent year, disaggregated into two
* panels (Female / Male), each a horizontal box-and-whisker across the 4
* race categories. Whisker = the model's reported 90% interval; box = the
* 25th-75th percentile of the group's 500 real posterior draws (or, if draws
* are unavailable, the fitSkewedInterval approximation); white tick =
* median; dark diamond = observed rate.
*/
const RACE_ORDER = ['WH', 'BL', 'HI', 'AM']
const SEX_PANELS = [{ sex: 'F', label: 'Female' }, { sex: 'M', label: 'Male' }]
const ROW_HEIGHT = 34
const MODEL = 'unified_m3_mod' // matches ChartPanel's WAVE_MODEL, which built `data`
function buildBox(d, draws) {
const median = Math.max(d.modeledMedian || 0, 0)
const lower = Math.max(Math.min(d.rateLower ?? median, median), 0)
const upper = Math.max(d.rateUpper ?? median, median)
if (draws && draws.length > 0) {
return { lower, upper, median, q1: Math.max(quantile(draws, 0.25), 0), q3: Math.max(quantile(draws, 0.75), median) }
}
const fit = fitSkewedInterval({ median, lower, upper })
return { lower, upper, median, q1: Math.max(fit.quantile(0.25), 0), q3: Math.max(fit.quantile(0.75), median) }
}
export default function RateByGroupBar({ data, leaid, state }) {
const groups = data.map((d) => ({ race: d.race, sex: d.sex, stuEnroll: d.enrollment }))
const { drawsByGroup } = useDrawDistribution({ leaid, state, model: MODEL, year: '21-22', groups })
// Only the RACE_ORDER × SEX_PANELS cells are actually drawn, so coverage is
// judged against those — not against everything the API returned. The hook's
// 'ready' status is an any-group signal and would over-claim here: individual
// boxes still fall back to fitSkewedInterval whenever their group is missing
// from the shard. A null map (loading, or the whole fetch failed) fails this
// check too, so the error path still shows the note.
const renderedGroups = data.filter(
(d) => RACE_ORDER.includes(d.race) && SEX_PANELS.some((p) => p.sex === d.sex),
)
const allGroupsHaveDraws = hasDrawsForAll(drawsByGroup, renderedGroups)
const maxRate = Math.max(
...data.map((d) => Math.max(d.observedRate, d.rateUpper ?? d.modeledMedian ?? 0)),
0.5
)
const { ticks, niceMax } = niceTicks(maxRate, 4)
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={{ fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }}>
Arrest rate by student group — 2021–22 (per 1,000)
</h3>
{!allGroupsHaveDraws && <ApproxNote />}
<div style={{ display: 'grid', gridTemplateColumns: '1fr 1fr', gap: 'var(--space-3)', marginTop: 'var(--space-2)' }}>
{SEX_PANELS.map(({ sex, label }) => (
<SexPanel
key={sex}
label={label}
rows={data.filter((d) => d.sex === sex)}
drawsByGroup={drawsByGroup}
ticks={ticks}
niceMax={niceMax}
/>
))}
</div>
<ChartLegend items={[
...RACE_ORDER.filter((race) => data.some((d) => d.race === race)).map((race) => ({
shape: 'swatch', color: raceColor(race), label: SHORT_RACE_LABEL[race],
})),
{ shape: 'diamond', color: OBSERVED_MARK_COLOR, label: 'Observed' },
]} />
<p style={{ fontSize: '0.72rem', color: 'var(--cv-ink-3)', marginTop: 'var(--space-1)' }}>
{allGroupsHaveDraws
? "Box = 25th–75th percentile of 500 real posterior draws; whisker = the model's reported 90% interval; white tick = median."
: "Box = modeled 25th–75th percentile (fitted approximation); whisker = the model's reported 90% interval; white tick = median."}
{' '}The dark diamond is the observed rate.
</p>
</div>
)
}
function SexPanel({ label, rows, drawsByGroup, ticks, niceMax }) {
const width = 300
const margin = { top: 30, right: 16, bottom: 34, left: 66 }
const innerWidth = width - margin.left - margin.right
const byRace = RACE_ORDER.map((race) => rows.find((d) => d.race === race)).filter(Boolean)
const bodyHeight = byRace.length * ROW_HEIGHT
const height = margin.top + bodyHeight + margin.bottom
const xScale = (val) => margin.left + (val / niceMax) * innerWidth
return (
<div style={{ border: '1px solid var(--cv-rule)', borderRadius: 'var(--radius-md)', padding: 'var(--space-1)' }}>
<svg width="100%" viewBox={`0 0 ${width} ${height}`} style={{ maxWidth: '100%' }}>
<text x={width / 2} y={16} textAnchor="middle" fontSize="0.78rem" fontWeight={700} fill="var(--cv-ink-2)">
{label}
</text>
<rect x={margin.left} y={margin.top} width={innerWidth} height={bodyHeight} fill="var(--cv-paper-2)" rx={4} />
{ticks.map((val) => {
const x = xScale(val)
return (
<g key={val}>
<line x1={x} y1={margin.top} x2={x} y2={margin.top + bodyHeight} stroke="var(--cv-rule)" strokeWidth={1} />
<text x={x} y={margin.top + bodyHeight + 14} textAnchor="middle" fontSize="0.6rem" fill="var(--cv-ink-3)">
{val}
</text>
</g>
)
})}
{byRace.map((d, i) => {
const rowY = margin.top + i * ROW_HEIGHT
const midY = rowY + ROW_HEIGHT / 2
const boxTop = midY - ROW_HEIGHT * 0.26
const boxBottom = midY + ROW_HEIGHT * 0.26
const color = raceColor(d.race)
const draws = drawsByGroup?.[groupKey(d.race, d.sex)]
const box = buildBox(d, draws)
const observedX = xScale(d.observedRate)
return (
<g key={d.race}>
<line x1={xScale(box.lower)} y1={midY} x2={xScale(box.upper)} y2={midY} stroke={color} strokeWidth={1.5} />
<line x1={xScale(box.lower)} y1={boxTop} x2={xScale(box.lower)} y2={boxBottom} stroke={color} strokeWidth={1.5} />
<line x1={xScale(box.upper)} y1={boxTop} x2={xScale(box.upper)} y2={boxBottom} stroke={color} strokeWidth={1.5} />
<rect
x={xScale(box.q1)} y={boxTop} width={Math.max(xScale(box.q3) - xScale(box.q1), 1)} height={boxBottom - boxTop}
fill={color} fillOpacity={0.55} stroke={color} strokeWidth={1} rx={1.5}
/>
<line x1={xScale(box.median)} y1={boxTop} x2={xScale(box.median)} y2={boxBottom} stroke="#fff" strokeWidth={1.5} />
<rect
x={observedX - 4} y={midY - 4} width={8} height={8}
fill={OBSERVED_MARK_COLOR} stroke="#fff" strokeWidth={1}
transform={`rotate(45 ${observedX} ${midY})`}
/>
<text x={margin.left - 8} y={midY + 4} textAnchor="end" fontSize="0.72rem" fill="var(--cv-ink)">
{SHORT_RACE_LABEL[d.race] || d.race}
</text>
</g>
)
})}
<text x={margin.left + innerWidth / 2} y={height - 6} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-ink-3)">
Rate per 1,000 students
</text>
</svg>
</div>
)
}
+532
View File
@@ -0,0 +1,532 @@
import { useMemo } from 'react'
import { area, curveMonotoneX, curveStep } from 'd3-shape'
import { scaleLinear } from 'd3-scale'
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
import { raceColor } from '../utils/colors.js'
import { agrestiCoull } from '../utils/agrestiCoull.js'
import { densityProfile } from '../utils/densityProfile.js'
import { displayDraws } from '../utils/districtGroups.js'
import { computeRateDomain } from '../utils/rateDomain.js'
import { toRates } from '../utils/pooling.js'
import { densityCurve, fitSkewedInterval } from '../utils/distributionApprox.js'
import ApproxNote from '../components/ApproxNote.jsx'
/**
* Chart A — "Arrest rate probability density".
*
* Port of the white paper's Fig 6 (`wp_fig_group_density` in
* crdc-arrests/R/paper_figures.R): each selected student group's posterior
* predictive arrest rate as a filled density, with that group's frequentist
* Agresti–Coull interval on a rail beneath, so model and observation can be
* read against each other.
*
* Layout follows the palette contract in utils/colors.js: **race is hue, sex is
* position**. Female and Male are two stacked sub-panels sharing one x-axis,
* never two hues; when sex pooling is on there is a single panel and the sex
* dimension disappears entirely rather than being recoloured.
*
* React owns the DOM here — d3 supplies scales and path generators only. No
* selections, no imperative mutation.
*/
const AGRESTI_COULL_LEVEL = 0.95
const FILL_OPACITY = 0.4
const STROKE_WIDTH = 2
const CHART_WIDTH = 760
const MARGIN = { top: 22, right: 18, bottom: 34, left: 16 }
const ROW_HEIGHT = 118
const COMPARE_ROW_HEIGHT = 74
const RAIL_HEIGHT = 26
export default function RateDensityPanel({
groups,
selectedKeys,
enrollByGroup,
byModel,
status,
nDraws,
selectedModel,
onSelectModel,
compareAll,
onCompareAllChange,
pooled,
}) {
const activeModels = useMemo(
() => (compareAll ? MODEL_QUADRANTS.map((q) => q.model) : [selectedModel]),
[compareAll, selectedModel],
)
const selected = useMemo(
() => groups.filter((g) => selectedKeys.includes(g.key)),
[groups, selectedKeys],
)
// Agresti–Coull is computed from observed counts, so it is identical across
// model specifications — the same rail is drawn on every model row, which is
// exactly what makes the rows comparable.
const acByKey = useMemo(() => {
const out = {}
for (const g of selected) {
if (!(g.enroll > 0)) continue
const ac = agrestiCoull(g.observed, g.enroll, AGRESTI_COULL_LEVEL)
out[g.key] = {
// The R original can return a negative lower bound for a single event;
// a negative arrest count is not drawable, so clamp here rather than in
// the port (see utils/agrestiCoull.js).
lower: Math.max(0, (ac.lower / g.enroll) * 1000),
upper: (ac.upper / g.enroll) * 1000,
point: (g.observed / g.enroll) * 1000,
}
}
return out
}, [selected])
// Draws for every active model, moved into display (pooled or unpooled) space.
const drawsByModel = useMemo(() => {
const out = {}
for (const model of activeModels) {
out[model] = displayDraws(byModel?.[model]?.counts, enrollByGroup, pooled)
}
return out
}, [activeModels, byModel, enrollByGroup, pooled])
const domain = useMemo(() => {
const rateArrays = []
for (const model of activeModels) {
const { counts, enroll } = drawsByModel[model] || {}
for (const g of selected) rateArrays.push(toRates(counts?.[g.key], enroll?.[g.key]))
}
return computeRateDomain(rateArrays, Object.values(acByKey).map((a) => a.upper))
}, [activeModels, drawsByModel, selected, acByKey])
// One entry per (model, group): the shape to draw and where it came from.
const rowsByModel = useMemo(() => {
const out = {}
for (const model of activeModels) {
const { counts, enroll } = drawsByModel[model] || {}
out[model] = selected.map((g) => ({
group: g,
...profileFor(g, counts?.[g.key], enroll?.[g.key], domain.niceMax),
}))
}
return out
}, [activeModels, drawsByModel, selected, domain.niceMax])
const anyApproximated = Object.values(rowsByModel)
.flat()
.some((r) => r.source !== 'draws')
const loading = status === 'loading'
const sexPanels = pooled
? [{ sex: null, label: null }]
: [{ sex: 'F', label: 'Female students' }, { sex: 'M', label: 'Male students' }]
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<div style={headerStyle}>
<div>
<h3 style={cardTitle}>Arrest rate probability density</h3>
<p style={subtitleStyle}>
Each curve is one student group&rsquo;s modelled arrest rate. The bar beneath each panel
is what was actually reported, with its {Math.round(AGRESTI_COULL_LEVEL * 100)}%
Agresti&ndash;Coull interval.
</p>
</div>
<SpecControls
selectedModel={selectedModel}
onSelectModel={onSelectModel}
compareAll={compareAll}
onCompareAllChange={onCompareAllChange}
loading={loading}
/>
</div>
{anyApproximated && <ApproxNote />}
{selected.length === 0 ? (
<p style={emptyStyle}>
No student groups are selected. Check a group in the table above to show its distribution.
</p>
) : loading ? (
<p style={emptyStyle}>
{compareAll
? 'Loading posterior draws for all four specifications…'
: 'Loading posterior draws…'}
</p>
) : (
sexPanels.map((panel) => (
<SexPanel
key={panel.sex || 'pooled'}
label={panel.label}
sex={panel.sex}
domain={domain}
rowsByModel={rowsByModel}
activeModels={activeModels}
acByKey={acByKey}
compareAll={compareAll}
/>
))
)}
{domain.clipped && (
<p style={noteStyle}>
The axis stops at {domain.niceMax} per 1,000; one or more groups extend beyond it. Those
curves and intervals are cut off at the right edge, marked &rsaquo;.
</p>
)}
<p style={captionStyle}>
Densities are drawn from {nDraws > 0 ? nDraws.toLocaleString() : '500'} posterior predictive
draws per group. Point ranges are the observed rate with its{' '}
{Math.round(AGRESTI_COULL_LEVEL * 100)}% Agresti&ndash;Coull interval. Source: CRDC School
Arrest Rate API (Knowles &amp; Miller 2025), 2021&ndash;22 Civil Rights Data Collection.
</p>
</div>
)
}
/**
* Picks the shape for one group: real draws when we have them, the fitted
* summary-interval approximation when we don't, and nothing at all for a pooled
* group with no draws — there is no honest pooled shape to fit (see
* `buildDisplayGroups`).
*/
function profileFor(group, counts, enroll, niceMax) {
const domain = { min: 0, max: niceMax, n: 60 }
const fromDraws = densityProfile(counts, enroll, domain)
if (fromDraws) return { profile: fromDraws, source: 'draws' }
if (!group.modeled || !(group.modeled.upper > 0)) return { profile: null, source: 'none' }
const fit = fitSkewedInterval({ ...group.modeled, intervalMass: 0.95 })
const points = densityCurve(fit, domain)
return {
profile: { kind: 'kde', points, maxY: Math.max(...points.map((p) => p.y)), step: 0 },
source: 'approx',
}
}
function SexPanel({ label, sex, domain, rowsByModel, activeModels, acByKey, compareAll }) {
const rows = activeModels.map((model) => ({
model,
label: MODEL_QUADRANTS.find((q) => q.model === model)?.label || model,
entries: (rowsByModel[model] || []).filter((r) => (sex ? r.group.sex === sex : true)),
}))
const hasAnything = rows.some((r) => r.entries.some((e) => e.profile))
const rowHeight = compareAll ? COMPARE_ROW_HEIGHT : ROW_HEIGHT
const innerWidth = CHART_WIDTH - MARGIN.left - MARGIN.right
const bandHeight = rowHeight + RAIL_HEIGHT
const height = MARGIN.top + rows.length * bandHeight + MARGIN.bottom
const x = scaleLinear().domain([0, domain.niceMax]).range([MARGIN.left, MARGIN.left + innerWidth])
return (
<figure style={{ margin: '0 0 var(--space-2)' }}>
{label && <figcaption style={panelLabelStyle}>{label}</figcaption>}
{!hasAnything ? (
<p style={emptyStyle}>No modelled distribution is available for these groups.</p>
) : (
<div style={{ overflowX: 'auto' }}>
<svg
width="100%"
viewBox={`0 0 ${CHART_WIDTH} ${height}`}
style={{ maxWidth: '100%', minWidth: '320px' }}
role="img"
aria-label={`Modelled arrest rate distributions${label ? ` for ${label.toLowerCase()}` : ''}`}
>
{domain.ticks.map((t) => (
<line
key={t}
x1={x(t)}
y1={MARGIN.top}
x2={x(t)}
y2={MARGIN.top + rows.length * bandHeight}
stroke="var(--cv-rule)"
strokeWidth={1}
/>
))}
{rows.map((row, i) => (
<ModelRow
key={row.model}
row={row}
compareAll={compareAll}
x={x}
top={MARGIN.top + i * bandHeight}
rowHeight={rowHeight}
acByKey={acByKey}
niceMax={domain.niceMax}
/>
))}
{domain.ticks.map((t) => (
<text
key={t}
x={x(t)}
y={MARGIN.top + rows.length * bandHeight + 15}
textAnchor="middle"
fontSize="0.62rem"
fill="var(--cv-ink-3)"
>
{t}
</text>
))}
<text
x={MARGIN.left + innerWidth / 2}
y={height - 4}
textAnchor="middle"
fontSize="0.64rem"
fill="var(--cv-ink-3)"
>
Arrests per 1,000 students
</text>
</svg>
</div>
)}
</figure>
)
}
function ModelRow({ row, compareAll, x, top, rowHeight, acByKey, niceMax }) {
const baselineY = top + rowHeight
const peakHeight = rowHeight * 0.86
const drawable = row.entries.filter((e) => e.profile)
return (
<g>
{compareAll && (
<text x={MARGIN.left + 2} y={top + 10} fontSize="0.62rem" fontWeight={700} fill="var(--cv-ink-2)">
{row.label}
</text>
)}
{drawable.map((entry, i) => (
<GroupArea
key={entry.group.key}
entry={entry}
x={x}
baselineY={baselineY}
// Each profile is normalized to its own peak. The two profile kinds
// carry different y units (probability mass vs. density), so a shared
// maximum would squash whichever kind happened to peak lower — and
// what the reader is comparing here is location and spread, not peak
// height.
peakHeight={peakHeight}
labelIndex={i}
compact={compareAll}
niceMax={niceMax}
/>
))}
<line
x1={MARGIN.left}
y1={baselineY}
x2={x(niceMax)}
y2={baselineY}
stroke="var(--cv-rule-strong)"
strokeWidth={1}
/>
{row.entries.map((entry, i) => {
const ac = acByKey[entry.group.key]
if (!ac) return null
return (
<PointRange
key={entry.group.key}
ac={ac}
color={raceColor(entry.group.race)}
x={x}
y={baselineY + 8 + (i % 3) * 5}
niceMax={niceMax}
/>
)
})}
</g>
)
}
function GroupArea({ entry, x, baselineY, peakHeight, labelIndex, compact, niceMax }) {
const { profile, group } = entry
const color = raceColor(group.race)
const y = scaleLinear().domain([0, profile.maxY || 1]).range([baselineY, baselineY - peakHeight])
const points =
profile.kind === 'mass'
? padMassPoints(profile.points, profile.step, niceMax)
: profile.points
const areaGen = area()
.x((p) => x(Math.min(p.x, niceMax)))
.y0(baselineY)
.y1((p) => y(p.y))
// A staircase for discrete mass, a smooth curve for a KDE. curveStep keeps
// the mass profile honest — it says "these values and no others" — while
// still visually rhyming with the filled areas beside it.
.curve(profile.kind === 'mass' ? curveStep : curveMonotoneX)
const peak = profile.points.reduce((best, p) => (p.y > best.y ? p : best), profile.points[0])
const peakX = x(Math.min(peak.x, niceMax))
const labelY = Math.max(baselineY - peakHeight - 2, y(peak.y) - 4 - (labelIndex % 2) * 11)
return (
<g>
<path d={areaGen(points)} fill={color} opacity={FILL_OPACITY} />
<path
d={areaGen.lineY1()(points)}
fill="none"
stroke={color}
strokeWidth={STROKE_WIDTH}
strokeLinejoin="round"
/>
{/* Direct-labelled at the peak so there is no legend to hunt through. */}
<text
x={peakX}
y={labelY}
textAnchor={peakX > x(niceMax) * 0.8 ? 'end' : 'middle'}
fontSize={compact ? '0.58rem' : '0.65rem'}
fontWeight={700}
fill={color}
stroke="var(--cv-paper)"
strokeWidth={3}
paintOrder="stroke"
>
{group.shortLabel}
</text>
</g>
)
}
/**
* A staircase needs a floor to close against on both sides: without the zero
* pads, curveStep leaves the first and last bars open and the fill bleeds to
* the panel edge.
*/
function padMassPoints(points, step, niceMax) {
const half = Math.max(step, 1e-6) / 2
const first = points[0]
const last = points[points.length - 1]
return [
{ x: Math.max(0, first.x - half), y: 0 },
...points,
{ x: Math.min(last.x + half, niceMax), y: 0 },
]
}
function PointRange({ ac, color, x, y, niceMax }) {
const lo = Math.min(ac.lower, niceMax)
const hi = Math.min(ac.upper, niceMax)
const point = Math.min(ac.point, niceMax)
const clipped = ac.upper > niceMax
return (
<g>
<line x1={x(lo)} y1={y} x2={x(hi)} y2={y} stroke={color} strokeWidth={2} strokeLinecap="round" />
<circle cx={x(point)} cy={y} r={3} fill={color} stroke="var(--cv-paper)" strokeWidth={1} />
{clipped && (
<text x={x(niceMax) + 3} y={y + 3} fontSize="0.6rem" fill={color}>
&rsaquo;
</text>
)}
</g>
)
}
function SpecControls({ selectedModel, onSelectModel, compareAll, onCompareAllChange, loading }) {
return (
<div style={{ display: 'flex', flexDirection: 'column', gap: '0.4rem', alignItems: 'flex-end' }}>
<div role="group" aria-label="Model specification" style={segmentedStyle}>
{MODEL_QUADRANTS.map((q) => {
const active = !compareAll && q.model === selectedModel
return (
<button
key={q.model}
type="button"
onClick={() => onSelectModel(q.model)}
aria-pressed={active}
disabled={compareAll}
style={{
...segmentStyle,
background: active ? 'var(--cv-navy-600)' : 'transparent',
color: active ? '#fff' : 'var(--cv-ink-2)',
opacity: compareAll ? 0.5 : 1,
cursor: compareAll ? 'not-allowed' : 'pointer',
}}
>
{q.label}
</button>
)
})}
</div>
<label style={switchLabel}>
<input
type="checkbox"
checked={compareAll}
onChange={(e) => onCompareAllChange(e.target.checked)}
/>
<span>Compare all four specifications</span>
{loading && compareAll && <Spinner />}
</label>
</div>
)
}
function Spinner() {
return (
<span
aria-label="Loading"
style={{
display: 'inline-block',
width: '11px',
height: '11px',
border: '2px solid var(--cv-rule-strong)',
borderTopColor: 'var(--cv-accent)',
borderRadius: '50%',
animation: 'cv-spin 700ms linear infinite',
}}
/>
)
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const headerStyle = {
display: 'flex',
justifyContent: 'space-between',
alignItems: 'flex-start',
flexWrap: 'wrap',
gap: 'var(--space-2)',
}
const subtitleStyle = { fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: 0, maxWidth: '34rem' }
const panelLabelStyle = {
fontSize: '0.75rem',
fontWeight: 700,
textTransform: 'uppercase',
letterSpacing: '0.06em',
color: 'var(--cv-ink-3)',
marginBottom: '0.1rem',
}
const emptyStyle = { color: 'var(--cv-ink-3)', fontSize: '0.85rem', padding: 'var(--space-2) 0' }
const noteStyle = { fontSize: '0.76rem', fontStyle: 'italic', color: 'var(--cv-ink-3)', margin: '0 0 0.4rem' }
const captionStyle = { fontSize: '0.74rem', color: 'var(--cv-ink-3)', margin: '0.4rem 0 0', lineHeight: 1.5 }
const segmentedStyle = {
display: 'inline-flex',
border: '1px solid var(--cv-rule)',
borderRadius: 'var(--radius-md)',
overflow: 'hidden',
}
const segmentStyle = {
padding: '0.25rem 0.55rem',
border: 'none',
borderRight: '1px solid var(--cv-rule)',
fontFamily: 'var(--font-sans)',
fontSize: '0.72rem',
fontWeight: 600,
}
const switchLabel = {
display: 'inline-flex',
alignItems: 'center',
gap: '0.4rem',
fontSize: '0.76rem',
color: 'var(--cv-ink-2)',
cursor: 'pointer',
}
-204
View File
@@ -1,204 +0,0 @@
import { useState } from 'react'
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
import { raceColor, OBSERVED_MARK_COLOR, RACE_LABELS, SHORT_RACE_LABEL } from '../utils/colors.js'
import { fitSkewedInterval, densityCurve } from '../utils/distributionApprox.js'
import { groupKey, hasDrawsForAll } from '../utils/drawGroups.js'
import { kdeCurve } from '../utils/kde.js'
import { niceTicks } from '../utils/niceTicks.js'
import { useDrawDistribution } from '../hooks/useDrawDistribution.js'
import ChartLegend from '../components/ChartLegend.jsx'
import ApproxNote from '../components/ApproxNote.jsx'
/**
* Modeled posterior density per race×sex group, for one selected model
* (dropdown, default three-year + referral rate), split into Female/Male
* columns. Each ridge is drawn from that group's 500 real posterior draws
* (Gaussian KDE) when available, falling back per-group to the analytic
* fitSkewedInterval/densityCurve approximation otherwise.
*/
const RACE_ORDER = ['WH', 'BL', 'HI', 'AM']
const SEX_COLUMNS = [{ sex: 'F', label: 'Female' }, { sex: 'M', label: 'Male' }]
const DEFAULT_MODEL = 'unified_m4_mod' // Three-year + referral rate
function buildGroupRow(row, draws) {
const enroll = row.stu_enroll || 0
const observedRate = enroll > 0 ? ((row.observed_arrests || 0) / enroll) * 1000 : 0
const rateMedian = (row.rate_median || 0) * 1000
const rateLower = (row.rate_lower || 0) * 1000
const rateUpper = (row.rate_upper || 0) * 1000
return {
race: row.race,
sex: row.sex,
observedRate,
draws,
fit: fitSkewedInterval({ median: rateMedian, lower: rateLower, upper: rateUpper }),
}
}
export default function RateDensityRidgeline({ quadData, leaid, state }) {
const [selectedModel, setSelectedModel] = useState(DEFAULT_MODEL)
const rows = (quadData && quadData[selectedModel]) || []
const groups = rows.map((r) => ({ race: r.race, sex: r.sex, stuEnroll: r.stu_enroll || 0 }))
const { drawsByGroup } = useDrawDistribution({ leaid, state, model: selectedModel, year: '21-22', groups })
// Only the RACE_ORDER × SEX_COLUMNS cells get a ridge, so coverage is judged
// against those. The hook's 'ready' status is an any-group signal: individual
// ridges still fall back to densityCurve when their group is missing from the
// shard, so gating the note on `status` alone would hide it while part of the
// chart is an approximation. A null map (loading or a failed fetch) also fails
// this check, so the error path still shows the note.
const renderedGroups = rows.filter(
(r) => RACE_ORDER.includes(r.race) && SEX_COLUMNS.some((c) => c.sex === r.sex),
)
const allGroupsHaveDraws = hasDrawsForAll(drawsByGroup, renderedGroups)
const modelSelect = (
<select
value={selectedModel}
onChange={(e) => setSelectedModel(e.target.value)}
style={{ padding: '0.25rem 0.5rem', fontFamily: 'var(--font-sans)', fontSize: '0.8rem' }}
>
{MODEL_QUADRANTS.map((q) => (
<option key={q.model} value={q.model}>{q.label}</option>
))}
</select>
)
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<div style={{ display: 'flex', justifyContent: 'space-between', alignItems: 'flex-start', flexWrap: 'wrap', gap: 'var(--space-2)' }}>
<div>
<h3 style={{ fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }}>
Predicted arrest rates by student group
</h3>
{!allGroupsHaveDraws && <ApproxNote />}
</div>
{modelSelect}
</div>
{rows.length === 0 ? (
<p style={{ color: 'var(--cv-ink-3)', marginTop: 'var(--space-2)' }}>No model data available.</p>
) : (
<>
<RidgeColumns rows={rows} drawsByGroup={drawsByGroup} />
<ChartLegend items={[
...RACE_ORDER.map((race) => ({ shape: 'swatch', color: raceColor(race), label: RACE_LABELS[race] })),
{ shape: 'diamond', color: OBSERVED_MARK_COLOR, label: 'Observed' },
]} />
</>
)}
</div>
)
}
function RidgeColumns({ rows, drawsByGroup }) {
// Shared x-domain across both columns, so Female/Male are directly comparable.
const allUpper = rows.map((r) => (r.rate_upper || 0) * 1000)
const allObserved = rows
.filter((r) => (r.stu_enroll || 0) > 0)
.map((r) => ((r.observed_arrests || 0) / r.stu_enroll) * 1000)
const rawMax = Math.min(Math.max(...allUpper, ...allObserved, 1) * 1.15, 30)
const { ticks, niceMax } = niceTicks(rawMax, 5)
return (
<div style={{ display: 'grid', gridTemplateColumns: '1fr 1fr', gap: 'var(--space-3)', marginTop: 'var(--space-2)' }}>
{SEX_COLUMNS.map(({ sex, label }) => (
<SexRidgeColumn
key={sex}
label={label}
rows={rows.filter((r) => r.sex === sex)}
drawsByGroup={drawsByGroup}
ticks={ticks}
maxRate={niceMax}
/>
))}
</div>
)
}
function SexRidgeColumn({ label, rows, drawsByGroup, ticks, maxRate }) {
const groups = RACE_ORDER
.map((race) => rows.find((r) => r.race === race))
.filter(Boolean)
.map((row) => buildGroupRow(row, drawsByGroup?.[groupKey(row.race, row.sex)]))
const width = 300
const rowHeight = 58
const margin = { top: 26, right: 16, bottom: 26, left: 56 }
const innerWidth = width - margin.left - margin.right
const height = margin.top + Math.max(groups.length, 1) * rowHeight + margin.bottom
const xScale = (val) => margin.left + (val / maxRate) * innerWidth
const curves = groups.map((g) =>
g.draws && g.draws.length > 0
? kdeCurve(g.draws, { min: 0, max: maxRate, n: 60 })
: densityCurve(g.fit, { min: 0, max: maxRate, n: 60 }),
)
const maxPdf = Math.max(...curves.flatMap((c) => c.map((p) => p.y)), 1e-9)
const peakHeight = rowHeight * 0.82
return (
<div style={{ border: '1px solid var(--cv-rule)', borderRadius: 'var(--radius-md)', padding: 'var(--space-1)' }}>
<svg width="100%" viewBox={`0 0 ${width} ${height}`} style={{ maxWidth: '100%' }}>
<text x={width / 2} y={16} textAnchor="middle" fontSize="0.78rem" fontWeight={700} fill="var(--cv-ink-2)">
{label}
</text>
{groups.length === 0 ? (
<text x={width / 2} y={height / 2} textAnchor="middle" fontSize="0.65rem" fill="var(--cv-ink-3)">
No data for this group
</text>
) : (
<>
{ticks.map((val) => {
const x = xScale(val)
return (
<g key={val}>
<line x1={x} y1={margin.top} x2={x} y2={margin.top + groups.length * rowHeight} stroke="var(--cv-rule)" strokeWidth={1} />
<text x={x} y={margin.top + groups.length * rowHeight + 14} textAnchor="middle" fontSize="0.58rem" fill="var(--cv-ink-3)">
{val}
</text>
</g>
)
})}
{groups.map((g, i) => {
const rowTop = margin.top + i * rowHeight
const baselineY = rowTop + rowHeight * 0.9
const curve = curves[i]
const color = raceColor(g.race)
const topPath = curve
.map((p, j) => `${j === 0 ? 'M' : 'L'}${xScale(p.x)},${baselineY - (p.y / maxPdf) * peakHeight}`)
.join(' ')
const areaPath = `${topPath} L${xScale(maxRate)},${baselineY} L${xScale(0)},${baselineY} Z`
const obsX = xScale(Math.min(g.observedRate, maxRate))
const obsY = baselineY - peakHeight * 0.15
return (
<g key={g.race}>
<text x={margin.left - 8} y={rowTop + rowHeight / 2 + 4} textAnchor="end" fontSize="0.68rem" fill="var(--cv-ink)">
{SHORT_RACE_LABEL[g.race] || g.race}
</text>
<path d={areaPath} fill={color} opacity={0.6} stroke="#fff" strokeWidth={0.5} />
<rect
x={obsX - 4} y={obsY - 4} width={8} height={8}
fill={OBSERVED_MARK_COLOR} stroke="#fff" strokeWidth={1}
transform={`rotate(45 ${obsX} ${obsY})`}
/>
</g>
)
})}
<text x={width / 2} y={height - 6} textAnchor="middle" fontSize="0.58rem" fill="var(--cv-ink-3)">
Arrests per 1,000 students
</text>
</>
)}
</svg>
</div>
)
}
+197 -65
View File
@@ -1,58 +1,163 @@
import { useState, useEffect } from 'react'
import { useCallback, useEffect, useMemo, useState } from 'react'
import * as api from '../hooks/useApi.js'
import { groupLabel } from '../utils/colors.js'
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
import { useDrawDistribution } from '../hooks/useDrawDistribution.js'
import {
buildDisplayGroups,
defaultDiffPair,
defaultSelectedKeys,
enrollByGroupKey,
} from '../utils/districtGroups.js'
import { shouldPoolBySex } from '../utils/pooling.js'
import ArrestsOverTime from '../charts/ArrestsOverTime.jsx'
import RateByGroupBar from '../charts/RateByGroupBar.jsx'
import RateDensityRidgeline from '../charts/RateDensityRidgeline.jsx'
import GroupDifference from '../charts/GroupDifference.jsx'
import RateDensityPanel from '../charts/RateDensityPanel.jsx'
import DistrictSummaryTable from './DistrictSummaryTable.jsx'
const ALL_WAVES = ['15-16', '17-18', '21-22']
const QUADRANT_MODELS = ['unified_m1_mod', 'unified_m2_mod', 'unified_m3_mod', 'unified_m4_mod']
const WAVE_MODEL = 'unified_m3_mod' // three-year, no covariate — powers Charts 1 & 2
const WAVE_MODEL = 'unified_m3_mod' // three-year, no covariate — powers the time series
const DEFAULT_SPEC = 'unified_m4_mod' // three-year + referral rate
const CURRENT_WAVE = '21-22'
const QUADRANT_MODELS = MODEL_QUADRANTS.map((q) => q.model)
/**
* ChartPanel — 3 charts: arrests over time (observed vs. modeled), arrest
* rate by student group (Female/Male panels), and the model-selectable
* posterior density ridge chart.
* The results page: what was observed (summary table), what the model says the
* rate is (density panel), and what it says about the gap between two groups
* (difference chart) — plus the district's arrest history.
*
* This component owns every piece of cross-chart state: the selected model
* specification, whether sexes are pooled, which groups are checked, and which
* pair is being differenced. The charts are given values and callbacks and hold
* none of it, so the table's checkboxes and the density panel can't drift apart.
*/
export default function ChartPanel({ district, state }) {
const [data, setData] = useState(null) // all fetched estimate rows keyed by year/model
const [waveData, setWaveData] = useState(null)
const [summaryByModel, setSummaryByModel] = useState({})
const [loading, setLoading] = useState(true)
const [selectedModel, setSelectedModel] = useState(DEFAULT_SPEC)
const [compareAll, setCompareAll] = useState(false)
const [poolOverride, setPoolOverride] = useState(null) // null = follow the rule
// Selections are stored with the pooling mode they were made in: pooled keys
// ('BL') and unpooled keys ('BL_F') are different namespaces, so a selection
// made in one mode must not be reapplied in the other.
const [selection, setSelection] = useState(null)
const [pairSelection, setPairSelection] = useState(null)
// ——— Fetch: the three waves for the time series ———
useEffect(() => {
async function fetchData() {
let cancelled = false
setLoading(true)
setWaveData(null)
setSummaryByModel({})
setPoolOverride(null)
setSelection(null)
setPairSelection(null)
setCompareAll(false)
async function run() {
const waves = {}
await Promise.all(
ALL_WAVES.map(async (year) => {
try {
waves[year] = await api.fetchDistrictEstimates(district.leaid, { model: WAVE_MODEL, year })
} catch (err) {
console.error(`ChartPanel: wave ${year} fetch failed:`, err)
}
}),
)
if (cancelled) return
setWaveData(waves)
setLoading(false)
}
run()
return () => {
cancelled = true
}
}, [district.leaid])
// ——— Fetch: the current wave's summary for whichever spec is selected ———
// Only the selected specification is fetched. Enrollment and observed arrest
// counts are district facts, not model outputs, and the modelled distributions
// now come from real draws rather than from these rows — so prefetching all
// four specs' summaries would be four requests for data three of which are
// never read.
useEffect(() => {
let cancelled = false
if (summaryByModel[selectedModel]) return
async function run() {
try {
// Chart 1 + 2: all 3 waves × three-year model (unified_m3_mod)
const waveData = {}
await Promise.all(ALL_WAVES.map(async (year) => {
try { waveData[year] = await api.fetchDistrictEstimates(district.leaid, { model: WAVE_MODEL, year }) } catch (e) {}
}))
// Chart 3: all 4 quadrant models, so the dropdown can switch between them
const quadData = {}
await Promise.all(QUADRANT_MODELS.map(async (model) => {
try { quadData[model] = await api.fetchDistrictEstimates(district.leaid, { model, year: '21-22' }) } catch (e) {}
}))
setData({ waveData, quadData })
const rows = await api.fetchDistrictEstimates(district.leaid, {
model: selectedModel,
year: CURRENT_WAVE,
})
if (!cancelled) setSummaryByModel((prev) => ({ ...prev, [selectedModel]: rows }))
} catch (err) {
console.error('ChartPanel data fetch failed:', err)
} finally {
setLoading(false)
console.error(`ChartPanel: summary fetch failed for ${selectedModel}:`, err)
if (!cancelled) setSummaryByModel((prev) => ({ ...prev, [selectedModel]: [] }))
}
}
fetchData()
}, [district.leaid])
run()
return () => {
cancelled = true
}
}, [district.leaid, selectedModel, summaryByModel])
if (loading || !data) {
return <LoadingCharts />
}
const rows = useMemo(() => summaryByModel[selectedModel] || [], [summaryByModel, selectedModel])
const mostRecent = data.waveData['21-22'] || []
const autoPooled = useMemo(() => shouldPoolBySex(rows), [rows])
const pooled = poolOverride ?? autoPooled
const groups = useMemo(() => buildDisplayGroups(rows, pooled), [rows, pooled])
const enrollByGroup = useMemo(() => enrollByGroupKey(rows), [rows])
const defaultKeys = useMemo(() => defaultSelectedKeys(groups), [groups])
const selectedKeys = selection?.pooled === pooled ? selection.keys : defaultKeys
const defaultPair = useMemo(() => defaultDiffPair(groups), [groups])
const pair = pairSelection?.pooled === pooled ? pairSelection.pair : defaultPair
const selectionIsEnrollmentFallback =
groups.length > 0 && groups.every((g) => g.observed === 0) && selection === null
const handleToggleGroup = useCallback(
(key) => {
const next = selectedKeys.includes(key)
? selectedKeys.filter((k) => k !== key)
: [...selectedKeys, key]
setSelection({ pooled, keys: next })
},
[selectedKeys, pooled],
)
const handlePooledChange = useCallback((next) => {
setPoolOverride(next)
// Both selections live in the other namespace now — drop them and let the
// defaults recompute for the new group set.
setSelection(null)
setPairSelection(null)
}, [])
const handlePairChange = useCallback((next) => setPairSelection({ pooled, pair: next }), [pooled])
const drawModels = useMemo(
() => (compareAll ? QUADRANT_MODELS : [selectedModel]),
[compareAll, selectedModel],
)
const { status, byModel, nDraws } = useDrawDistribution({
leaid: district.leaid,
state,
models: drawModels,
year: CURRENT_WAVE,
})
if (loading || !waveData) return <LoadingCharts />
// Chart 1: Arrests over time by wave — observed total + modeled (three-year model) point-range
const timeSeriesData = ALL_WAVES.map((year) => {
const yearRows = data.waveData[year] || []
const yearRows = waveData[year] || []
return {
year,
label: `20${year.replace('-', '-')}`,
@@ -64,20 +169,8 @@ export default function ChartPanel({ district, state }) {
}
})
// Chart 2: rate by student group (most recent wave, per 1k), with modeled + observed + interval
const rateByGroup = mostRecent.map((r) => ({
race: r.race, sex: r.sex, label: groupLabel(r.race, r.sex),
observedRate: (r.observed_arrests || 0) / ((r.stu_enroll || 1) / 1000),
modeledMedian: (r.count_median || 0) / ((r.stu_enroll || 1) / 1000),
rateLower: (r.rate_lower || 0) * 1000,
rateUpper: (r.rate_upper || 0) * 1000,
observedArrests: r.observed_arrests || 0,
enrollment: r.stu_enroll || 0,
}))
return (
<div style={{ padding: 'var(--space-3) 0 var(--space-7)' }}>
{/* Chart panel header */}
<div style={{ marginBottom: 'var(--space-4)' }}>
<span className="eyebrow">District estimates — Bayesian model comparison</span>
<h2 style={{ marginTop: 'var(--space-1)', marginBottom: 0 }}>
@@ -85,29 +178,68 @@ export default function ChartPanel({ district, state }) {
</h2>
</div>
<div style={{
display: 'flex',
flexDirection: 'column',
gap: 'var(--space-5)',
maxWidth: '70rem',
marginLeft: 'auto',
marginRight: 'auto'
}}>
<div
style={{
display: 'flex',
flexDirection: 'column',
gap: 'var(--space-5)',
maxWidth: '70rem',
marginLeft: 'auto',
marginRight: 'auto',
}}
>
<DistrictSummaryTable
groups={groups}
selectedKeys={selectedKeys}
onToggleGroup={handleToggleGroup}
pooled={pooled}
autoPooled={autoPooled}
onPooledChange={handlePooledChange}
selectionIsEnrollmentFallback={selectionIsEnrollmentFallback}
/>
<RateDensityPanel
groups={groups}
selectedKeys={selectedKeys}
enrollByGroup={enrollByGroup}
byModel={byModel}
status={status}
nDraws={nDraws}
selectedModel={selectedModel}
onSelectModel={setSelectedModel}
compareAll={compareAll}
onCompareAllChange={setCompareAll}
pooled={pooled}
/>
<GroupDifference
groups={groups}
enrollByGroup={enrollByGroup}
byModel={byModel}
selectedModel={selectedModel}
status={status}
pooled={pooled}
pair={pair}
onPairChange={handlePairChange}
/>
<ArrestsOverTime data={timeSeriesData} modelId={WAVE_MODEL} />
<RateByGroupBar data={rateByGroup} leaid={district.leaid} state={state} />
<RateDensityRidgeline quadData={data.quadData} leaid={district.leaid} state={state} />
</div>
{/* Methodology footer */}
<div style={{ marginTop: 'var(--space-6)', padding: 'var(--space-3) 0', borderTop: '1px solid var(--cv-rule)' }}>
<div
style={{
marginTop: 'var(--space-6)',
padding: 'var(--space-3) 0',
borderTop: '1px solid var(--cv-rule)',
}}
>
<span className="eyebrow">Methodology</span>
<p style={{ fontSize: '0.85rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-1)' }}>
Estimates are from the CRDC School Arrest Rate API (Knowles & Miller 2025). Data shown spans three waves of
the Civil Rights Data Collection (2015–16, 2017–18, 2021–22) and lets you explore four Bayesian model
specifications: one-year vs. three-year models with and without referral-rate covariates. All rates are
per 1,000 students.
Estimates are from the CRDC School Arrest Rate API (Knowles &amp; Miller 2025). Data shown
spans three waves of the Civil Rights Data Collection (2015–16, 2017–18, 2021–22) and lets
you explore four Bayesian model specifications: one-year vs. three-year models with and
without referral-rate covariates. All rates are per 1,000 students. Modelled distributions
are drawn from the published posterior predictive draws, fetched in the browser.
</p>
</div>
</div>
+54 -33
View File
@@ -1,10 +1,42 @@
import { useState, useEffect } from 'react'
import { searchDistricts, fetchStateDistricts } from '../hooks/useApi.js'
import { searchDistricts } from '../hooks/useApi.js'
const TOP_DISTRICTS_URL = `${import.meta.env.BASE_URL}data/top_districts.json`
const SUGGESTION_COUNT = 8
// Module-level cache: the fixture covers every state, so it is fetched at most
// once per page load no matter how many states the visitor browses through.
let topDistrictsPromise = null
function loadTopDistricts() {
if (!topDistrictsPromise) {
topDistrictsPromise = fetch(TOP_DISTRICTS_URL)
.then((res) => {
if (!res.ok) throw new Error(`HTTP ${res.status} loading ${TOP_DISTRICTS_URL}`)
return res.json()
})
.catch((err) => {
// Don't poison the cache with a rejected promise — let a later visit retry.
topDistrictsPromise = null
throw err
})
}
return topDistrictsPromise
}
/**
* District search screen with:
* 1. "Interesting" suggestions — districts with the most arrests in the selected state (fetched once)
* 1. "Interesting" suggestions — the districts with the most arrests in the
* selected state, read from the committed `public/data/top_districts.json`
* fixture (built by `scripts/build-top-districts.mjs`).
* 2. Live-search as you type → /api/v1/districts?q=...&state=XX
*
* The suggestions used to be computed at runtime from
* `/estimates?state=XX&year=21-22&limit=500`. That endpoint returns rows
* `ORDER BY LEAID, RACE, SEX` at eight rows per district, so the cap selected
* the ~62 lowest-LEAID districts in the state rather than the busiest ones —
* California has 11,488 rows. The fixture is ranked over the complete result
* set, and it also removes a multi-second fetch from this screen.
*/
export default function DistrictSearch({ state, onSelect, onBack }) {
const [query, setQuery] = useState('')
@@ -13,38 +45,32 @@ export default function DistrictSearch({ state, onSelect, onBack }) {
const [loadingSugg, setLoadingSugg] = useState(true)
const [loadingSearch, setLoadingSearch] = useState(false)
// ——— Fetch interesting suggestions (top arrests in this state) once on mount ———
// ——— Read the suggestion fixture for this state ———
useEffect(() => {
let cancelled = false
async function loadSuggestions() {
try {
// /estimates?state=XX&year=21-22 returns all districts; sort by observed_arrests desc client-side
const data = await fetchStateDistricts(state, '21-22', 500)
setLoadingSugg(true)
// Aggregate arrests per district (sum across race×sex groups), then sort
const byLeaid = {}
for (const row of data) {
if (!byLeaid[row.leaid]) {
byLeaid[row.leaid] = { leaid: row.leaid, lea_name: row.lea_name, state: row.state, observed_arrests: 0 }
}
byLeaid[row.leaid].observed_arrests += (row.observed_arrests || 0)
}
const sorted = Object.values(byLeaid).sort((a, b) => b.observed_arrests - a.observed_arrests)
if (!cancelled) {
// Take top 8 for suggestions; filter to those with >0 arrests
setSuggestions(sorted.filter(d => d.observed_arrests > 0).slice(0, 8))
}
} catch (err) {
console.error('Failed to load interesting districts:', err)
loadTopDistricts()
.then((fixture) => {
if (cancelled) return
const forState = fixture?.states?.[state] || []
setSuggestions(
forState.slice(0, SUGGESTION_COUNT).map((d) => ({
leaid: d.leaid,
lea_name: d.name,
state,
observed_arrests: d.arrests,
})),
)
})
.catch((err) => {
console.error('Failed to load suggested districts:', err)
if (!cancelled) setSuggestions([]) // degrade gracefully — just show search box
} finally {
})
.finally(() => {
if (!cancelled) setLoadingSugg(false)
}
}
})
loadSuggestions()
return () => { cancelled = true }
}, [state])
@@ -158,11 +184,6 @@ export default function DistrictSearch({ state, onSelect, onBack }) {
</div>
)}
</div>
{/* Hint */}
<p style={{ fontSize: '0.85rem', color: 'var(--cv-ink-3)', marginTop: 'var(--space-4)' }}>
Tip: Try districts like Derby (KS), Paterson (NJ), or Mobile County (AL) — they have notable arrest rates.
</p>
</div>
)
}
+234
View File
@@ -0,0 +1,234 @@
import { raceColor } from '../utils/colors.js'
import { POOL_BY_SEX_ARREST_THRESHOLD } from '../utils/pooling.js'
/**
* The results page's first block: what was actually observed, before any
* modelling. One row per student group plus a district total.
*
* The table doubles as the legend and the control for the density panel — each
* row carries the checkbox that adds or removes that group's curve, and the
* colour swatch is the same hue the curve is drawn in. That's deliberate:
* a separate legend plus a separate group picker would make the reader hold
* three mappings in their head instead of one.
*
* @param {{
* groups: Array<{key:string, race:string, sex:string|null, label:string,
* enroll:number, observed:number, rate:number}>,
* selectedKeys: string[],
* onToggleGroup: (key: string) => void,
* pooled: boolean,
* autoPooled: boolean,
* onPooledChange: (pooled: boolean) => void,
* selectionIsEnrollmentFallback: boolean,
* }} props
*/
export default function DistrictSummaryTable({
groups,
selectedKeys,
onToggleGroup,
pooled,
autoPooled,
onPooledChange,
selectionIsEnrollmentFallback,
}) {
if (!groups?.length) {
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={cardTitle}>Reported arrests and enrollment</h3>
<p style={{ color: 'var(--cv-ink-3)', margin: 0 }}>
No student-group data is available for this district in 2021–22.
</p>
</div>
)
}
const selected = new Set(selectedKeys)
const totalEnroll = groups.reduce((sum, g) => sum + g.enroll, 0)
const totalObserved = groups.reduce((sum, g) => sum + g.observed, 0)
const totalRate = totalEnroll > 0 ? (totalObserved / totalEnroll) * 1000 : 0
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={cardTitle}>Reported arrests and enrollment, 2021–22</h3>
<p style={{ fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: '0 0 var(--space-2)' }}>
Check a group to show its modelled arrest-rate distribution in the chart below.
</p>
{pooled && (
<PoolingBanner autoPooled={autoPooled} totalObserved={totalObserved} onPooledChange={onPooledChange} />
)}
{selectionIsEnrollmentFallback && (
<p style={noteStyle}>
No group in this district reported an arrest in 2021–22, so the two largest groups by
enrollment are shown by default.
</p>
)}
<div style={{ overflowX: 'auto' }}>
<table style={tableStyle}>
<caption style={{ captionSide: 'bottom', textAlign: 'left', paddingTop: 'var(--space-1)', fontSize: '0.75rem', color: 'var(--cv-ink-3)' }}>
Students counted are those in the four modelled race groups (American Indian / Alaska
Native, Black, Hispanic, White) — not the district&rsquo;s total enrollment, which also
includes groups this model does not estimate.
</caption>
<thead>
<tr>
<th scope="col" style={{ ...thStyle, textAlign: 'left' }}>Student group</th>
<th scope="col" style={thStyle}>Students</th>
<th scope="col" style={thStyle}>Observed arrests</th>
<th scope="col" style={thStyle}>Rate per 1,000</th>
</tr>
</thead>
<tbody>
{groups.map((g) => (
<GroupRow
key={g.key}
group={g}
checked={selected.has(g.key)}
onToggle={() => onToggleGroup(g.key)}
/>
))}
</tbody>
<tfoot>
<tr>
<th scope="row" style={{ ...tdStyle, textAlign: 'left', fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
All modelled groups
</th>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{totalEnroll.toLocaleString()}
</td>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{totalObserved.toLocaleString()}
</td>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{formatRate(totalRate)}
</td>
</tr>
</tfoot>
</table>
</div>
{!pooled && (
<label style={{ ...toggleLabel, marginTop: 'var(--space-2)' }}>
<input type="checkbox" checked={false} onChange={() => onPooledChange(true)} />
<span>Combine Female and Male within each race</span>
</label>
)}
</div>
)
}
function GroupRow({ group, checked, onToggle }) {
const zero = group.observed === 0
const inputId = `group-toggle-${group.key}`
return (
<tr style={{ opacity: zero ? 0.62 : 1 }}>
<td style={{ ...tdStyle, textAlign: 'left' }}>
<label htmlFor={inputId} style={{ display: 'flex', alignItems: 'center', gap: '0.5rem', cursor: 'pointer' }}>
<input id={inputId} type="checkbox" checked={checked} onChange={onToggle} />
<span
aria-hidden="true"
style={{
width: '11px',
height: '11px',
borderRadius: '3px',
flexShrink: 0,
// Hollow when deselected: the swatch tracks whether that curve is
// on screen, so the table stays a truthful legend.
background: checked ? raceColor(group.race) : 'transparent',
border: `2px solid ${raceColor(group.race)}`,
}}
/>
<span>{group.label}</span>
</label>
</td>
<td style={numStyle}>{group.enroll.toLocaleString()}</td>
<td style={numStyle}>{group.observed.toLocaleString()}</td>
<td style={numStyle}>{formatRate(group.rate)}</td>
</tr>
)
}
function PoolingBanner({ autoPooled, totalObserved, onPooledChange }) {
return (
<div style={bannerStyle}>
<p style={{ margin: 0, fontSize: '0.82rem' }}>
{autoPooled ? (
<>
This district reported {totalObserved.toLocaleString()} arrests in total — fewer than{' '}
{POOL_BY_SEX_ARREST_THRESHOLD} — so Female and Male students are combined within each
race to give each estimate more data to stand on.
</>
) : (
<>Female and Male students are combined within each race.</>
)}
</p>
<label style={toggleLabel}>
<input type="checkbox" checked onChange={() => onPooledChange(false)} />
<span>Combined — uncheck to show Female and Male separately</span>
</label>
</div>
)
}
function formatRate(rate) {
if (!(rate > 0)) return '0.0'
return rate.toLocaleString(undefined, { minimumFractionDigits: 1, maximumFractionDigits: 1 })
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const tableStyle = {
width: '100%',
borderCollapse: 'collapse',
fontSize: '0.85rem',
fontVariantNumeric: 'tabular-nums',
}
const thStyle = {
textAlign: 'right',
padding: '0.35rem 0.5rem',
borderBottom: '1px solid var(--cv-rule-strong)',
fontSize: '0.72rem',
fontWeight: 600,
textTransform: 'uppercase',
letterSpacing: '0.06em',
color: 'var(--cv-ink-3)',
whiteSpace: 'nowrap',
}
const tdStyle = {
padding: '0.35rem 0.5rem',
borderBottom: '1px solid var(--cv-rule)',
}
const numStyle = { ...tdStyle, textAlign: 'right', whiteSpace: 'nowrap' }
const bannerStyle = {
background: 'var(--cv-paper-2)',
border: '1px solid var(--cv-rule)',
borderLeft: '3px solid var(--cv-accent)',
borderRadius: 'var(--radius-md)',
padding: 'var(--space-1) var(--space-2)',
marginBottom: 'var(--space-2)',
display: 'flex',
flexDirection: 'column',
gap: '0.4rem',
}
const toggleLabel = {
display: 'inline-flex',
alignItems: 'center',
gap: '0.45rem',
fontSize: '0.78rem',
color: 'var(--cv-ink-2)',
cursor: 'pointer',
}
const noteStyle = {
fontSize: '0.78rem',
fontStyle: 'italic',
color: 'var(--cv-ink-3)',
margin: '0 0 var(--space-2)',
}
+10 -1
View File
@@ -119,7 +119,16 @@ export async function fetchStateEstimates(state, options = {}) {
return apiFetch(`/states/${state}${qs ? `?${qs}` : ''}`)
}
/** GET /estimates?state=XX&year=Y — all districts in a state for "interesting" suggestions */
/**
* GET /estimates?state=XX&year=Y — every estimate row in a state.
*
* Not used by the running app. The rows come back `ORDER BY LEAID, RACE, SEX`
* at eight per district, so any `limit` short of the state's full row count
* selects the lowest-LEAID districts rather than a meaningful sample —
* `DistrictSearch` reads the pre-ranked `public/data/top_districts.json`
* fixture instead. Kept for scripts and ad-hoc use; page it with `meta.total`
* as `scripts/build-top-districts.mjs` does.
*/
export async function fetchStateDistricts(state, year = '21-22', limit = 500) {
return apiFetch(`/estimates?state=${state}&year=${year}&limit=${limit}`)
}
+192 -83
View File
@@ -1,19 +1,68 @@
import { useEffect, useState } from 'react'
import { useEffect, useMemo, useState } from 'react'
import { getDb } from '../utils/duckdbClient.js'
import { groupKey } from '../utils/drawGroups.js'
import { groupKey, isCompleteDrawSet } from '../utils/drawGroups.js'
const HF_BASE = 'https://huggingface.co/datasets/civilytics/crdc-school-arrest-rates/resolve/main/parquet'
const HF_DATASET = 'civilytics/crdc-school-arrest-rates'
const HF_BASE = `https://huggingface.co/datasets/${HF_DATASET}/resolve/main/parquet`
const HF_TREE = `https://huggingface.co/api/datasets/${HF_DATASET}/tree/main/parquet`
// Module-level cache: one registered duckdb-wasm file buffer per
// Module-level cache: the registered duckdb-wasm file buffers for one
// (model, year, state) shard, shared across every component instance and
// district navigated to in this browser session. See Global Constraints —
// in-memory only, no persistence across page loads.
// in-memory only, no persistence across page loads. Four models is simply four
// cache entries; nothing else about this cache changes when comparing specs.
const shardCache = new Map()
// Runaway guard on a malformed directory listing, not an expected limit.
const MAX_SHARD_PARTS = 64
function shardKey(model, year, state) {
return `${model}__${year}__${state}`
}
function shardDir(model, year, state) {
return `model_id=${model}/YEAR=${year}/LEA_STATE=${state}`
}
/**
* Lists the parquet parts published for one (model, year, state).
*
* A state's draws are split across `data_0.parquet`, `data_1.parquet`, … and
* the part count varies by state: Nevada is one file, California is eight
* (6.2MB in total, of which data_0 is 37KB). Reading only data_0 covers 11 of
* California's 1,715 districts and makes every other CA district look absent
* from the published data, silently falling back to the approximation.
*
* The directory listing is used rather than probing `data_N` until a 404
* because a 404 is logged as a console error by the browser's network layer no
* matter how cleanly the fetch handles it — and a red error on every load is
* indistinguishable from a real one. Probing remains the fallback if the
* listing API is unavailable or changes shape.
*/
async function listShardParts(model, year, state) {
const dir = shardDir(model, year, state)
try {
const res = await fetch(`${HF_TREE}/${dir}`)
if (!res.ok) throw new Error(`tree listing HTTP ${res.status}`)
const entries = await res.json()
const parts = entries
.filter((e) => e?.type === 'file' && /\/data_\d+\.parquet$/.test(e.path || ''))
.map((e) => ({ part: Number(e.path.match(/data_(\d+)\.parquet$/)[1]), path: e.path }))
.sort((a, b) => a.part - b.part)
if (parts.length > 0) return parts.map((p) => p.part)
throw new Error('tree listing contained no parquet parts')
} catch (err) {
console.warn(`useDrawDistribution: falling back to sequential part probing for ${dir}:`, err.message)
const parts = []
for (let part = 0; part < MAX_SHARD_PARTS; part++) {
const res = await fetch(`${HF_BASE}/${dir}/data_${part}.parquet`, { method: 'HEAD' })
if (!res.ok) break
parts.push(part)
}
return parts
}
}
function ensureShardRegistered(db, model, year, state) {
const key = shardKey(model, year, state)
if (!shardCache.has(key)) {
@@ -21,13 +70,24 @@ function ensureShardRegistered(db, model, year, state) {
key,
(async () => {
try {
const url = `${HF_BASE}/model_id=${model}/YEAR=${year}/LEA_STATE=${state}/data_0.parquet`
const res = await fetch(url)
if (!res.ok) throw new Error(`Failed to fetch draw shard: HTTP ${res.status}`)
const buffer = new Uint8Array(await res.arrayBuffer())
const fileName = `${key}.parquet`
await db.registerFileBuffer(fileName, buffer)
return fileName
const dir = shardDir(model, year, state)
const parts = await listShardParts(model, year, state)
if (parts.length === 0) {
throw new Error(`No draw shard published for ${state}/${year}/${model}`)
}
// Parts are independent files, so fetch them together rather than
// walking them one at a time — California is eight round trips.
const fileNames = await Promise.all(
parts.map(async (part) => {
const res = await fetch(`${HF_BASE}/${dir}/data_${part}.parquet`)
if (!res.ok) throw new Error(`Failed to fetch draw shard part ${part}: HTTP ${res.status}`)
const buffer = new Uint8Array(await res.arrayBuffer())
const fileName = `${key}__${part}.parquet`
await db.registerFileBuffer(fileName, buffer)
return fileName
}),
)
return fileNames
} catch (err) {
// Don't let a transient failure (network blip, HF outage) poison the
// cache forever — remove the rejected entry so the next caller for
@@ -42,100 +102,145 @@ function ensureShardRegistered(db, model, year, state) {
}
/**
* Fetches real posterior draws for one district/model/year from the Hugging
* Face parquet dataset via duckdb-wasm, converts predicted counts to
* rate-per-1,000 using each group's stu_enroll (not present in the draws
* table itself — joined here from data this app already has), and returns
* them keyed by "RACE_SEX" (see `groupKey` in utils/drawGroups.js).
* Reads one model's shard and returns that district's predicted counts keyed by
* group and indexed by draw.
*
* `status` is an ANY-group signal: 'ready' means at least one group came back
* with real draws, not that every requested group did. Groups can be missing
* individually (falsy enrollment, or absent from the parquet shard), so a caller
* that wants to claim "these are all real draws" must check its own rendered
* groups against `drawsByGroup` — use `hasDrawsForAll` from utils/drawGroups.js.
*
* @param {{leaid: string, state: string, model: string, year: string,
* groups: Array<{race: string, sex: string, stuEnroll: number}>}} params
* @returns {{status: 'loading'|'ready'|'error', drawsByGroup: Record<string, number[]> | null}}
* Indexing by `draw_id - 1` rather than by push order means DuckDB's row
* ordering is irrelevant, and it makes a missing draw a hole instead of a
* silently shorter array — which `isCompleteDrawSet` then rejects.
*/
export function useDrawDistribution({ leaid, state, model, year, groups }) {
const [status, setStatus] = useState('loading')
const [drawsByGroup, setDrawsByGroup] = useState(null)
async function fetchModelCounts(db, { leaid, state, model, year }) {
const fileNames = await ensureShardRegistered(db, model, year, state)
let conn
try {
conn = await db.connect()
// read_parquet over the full part list — a district lives in exactly one
// part, and which one is not predictable from its LEAID.
const fileList = fileNames.map((f) => `'${f}'`).join(', ')
const stmt = await conn.prepare(
`SELECT RACE, SEX, draw_id, pred FROM read_parquet([${fileList}]) WHERE LEAID = ?`,
)
const table = await stmt.query(leaid)
await stmt.close()
const rows = table.toArray().map((r) => r.toJSON())
// groups is typically a fresh array literal every render; derive a stable
const raw = {}
let nDraws = 0
for (const row of rows) {
const drawId = Number(row.draw_id)
const pred = Number(row.pred)
if (!Number.isFinite(drawId) || drawId < 1 || !Number.isFinite(pred)) continue
const key = groupKey(row.RACE, row.SEX)
;(raw[key] ??= [])[drawId - 1] = pred
if (drawId > nDraws) nDraws = drawId
}
// A group present for only part of the draw set must fall back, not render
// a short draw set: its density and interval would be computed off a
// biased subsample and look identical on screen to a complete one.
const counts = {}
for (const [key, arr] of Object.entries(raw)) {
if (isCompleteDrawSet(arr, nDraws)) counts[key] = arr
else console.warn('useDrawDistribution: dropping incomplete draw set for', { leaid, model, key, got: arr.length, want: nDraws })
}
return Object.keys(counts).length > 0 ? { counts, nDraws } : null
} finally {
if (conn) await conn.close()
}
}
/**
* Fetches real posterior predictive draws for one district from the Hugging
* Face parquet dataset via duckdb-wasm, for one or more model specifications at
* once, and returns **predicted counts indexed by draw** — not rates.
*
* Counts rather than rates is what unlocks the rest of the app: sex pooling has
* to sum numerators and denominators separately (see `utils/pooling.js`), and
* a between-group difference has to be taken at a common draw index. Callers
* divide by their own enrollment, which is why this hook no longer takes a
* `groups` argument at all.
*
* A caveat worth carrying into any caption: `draw_id` is renumbered 1–500 per
* write batch upstream, and a district's groups land in different batches, so
* draws are **not** paired parameter draws across groups — the pairing is
* effectively independent, dominated by `posterior_predict` observation noise.
* Published Fig 7 has the same property. Say "posterior predictive draws".
*
* `status` is an ANY-model, ANY-group signal: 'ready' means at least one model
* returned at least one complete group. A caller claiming "these are all real
* draws" must check its own rendered groups with `hasDrawsForAll`.
*
* @param {{leaid: string, state: string, models: string[], year: string}} params
* @returns {{status: 'loading'|'ready'|'error',
* byModel: Record<string, {counts: Record<string, number[]>, nDraws: number} | null> | null,
* nDraws: number}} `nDraws` is the largest draw count among models that came
* back; per-model counts live in `byModel[model].nDraws`.
*/
export function useDrawDistribution({ leaid, state, models, year }) {
const [status, setStatus] = useState('loading')
const [byModel, setByModel] = useState(null)
// `models` is typically a fresh array literal every render; derive a stable
// primitive so the effect only re-runs when its actual content changes.
const groupsSignature = (groups || []).map((g) => `${g.race}:${g.sex}:${g.stuEnroll}`).join(',')
const modelsSignature = (models || []).filter(Boolean).join(',')
useEffect(() => {
let cancelled = false
// Reset BEFORE the input guard below, not after. Clearing any previous
// model/district's draws has to happen on every input change, including the
// ones that have nothing to fetch. Without this ordering, a chart that
// varies `model` across renders (e.g. RateDensityRidgeline's dropdown) and
// lands on a model whose `groups` is empty (that model's upstream fetch
// failed) would keep reporting 'ready' and keep handing back the *previous*
// model's real draws — keyed by the same RACE_SEX strings — under the newly
// selected model's summary stats, silently mixing two models' data.
// varies its model selection across renders and lands on a model whose
// fetch fails would keep reporting 'ready' and keep handing back the
// *previous* model's real draws — keyed by the same group strings — under
// the newly selected model's label, silently mixing two models' data.
setStatus('loading')
setDrawsByGroup(null)
setByModel(null)
const modelList = modelsSignature ? modelsSignature.split(',') : []
// Nothing to fetch: stay in 'loading' with no draws, which every consumer
// already treats as "fall back to the approximation". No cleanup needed —
// nothing async was started.
if (!leaid || !state || !model || !year || !groups?.length) return
if (!leaid || !state || !year || modelList.length === 0) return
async function run() {
let conn
try {
const db = await getDb()
const fileName = await ensureShardRegistered(db, model, year, state)
conn = await db.connect()
const stmt = await conn.prepare(`SELECT RACE, SEX, pred FROM read_parquet('${fileName}') WHERE LEAID = ?`)
const table = await stmt.query(leaid)
await stmt.close()
const rows = table.toArray().map((r) => r.toJSON())
// Per-model try/catch: one shard 404ing (or one model missing for this
// state) must not blank out the models that did load.
const results = await Promise.all(
modelList.map(async (model) => {
try {
return [model, await fetchModelCounts(db, { leaid, state, model, year })]
} catch (err) {
console.error(`useDrawDistribution: model ${model} failed:`, err)
return [model, null]
}
}),
)
const enrollByGroup = {}
for (const g of groups) enrollByGroup[groupKey(g.race, g.sex)] = g.stuEnroll || 0
const byGroup = {}
for (const row of rows) {
const key = groupKey(row.RACE, row.SEX)
const enroll = enrollByGroup[key]
if (!enroll) continue
const rate = (Number(row.pred) / enroll) * 1000
;(byGroup[key] ??= []).push(rate)
if (cancelled) return
const next = Object.fromEntries(results)
const anyReady = Object.values(next).some((v) => v !== null)
if (anyReady) {
setByModel(next)
setStatus('ready')
} else {
console.warn('useDrawDistribution: no usable draws for', { leaid, state, year, models: modelList })
setByModel(null)
setStatus('error')
}
// A successful query with zero matching rows (this district isn't in
// the draws shard, or none of its rows matched a known group) is not
// "ready" — there's no real data to show, so treat it like a failure
// and let the caller fall back, rather than silently claiming real
// draws while every group actually uses the fitted approximation.
const hasDraws = Object.values(byGroup).some((draws) => draws.length > 0)
if (!cancelled) {
if (hasDraws) {
setDrawsByGroup(byGroup)
setStatus('ready')
} else {
console.warn('useDrawDistribution: query succeeded but returned no usable draws for', { leaid, model, year, state })
setDrawsByGroup(null)
setStatus('error')
}
}
} catch (err) {
// Only reached when the shared duckdb engine itself fails to load.
console.error('useDrawDistribution failed:', err)
if (!cancelled) {
// Belt-and-suspenders alongside the setDrawsByGroup(null) at the
// top of this effect: a failed fetch/query must never leave a
// *previous* model's real draws in place under the newly-selected
// model's status/summary stats.
setDrawsByGroup(null)
// Belt-and-suspenders alongside the setByModel(null) at the top of
// this effect: a failed load must never leave a *previous* model's
// real draws in place under the newly-selected model's label.
setByModel(null)
setStatus('error')
}
} finally {
if (conn) await conn.close()
}
}
@@ -143,8 +248,12 @@ export function useDrawDistribution({ leaid, state, model, year, groups }) {
return () => {
cancelled = true
}
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [leaid, state, model, year, groupsSignature])
}, [leaid, state, year, modelsSignature])
return { status, drawsByGroup }
const nDraws = useMemo(
() => Math.max(0, ...Object.values(byModel || {}).map((v) => v?.nDraws || 0)),
[byModel],
)
return { status, byModel, nDraws }
}
+6
View File
@@ -270,3 +270,9 @@ code, .mono { font-family: var(--font-mono); font-feature-settings: "tnum" 1, "z
/* Responsive */
@media (max-width: 680px) { .cv-footer { flex-direction: column; text-align: center; } }
/* Spinner used while opt-in multi-shard draw fetches are in flight */
@keyframes cv-spin { to { transform: rotate(360deg); } }
@media (prefers-reduced-motion: reduce) {
@keyframes cv-spin { to { transform: none; } }
}
+66
View File
@@ -0,0 +1,66 @@
/**
* Agresti–Coull approximate interval for a rare-event count.
*
* Direct port of `agresti_coull()` in crdc-arrests/R/paper_figures.R:219-237 —
* the frequentist point range drawn beside the posterior densities in the white
* paper's Figs 6 and 7. Keeping this a faithful port is the whole point: the
* app's error bars have to be the *same* interval the paper published, so the
* two can be compared directly.
*
* Two consequences of that faithfulness, both deliberate:
*
* 1. The bounds are on the **count** scale, not the proportion scale — the R
* function multiplies back up by the adjusted denominator. Divide by
* enrollment yourself to plot per-1,000.
* 2. `lower` can be **negative** for very small numerators (e.g. 1 arrest in 53
* students gives -0.32). Clamp for display at the call site; do not clamp
* here, or this stops matching the published figures.
*
* The R original returns an unnamed vector in the order
* `c(ci_upper, ci_lower, sd, phat_se, phat)` — upper *first*, which is easy to
* transcribe backwards. This returns a named object instead.
*/
import { probit } from './distributionApprox.js'
/**
* @param {number} numerator - observed events (arrests)
* @param {number} denominator - trials (students enrolled)
* @param {number} [confidenceLevel=0.95] - e.g. 0.95 for a 95% interval
* @returns {{upper: number, lower: number, sd: number, se: number, phat: number}}
* `upper`/`lower`/`sd` are counts. In the zero-numerator branch the R
* original also reports `phat` as a count (the interval midpoint) rather than
* a proportion; that quirk is preserved.
*/
export function agrestiCoull(numerator, denominator, confidenceLevel = 0.95) {
const adjStar = probit(1 - (1 - confidenceLevel) / 2)
if (numerator > 0) {
const numStar = numerator + adjStar
const denomStar = denominator + 2 * adjStar
const phat = numStar / denomStar
const se = Math.sqrt((phat / denomStar) * (1 - phat))
return {
upper: (phat + adjStar * se) * denomStar,
lower: (phat - adjStar * se) * denomStar,
sd: se * denomStar,
se,
phat,
}
}
// Zero events: the rule of three. The R original writes the bound as
// denominator * (-log(1 - level) / denominator), where the denominator
// cancels — so the upper bound is -log(1 - level) ≈ 3 at 95% regardless of
// how many students were enrolled. Kept in the cancelled form so the value
// is identical rather than merely close.
const upper = -Math.log(1 - confidenceLevel)
const midpoint = (upper + 0) / 2
return {
upper,
lower: 0,
sd: midpoint / confidenceLevel,
se: 0,
phat: midpoint,
}
}
+109
View File
@@ -0,0 +1,109 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { agrestiCoull } from './agrestiCoull.js'
/**
* Reference values produced by the R original, `agresti_coull()` in
* crdc-arrests/R/paper_figures.R:219-237, printed at 15 significant digits:
*
* agresti_coull(15, 499, 0.95) c(24.89431440404100, 9.02561356503911,
* 4.04821235598516, 0.00804941727469993,
* 0.03372299056239180)
* agresti_coull(0, 53, 0.95) c(2.99573227355399, 0, 1.57670119660736,
* 0, 1.49786613677699)
* agresti_coull(1, 53, 0.95) c(6.24314599198645, -0.32321802290634,
* 1.67512364173205, 0.02942947578293,
* 0.05200224403214)
* agresti_coull(100, 148928, 0.95) c(121.743969120453, 82.1759588486273,
* 10.0940656522091, 6.77763749845643e-05,
* 6.84607866694077e-04)
* agresti_coull(3, 200, 0.90) c(8.14909680812749, 1.14061044577545,
* 2.13042858267619, 0.01047976610058,
* 0.02284844466400)
*
* R's qnorm is exact to double precision; this port uses Acklam's rational
* probit (relative error < 1.15e-9), so equality is asserted to 1e-7 relative.
*/
const REL_TOL = 1e-7
function assertClose(actual, expected, label) {
const scale = Math.max(Math.abs(expected), 1e-9)
assert.ok(
Math.abs(actual - expected) / scale < REL_TOL,
`${label}: expected ${expected}, got ${actual}`,
)
}
function assertMatchesR(result, [upper, lower, sd, se, phat], label) {
assertClose(result.upper, upper, `${label} upper`)
assertClose(result.lower, lower, `${label} lower`)
assertClose(result.sd, sd, `${label} sd`)
assertClose(result.se, se, `${label} se`)
assertClose(result.phat, phat, `${label} phat`)
}
test('agrestiCoull: matches R for a typical rare-event cell', () => {
assertMatchesR(
agrestiCoull(15, 499),
[24.894314404041, 9.02561356503911, 4.04821235598516, 0.00804941727469993, 0.0337229905623918],
'ac(15, 499, 0.95)',
)
})
test('agrestiCoull: matches R for a large district cell', () => {
assertMatchesR(
agrestiCoull(100, 148928),
[121.743969120453, 82.1759588486273, 10.0940656522091, 6.77763749845643e-5, 6.84607866694077e-4],
'ac(100, 148928, 0.95)',
)
})
test('agrestiCoull: matches R at a non-default confidence level', () => {
assertMatchesR(
agrestiCoull(3, 200, 0.9),
[8.14909680812749, 1.14061044577545, 2.13042858267619, 0.0104797661005796, 0.0228484446639996],
'ac(3, 200, 0.90)',
)
})
test('agrestiCoull: zero numerator uses the rule of three', () => {
// -log(1 - 0.95) = 2.9957…, the classic "rule of three" upper bound for zero
// events. The R original writes it as denominator * (-log(1-cl)/denominator),
// which cancels — the bound does not depend on the denominator at all.
const r = agrestiCoull(0, 53)
assertMatchesR(r, [2.99573227355399, 0, 1.57670119660736, 0, 1.49786613677699], 'ac(0, 53, 0.95)')
assert.equal(r.lower, 0)
})
test('agrestiCoull: the zero-numerator upper bound ignores the denominator', () => {
assert.equal(agrestiCoull(0, 53).upper, agrestiCoull(0, 500000).upper)
})
test('agrestiCoull: keeps R\'s negative lower bound for a single event', () => {
// Faithful to the R original: with numerator = 1 the lower bound goes below
// zero. Callers clamp for display — do NOT clamp here, or this port silently
// stops matching the published figures.
const r = agrestiCoull(1, 53)
assertMatchesR(
r,
[6.24314599198645, -0.323218022906343, 1.67512364173205, 0.029429475782928, 0.0520022440321421],
'ac(1, 53, 0.95)',
)
assert.ok(r.lower < 0, 'lower bound should be negative for numerator = 1, n = 53')
})
test('agrestiCoull: bounds are on the count scale, not the proportion scale', () => {
const r = agrestiCoull(15, 499)
assert.ok(r.upper > 15 && r.lower < 15, 'the interval should bracket the observed count')
})
test('agrestiCoull: defaults to 95%', () => {
assert.deepEqual(agrestiCoull(15, 499), agrestiCoull(15, 499, 0.95))
})
test('agrestiCoull: a wider confidence level gives a wider interval', () => {
const narrow = agrestiCoull(15, 499, 0.8)
const wide = agrestiCoull(15, 499, 0.99)
assert.ok(wide.upper > narrow.upper)
assert.ok(wide.lower < narrow.lower)
})
+18 -2
View File
@@ -31,10 +31,26 @@ export function raceColor(race) {
return RACE_COLORS[race] || REFERENCE_GRAY
}
// Both label helpers take an optional `sex`: omitting it names a sex-pooled
// group (race alone), mirroring `groupKey(race)` in utils/drawGroups.js. Don't
// let a missing sex fall through to "Male".
export function groupLabel(race, sex) {
return `${RACE_LABELS[race] || race} ${sex === 'F' ? 'Female' : 'Male'}`
const race_ = RACE_LABELS[race] || race
if (!sex) return race_
return `${race_} ${sex === 'F' ? 'Female' : 'Male'}`
}
export function shortGroupLabel(race, sex) {
return `${SHORT_RACE_LABEL[race] || race} ${sex}`
const race_ = SHORT_RACE_LABEL[race] || race
return sex ? `${race_} ${sex}` : race_
}
// For labels that appear mid-sentence ("…the Black male arrest rate exceeds the
// White male rate"). The race is a proper noun and keeps its capital; only the
// sex word is lowercased. Don't reach for .toLowerCase() on groupLabel() — it
// turns "Black Male" into "black male".
export function sentenceGroupLabel(race, sex) {
const race_ = RACE_LABELS[race] || race
if (!sex) return race_
return `${race_} ${sex === 'F' ? 'female' : 'male'}`
}
+114
View File
@@ -0,0 +1,114 @@
/**
* Chooses an honest visual profile for one group's posterior predictive draws.
*
* In a sparse district the posterior predictive is a *discrete count*
* distribution, not a smooth one. Carson City NV (3200390) has a 53-student
* AI/AN female cell where a single arrest is 18.9 per 1,000: the draws take
* four distinct values, and a Gaussian KDE renders those four spikes as a lumpy
* smear that reads as a rendering bug rather than as a finding about the data.
*
* So: few distinct values → draw the actual probability mass at each achievable
* rate. Many → the smooth KDE, delegated to `kde.js` unchanged (its bandwidth
* clamp was tuned for exactly these zero-inflated posteriors — see
* BANDWIDTH_FLOOR_DIVISOR there; do not re-derive it here).
*/
import { kdeCurve } from './kde.js'
import { toRates } from './pooling.js'
/**
* At or below this many distinct predicted counts, the draws are shown as
* discrete mass rather than smoothed.
*
* 12 is comfortably above the 3–5 distinct values a genuinely sparse cell
* produces and comfortably below the ~40+ a district with real arrest volume
* produces, so the switch happens well away from either regime rather than
* flickering at the boundary.
*/
export const MASS_MAX_DISTINCT = 12
/**
* @param {number[] | null | undefined} counts - predicted counts, indexed by draw
* @param {number | null | undefined} enroll - students in the group
* @param {{min?: number, max?: number, n?: number}} [domain] - the x-range the
* profile will be drawn over, in rate per 1,000. Passed straight to
* `kdeCurve`, whose bandwidth clamp is relative to this width.
* @returns {{kind: 'kde'|'mass', points: Array<{x: number, y: number}>,
* maxY: number, step: number} | null}
* `null` when there is nothing to draw (no draws, or no denominator).
* For `kind: 'mass'`, `y` is a probability and `step` is the spacing between
* achievable rates — the bar width for a filled staircase. For `kind: 'kde'`,
* `y` is a density and `step` is 0.
* `maxY` is the profile's own peak: because the two kinds carry different y
* units, a chart overlaying several groups must normalize each profile by its
* own `maxY` rather than by a shared maximum.
*/
export function densityProfile(counts, enroll, domain = {}) {
const rates = toRates(counts, enroll)
if (!rates.length) return null
// Counts are integers, so consecutive achievable rates are exactly one
// student-rate apart — pass that explicitly rather than inferring it from the
// gaps actually observed, which overstates the bar width when (say) only
// counts 0 and 3 appear.
return rateProfile(rates, domain, 1000 / enroll)
}
/**
* The same choice made directly on a set of rates, for quantities that are
* already differences rather than count/denominator pairs (see
* `utils/groupDifference.js`). A difference of two discrete posteriors is
* itself discrete, and deserves the same honesty.
*
* @param {number[] | null | undefined} rates - values on the plotted scale
* @param {{min?: number, max?: number, n?: number}} [domain]
* @param {number} [explicitStep] - known spacing between achievable values;
* inferred from the smallest observed gap when omitted.
* @returns {{kind: 'kde'|'mass', points: Array<{x: number, y: number}>,
* maxY: number, step: number} | null}
*/
export function rateProfile(rates, domain = {}, explicitStep = 0) {
if (!rates?.length) return null
const { min = 0, max, n = 60 } = domain
// Round before tallying so two draws that differ only in floating-point noise
// count as one achievable value rather than two.
const tally = new Map()
for (const r of rates) {
const key = Math.round(r * 1e9) / 1e9
tally.set(key, (tally.get(key) || 0) + 1)
}
if (tally.size <= MASS_MAX_DISTINCT) {
const total = rates.length
const points = [...tally.entries()]
.map(([x, freq]) => ({ x, y: freq / total }))
.sort((a, b) => a.x - b.x)
return {
kind: 'mass',
points,
maxY: Math.max(...points.map((p) => p.y)),
step: explicitStep > 0 ? explicitStep : inferStep(points, max, min),
}
}
const points = kdeCurve(rates, { min, max, n })
return {
kind: 'kde',
points,
maxY: Math.max(...points.map((p) => p.y)),
step: 0,
}
}
/** Smallest gap between achievable values; a visible default for a single point. */
function inferStep(points, max, min) {
let smallest = Infinity
for (let i = 1; i < points.length; i++) {
const gap = points[i].x - points[i - 1].x
if (gap > 0 && gap < smallest) smallest = gap
}
if (Number.isFinite(smallest)) return smallest
const width = Number.isFinite(max) && max > min ? max - min : 1
return width / 40
}
+135
View File
@@ -0,0 +1,135 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { MASS_MAX_DISTINCT, densityProfile, rateProfile } from './densityProfile.js'
import { kdeCurve } from './kde.js'
import { toRates } from './pooling.js'
/** counts whose distinct-value count is exactly `k` (values 0…k-1, padded). */
function countsWithDistinct(k, total = 500) {
const out = []
for (let i = 0; i < total; i++) out.push(i % k)
return out
}
test('densityProfile: discrete posterior returns a mass profile', () => {
// 500 draws over 4 achievable counts is the Carson City AI/AN case: a KDE
// renders it as a lumpy smear that reads as a rendering bug.
const profile = densityProfile(countsWithDistinct(4), 53, { max: 100 })
assert.equal(profile.kind, 'mass')
})
test('densityProfile: mass points are probabilities at achievable rates', () => {
const profile = densityProfile([0, 0, 0, 1], 500, { max: 10 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [
{ x: 0, y: 0.75 },
{ x: 2, y: 0.25 },
])
})
test('densityProfile: mass probabilities sum to 1', () => {
const profile = densityProfile([0, 0, 1, 2, 2, 5], 1000, { max: 10 })
const total = profile.points.reduce((sum, p) => sum + p.y, 0)
assert.ok(Math.abs(total - 1) < 1e-12, `probabilities summed to ${total}`)
})
test('densityProfile: mass points are sorted by rate ascending', () => {
const profile = densityProfile([5, 0, 3, 1], 1000, { max: 10 })
const xs = profile.points.map((p) => p.x)
assert.deepEqual(xs, [...xs].sort((a, b) => a - b))
})
test('densityProfile: mass carries the achievable-rate step for staircase width', () => {
// Counts are integers, so achievable rates are spaced 1000/enroll apart.
// The chart needs that width to draw a bar rather than a hairline.
const profile = densityProfile([0, 1], 250, { max: 10 })
assert.equal(profile.step, 4)
})
test('densityProfile: a fully degenerate draw set is a single mass point', () => {
// 500 identical zeros is common in small districts. This is the case a
// Gaussian KDE turns into a delta spike.
const profile = densityProfile(new Array(500).fill(0), 4073, { max: 20 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [{ x: 0, y: 1 }])
})
test('densityProfile: switches to KDE above the distinct-value cutoff', () => {
const atCutoff = densityProfile(countsWithDistinct(MASS_MAX_DISTINCT), 1000, { max: 50 })
const aboveCutoff = densityProfile(countsWithDistinct(MASS_MAX_DISTINCT + 1), 1000, { max: 50 })
assert.equal(atCutoff.kind, 'mass')
assert.equal(aboveCutoff.kind, 'kde')
})
test('densityProfile: KDE branch delegates to kdeCurve over the given domain', () => {
// The bandwidth clamp in kde.js was tuned for exactly these zero-inflated
// posteriors — this must delegate, not re-derive.
const counts = countsWithDistinct(40)
const enroll = 1000
const profile = densityProfile(counts, enroll, { min: 0, max: 50, n: 60 })
const expected = kdeCurve(toRates(counts, enroll), { min: 0, max: 50, n: 60 })
assert.equal(profile.kind, 'kde')
assert.deepEqual(profile.points, expected)
})
test('densityProfile: reports the profile peak for per-group normalization', () => {
const profile = densityProfile([0, 0, 0, 1], 500, { max: 10 })
assert.equal(profile.maxY, 0.75)
const kde = densityProfile(countsWithDistinct(40), 1000, { max: 50 })
assert.equal(kde.maxY, Math.max(...kde.points.map((p) => p.y)))
})
test('densityProfile: null for unusable input rather than an empty curve', () => {
assert.equal(densityProfile([], 500, { max: 10 }), null)
assert.equal(densityProfile(undefined, 500, { max: 10 }), null)
assert.equal(densityProfile([0, 1], 0, { max: 10 }), null)
assert.equal(densityProfile([0, 1], undefined, { max: 10 }), null)
})
test('densityProfile: does not mutate its input counts', () => {
const counts = [3, 1, 2]
densityProfile(counts, 1000, { max: 10 })
assert.deepEqual(counts, [3, 1, 2])
})
// ——— rateProfile (used for already-differenced quantities) ———
test('rateProfile: mass profile straight from rate values', () => {
const profile = rateProfile([-1, -1, 0, 2], { min: -5, max: 5 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [
{ x: -1, y: 0.5 },
{ x: 0, y: 0.25 },
{ x: 2, y: 0.25 },
])
})
test('rateProfile: infers the staircase step from the smallest observed gap', () => {
assert.equal(rateProfile([0, 0.5, 2], { min: 0, max: 5 }).step, 0.5)
})
test('rateProfile: falls back to a visible step for a single achievable value', () => {
const profile = rateProfile([3, 3, 3], { min: 0, max: 40 })
assert.ok(profile.step > 0, 'a single mass point still needs a drawable width')
})
test('rateProfile: an explicit step wins over the inferred one', () => {
assert.equal(rateProfile([0, 3], { min: 0, max: 5 }, 0.25).step, 0.25)
})
test('rateProfile: merges values differing only by floating-point noise', () => {
const profile = rateProfile([0.1 + 0.2, 0.3, 0.3], { min: 0, max: 1 })
assert.equal(profile.points.length, 1)
assert.equal(profile.points[0].y, 1)
})
test('rateProfile: switches to KDE above the distinct-value cutoff', () => {
const many = Array.from({ length: 300 }, (_, i) => (i % (MASS_MAX_DISTINCT + 1)) * 0.7)
assert.equal(rateProfile(many, { min: 0, max: 10 }).kind, 'kde')
})
test('rateProfile: null for an empty or missing set', () => {
assert.equal(rateProfile([], { max: 5 }), null)
assert.equal(rateProfile(undefined, { max: 5 }), null)
})
+19 -6
View File
@@ -42,9 +42,18 @@ function standardNormalCdf(z) {
return 0.5 * (1 + erf(z / Math.SQRT2))
}
// Peter Acklam's rational approximation of the inverse standard normal CDF
// (probit), relative error < 1.15e-9. Supports arbitrary intervalMass.
function probit(p) {
/**
* Peter Acklam's rational approximation of the inverse standard normal CDF
* (probit / R's `qnorm`), relative error < 1.15e-9. Supports arbitrary
* intervalMass.
*
* Exported so `agrestiCoull.js` can reuse it — the app should carry exactly one
* probit implementation.
*
* @param {number} p - probability in (0, 1)
* @returns {number}
*/
export function probit(p) {
const a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02, 1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00]
const b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02, 6.680131188771972e+01, -1.328068155288572e+01]
const c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00, -2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00]
@@ -70,11 +79,15 @@ function probit(p) {
/**
* @param {{median:number, lower:number, upper:number, intervalMass?:number, floorAtZero?:boolean}} p
* intervalMass: fraction of probability covered by [lower, upper] — the
* API's rate/count interval bounds are a 90% interval, so default 0.90.
* intervalMass: fraction of probability covered by [lower, upper]. The API's
* rate/count bounds are a **95%** interval — `validate_interval()` in
* crdc-arrests/api/R/validate.R defaults to 95 and the app never passes
* `interval=` — so the default is 0.95. Fitting 95% bounds as if they were
* 90% understates sigma by ~16% and draws a distribution narrower than the
* model's own.
* @returns {{ median:number, sigmaLeft:number, sigmaRight:number, pdf:(x:number)=>number, cdf:(x:number)=>number }}
*/
export function fitSkewedInterval({ median, lower, upper, intervalMass = 0.90, floorAtZero = true }) {
export function fitSkewedInterval({ median, lower, upper, intervalMass = 0.95, floorAtZero = true }) {
const z = probit((1 + intervalMass) / 2)
let sigmaLeft = (median - lower) / z
+128
View File
@@ -0,0 +1,128 @@
/**
* Turns the API's per-race×sex estimate rows into the row set the results page
* actually displays, in one place, so the summary table, the density panel and
* the difference chart can never disagree about which groups exist, what they
* are called, or which are selected by default.
*
* "Display space" is either eight race×sex groups or — when sex pooling is on —
* four race groups. Keys come from `groupKey`, so the same map lookups work in
* both modes.
*/
import { groupLabel, sentenceGroupLabel, shortGroupLabel } from './colors.js'
import { groupKey } from './drawGroups.js'
import { poolBySex } from './pooling.js'
/** The four modeled race groups, in the fixed palette order (see colors.js). */
export const RACE_ORDER = ['WH', 'BL', 'HI', 'AM']
const SEXES = ['F', 'M']
/**
* @param {Array<object>} rows - estimate rows for one district/model/year
* @param {boolean} pooled - collapse Female + Male into one row per race
* @returns {Array<{key: string, race: string, sex: string|null, label: string,
* shortLabel: string, enroll: number, observed: number, rate: number}>}
* sorted by observed arrests descending.
*/
export function buildDisplayGroups(rows, pooled) {
const usable = (rows || []).filter(
(r) => RACE_ORDER.includes(r.race) && SEXES.includes(r.sex),
)
const acc = new Map()
for (const r of usable) {
const sex = pooled ? null : r.sex
const key = groupKey(r.race, sex)
const prev = acc.get(key) || { key, race: r.race, sex, enroll: 0, observed: 0 }
acc.set(key, {
...prev,
enroll: prev.enroll + (r.stu_enroll || 0),
observed: prev.observed + (r.observed_arrests || 0),
// The model's own summary interval, carried per 1,000 so a chart can fall
// back to the fitted approximation when real draws can't be fetched.
// Null when pooled: adding two groups' interval *bounds* together is not
// a pooled interval, and there is no honest way to fake one without the
// draws. A pooled group with no draws simply has no modelled shape.
modeled: pooled
? null
: {
median: (r.rate_median || 0) * 1000,
lower: (r.rate_lower || 0) * 1000,
upper: (r.rate_upper || 0) * 1000,
},
})
}
return [...acc.values()]
.map((g) => ({
...g,
label: groupLabel(g.race, g.sex),
shortLabel: shortGroupLabel(g.race, g.sex),
sentenceLabel: sentenceGroupLabel(g.race, g.sex),
// A rate with no denominator is not a large rate — it is no rate. Report
// 0 and let the table's enrollment column show why.
rate: g.enroll > 0 ? (g.observed / g.enroll) * 1000 : 0,
}))
.sort((a, b) => b.observed - a.observed || b.enroll - a.enroll || a.key.localeCompare(b.key))
}
/**
* Groups checked on first render: every group with at least one observed
* arrest. When a district reports none at all, falls back to the two largest by
* enrollment so the chart still shows something explainable rather than an
* empty panel (the caller says so in the UI).
*
* @param {ReturnType<typeof buildDisplayGroups>} groups
* @returns {string[]} group keys
*/
export function defaultSelectedKeys(groups) {
if (!groups?.length) return []
const withArrests = groups.filter((g) => g.observed > 0)
if (withArrests.length > 0) return withArrests.map((g) => g.key)
return [...groups]
.sort((a, b) => b.enroll - a.enroll)
.slice(0, 2)
.map((g) => g.key)
}
/**
* The difference chart's opening pair: the two groups with the most observed
* arrests. Null when there aren't two groups to compare.
*
* @param {ReturnType<typeof buildDisplayGroups>} groups
* @returns {[string, string] | null}
*/
export function defaultDiffPair(groups) {
if (!groups || groups.length < 2) return null
return [groups[0].key, groups[1].key]
}
/**
* Enrollment for every race×sex cell, keyed for `poolBySex`/`displayDraws`.
* Always unpooled — pooling sums these itself.
*
* @param {Array<object>} rows
* @returns {Record<string, number>}
*/
export function enrollByGroupKey(rows) {
const out = {}
for (const r of rows || []) {
if (!RACE_ORDER.includes(r.race) || !SEXES.includes(r.sex)) continue
out[groupKey(r.race, r.sex)] = r.stu_enroll || 0
}
return out
}
/**
* Moves one model's raw race×sex draw counts into display space.
*
* @param {Record<string, number[]> | null | undefined} counts - keyed race×sex
* @param {Record<string, number>} enrollByGroup - keyed race×sex
* @param {boolean} pooled
* @returns {{counts: Record<string, number[]>, enroll: Record<string, number>}}
*/
export function displayDraws(counts, enrollByGroup, pooled) {
if (!counts) return { counts: {}, enroll: {} }
if (pooled) return poolBySex(counts, enrollByGroup)
return { counts, enroll: enrollByGroup || {} }
}
+169
View File
@@ -0,0 +1,169 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
buildDisplayGroups,
defaultDiffPair,
defaultSelectedKeys,
displayDraws,
enrollByGroupKey,
} from './districtGroups.js'
const row = (race, sex, stu_enroll, observed_arrests) => ({ race, sex, stu_enroll, observed_arrests })
const CLARK = [
row('WH', 'F', 30000, 8),
row('WH', 'M', 31000, 20),
row('BL', 'F', 12000, 18),
row('BL', 'M', 12500, 40),
row('HI', 'F', 30000, 5),
row('HI', 'M', 31000, 8),
row('AM', 'F', 200, 0),
row('AM', 'M', 228, 1),
]
// ——— buildDisplayGroups ———
test('buildDisplayGroups: one row per race×sex when not pooled', () => {
const groups = buildDisplayGroups(CLARK, false)
assert.equal(groups.length, 8)
assert.deepEqual(
groups.map((g) => g.key).sort(),
['AM_F', 'AM_M', 'BL_F', 'BL_M', 'HI_F', 'HI_M', 'WH_F', 'WH_M'],
)
})
test('buildDisplayGroups: collapses to four race rows when pooled', () => {
const groups = buildDisplayGroups(CLARK, true)
assert.equal(groups.length, 4)
const bl = groups.find((g) => g.key === 'BL')
assert.equal(bl.enroll, 24500)
assert.equal(bl.observed, 58)
assert.equal(bl.sex, null)
})
test('buildDisplayGroups: sorted by observed arrests descending', () => {
const observed = buildDisplayGroups(CLARK, false).map((g) => g.observed)
assert.deepEqual(observed, [...observed].sort((a, b) => b - a))
})
test('buildDisplayGroups: rate is per 1,000 students', () => {
const bl = buildDisplayGroups(CLARK, false).find((g) => g.key === 'BL_M')
assert.ok(Math.abs(bl.rate - (40 / 12500) * 1000) < 1e-12)
})
test('buildDisplayGroups: rate is 0 rather than Infinity with no enrollment', () => {
const groups = buildDisplayGroups([row('BL', 'F', 0, 3)], false)
assert.equal(groups[0].rate, 0)
})
test('buildDisplayGroups: labels a pooled group by race alone', () => {
const bl = buildDisplayGroups(CLARK, true).find((g) => g.key === 'BL')
assert.equal(bl.label, 'Black')
const blf = buildDisplayGroups(CLARK, false).find((g) => g.key === 'BL_F')
assert.equal(blf.label, 'Black Female')
})
test('buildDisplayGroups: the sentence label keeps the race capitalized', () => {
// Guards the readout in GroupDifference: lowercasing the whole label turned
// "Black Male" into "black male" mid-sentence.
const groups = buildDisplayGroups(CLARK, false)
assert.equal(groups.find((g) => g.key === 'BL_M').sentenceLabel, 'Black male')
assert.equal(groups.find((g) => g.key === 'WH_F').sentenceLabel, 'White female')
assert.equal(buildDisplayGroups(CLARK, true).find((g) => g.key === 'HI').sentenceLabel, 'Hispanic')
})
test('buildDisplayGroups: carries the modeled interval per 1,000 when not pooled', () => {
const groups = buildDisplayGroups(
[{ race: 'BL', sex: 'M', stu_enroll: 1000, observed_arrests: 4, rate_median: 0.004, rate_lower: 0.001, rate_upper: 0.009 }],
false,
)
assert.deepEqual(groups[0].modeled, { median: 4, lower: 1, upper: 9 })
})
test('buildDisplayGroups: pooled groups carry no modeled interval', () => {
// Summing two groups' interval bounds is not a pooled interval — there is no
// honest fallback shape without the draws, so don't invent one.
const groups = buildDisplayGroups(CLARK, true)
assert.ok(groups.every((g) => g.modeled === null))
})
test('buildDisplayGroups: ignores races and sexes outside the modeled set', () => {
const groups = buildDisplayGroups([...CLARK, row('AS', 'F', 900, 4), row('WH', 'X', 5, 5)], false)
assert.equal(groups.length, 8)
assert.ok(!groups.some((g) => g.race === 'AS'))
})
test('buildDisplayGroups: missing counts default to zero', () => {
const groups = buildDisplayGroups([{ race: 'BL', sex: 'F' }], false)
assert.equal(groups[0].observed, 0)
assert.equal(groups[0].enroll, 0)
})
test('buildDisplayGroups: empty input gives an empty list', () => {
assert.deepEqual(buildDisplayGroups([], false), [])
assert.deepEqual(buildDisplayGroups(undefined, true), [])
})
// ——— defaultSelectedKeys ———
test('defaultSelectedKeys: every group with at least one observed arrest', () => {
const groups = buildDisplayGroups(CLARK, false)
const keys = defaultSelectedKeys(groups)
assert.ok(!keys.includes('AM_F'))
assert.ok(keys.includes('AM_M'))
assert.equal(keys.length, 7)
})
test('defaultSelectedKeys: falls back to the two largest by enrollment', () => {
// A district with no arrests anywhere still has to show something, or the
// chart renders empty with no explanation.
const groups = buildDisplayGroups(
[row('WH', 'F', 900, 0), row('WH', 'M', 1000, 0), row('AM', 'F', 20, 0)],
false,
)
assert.deepEqual(defaultSelectedKeys(groups).sort(), ['WH_F', 'WH_M'])
})
test('defaultSelectedKeys: empty for no groups', () => {
assert.deepEqual(defaultSelectedKeys([]), [])
})
// ——— defaultDiffPair ———
test('defaultDiffPair: the two groups with the most observed arrests', () => {
assert.deepEqual(defaultDiffPair(buildDisplayGroups(CLARK, false)), ['BL_M', 'WH_M'])
})
test('defaultDiffPair: null when fewer than two groups exist', () => {
assert.equal(defaultDiffPair(buildDisplayGroups([row('BL', 'F', 10, 1)], false)), null)
assert.equal(defaultDiffPair([]), null)
})
// ——— enrollByGroupKey ———
test('enrollByGroupKey: maps race×sex keys to enrollment', () => {
const map = enrollByGroupKey(CLARK)
assert.equal(map.BL_M, 12500)
assert.equal(Object.keys(map).length, 8)
})
// ——— displayDraws ———
test('displayDraws: passes raw counts straight through when not pooled', () => {
const counts = { BL_F: [1, 2], BL_M: [3, 4] }
const out = displayDraws(counts, { BL_F: 100, BL_M: 200 }, false)
assert.deepEqual(out.counts, counts)
assert.deepEqual(out.enroll, { BL_F: 100, BL_M: 200 })
})
test('displayDraws: pools by sex when pooling is on', () => {
const out = displayDraws({ BL_F: [1, 2], BL_M: [3, 4] }, { BL_F: 100, BL_M: 200 }, true)
assert.deepEqual(out.counts, { BL: [4, 6] })
assert.deepEqual(out.enroll, { BL: 300 })
})
test('displayDraws: empty structures for missing counts', () => {
const out = displayDraws(null, { BL_F: 100 }, false)
assert.deepEqual(out.counts, {})
assert.deepEqual(out.enroll, {})
})
+48 -9
View File
@@ -1,33 +1,72 @@
/**
* Shared key format and coverage check for the per-group posterior draws
* Shared key format and coverage checks for the per-group posterior draws
* returned by `useDrawDistribution`. Both the hook (which builds the map) and
* the charts (which read it, and decide whether to claim "real draws") go
* through here so the key format lives in exactly one place.
*
* Draw arrays are *counts indexed by draw* — `counts[draw_id - 1]` — not rates
* and not push-ordered. That makes row order out of DuckDB irrelevant, and it
* makes a missing draw a hole rather than a silently shorter array, which is
* why `isCompleteDrawSet` checks every index instead of trusting `length`.
*/
/** Canonical key for one race×sex group in a `drawsByGroup` map. */
/**
* Canonical key for one group in a draws map.
*
* With a sex, this is one race×sex cell (`'BL_F'`). Without one — the pooled
* mode sparse districts fall back to — it is the race alone (`'BL'`). The two
* shapes can never collide, so a single map type serves both modes.
*
* @param {string} race
* @param {string} [sex] - omit (or pass null/'') for a sex-pooled group
* @returns {string}
*/
export function groupKey(race, sex) {
return `${race}_${sex}`
return sex ? `${race}_${sex}` : `${race}`
}
/**
* True only when *every* group in `groups` has a non-empty draw array.
* True when `counts` holds a finite value at every draw index 0…nDraws-1.
*
* A partial group must fall back rather than render a short draw set: a group
* present for only 300 of 500 draws would otherwise get a density and an
* interval computed off a biased subsample, indistinguishable on screen from
* a complete one.
*
* @param {number[] | null | undefined} counts
* @param {number} nDraws
* @returns {boolean}
*/
export function isCompleteDrawSet(counts, nDraws) {
if (!Array.isArray(counts) || !(nDraws > 0) || counts.length !== nDraws) return false
for (let i = 0; i < nDraws; i++) {
if (!Number.isFinite(counts[i])) return false
}
return true
}
/**
* True only when *every* group in `groups` has a complete draw array.
*
* `useDrawDistribution`'s `status` is an any-group signal: it reports 'ready'
* as soon as one group has real draws. Charts fall back per-group, so the
* chart-wide "these are real posterior draws" note/caption must be gated on
* complete coverage instead — a group can be missing because its enrollment is
* falsy or because its (LEAID, RACE, SEX) isn't in the parquet shard.
* complete coverage instead — a group can be missing because its (LEAID, RACE,
* SEX) isn't in the parquet shard.
*
* Returns false for an empty group list (nothing rendered means nothing to
* claim) and for a null map (loading, or the fetch failed outright).
*
* @param {Record<string, number[]> | null | undefined} drawsByGroup
* @param {Array<{race: string, sex: string}>} groups - the groups a chart is
* actually rendering, not everything the API returned.
* @param {Array<{race: string, sex?: string}>} groups - the groups a chart is
* actually rendering, not everything the API returned. Omit `sex` for pooled
* groups.
* @returns {boolean}
*/
export function hasDrawsForAll(drawsByGroup, groups) {
if (!drawsByGroup || !groups?.length) return false
return groups.every((g) => (drawsByGroup[groupKey(g.race, g.sex)]?.length ?? 0) > 0)
return groups.every((g) => {
const counts = drawsByGroup[groupKey(g.race, g.sex)]
return isCompleteDrawSet(counts, counts?.length ?? 0)
})
}
+60 -1
View File
@@ -1,11 +1,24 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { groupKey, hasDrawsForAll } from './drawGroups.js'
import { groupKey, hasDrawsForAll, isCompleteDrawSet } from './drawGroups.js'
test('groupKey: joins race and sex with an underscore', () => {
assert.equal(groupKey('BL', 'F'), 'BL_F')
})
test('groupKey: returns the race alone for a pooled group', () => {
// Pooled rows are keyed by race only, so the same helper builds both key
// shapes and nothing downstream has to know which mode it is in.
assert.equal(groupKey('BL'), 'BL')
assert.equal(groupKey('BL', null), 'BL')
assert.equal(groupKey('BL', ''), 'BL')
})
test('groupKey: a pooled key never collides with an unpooled one', () => {
assert.notEqual(groupKey('BL'), groupKey('BL', 'F'))
assert.notEqual(groupKey('BL'), groupKey('BL', 'M'))
})
test('hasDrawsForAll: true when every rendered group has draws', () => {
const map = { WH_F: [1, 2], BL_F: [3] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), true)
@@ -39,3 +52,49 @@ test('hasDrawsForAll: ignores groups the chart is not rendering', () => {
const map = { WH_F: [1], BL_F: [2], AS_F: [3] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), true)
})
test('hasDrawsForAll: works on pooled (race-only) groups', () => {
const map = { WH: [1], BL: [2] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH' }, { race: 'BL' }]), true)
assert.equal(hasDrawsForAll(map, [{ race: 'WH' }, { race: 'HI' }]), false)
})
test('hasDrawsForAll: false when a count array has a hole', () => {
// Counts are indexed by draw_id, so a group missing draw 2 leaves a hole
// rather than a short array. length alone would call this complete.
const holey = [1, 2, 3]
delete holey[1]
assert.equal(hasDrawsForAll({ BL_F: holey }, [{ race: 'BL', sex: 'F' }]), false)
})
// ——— isCompleteDrawSet ———
test('isCompleteDrawSet: true for a dense array of the expected length', () => {
assert.equal(isCompleteDrawSet([0, 1, 2], 3), true)
})
test('isCompleteDrawSet: false when the array is shorter than nDraws', () => {
assert.equal(isCompleteDrawSet([0, 1], 3), false)
})
test('isCompleteDrawSet: false when the array is longer than nDraws', () => {
assert.equal(isCompleteDrawSet([0, 1, 2, 3], 3), false)
})
test('isCompleteDrawSet: false when a draw index was never filled', () => {
const holey = new Array(3)
holey[0] = 1
holey[2] = 3
assert.equal(isCompleteDrawSet(holey, 3), false)
})
test('isCompleteDrawSet: false for a non-finite entry', () => {
assert.equal(isCompleteDrawSet([1, NaN, 3], 3), false)
assert.equal(isCompleteDrawSet([1, Infinity, 3], 3), false)
})
test('isCompleteDrawSet: false for a missing array or a zero-draw expectation', () => {
assert.equal(isCompleteDrawSet(null, 3), false)
assert.equal(isCompleteDrawSet(undefined, 3), false)
assert.equal(isCompleteDrawSet([], 0), false)
})
+57
View File
@@ -0,0 +1,57 @@
/**
* The difference in modelled arrest rate between two student groups, taken
* draw by draw — the quantity behind the white paper's Fig 7
* (`wp_fig_group_difference`).
*
* Read the wording carefully before writing a caption from these numbers.
* Upstream, `draw_id` is renumbered 1–500 per write batch, and a district's
* groups land in different batches, so draw *k* of group A and draw *k* of
* group B are not the same parameter draw. Measured correlation between two
* groups' `pred` in Clark County was ≈ 0.02 even within a batch — observation
* noise from `posterior_predict` dominates. The published figure has exactly
* the same property, so the app matches the paper; what neither can claim is a
* paired-parameter contrast. Always say **posterior predictive draws**, never
* "paired parameter draws".
*/
import { quantile } from './kde.js'
/**
* Per-draw difference in rate per 1,000: group A minus group B.
*
* Returns an empty array rather than a truncated one when the two draw sets
* disagree in length — pairing draw 3 of one group with draw 7 of another would
* fabricate a difference distribution out of unrelated draws.
*
* @param {number[] | null | undefined} countsA - predicted counts, indexed by draw
* @param {number | null | undefined} enrollA
* @param {number[] | null | undefined} countsB
* @param {number | null | undefined} enrollB
* @returns {number[]}
*/
export function differenceRates(countsA, enrollA, countsB, enrollB) {
if (!countsA?.length || !countsB?.length) return []
if (countsA.length !== countsB.length) return []
if (!(enrollA > 0) || !(enrollB > 0)) return []
return countsA.map((a, i) => (a / enrollA) * 1000 - (countsB[i] / enrollB) * 1000)
}
/**
* @param {number[] | null | undefined} deltas
* @returns {{n: number, prGreater: number, median: number, lower80: number,
* upper80: number, lower95: number, upper95: number} | null}
* `prGreater` counts draws **strictly** above zero, so a group whose every
* draw ties reads as 0%, not 100%.
*/
export function differenceSummary(deltas) {
if (!deltas?.length) return null
return {
n: deltas.length,
prGreater: deltas.filter((d) => d > 0).length / deltas.length,
median: quantile(deltas, 0.5),
lower80: quantile(deltas, 0.1),
upper80: quantile(deltas, 0.9),
lower95: quantile(deltas, 0.025),
upper95: quantile(deltas, 0.975),
}
}
+77
View File
@@ -0,0 +1,77 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { differenceRates, differenceSummary } from './groupDifference.js'
// ——— differenceRates ———
test('differenceRates: subtracts rates at the same draw index', () => {
// A: 2 and 4 arrests in 1,000 students → 2 and 4 per 1,000
// B: 1 and 1 arrests in 500 students → 2 and 2 per 1,000
const deltas = differenceRates([2, 4], 1000, [1, 1], 500)
assert.deepEqual(deltas, [0, 2])
})
test('differenceRates: differences can be negative', () => {
assert.deepEqual(differenceRates([0], 1000, [5], 1000), [-5])
})
test('differenceRates: empty when the draw sets are different lengths', () => {
// Pairing draw 3 of one group with draw 7 of another would fabricate a
// difference distribution out of unrelated draws.
assert.deepEqual(differenceRates([1, 2, 3], 1000, [1, 2], 1000), [])
})
test('differenceRates: empty when either denominator is missing', () => {
assert.deepEqual(differenceRates([1], 0, [1], 1000), [])
assert.deepEqual(differenceRates([1], 1000, [1], undefined), [])
})
test('differenceRates: empty when either draw set is missing', () => {
assert.deepEqual(differenceRates(null, 1000, [1], 1000), [])
assert.deepEqual(differenceRates([1], 1000, [], 1000), [])
})
test('differenceRates: does not mutate its inputs', () => {
const a = [1, 2]
const b = [3, 4]
differenceRates(a, 1000, b, 1000)
assert.deepEqual(a, [1, 2])
assert.deepEqual(b, [3, 4])
})
// ——— differenceSummary ———
test('differenceSummary: Pr(delta > 0) is the share of draws strictly above zero', () => {
const summary = differenceSummary([-1, 0, 1, 2])
assert.equal(summary.prGreater, 0.5)
})
test('differenceSummary: a draw of exactly zero does not count as greater', () => {
assert.equal(differenceSummary([0, 0, 0, 0]).prGreater, 0)
})
test('differenceSummary: reports the median and both interval widths', () => {
const deltas = Array.from({ length: 101 }, (_, i) => i) // 0…100
const summary = differenceSummary(deltas)
assert.equal(summary.median, 50)
assert.equal(summary.lower80, 10)
assert.equal(summary.upper80, 90)
assert.equal(summary.lower95, 2.5)
assert.equal(summary.upper95, 97.5)
})
test('differenceSummary: the 95% interval contains the 80% interval', () => {
const deltas = Array.from({ length: 500 }, (_, i) => Math.sin(i) * 4)
const s = differenceSummary(deltas)
assert.ok(s.lower95 <= s.lower80)
assert.ok(s.upper95 >= s.upper80)
})
test('differenceSummary: reports the draw count it summarized', () => {
assert.equal(differenceSummary([1, 2, 3]).n, 3)
})
test('differenceSummary: null for an empty or missing set', () => {
assert.equal(differenceSummary([]), null)
assert.equal(differenceSummary(undefined), null)
})
+115
View File
@@ -0,0 +1,115 @@
/**
* Sex pooling for sparse districts.
*
* This is the project's own house method applied on a different axis:
* `build_state_summary()` in `crdc-arrests/R/summarize_draws.R` pools across
* LEAs by summing `pred` and `stu_enroll` *within each draw* and only then
* summarizing across draws. Pooling Female + Male inside one district is the
* identical operation — sum the numerators draw-by-draw, sum the denominators,
* divide once. Summing rates instead would weight a 53-student cell the same as
* a 3,000-student one.
*
* Pure functions; no React, no fetching.
*/
import { groupKey } from './drawGroups.js'
/**
* Whole-district trigger for sex pooling: fewer than this many observed arrests
* across all eight race×sex cells.
*
* 20 is chosen so that the average cell carries at least ~2 arrests before the
* app is willing to show eight separate posteriors. Below it, most cells are
* zero-count and their posteriors are dominated by the prior and by
* observation noise, so eight ridges read as eight findings when they are
* really one. The threshold is on the *district* total rather than per cell so
* the table's row set doesn't change shape group by group.
*/
export const POOL_BY_SEX_ARREST_THRESHOLD = 20
/**
* True when a district is sparse enough that its eight race×sex cells should
* collapse to four race rows.
*
* Returns false for an empty row set: no rows is not "a sparse district", it is
* no district data at all, and firing the pooling banner there would explain a
* rule that never applied.
*
* @param {Array<{observed_arrests?: number|null}>} rows - estimate rows for one
* district/model/year, one per race×sex cell.
* @returns {boolean}
*/
export function shouldPoolBySex(rows) {
if (!rows?.length) return false
const total = rows.reduce((sum, r) => sum + (r.observed_arrests || 0), 0)
return total < POOL_BY_SEX_ARREST_THRESHOLD
}
/**
* Per-1,000 rate array from an array of predicted counts.
*
* Returns an empty array when the denominator is missing or non-positive: a
* rate with no denominator is undefined, not zero, and every consumer already
* treats an empty draw array as "nothing to draw here".
*
* @param {number[] | null | undefined} counts - predicted counts, indexed by draw
* @param {number | null | undefined} enroll - students in the group
* @returns {number[]}
*/
export function toRates(counts, enroll) {
if (!counts?.length || !(enroll > 0)) return []
return counts.map((c) => (c / enroll) * 1000)
}
/**
* Pools Female + Male within each race, summing predicted counts at the same
* draw index and summing enrollment.
*
* Two refusals keep the arithmetic honest:
*
* - A race whose sexes have different draw-array lengths is dropped entirely.
* Element-wise addition would pair unrelated draws and truncate to the
* shorter array, producing a narrower pooled interval than the data supports.
* - Numerator and denominator are built from the same set of sexes. Adding both
* sexes' enrollment to one sex's counts would roughly halve the rate — an
* invented improvement, not a pooled estimate — and adding a sex's counts
* without its enrollment inflates it the same way.
*
* @param {Record<string, number[]>} countsByGroup - keyed by `groupKey(race, sex)`
* @param {Record<string, number>} enrollByGroup - keyed by `groupKey(race, sex)`
* @returns {{counts: Record<string, number[]>, enroll: Record<string, number>}}
* both keyed by race alone (see `groupKey(race)`).
*/
export function poolBySex(countsByGroup, enrollByGroup) {
const counts = {}
const enroll = {}
if (!countsByGroup) return { counts, enroll }
const byRace = {}
for (const key of Object.keys(countsByGroup)) {
const [race] = key.split('_')
;(byRace[race] ??= []).push(key)
}
for (const [race, keys] of Object.entries(byRace)) {
// A group must bring both halves of the fraction. Counts with no matching
// enrollment would land in the numerator while contributing nothing to the
// denominator, inflating the pooled rate.
const present = keys.filter((k) => countsByGroup[k]?.length > 0 && enrollByGroup?.[k] > 0)
if (present.length === 0) continue
const nDraws = countsByGroup[present[0]].length
if (present.some((k) => countsByGroup[k].length !== nDraws)) continue
const summed = new Array(nDraws).fill(0)
for (const k of present) {
const arr = countsByGroup[k]
for (let i = 0; i < nDraws; i++) summed[i] += arr[i]
}
counts[groupKey(race)] = summed
enroll[groupKey(race)] = present.reduce((sum, k) => sum + (enrollByGroup?.[k] || 0), 0)
}
return { counts, enroll }
}
+110
View File
@@ -0,0 +1,110 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
POOL_BY_SEX_ARREST_THRESHOLD,
poolBySex,
shouldPoolBySex,
toRates,
} from './pooling.js'
// ——— toRates ———
test('toRates: converts counts to a per-1,000 rate array', () => {
assert.deepEqual(toRates([0, 1, 2], 500), [0, 2, 4])
})
test('toRates: does not mutate its input', () => {
const counts = [1, 2]
toRates(counts, 1000)
assert.deepEqual(counts, [1, 2])
})
test('toRates: returns an empty array for non-positive enrollment', () => {
// A rate with no denominator is not zero, it is undefined — return nothing
// rather than a column of Infinity that a chart would happily plot.
assert.deepEqual(toRates([1, 2], 0), [])
assert.deepEqual(toRates([1, 2], -5), [])
assert.deepEqual(toRates([1, 2], undefined), [])
})
test('toRates: returns an empty array when counts are missing', () => {
assert.deepEqual(toRates(undefined, 100), [])
assert.deepEqual(toRates([], 100), [])
})
// ——— shouldPoolBySex ———
test('shouldPoolBySex: true when total observed arrests is under the threshold', () => {
const rows = [{ observed_arrests: 4 }, { observed_arrests: 2 }]
assert.equal(shouldPoolBySex(rows), true)
})
test('shouldPoolBySex: false at exactly the threshold', () => {
const rows = [{ observed_arrests: POOL_BY_SEX_ARREST_THRESHOLD }]
assert.equal(shouldPoolBySex(rows), false)
})
test('shouldPoolBySex: treats missing observed_arrests as zero', () => {
assert.equal(shouldPoolBySex([{ observed_arrests: null }, {}]), true)
})
test('shouldPoolBySex: false for an empty or missing row set', () => {
// No rows is not "a sparse district" — it is no district data at all. Firing
// the pooling banner there would explain a rule that never applied.
assert.equal(shouldPoolBySex([]), false)
assert.equal(shouldPoolBySex(undefined), false)
})
// ——— poolBySex ———
test('poolBySex: sums counts draw-by-draw and enrollment across F and M', () => {
const counts = { BL_F: [1, 2, 3], BL_M: [10, 20, 30] }
const enroll = { BL_F: 100, BL_M: 300 }
const pooled = poolBySex(counts, enroll)
assert.deepEqual(pooled.counts, { BL: [11, 22, 33] })
assert.deepEqual(pooled.enroll, { BL: 400 })
})
test('poolBySex: keys the result by race alone', () => {
const pooled = poolBySex(
{ WH_F: [1], WH_M: [1], HI_F: [2], HI_M: [2] },
{ WH_F: 10, WH_M: 10, HI_F: 20, HI_M: 20 },
)
assert.deepEqual(Object.keys(pooled.counts).sort(), ['HI', 'WH'])
})
test('poolBySex: pools a race present for only one sex', () => {
const pooled = poolBySex({ AM_F: [1, 1] }, { AM_F: 53, AM_M: 61 })
// Only F has draws, so only F's enrollment may go in the denominator —
// summing both sexes' enrollment against one sex's counts would halve the rate.
assert.deepEqual(pooled.counts, { AM: [1, 1] })
assert.deepEqual(pooled.enroll, { AM: 53 })
})
test('poolBySex: drops a race whose sexes have mismatched draw counts', () => {
// Adding a 500-draw array to a 400-draw array element-wise would silently
// pair unrelated draws and truncate; refuse instead.
const pooled = poolBySex({ BL_F: [1, 2, 3], BL_M: [1, 2] }, { BL_F: 10, BL_M: 10 })
assert.deepEqual(pooled.counts, {})
assert.deepEqual(pooled.enroll, {})
})
test('poolBySex: ignores a group with no enrollment entry', () => {
const pooled = poolBySex({ BL_F: [1], BL_M: [2] }, { BL_F: 100 })
assert.deepEqual(pooled.counts, { BL: [1] })
assert.deepEqual(pooled.enroll, { BL: 100 })
})
test('poolBySex: returns empty structures for empty input', () => {
const pooled = poolBySex({}, {})
assert.deepEqual(pooled.counts, {})
assert.deepEqual(pooled.enroll, {})
})
test('poolBySex: does not mutate its inputs', () => {
const counts = { BL_F: [1], BL_M: [2] }
const enroll = { BL_F: 10, BL_M: 20 }
poolBySex(counts, enroll)
assert.deepEqual(counts, { BL_F: [1], BL_M: [2] })
assert.deepEqual(enroll, { BL_F: 10, BL_M: 20 })
})
+52
View File
@@ -0,0 +1,52 @@
/**
* Shared x-domain for the arrest-rate charts, in rate per 1,000.
*
* Every panel and every model row has to sit on the same axis or the reader
* will compare curves that aren't comparable, so the domain is computed once
* from everything that will be drawn.
*/
import { quantile } from './kde.js'
import { niceTicks } from './niceTicks.js'
/**
* Hard ceiling on the axis, inherited from the chart this replaces.
*
* A four-student cell can produce posterior draws in the hundreds per 1,000;
* letting those set the axis squashes every other group into the leftmost few
* pixels. `computeRateDomain` reports `clipped` whenever the cap actually
* binds, so the chart can say so rather than cropping data off the edge in
* silence.
*/
export const MAX_RATE_DOMAIN = 30
/**
* Ignore the very top of each draw set when sizing the axis. The posterior
* predictive has a long right tail by construction; one draw in 500 should not
* decide the axis for all of them.
*/
const DOMAIN_QUANTILE = 0.995
const HEADROOM = 1.15
/**
* @param {Array<number[]>} rateArrays - per-group draw rates, per 1,000
* @param {number[]} acUpperRates - Agresti–Coull upper bounds, per 1,000
* @returns {{ticks: number[], niceMax: number, clipped: boolean}}
* `clipped` is true when something that will be drawn extends past `niceMax`.
*/
export function computeRateDomain(rateArrays, acUpperRates) {
const candidates = []
for (const rates of rateArrays || []) {
if (rates?.length) candidates.push(quantile(rates, DOMAIN_QUANTILE))
}
for (const upper of acUpperRates || []) {
if (Number.isFinite(upper)) candidates.push(upper)
}
const dataMax = candidates.length ? Math.max(...candidates) : 0
const rawMax = Math.max(dataMax, 1) * HEADROOM
const { ticks, niceMax } = niceTicks(Math.min(rawMax, MAX_RATE_DOMAIN), 5)
return { ticks, niceMax, clipped: dataMax > niceMax }
}
+51
View File
@@ -0,0 +1,51 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { MAX_RATE_DOMAIN, computeRateDomain } from './rateDomain.js'
test('computeRateDomain: covers the draws and the frequentist bounds', () => {
const domain = computeRateDomain([[0, 1, 2, 3, 4]], [6])
assert.ok(domain.niceMax >= 6, `niceMax ${domain.niceMax} should cover the AC upper bound`)
assert.equal(domain.clipped, false)
})
test('computeRateDomain: ticks start at zero and end at niceMax', () => {
const domain = computeRateDomain([[0, 1, 2]], [4])
assert.equal(domain.ticks[0], 0)
assert.equal(domain.ticks[domain.ticks.length - 1], domain.niceMax)
})
test('computeRateDomain: one extreme draw does not blow out the axis', () => {
// A single 400-per-1,000 draw in a 53-student cell would otherwise squash
// every other group into the leftmost pixel.
const bulk = new Array(500).fill(2)
bulk[0] = 400
const domain = computeRateDomain([bulk], [3])
assert.ok(domain.niceMax < 10, `niceMax was ${domain.niceMax}`)
})
test('computeRateDomain: caps the domain and reports the clip', () => {
// The cap keeps a tiny-denominator group from flattening every other curve,
// but the caller has to be able to say so on screen rather than silently
// cropping data off the right edge.
const wide = new Array(500).fill(0).map((_, i) => (i < 490 ? 60 : 0))
const domain = computeRateDomain([wide], [80])
assert.equal(domain.niceMax, MAX_RATE_DOMAIN)
assert.equal(domain.clipped, true)
})
test('computeRateDomain: never returns a zero-width domain', () => {
const domain = computeRateDomain([[0, 0, 0]], [0])
assert.ok(domain.niceMax > 0)
assert.equal(domain.clipped, false)
})
test('computeRateDomain: handles no data at all', () => {
const domain = computeRateDomain([], [])
assert.ok(domain.niceMax > 0)
assert.ok(domain.ticks.length > 1)
})
test('computeRateDomain: ignores empty rate arrays', () => {
const domain = computeRateDomain([[], [1, 2, 3]], [])
assert.ok(domain.niceMax > 0 && domain.niceMax < MAX_RATE_DOMAIN)
})