Ports Figs 6 and 7 from crdc-arrests/R/paper_figures.R (wp_fig_group_density and
wp_fig_group_difference) so someone who has read the paper sees the paper. The
three old charts become a summary table and two charts; ArrestsOverTime stays.
Draws pipeline. useDrawDistribution now selects draw_id and returns predicted
counts indexed by draw rather than pre-divided rates. That one column is what
unlocks the rest: pooling has to sum numerators and denominators separately, and
a between-group difference has to be taken at a common draw index. Indexing by
draw_id rather than push order makes DuckDB's row ordering irrelevant and turns
a missing draw into a hole, which isCompleteDrawSet then rejects — a group
present for 300 of 500 draws would otherwise get an interval computed off a
biased subsample that looks identical on screen to a complete one. It also takes
models[] instead of a single model, so "compare all four" needs no conditional
hooks. The reset-before-guard ordering is preserved.
Also fixes a pre-existing bug in that pipeline: a state's draws are split across
data_0.parquet, data_1.parquet, … and the part count varies by state. The app
only ever fetched data_0. Nevada has one part, so this was invisible in every
Nevada test; California has eight, totalling 6.2MB, of which data_0 is 37KB and
holds 11 of California's 1,715 districts. Every other CA district looked absent
from the published data and silently fell back to the approximation. Parts are
now discovered from the Hugging Face tree listing API and fetched in parallel —
listing rather than probing data_N until a 404, because the browser logs a 404
as a console error however cleanly the fetch handles it, and a red error on
every load is indistinguishable from a real one. HEAD probing is the fallback.
New pure utils, all written against tests first:
agrestiCoull faithful port incl. the zero-numerator rule of three and the
negative lower bound at (1, 53); pinned to five R outputs
pooling sex pooling for sparse districts; numerator and denominator
are always drawn from the same set of groups
densityProfile discrete probability mass below 12 distinct values, KDE above;
KDE delegates to kde.js, whose bandwidth clamp is untouched
districtGroups display-row derivation, defaults, pooled vs unpooled keys
groupDifference per-draw delta; refuses to pair mismatched draw sets
rateDomain shared x-axis, with a clip flag so the cap is never silent
Chart A keeps the palette contract: race is hue, sex is position. Female and
Male are stacked panels sharing one axis, collapsing to one panel when pooled —
never a second hue. Chart B uses a diverging ramp centred at zero rather than
the paper's sequential YlOrRd, because delta is signed and a sequential ramp
encodes "more" where the data means "which direction".
Captions say "posterior predictive draws", never "paired parameter draws":
draw_id is renumbered per write batch upstream and a district's groups land in
different batches, so cross-group pairing is effectively independent (measured
cor ~= 0.02). The published figure has the same property; what neither can claim
is a paired-parameter contrast.
Sparse districts (under 20 arrests district-wide) pool Female and Male within
each race, with a banner stating the rule and a switch to override it. A pooled
group carries no modelled interval — summing two groups' interval bounds is not
a pooled interval, and there is no honest way to fake one without the draws.
React still owns the DOM. d3-scale/shape/array/interpolate supply scales, path
generators and colour interpolation; no selections, no useEffect DOM mutation.
Deletes RateByGroupBar and RateDensityRidgeline. Keeps distributionApprox.js and
ApproxNote.jsx — still the per-group fallback when draws can't be fetched.
112 tests pass; npm run build clean.
Adds src/utils/duckdbClient.js, a lazy-initialized getDb() singleton
wrapping the AsyncDuckDB MVP (single-threaded) wasm bundle. Pins
@duckdb/duckdb-wasm to ^1.32.0 (npm's latest dist-tag currently points
at a -dev prerelease). Excludes the package from Vite's dev-server
dependency pre-bundling so its worker/wasm ?url imports resolve
correctly.
Verified live in the browser under the /crdc-demo/ base path: wasm and
worker assets load with 200, and a SELECT 42 query round-trips
correctly through the singleton, confirming the risk flagged in the
design spec (duckdb-wasm loading correctly from a subpath-served Vite
dev server) does not materialize.
Rebuilds all 6 charts around a validated categorical race palette, row-based
sex encoding, and a consistent observed-vs-modeled mark convention (diamond
vs. filled bar/density) instead of ad hoc per-chart color schemes. Adds a
shared student-group filter (defaults to all 8 groups) that scopes every
chart's data from one place in ChartPanel.
Drops the D3 dependency entirely in favor of plain SVG, removing the
imperative-DOM bug class behind this app's repeated "fix the fix" commits.
Replaces the fake symmetric-normal posterior approximation with a skewed,
median-preserving fit to the API's interval bounds, clearly labeled as an
approximation. Fixes two broken SVG fill attributes, a decorative model
dropdown that never affected its chart, dead code, an orphaned component,
and a broken CSS token reference.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Install d3@7
- Add fetchDistrictDraws API function (for future use with raw draws)
- Create RateDensityRidgeline component using D3.js
- Density ridges showing posterior distributions per race group
- Diamond markers for observed rates (matching R design)
- Dropdown to switch between model specifications
- Matches Civilytics color palette (navy fill, danger diamonds)
React + Vite static site demonstrating the CRDC School Arrest Rate API.
Features 6 interactive charts comparing observed arrest data against Bayesian
model estimates across U.S. school districts, with Civilytics visual identity.
- State selector and district search with 'interesting' suggestions (top arrests)
- Animated histogram loading grid showing posterior draw progress
- Charts: time series, rate by group, district vs national, model comparison quadrants, density proxy, exceedance probability
- Static JSON fixture for national rates (no API changes needed)
- Dockerfile + docker-compose.yml for self-hosted deployment
- GitHub Actions workflow and _config.yml for Git Pages