Ports Figs 6 and 7 from crdc-arrests/R/paper_figures.R (wp_fig_group_density and
wp_fig_group_difference) so someone who has read the paper sees the paper. The
three old charts become a summary table and two charts; ArrestsOverTime stays.
Draws pipeline. useDrawDistribution now selects draw_id and returns predicted
counts indexed by draw rather than pre-divided rates. That one column is what
unlocks the rest: pooling has to sum numerators and denominators separately, and
a between-group difference has to be taken at a common draw index. Indexing by
draw_id rather than push order makes DuckDB's row ordering irrelevant and turns
a missing draw into a hole, which isCompleteDrawSet then rejects — a group
present for 300 of 500 draws would otherwise get an interval computed off a
biased subsample that looks identical on screen to a complete one. It also takes
models[] instead of a single model, so "compare all four" needs no conditional
hooks. The reset-before-guard ordering is preserved.
Also fixes a pre-existing bug in that pipeline: a state's draws are split across
data_0.parquet, data_1.parquet, … and the part count varies by state. The app
only ever fetched data_0. Nevada has one part, so this was invisible in every
Nevada test; California has eight, totalling 6.2MB, of which data_0 is 37KB and
holds 11 of California's 1,715 districts. Every other CA district looked absent
from the published data and silently fell back to the approximation. Parts are
now discovered from the Hugging Face tree listing API and fetched in parallel —
listing rather than probing data_N until a 404, because the browser logs a 404
as a console error however cleanly the fetch handles it, and a red error on
every load is indistinguishable from a real one. HEAD probing is the fallback.
New pure utils, all written against tests first:
agrestiCoull faithful port incl. the zero-numerator rule of three and the
negative lower bound at (1, 53); pinned to five R outputs
pooling sex pooling for sparse districts; numerator and denominator
are always drawn from the same set of groups
densityProfile discrete probability mass below 12 distinct values, KDE above;
KDE delegates to kde.js, whose bandwidth clamp is untouched
districtGroups display-row derivation, defaults, pooled vs unpooled keys
groupDifference per-draw delta; refuses to pair mismatched draw sets
rateDomain shared x-axis, with a clip flag so the cap is never silent
Chart A keeps the palette contract: race is hue, sex is position. Female and
Male are stacked panels sharing one axis, collapsing to one panel when pooled —
never a second hue. Chart B uses a diverging ramp centred at zero rather than
the paper's sequential YlOrRd, because delta is signed and a sequential ramp
encodes "more" where the data means "which direction".
Captions say "posterior predictive draws", never "paired parameter draws":
draw_id is renumbered per write batch upstream and a district's groups land in
different batches, so cross-group pairing is effectively independent (measured
cor ~= 0.02). The published figure has the same property; what neither can claim
is a paired-parameter contrast.
Sparse districts (under 20 arrests district-wide) pool Female and Male within
each race, with a banner stating the rule and a switch to override it. A pooled
group carries no modelled interval — summing two groups' interval bounds is not
a pooled interval, and there is no honest way to fake one without the draws.
React still owns the DOM. d3-scale/shape/array/interpolate supply scales, path
generators and colour interpolation; no selections, no useEffect DOM mutation.
Deletes RateByGroupBar and RateDensityRidgeline. Keeps distributionApprox.js and
ApproxNote.jsx — still the per-group fallback when draws can't be fetched.
112 tests pass; npm run build clean.
DistrictSearch built its "most arrests" list from
/estimates?state=XX&year=21-22&limit=500. That endpoint returns rows
ORDER BY LEAID, RACE, SEX at eight rows per district, so a 500-row cap is not a
sample of the state — it is the ~62 lowest-LEAID districts in it.
Measured against California (11,488 rows, 1,715 districts): the old read covered
68 districts, and 6 of the true top 8 were invisible to it. It suggested
districts with 1 and 2 arrests as the state's most notable, while San Diego
Unified (178), Fresno Unified (77) and Kern High (69) never appeared.
Replaced with a committed fixture, public/data/top_districts.json, generated by
scripts/build-top-districts.mjs. The script pages each state to completion using
meta.total from the response envelope and fails loudly on a short read, since a
silent truncation there would reintroduce exactly this bug. 51 states, 135
requests, ~104KB, following the national_rates.json precedent. Re-run it only
when a new CRDC wave lands.
The search screen also loses a multi-second fetch on every visit, and the
hardcoded "Try Derby (KS), Paterson (NJ)" hint goes with it — the real list
supersedes it. Degrades to search-only if the fixture is missing.
fetchStateDistricts() is kept for scripts and ad-hoc use, with its JSDoc now
warning that any short read ranks by LEAID.
pages.yml gains public/** in its paths filter: the fixture ships with the build,
so regenerating it has to be able to trigger a deploy on its own.
The API returns 95% intervals: validate_interval() in
crdc-arrests/api/R/validate.R defaults to 95L and this app never passes
`interval=`. Two places claimed 90% anyway.
fitSkewedInterval's `intervalMass` defaulted to 0.90, so the analytic fallback
fitted 95% bounds as if they covered 90% of the mass. That divides each
half-interval by 1.645 instead of 1.960 and understates sigma by ~16% — the
fallback drew a distribution visibly narrower than the model's own, in the one
code path where we have no draws to check it against.
ArrestsOverTime's legend read "Modeled (median + 90% interval)" while plotting
count_lower/count_upper, which are the same 95% bounds.
Also exports probit() from distributionApprox.js so the Agresti-Coull port can
reuse it rather than carrying a second qnorm implementation.
AGENTS.md's "Error Bar Convention" paragraph was wrong in four ways after the
last docs pass: it pointed "above" at a section that is below it, described the
fallback as "synthetic draw generation" when it generates no draws at all, cited
an `intervalWidth / 3.29` expression in RateDensityRidgeline.jsx that does not
exist, and listed ModelDrawsComparison.jsx, a file that was deleted when the
demo was reduced to 3 charts. Rewritten against the current code: the fallback
is fitSkewedInterval's analytic two-piece normal, whose divisor is
probit((1 + intervalMass) / 2) with intervalMass defaulting to 0.90, and the
"if you change this" list now names the three files that actually encode 90%.
README: 6 charts -> 3, the ridgeline is Chart 3 not Chart 5, Chart 2's use of
real draws is now mentioned, the D3.js v7 claim is dropped (d3 is not a
dependency), @duckdb/duckdb-wasm is listed in the tech stack, the file tree
matches src/, and the container-size note accounts for the wasm engine.
The wasm engine is the single largest asset in the build (39,362,651 bytes) but
the immutable-cache location regex did not include `wasm`, so on the
Docker/nginx path it got no Cache-Control at all, and no gzip directive existed
anywhere -- nginx's default gzip_types is text/html only, so it shipped
uncompressed on every cold load.
Verified against nginx:alpine (1.31.3) with the real dist/ mounted:
`nginx -t` passes, and the wasm now returns Content-Type: application/wasm,
Content-Encoding: gzip, Cache-Control: public, max-age=31536000, immutable --
8,766,496 bytes on the wire instead of 39,362,651.
Dynamic gzip rather than gzip_static because the Vite build emits no
pre-compressed .gz files.
?state= was used verbatim and ends up interpolated into a fetch URL and into
the Hugging Face parquet shard path duckdb-wasm reads. Impact is low -- it is a
client-only fetch to a public HTTPS URL, and the wasm sandbox loads no httpfs
-- but validating at the boundary is cheap and correct. Deep links now require
/^[A-Z]{2}$/ (after trim + uppercase, so ?state=az still works) and fall back
to the state selector with a warning when they do not match.
useDrawDistribution's status is an any-group signal -- 'ready' means at least
one group came back with draws -- but both charts used it chart-wide to hide
<ApproxNote /> and print a caption claiming every box/ridge is real draws.
Individual boxes and ridges already fall back per-group, so the note and
caption were the only things over-claiming.
This is reachable with real data, not just in theory: for LEAID 0400311 (AZ,
unified_m4_mod) the estimates API returns all 8 race x sex groups but the
parquet shard contains draws for WH_M only. The chart reported 'ready' and hid
the note while 7 of its 8 ridges were the analytic approximation.
New src/utils/drawGroups.js owns the "RACE_SEX" key format (previously
duplicated across the hook and both charts) and a hasDrawsForAll() coverage
check. Each chart now checks the groups it actually renders -- RACE_ORDER x its
sex panels/columns -- rather than everything the API returned. A null map
(loading, or a fetch that failed outright) fails the check, so the error path
still shows the note.
Also fixes a stale-state hazard in the same hook: the input guard ran before
setStatus('loading')/setDrawsByGroup(null), so switching to a model whose
groups list is empty (that model's upstream fetch failed) left status at
'ready' with the previous model's draws still in state. Verified by
instrumenting the hook: with the old ordering, switching from a ready model to
an empty one kept status 'ready' and all 8 previous draw keys; with the reset
moved above the guard it correctly resets to 'loading' with no draws.
Small districts produce rare-event count posteriors that are heavily
zero-inflated, and the absolute 1e-3 floor / raw-sd fallback broke on both
extremes of that data:
- All-identical draws (e.g. 500 zeros, the norm for a group with a handful of
students) collapsed to h = 1e-3, a near-delta spike of density ~399. Because
SexRidgeColumn shares one maxPdf per column, that single spike flattened every
other ridge in the column to sub-pixel height. Measured on a real district
(0400315 AZ, unified_m4_mod): the three informative ridges rendered at
0.07-0.13px of a 47.56px row.
- When IQR is 0 -- the normal case when most draws are 0 -- min(sd, iqr/1.34)
was falsy and the rule fell all the way back to raw sd, which a few extreme
draws inflate until the ridge is a flat line claiming maximal uncertainty.
Bandwidth is now clamped to [domainWidth/50, domainWidth/6] and the robust rule
degrades by picking the smallest *positive* spread estimate instead of
discarding robustness entirely. The floor is slightly wider than one render
step at n = 60, so a degenerate draw set resolves as a narrow bump; it does not
bind on an ordinary posterior (spread wider than ~8% of the domain keeps its
own Silverman bandwidth). On the district above the informative ridges now
render at 5.66-5.70px, a ~60x improvement.
Five new tests cover both failure modes; all five fail against the old formula.
AGENTS.md Error Bar Convention section now explicitly states that synthetic
draw generation from normal approximation is only used when real draws can't
be fetched (network error, unsupported browser, HF outage). Makes clear the
fallback is secondary, not primary, to the real-draw mechanism described in
the new §5.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- AGENTS.md: Replace §4 (API Endpoint Availability) with updated text; insert
new §5 (Real posterior draws via duckdb-wasm) documenting the shift from
synthetic normal-approximation draws to client-side fetches via duckdb-wasm
against the public Hugging Face parquet dataset. Include actual payload size
(~39MB uncompressed / ~8.86MB gzipped). Renumber subsequent items.
- README.md: Update Chart 5 description in "What It Does" to reflect real draws
+ fallback behavior. Update API table to clarify that /api/v1/draws is not
called from app but informs the Hugging Face URL the app fetches directly.
- HANDOFF.md: Mark "Raw posterior draws" as done (2026-08-11) with reference
to the design spec.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
useDrawDistribution never cleared drawsByGroup on a new fetch or on
error, only status. A chart that varies `model` across renders (Chart
3's dropdown) could switch from a model that had succeeded to one
that then failed, and keep rendering the *previous* model's real
posterior draws — keyed by the same RACE_SEX strings — under the
newly-selected model's summary stats, with ApproxNote visible
suggesting (wrongly) that the fallback approximation was in use.
Clear drawsByGroup to null both when a new fetch starts and in the
catch branch, so a failed model switch never mixes draws from two
different models.
Wire useDrawDistribution into RateDensityRidgeline so each race×sex ridge
uses a Gaussian KDE over that group's 500 real posterior draws when the
selected model's shard has loaded, falling back per-group to the analytic
fitSkewedInterval approximation otherwise. ApproxNote now only shows when
status !== 'ready'. Pass leaid/state through from ChartPanel.
Two Important review findings on useDrawDistribution:
- ensureShardRegistered cached the rejected promise on fetch/registration
failure, permanently stuck for that (model,year,state) key until a full
page reload. Now deletes the cache entry on failure so the next caller
retries fresh.
- status could become 'ready' even when the LEAID-bound query returned zero
usable rows (district not present in the draws shard), silently mislabeling
a fitSkewedInterval-only render as real posterior draws. Now requires at
least one group with actual draws before reporting 'ready'.
Adds src/utils/duckdbClient.js, a lazy-initialized getDb() singleton
wrapping the AsyncDuckDB MVP (single-threaded) wasm bundle. Pins
@duckdb/duckdb-wasm to ^1.32.0 (npm's latest dist-tag currently points
at a -dev prerelease). Excludes the package from Vite's dev-server
dependency pre-bundling so its worker/wasm ?url imports resolve
correctly.
Verified live in the browser under the /crdc-demo/ base path: wasm and
worker assets load with 200, and a SELECT 42 query round-trips
correctly through the singleton, confirming the risk flagged in the
design spec (duckdb-wasm loading correctly from a subpath-served Vite
dev server) does not materialize.
Scopes replacing the analytic distribution approximation with real
posterior draws, fetched client-side from the Hugging Face parquet
dataset using duckdb-wasm — no new backend endpoint required.
The modeled point-range was dodged 18px to the side of its observed point,
which read as misaligned rather than paired. Removes the dodge so both sit
on the same x column - directly comparable at a glance, matching the
reference whitepaper figures' own convention.
Observed points were plain circles, inconsistent with every other chart in
the app where a diamond means "observed." Switches them to diamonds, drawn
last so they stay on top of the modeled marks sharing their column.
Adds a caption naming the model the modeled interval is drawn from (three-
year, no covariate / unified_m3_mod), sourced from ChartPanel's existing
WAVE_MODEL constant via the shared MODEL_QUADRANT_LABEL lookup rather than
hardcoded, so it can't drift if the model choice changes.
Also gives the rate-per-1k labels a paper-colored text halo (paintOrder:
stroke) so they stay legible now that they can sit directly over the
modeled whisker line instead of needing to dodge around it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The district-name watermark inside the arrests-over-time chart was
redundant with the page's own heading and was the direct cause of a label
collision (a high point's rate label rendered on top of it). Removes it.
Two more collisions, both edge cases the screenshot happened to hit: the
first wave's label sat close enough to the plot's top that it could
overlap the topmost gridline text, and a point sitting flush on the plot's
left edge had its center-anchored label bleed into the y-axis tick labels.
Fixes both with more vertical clearance and edge-aware text anchoring
(first point anchors right, last point anchors left, matching the same
fix already applied to chart 4's wave ticks before that chart was cut).
Adds a shared niceTicks() utility (Heckbert's nice-numbers algorithm) so
axis ticks read as round numbers (0/100/200) instead of arbitrary
fractions of the data max (0/113/226/339) — applied to all three charts
for consistency. Also fixes a legend/mark mismatch in the arrests-over-time
chart (the modeled series legend showed a diamond; the actual mark is a
dot) by adding a proper 'dot' shape to ChartLegend.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Fixes the deep-link/back-button bug where the district name showed
"Unknown District" on return: App.jsx was passing the LEAID as a text
search query to searchDistricts() (a name search), which never matches a
numeric ID. Resolves it instead via fetchDistrictEstimates(leaid, ...) and
reads lea_name/state directly from the returned row - correct regardless
of how the page was reached (fresh load, refresh, or browser back/forward).
Cuts the demo from 6 charts to 3, per request: arrests over time (kept),
arrest rate by student group restructured into Female/Male box-and-whisker
panels (kept), and the posterior density ridge chart restructured from a
2x2 model-quadrant grid into a single selected model (dropdown, default
three-year + referral rate) with Female/Male ridge columns. Removes
DistrictVsNational, ModelDrawsComparison, and ExceedanceProbability
entirely, along with the student-group filter (no longer needed - the
remaining charts always show the full breakdown) and the national-rates
fetch/plumbing that only those removed charts used.
Also updates LoadingAnimation's copy and dedupes its STUDENT_GROUPS/
MODEL_QUADRANTS constants against the shared ones in useApi.js.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Chart 4 (model predictions vs. observed) was picking the first row for a
given wave/model instead of summing across selected groups, so any district
whose first-returned group had zero counts (e.g. a suppressed race group)
showed a collapsed, flat modeled marker instead of the real aggregate
interval. Also fixes the rightmost wave label ("2021-22") clipping off the
edge of the small-multiple panels by anchoring edge ticks away from the
viewBox boundary instead of centering on it.
Chart 6's y-axis group labels (e.g. "American Indian / Alaska Native
Female") were wider than their margin and clipped past the left edge of the
SVG. Adds a shared shortGroupLabel() to colors.js and uses it here, matching
the convention already used in charts 2 and 5.
Chart 2 previously drew a plain bar to the modeled median next to an
observed diamond that often sat well past the bar's end, reading as if the
bar itself were an uncertainty range when it wasn't. Replaces it with an
actual horizontal box-and-whisker: whisker = the API's reported 90%
interval, box = the fitted approximation's 25th/75th percentiles (via a new
fit.quantile() inverse-CDF, exact round-trip of fitSkewedInterval's own
construction), median tick, observed diamond overlaid.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Rebuilds all 6 charts around a validated categorical race palette, row-based
sex encoding, and a consistent observed-vs-modeled mark convention (diamond
vs. filled bar/density) instead of ad hoc per-chart color schemes. Adds a
shared student-group filter (defaults to all 8 groups) that scopes every
chart's data from one place in ChartPanel.
Drops the D3 dependency entirely in favor of plain SVG, removing the
imperative-DOM bug class behind this app's repeated "fix the fix" commits.
Replaces the fake symmetric-normal posterior approximation with a skewed,
median-preserving fit to the API's interval bounds, clearly labeled as an
approximation. Fixes two broken SVG fill attributes, a decorative model
dropdown that never affected its chart, dead code, an orphaned component,
and a broken CSS token reference.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Create AGENTS.md: comprehensive agent guide covering architecture,
data flow, error bar convention (90% intervals), D3 usage patterns,
null safety pitfalls, deployment checklist, and related repos
- Update HANDOFF.md: mark CORS as resolved, document chart improvements
(Chart 5 ridgeline rewrite with synthetic draws + smooth rendering)
- Refresh README.md: accurate tech stack (D3 v7), updated file structure
with new chart files, corrected API endpoint table
- Rewrite RateDensityRidgeline.jsx with proper histogram-to-density conversion
- Build top-edge points from binned draws, then smooth with d3.curveBasis
- Construct complete area path: bottom edge + smoothed top + close
- Position observed rate diamonds above the ridge peak
- Add fetchDistrictDraws call in ChartPanel for all 4 quadrant models
- Pass modelDraws to ModelDrawsComparison component
- Rewrite Chart 4 with D3 showing:
- Density histogram from raw draws when available (per-year distribution)
- Fall back to interval-based rendering otherwise
- 90% credible intervals from quantile calculation
- Legend explaining colors (red=observed, teal=one-year, navy=three-year)
- Install d3@7
- Add fetchDistrictDraws API function (for future use with raw draws)
- Create RateDensityRidgeline component using D3.js
- Density ridges showing posterior distributions per race group
- Diamond markers for observed rates (matching R design)
- Dropdown to switch between model specifications
- Matches Civilytics color palette (navy fill, danger diamonds)
- Increase SVG dimensions in all 6 charts to prevent label truncation
- Fix ModelDrawsComparison: pass quadData so predictions render for three-year models
- Change error bars from 95% to 90% (legend text + SD calculation)
- Widen chart viewport: max-width 70rem vs default 60rem text width
- Fix deep link: fetch district name from API instead of showing 'Loading...'
Charts and their titles/captions need more horizontal space for readability. Single column layout ensures each chart renders at full container width without cramped side-by-side placement.
Source maps cause harmless but noisy 404 errors in browser dev tools when deployed to git-pages, since the hashed filenames change between local and CI builds. Disabling them eliminates these warnings without affecting app functionality.
The obsWidth, lowerX, upperWidth, and medX variables were declared inline within the SVG <g> element's children, which is invalid JavaScript/JSX. This caused 'ReferenceError: obsWidth is not defined' at runtime when rendering Chart 5.
Fix: Moved all const declarations to the top of the map callback, before the return statement.
- Add missing data/national_rates.json fixture (was only in src/data/, not served by Vite)
- LoadingAnimation.jsx: Use import.meta.env.BASE_URL for national rates path
- Ensures file is available at /crdc-demo/data/national_rates.json on git-pages
- App.jsx: Use import.meta.env.BASE_URL prefix for civilytics-logo.svg (was hardcoded as '/civilytics-logo.svg' → 404 on pages.civilytics.org/crdc-demo/)
- ChartPanel.jsx: Same fix for /data/national_rates.json path — was causing JSON.parse errors because the 404 HTML response couldn't be parsed
- App.jsx: Fix 'search another district' button href to use /crdc-demo/ instead of / (root)
- LoadingAnimation.jsx: Replace broken inline fetch helper with centralized api client from useApi.js — was bypassing envelope unwrapping ({status:'success',data:[...]}), causing all 43 API calls to fail silently
- ChartPanel.jsx: Same fix — replace raw fetch() calls with api.fetchDistrictEstimates(), fix national_rates.json path (remove hardcoded /crdc-demo/ prefix that breaks in dev), parallelize wave/model fetches with Promise.all for faster loading
Root causes fixed:
1. Navigation button linked to site root instead of app subdirectory → users lost their way after selecting a district
2. Dual API client implementations — inline fetch helpers didn't unwrap the JSON envelope structure, causing data parsing failures that made the page appear stuck on 'Organizing data...'
3. Sequential fetches in ChartPanel caused slow loading; parallelized for better UX
- Add VITE_PROXY_URL env var to route requests through a CORS proxy when
deployed as static site (CRDC API doesn't send Access-Control-Allow-Origin)
- Update useApi.js, LoadingAnimation.jsx, and ChartPanel.jsx to respect proxy
- Add nginx.conf with /api/v1/ reverse proxy + CORS headers for Docker deployment
- Include proxy.php — simple PHP CORS proxy for git-pages hosts that support PHP
- Document CORS workaround in README
Without this fix, browser fetch() calls are blocked by same-origin policy and
the app shows a blank white screen despite serving correct HTML.
- Set vite.config.mjs base to '/crdc-demo/' so asset paths resolve correctly
when served from a subdirectory (not root)
- Add favicon.svg in public/ folder
- Update index.html with correct absolute paths including /crdc-demo/ prefix
Replaces the GitHub Actions reference with .gitea/workflows/pages.yml,
following the sln-school-comparison deployment pattern. Documents the
git-pages.civilytics.org/crdc-demo/ URL and token configuration.
Adapts the sln-school-comparison pattern: builds Vite static site in CI,
then copies dist/ to a pages branch served by the git-pages webhook at
pages.civilytics.org/crdc-demo/
React + Vite static site demonstrating the CRDC School Arrest Rate API.
Features 6 interactive charts comparing observed arrest data against Bayesian
model estimates across U.S. school districts, with Civilytics visual identity.
- State selector and district search with 'interesting' suggestions (top arrests)
- Animated histogram loading grid showing posterior draw progress
- Charts: time series, rate by group, district vs national, model comparison quadrants, density proxy, exceedance probability
- Static JSON fixture for national rates (no API changes needed)
- Dockerfile + docker-compose.yml for self-hosted deployment
- GitHub Actions workflow and _config.yml for Git Pages