Compare commits

56 Commits
Author SHA1 Message Date
jared 587e74ca80 chore: add CLAUDE.md that imports AGENTS.md
Claude Code loads a project's AGENTS.md only when the project has no CLAUDE.md, and then
also loads every AGENTS.md and .claude/AGENTS.md above it, including ~/AGENTS.md and
~/.claude/AGENTS.md. This one-line import keeps the project's instructions to its own
AGENTS.md, matching the other Civilytics repos.
2026-10-11 16:57:01 -04:00
jared 654b42ca71 fix: raise MAX_RATE_DOMAIN from 30 to 100 per 1,000 students
Deploy to git-pages / deploy (push) Successful in 19s
The previous cap of 30 was exceeded by over 60% of districts — mostly sparse
groups with small enrollment cells whose Agresti-Coull upper bounds genuinely
extend that far. Analysis across all 51 states (16,279 districts) showed the
median max x-value is already ~56 per 1,000; only truly degenerate cases like a
single predicted arrest in a four-student cell (~250/1000) need capping.

A cap of 100 still guards against these outliers while letting realistic data
drive the axis for the vast majority of districts. The existing 'clipped' flag
and note mechanism remain unchanged — they activate only when extreme values are
encountered.
2026-08-22 20:29:32 -04:00
jared fb0beaafc8 fix: compute the district total's interval as HPD, matching the API
Deploy to git-pages / deploy (push) Successful in 16s
The API's count_lower/count_upper are highest-density bounds — hpd_bounds_sql()
in crdc-arrests/R/summarize_draws.R — not equal-tailed quantiles, and the white
paper's figures use the same. The draws-based total introduced in the previous
commit used 2.5/97.5 quantiles, so the chart's two modes labelled "95% interval"
meant two different things depending on whether the draws had been fetched.

Measured on Clark County, the difference is small — 2 to 3 arrests on a band of
roughly 50, about a pixel — but it is a distinction the page offers no
explanation for, and the fallback mode's bounds are already HPD.

  wave     equal-tailed   HPD        delta
  2015-16  215-276        219-277    width -3
  2017-18  175-230        173-226    width -2
  2021-22  116-164        115-161    width -2

hpdBounds() takes the narrowest window covering the mass, contiguous as the R
original is. Rendered values read back off the chart geometry now match an
offline DuckDB computation exactly.
2026-08-12 09:34:53 -04:00
jared 15f8a86409 feat: take the arrests-over-time interval from the draws, not summed bounds
Deploy to git-pages / deploy (push) Successful in 16s
The band on the time-series chart summed each student group's own
count_lower/count_upper. That is not the interval of the district total: the sum
of eight 97.5th percentiles is the case where every group lands at its extreme
in the same draw, which is far less likely than any one group doing so. The
chart was overstating uncertainty by roughly a factor of two.

Now summed within each draw and summarized across draws — build_state_summary()'s
method from crdc-arrests/R/summarize_draws.R on a different axis, the same
operation poolBySex already performs. Clark County NV, 95% interval of the total:

  wave     draws-based        summed bounds     change
  2015-16  215-276 (w 61)     172-322 (w 150)   -60%
  2017-18  175-230 (w 55)     134-269 (w 135)   -60%
  2021-22  116-164 (w 48)      78-188 (w 110)   -57%

Medians barely move (245->246, 200->202, 137->138), which is the expected
signature: the sum of medians already approximates the median of the sum. It is
the tails that were wrong. Values verified against an offline DuckDB computation
over the same shards and then read back off the rendered SVG geometry.

Cost is very uneven, so it is not unconditional. Three waves is 0.31MB for
Nevada but 13.9MB for California and 14.3MB for Texas. Directory listings
(~1KB/year, and already fetched for part discovery) carry each part's byte size,
so useDistrictTotalDraws probes the total up front: under 3MB it fetches
automatically, above it the chart offers a button naming the real size
("Compute the exact interval from the draws (13.9 MB)"). Measured: Nevada adds
no perceptible time at all — the totals resolve in the same frame as the density
panel, since both wait on the duckdb-wasm engine that dominates a cold load —
and California's opt-in completes in 1.3s.

totalPerDraw refuses to sum unless every group has a complete draw set, and
fetchModelCounts now reports how many groups it dropped: for a per-group density
a missing group is a gap, but for a total it is an undercount that would shift
the whole interval down while looking entirely plausible.

The caption states which interval is on screen in either case, and names the
assumption behind the draws-based one — draw_id is renumbered per write batch
upstream, so summing at a draw index convolves what are effectively independent
draws. That is defensible for a posterior *predictive* total, whose observation
noise is independent across groups by construction and dominates the
parameter-level covariance, but it should be read as a predictive total rather
than a correlation-preserving contrast.
2026-08-12 09:31:04 -04:00
jared 9df4f4b5aa Tweak footnote wording
Deploy to git-pages / deploy (push) Successful in 17s
2026-08-12 09:22:07 -04:00
jared fd7acddeb2 docs: note the student groups the research does not model
Deploy to git-pages / deploy (push) Successful in 17s
The CRDC also reports arrests for Asian, Hawaiian/Pacific Islander, and
multiracial students. Nothing on the results page said so, which leaves a reader
to assume the four race groups shown are the whole picture.

The methodology block now names the three groups and gives the paper's own
reason. Worth stating precisely, because it is not a small-sample argument:
white_paper.qmd footnote 4 explains the focus on White, Black, Hispanic and
American Indian students as an analytic choice — prior work documents the
starkest disparities for Black and American Indian students relative to White
students, with Hispanic students included for their share of enrollment — and
line 309 gives the second reason, that the stratified specifications are fit per
group, so covering all 14 reported combinations would mean 70 models instead of
40. Small sample size is the stated reason for excluding *disability* status
(line 168), not race groups.

The note says plainly that the absence reflects scope rather than a finding, and
that no estimate for those students should be inferred from this page. Also
mentions the two upstream exclusions a reader of the totals would otherwise
trip over: arrests not disaggregable by both race and sex (including Section 504
students), and schools not enrolling grade 7 or above.
2026-08-12 09:00:23 -04:00
jared c8757cd786 fix: pack density labels into lanes and match the line chart's scale
Two visual defects reported against the deployed site.

Direct labels collided. Each density is normalized to its own peak so the
curves stay comparable in shape, which means every curve peaks at the *same*
height — so labelling "at the apex" put every label on one line, and the
alternating two-row offset only ever separated two of them. Three groups with
similar rates rendered as "HiWhite:F: Black F".

Labels now live in a reserved band above each row and are packed into lanes by
utils/labelLayout.js: first-fit by x, dropping to a new lane only where the
previous one is occupied, so well-separated groups still share a lane and the
common case stays compact. Rows size themselves from the lane count. Widths are
estimated from character count — SVG text can't be measured before render — and
the estimate is deliberately generous so packing errs toward separation.
Verified by measuring rendered getBBox rects in the browser: zero overlaps for
Clark County unpooled, pooled with all four races checked, and compare-all-four
(16 labels per panel), with nothing outside the viewBox.

The line chart looked like it came from a different app because it did: its
viewBox was 360 wide where the other charts are 760. Both render at width="100%"
in the same card, so its 0.7rem text was scaled up roughly twice as far. Now on
the same 760 grid with matching type sizes (ticks 0.62rem, axis titles 0.64rem),
the beige plot fill dropped to match the other cards, and a baseline under the
waves.

Series ends had no room: the first and last waves sat flush against the plot
edges, clipping half of each end diamond and forcing their labels to be anchored
outward to avoid overflowing. X_PAD insets the scale so every marker has 88px of
clearance and every label centres over its own point. The rate label is one line
("1.71 per 1,000") instead of a stacked number and unit, and the modeled
point-range is thinner and slightly transparent so the observed series reads as
primary.

Also fixes a note that fired too early: ApproxNote rendered while the draws
fetch was still in flight, because with no draws yet every group counts as
approximated — claiming a fallback that hadn't happened, directly above a
spinner saying the real draws were still coming.
2026-08-12 09:00:22 -04:00
jared 4396a56a4a docs: rewrite AGENTS.md and the README chart list for the rebuilt app
Deploy to git-pages / deploy (push) Successful in 23s
AGENTS.md described 6 charts, D3 selections, and DistrictVsNational /
ModelDrawsComparison / ExceedanceProbability — none of which exist. An agent
reading it as authoritative would have been actively misled, so it now opens by
saying src/ wins any disagreement.

Rewritten: the real component tree and data flow, ChartPanel as the owner of all
cross-chart state, the pooled/unpooled key namespace trap, the four properties
of the draws pipeline that are easy to break, the multi-part shard layout and
why discovery uses the tree listing, the "posterior predictive draws" wording
rule and the reason for it, the 95% convention, and a table of which tuning
decisions carry stated rationale and should not be re-derived (the CVD-validated
palette, the KDE bandwidth clamp, the pooling threshold, the mass/KDE cutoff,
the axis cap).

Adds a testing section — there are automated tests now — and notes that the
deployment check should use a multi-part state like California, since a
single-part state cannot catch a regression in part discovery. Records that
App.jsx's "Search another district" is a full page reload that discards the
shard cache.

README: the chart list becomes the summary table plus the two ported figures,
the file tree matches src/, the d3 role is stated precisely (scales and paths,
not selections), and the endpoint table warns that /estimates?state= ranks by
LEAID on a short read.

Also commits the plan this work followed.
2026-08-12 08:41:01 -04:00
jared c1920c11fb feat: rebuild the results page around the white-paper figures
Ports Figs 6 and 7 from crdc-arrests/R/paper_figures.R (wp_fig_group_density and
wp_fig_group_difference) so someone who has read the paper sees the paper. The
three old charts become a summary table and two charts; ArrestsOverTime stays.

Draws pipeline. useDrawDistribution now selects draw_id and returns predicted
counts indexed by draw rather than pre-divided rates. That one column is what
unlocks the rest: pooling has to sum numerators and denominators separately, and
a between-group difference has to be taken at a common draw index. Indexing by
draw_id rather than push order makes DuckDB's row ordering irrelevant and turns
a missing draw into a hole, which isCompleteDrawSet then rejects — a group
present for 300 of 500 draws would otherwise get an interval computed off a
biased subsample that looks identical on screen to a complete one. It also takes
models[] instead of a single model, so "compare all four" needs no conditional
hooks. The reset-before-guard ordering is preserved.

Also fixes a pre-existing bug in that pipeline: a state's draws are split across
data_0.parquet, data_1.parquet, … and the part count varies by state. The app
only ever fetched data_0. Nevada has one part, so this was invisible in every
Nevada test; California has eight, totalling 6.2MB, of which data_0 is 37KB and
holds 11 of California's 1,715 districts. Every other CA district looked absent
from the published data and silently fell back to the approximation. Parts are
now discovered from the Hugging Face tree listing API and fetched in parallel —
listing rather than probing data_N until a 404, because the browser logs a 404
as a console error however cleanly the fetch handles it, and a red error on
every load is indistinguishable from a real one. HEAD probing is the fallback.

New pure utils, all written against tests first:
  agrestiCoull    faithful port incl. the zero-numerator rule of three and the
                  negative lower bound at (1, 53); pinned to five R outputs
  pooling         sex pooling for sparse districts; numerator and denominator
                  are always drawn from the same set of groups
  densityProfile  discrete probability mass below 12 distinct values, KDE above;
                  KDE delegates to kde.js, whose bandwidth clamp is untouched
  districtGroups  display-row derivation, defaults, pooled vs unpooled keys
  groupDifference per-draw delta; refuses to pair mismatched draw sets
  rateDomain      shared x-axis, with a clip flag so the cap is never silent

Chart A keeps the palette contract: race is hue, sex is position. Female and
Male are stacked panels sharing one axis, collapsing to one panel when pooled —
never a second hue. Chart B uses a diverging ramp centred at zero rather than
the paper's sequential YlOrRd, because delta is signed and a sequential ramp
encodes "more" where the data means "which direction".

Captions say "posterior predictive draws", never "paired parameter draws":
draw_id is renumbered per write batch upstream and a district's groups land in
different batches, so cross-group pairing is effectively independent (measured
cor ~= 0.02). The published figure has the same property; what neither can claim
is a paired-parameter contrast.

Sparse districts (under 20 arrests district-wide) pool Female and Male within
each race, with a banner stating the rule and a switch to override it. A pooled
group carries no modelled interval — summing two groups' interval bounds is not
a pooled interval, and there is no honest way to fake one without the draws.

React still owns the DOM. d3-scale/shape/array/interpolate supply scales, path
generators and colour interpolation; no selections, no useEffect DOM mutation.

Deletes RateByGroupBar and RateDensityRidgeline. Keeps distributionApprox.js and
ApproxNote.jsx — still the per-group fallback when draws can't be fetched.

112 tests pass; npm run build clean.
2026-08-12 08:40:46 -04:00
jared 63413f9eb7 fix: rank district suggestions over the full state, not the first 500 rows
DistrictSearch built its "most arrests" list from
/estimates?state=XX&year=21-22&limit=500. That endpoint returns rows
ORDER BY LEAID, RACE, SEX at eight rows per district, so a 500-row cap is not a
sample of the state — it is the ~62 lowest-LEAID districts in it.

Measured against California (11,488 rows, 1,715 districts): the old read covered
68 districts, and 6 of the true top 8 were invisible to it. It suggested
districts with 1 and 2 arrests as the state's most notable, while San Diego
Unified (178), Fresno Unified (77) and Kern High (69) never appeared.

Replaced with a committed fixture, public/data/top_districts.json, generated by
scripts/build-top-districts.mjs. The script pages each state to completion using
meta.total from the response envelope and fails loudly on a short read, since a
silent truncation there would reintroduce exactly this bug. 51 states, 135
requests, ~104KB, following the national_rates.json precedent. Re-run it only
when a new CRDC wave lands.

The search screen also loses a multi-second fetch on every visit, and the
hardcoded "Try Derby (KS), Paterson (NJ)" hint goes with it — the real list
supersedes it. Degrades to search-only if the fixture is missing.

fetchStateDistricts() is kept for scripts and ad-hoc use, with its JSDoc now
warning that any short read ranks by LEAID.

pages.yml gains public/** in its paths filter: the fixture ships with the build,
so regenerating it has to be able to trigger a deploy on its own.
2026-08-12 08:40:23 -04:00
jared c62d1e3068 fix: report the API's 95% intervals as 95%, not 90%
The API returns 95% intervals: validate_interval() in
crdc-arrests/api/R/validate.R defaults to 95L and this app never passes
`interval=`. Two places claimed 90% anyway.

fitSkewedInterval's `intervalMass` defaulted to 0.90, so the analytic fallback
fitted 95% bounds as if they covered 90% of the mass. That divides each
half-interval by 1.645 instead of 1.960 and understates sigma by ~16% — the
fallback drew a distribution visibly narrower than the model's own, in the one
code path where we have no draws to check it against.

ArrestsOverTime's legend read "Modeled (median + 90% interval)" while plotting
count_lower/count_upper, which are the same 95% bounds.

Also exports probit() from distributionApprox.js so the Agresti-Coull port can
reuse it rather than carrying a second qnorm implementation.
2026-08-12 08:40:02 -04:00
jared db5111dab9 docs: correct the error-bar convention and README to match the actual app
Deploy to git-pages / deploy (push) Successful in 22s
AGENTS.md's "Error Bar Convention" paragraph was wrong in four ways after the
last docs pass: it pointed "above" at a section that is below it, described the
fallback as "synthetic draw generation" when it generates no draws at all, cited
an `intervalWidth / 3.29` expression in RateDensityRidgeline.jsx that does not
exist, and listed ModelDrawsComparison.jsx, a file that was deleted when the
demo was reduced to 3 charts. Rewritten against the current code: the fallback
is fitSkewedInterval's analytic two-piece normal, whose divisor is
probit((1 + intervalMass) / 2) with intervalMass defaulting to 0.90, and the
"if you change this" list now names the three files that actually encode 90%.

README: 6 charts -> 3, the ridgeline is Chart 3 not Chart 5, Chart 2's use of
real draws is now mentioned, the D3.js v7 claim is dropped (d3 is not a
dependency), @duckdb/duckdb-wasm is listed in the tech stack, the file tree
matches src/, and the container-size note accounts for the wasm engine.
2026-08-11 10:37:37 -04:00
jared 096455944e fix(nginx): cache and compress the duckdb-wasm asset
The wasm engine is the single largest asset in the build (39,362,651 bytes) but
the immutable-cache location regex did not include `wasm`, so on the
Docker/nginx path it got no Cache-Control at all, and no gzip directive existed
anywhere -- nginx's default gzip_types is text/html only, so it shipped
uncompressed on every cold load.

Verified against nginx:alpine (1.31.3) with the real dist/ mounted:
`nginx -t` passes, and the wasm now returns Content-Type: application/wasm,
Content-Encoding: gzip, Cache-Control: public, max-age=31536000, immutable --
8,766,496 bytes on the wire instead of 39,362,651.

Dynamic gzip rather than gzip_static because the Vite build emits no
pre-compressed .gz files.
2026-08-11 10:37:27 -04:00
jared 420fc41571 fix: validate the deep-link state param before propagating it
?state= was used verbatim and ends up interpolated into a fetch URL and into
the Hugging Face parquet shard path duckdb-wasm reads. Impact is low -- it is a
client-only fetch to a public HTTPS URL, and the wasm sandbox loads no httpfs
-- but validating at the boundary is cheap and correct. Deep links now require
/^[A-Z]{2}$/ (after trim + uppercase, so ?state=az still works) and fall back
to the state selector with a warning when they do not match.
2026-08-11 10:37:26 -04:00
jared 70a12cad22 fix: gate the "real posterior draws" claim on per-group coverage
useDrawDistribution's status is an any-group signal -- 'ready' means at least
one group came back with draws -- but both charts used it chart-wide to hide
<ApproxNote /> and print a caption claiming every box/ridge is real draws.
Individual boxes and ridges already fall back per-group, so the note and
caption were the only things over-claiming.

This is reachable with real data, not just in theory: for LEAID 0400311 (AZ,
unified_m4_mod) the estimates API returns all 8 race x sex groups but the
parquet shard contains draws for WH_M only. The chart reported 'ready' and hid
the note while 7 of its 8 ridges were the analytic approximation.

New src/utils/drawGroups.js owns the "RACE_SEX" key format (previously
duplicated across the hook and both charts) and a hasDrawsForAll() coverage
check. Each chart now checks the groups it actually renders -- RACE_ORDER x its
sex panels/columns -- rather than everything the API returned. A null map
(loading, or a fetch that failed outright) fails the check, so the error path
still shows the note.

Also fixes a stale-state hazard in the same hook: the input guard ran before
setStatus('loading')/setDrawsByGroup(null), so switching to a model whose
groups list is empty (that model's upstream fetch failed) left status at
'ready' with the previous model's draws still in state. Verified by
instrumenting the hook: with the old ordering, switching from a ready model to
an empty one kept status 'ready' and all 8 previous draw keys; with the reset
moved above the guard it correctly resets to 'loading' with no draws.
2026-08-11 10:37:15 -04:00
jared 138a083c6f fix: clamp KDE bandwidth to the plotting domain, not absolute units
Small districts produce rare-event count posteriors that are heavily
zero-inflated, and the absolute 1e-3 floor / raw-sd fallback broke on both
extremes of that data:

- All-identical draws (e.g. 500 zeros, the norm for a group with a handful of
  students) collapsed to h = 1e-3, a near-delta spike of density ~399. Because
  SexRidgeColumn shares one maxPdf per column, that single spike flattened every
  other ridge in the column to sub-pixel height. Measured on a real district
  (0400315 AZ, unified_m4_mod): the three informative ridges rendered at
  0.07-0.13px of a 47.56px row.
- When IQR is 0 -- the normal case when most draws are 0 -- min(sd, iqr/1.34)
  was falsy and the rule fell all the way back to raw sd, which a few extreme
  draws inflate until the ridge is a flat line claiming maximal uncertainty.

Bandwidth is now clamped to [domainWidth/50, domainWidth/6] and the robust rule
degrades by picking the smallest *positive* spread estimate instead of
discarding robustness entirely. The floor is slightly wider than one render
step at n = 60, so a degenerate draw set resolves as a narrow bump; it does not
bind on an ordinary posterior (spread wider than ~8% of the domain keeps its
own Silverman bandwidth). On the district above the informative ridges now
render at 5.66-5.70px, a ~60x improvement.

Five new tests cover both failure modes; all five fail against the old formula.
2026-08-11 10:37:01 -04:00
jaredandClaude Sonnet 5 b5ebf6bec2 docs: clarify that synthetic draw generation is the fallback mechanism
AGENTS.md Error Bar Convention section now explicitly states that synthetic
draw generation from normal approximation is only used when real draws can't
be fetched (network error, unsupported browser, HF outage). Makes clear the
fallback is secondary, not primary, to the real-draw mechanism described in
the new §5.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-11 10:02:19 -04:00
jaredandClaude Sonnet 5 523c21d78c docs: reflect real posterior-draw architecture in AGENTS/README/HANDOFF
- AGENTS.md: Replace §4 (API Endpoint Availability) with updated text; insert
  new §5 (Real posterior draws via duckdb-wasm) documenting the shift from
  synthetic normal-approximation draws to client-side fetches via duckdb-wasm
  against the public Hugging Face parquet dataset. Include actual payload size
  (~39MB uncompressed / ~8.86MB gzipped). Renumber subsequent items.
- README.md: Update Chart 5 description in "What It Does" to reflect real draws
  + fallback behavior. Update API table to clarify that /api/v1/draws is not
  called from app but informs the Hugging Face URL the app fetches directly.
- HANDOFF.md: Mark "Raw posterior draws" as done (2026-08-11) with reference
  to the design spec.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-11 09:57:12 -04:00
jared 20c08ae198 fix: clear stale drawsByGroup when a model switch starts or fails
useDrawDistribution never cleared drawsByGroup on a new fetch or on
error, only status. A chart that varies `model` across renders (Chart
3's dropdown) could switch from a model that had succeeded to one
that then failed, and keep rendering the *previous* model's real
posterior draws — keyed by the same RACE_SEX strings — under the
newly-selected model's summary stats, with ApproxNote visible
suggesting (wrongly) that the fallback approximation was in use.

Clear drawsByGroup to null both when a new fetch starts and in the
catch branch, so a failed model switch never mixes draws from two
different models.
2026-08-11 09:43:09 -04:00
jared b5941a78d1 feat: drive Chart 3's ridgelines from real posterior draws, with fallback
Wire useDrawDistribution into RateDensityRidgeline so each race×sex ridge
uses a Gaussian KDE over that group's 500 real posterior draws when the
selected model's shard has loaded, falling back per-group to the analytic
fitSkewedInterval approximation otherwise. ApproxNote now only shows when
status !== 'ready'. Pass leaid/state through from ChartPanel.
2026-08-11 09:30:46 -04:00
jared ea6d554e3d fix: don't poison shard cache on failure; require real draws for ready status
Two Important review findings on useDrawDistribution:
- ensureShardRegistered cached the rejected promise on fetch/registration
  failure, permanently stuck for that (model,year,state) key until a full
  page reload. Now deletes the cache entry on failure so the next caller
  retries fresh.
- status could become 'ready' even when the LEAID-bound query returned zero
  usable rows (district not present in the draws shard), silently mislabeling
  a fitSkewedInterval-only render as real posterior draws. Now requires at
  least one group with actual draws before reporting 'ready'.
2026-08-11 09:22:01 -04:00
jared d5bf18482a feat: drive Chart 2's box plot from real posterior draws, with fallback 2026-08-11 09:07:53 -04:00
jared 91e21feb74 feat: add empirical KDE/quantile utilities for real posterior draws 2026-08-11 08:55:06 -04:00
jared 88b09a5d2e feat: add duckdb-wasm singleton for client-side parquet queries
Adds src/utils/duckdbClient.js, a lazy-initialized getDb() singleton
wrapping the AsyncDuckDB MVP (single-threaded) wasm bundle. Pins
@duckdb/duckdb-wasm to ^1.32.0 (npm's latest dist-tag currently points
at a -dev prerelease). Excludes the package from Vite's dev-server
dependency pre-bundling so its worker/wasm ?url imports resolve
correctly.

Verified live in the browser under the /crdc-demo/ base path: wasm and
worker assets load with 200, and a SELECT 42 query round-trips
correctly through the singleton, confirming the risk flagged in the
design spec (duckdb-wasm loading correctly from a subpath-served Vite
dev server) does not materialize.
2026-08-11 08:49:05 -04:00
jared 991a6912f8 docs: add implementation plan for empirical draw distributions via duckdb-wasm 2026-08-11 08:42:01 -04:00
jared 2cb04ad9c1 docs: add design spec for empirical draw distributions via duckdb-wasm
Scopes replacing the analytic distribution approximation with real
posterior draws, fetched client-side from the Hugging Face parquet
dataset using duckdb-wasm — no new backend endpoint required.
2026-08-11 07:41:21 -04:00
jaredandClaude Sonnet 5 466ebdb9a6 Align modeled intervals with their observed points, use diamonds, name the model
Deploy to git-pages / deploy (push) Successful in 11s
The modeled point-range was dodged 18px to the side of its observed point,
which read as misaligned rather than paired. Removes the dodge so both sit
on the same x column - directly comparable at a glance, matching the
reference whitepaper figures' own convention.

Observed points were plain circles, inconsistent with every other chart in
the app where a diamond means "observed." Switches them to diamonds, drawn
last so they stay on top of the modeled marks sharing their column.

Adds a caption naming the model the modeled interval is drawn from (three-
year, no covariate / unified_m3_mod), sourced from ChartPanel's existing
WAVE_MODEL constant via the shared MODEL_QUADRANT_LABEL lookup rather than
hardcoded, so it can't drift if the model choice changes.

Also gives the rate-per-1k labels a paper-colored text halo (paintOrder:
stroke) so they stay legible now that they can sit directly over the
modeled whisker line instead of needing to dodge around it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 19:44:28 -04:00
jaredandClaude Sonnet 5 fa35d97269 Clean up the hero chart: remove watermark, fix label collisions, round axis ticks
Deploy to git-pages / deploy (push) Successful in 12s
The district-name watermark inside the arrests-over-time chart was
redundant with the page's own heading and was the direct cause of a label
collision (a high point's rate label rendered on top of it). Removes it.

Two more collisions, both edge cases the screenshot happened to hit: the
first wave's label sat close enough to the plot's top that it could
overlap the topmost gridline text, and a point sitting flush on the plot's
left edge had its center-anchored label bleed into the y-axis tick labels.
Fixes both with more vertical clearance and edge-aware text anchoring
(first point anchors right, last point anchors left, matching the same
fix already applied to chart 4's wave ticks before that chart was cut).

Adds a shared niceTicks() utility (Heckbert's nice-numbers algorithm) so
axis ticks read as round numbers (0/100/200) instead of arbitrary
fractions of the data max (0/113/226/339) — applied to all three charts
for consistency. Also fixes a legend/mark mismatch in the arrests-over-time
chart (the modeled series legend showed a diamond; the actual mark is a
dot) by adding a proper 'dot' shape to ChartLegend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 19:37:02 -04:00
jaredandClaude Sonnet 5 a1ffa600b3 Reduce demo to 3 charts, fix back-button district name bug
Deploy to git-pages / deploy (push) Successful in 26s
Fixes the deep-link/back-button bug where the district name showed
"Unknown District" on return: App.jsx was passing the LEAID as a text
search query to searchDistricts() (a name search), which never matches a
numeric ID. Resolves it instead via fetchDistrictEstimates(leaid, ...) and
reads lea_name/state directly from the returned row - correct regardless
of how the page was reached (fresh load, refresh, or browser back/forward).

Cuts the demo from 6 charts to 3, per request: arrests over time (kept),
arrest rate by student group restructured into Female/Male box-and-whisker
panels (kept), and the posterior density ridge chart restructured from a
2x2 model-quadrant grid into a single selected model (dropdown, default
three-year + referral rate) with Female/Male ridge columns. Removes
DistrictVsNational, ModelDrawsComparison, and ExceedanceProbability
entirely, along with the student-group filter (no longer needed - the
remaining charts always show the full breakdown) and the national-rates
fetch/plumbing that only those removed charts used.

Also updates LoadingAnimation's copy and dedupes its STUDENT_GROUPS/
MODEL_QUADRANTS constants against the shared ones in useApi.js.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 19:27:24 -04:00
jaredandClaude Sonnet 5 cf8e65bb8e Fix chart bugs found on deployed site, redesign Chart 2 as box-and-whisker
Deploy to git-pages / deploy (push) Successful in 12s
Chart 4 (model predictions vs. observed) was picking the first row for a
given wave/model instead of summing across selected groups, so any district
whose first-returned group had zero counts (e.g. a suppressed race group)
showed a collapsed, flat modeled marker instead of the real aggregate
interval. Also fixes the rightmost wave label ("2021-22") clipping off the
edge of the small-multiple panels by anchoring edge ticks away from the
viewBox boundary instead of centering on it.

Chart 6's y-axis group labels (e.g. "American Indian / Alaska Native
Female") were wider than their margin and clipped past the left edge of the
SVG. Adds a shared shortGroupLabel() to colors.js and uses it here, matching
the convention already used in charts 2 and 5.

Chart 2 previously drew a plain bar to the modeled median next to an
observed diamond that often sat well past the bar's end, reading as if the
bar itself were an uncertainty range when it wasn't. Replaces it with an
actual horizontal box-and-whisker: whisker = the API's reported 90%
interval, box = the fitted approximation's 25th/75th percentiles (via a new
fit.quantile() inverse-CDF, exact round-trip of fitSkewedInterval's own
construction), median tick, observed diamond overlaid.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 18:58:35 -04:00
jaredandClaude Sonnet 5 67caa02fe2 Redesign chart visualizations to match Civilytics white-paper style
Deploy to git-pages / deploy (push) Successful in 18s
Rebuilds all 6 charts around a validated categorical race palette, row-based
sex encoding, and a consistent observed-vs-modeled mark convention (diamond
vs. filled bar/density) instead of ad hoc per-chart color schemes. Adds a
shared student-group filter (defaults to all 8 groups) that scopes every
chart's data from one place in ChartPanel.

Drops the D3 dependency entirely in favor of plain SVG, removing the
imperative-DOM bug class behind this app's repeated "fix the fix" commits.
Replaces the fake symmetric-normal posterior approximation with a skewed,
median-preserving fit to the API's interval bounds, clearly labeled as an
approximation. Fixes two broken SVG fill attributes, a decorative model
dropdown that never affected its chart, dead code, an orphaned component,
and a broken CSS token reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 18:45:37 -04:00
jared 6c8e1e8409 docs: refresh AGENTS.md, HANDOFF.md, and README with recent work
- Create AGENTS.md: comprehensive agent guide covering architecture,
  data flow, error bar convention (90% intervals), D3 usage patterns,
  null safety pitfalls, deployment checklist, and related repos
- Update HANDOFF.md: mark CORS as resolved, document chart improvements
  (Chart 5 ridgeline rewrite with synthetic draws + smooth rendering)
- Refresh README.md: accurate tech stack (D3 v7), updated file structure
  with new chart files, corrected API endpoint table
2026-08-10 17:10:35 -04:00
jared 3fe004356e Fix ridgeline chart to render smooth density curves instead of individual columns
Deploy to git-pages / deploy (push) Successful in 13s
- Rewrite RateDensityRidgeline.jsx with proper histogram-to-density conversion
- Build top-edge points from binned draws, then smooth with d3.curveBasis
- Construct complete area path: bottom edge + smoothed top + close
- Position observed rate diamonds above the ridge peak
2026-08-10 17:05:38 -04:00
jared fc0b7e9b6d Fix: remove broken .replace() call causing TypeError in RateDensityRidgeline
Deploy to git-pages / deploy (push) Successful in 12s
2026-08-10 16:56:29 -04:00
jared c32aa46f7c Fix: null check before accessing modelDraws[q.key]
Deploy to git-pages / deploy (push) Successful in 12s
2026-08-10 16:51:53 -04:00
jared c64682228d Fix: null checks for quadData in ChartPanel and child components
Deploy to git-pages / deploy (push) Successful in 14s
2026-08-10 16:43:30 -04:00
jared 83417368f2 Fix: handle null modelDraws gracefully
Deploy to git-pages / deploy (push) Successful in 13s
2026-08-10 16:37:40 -04:00
jared be647d0882 Fix: remove draws fetch (API endpoint not available for browser)
Deploy to git-pages / deploy (push) Successful in 12s
The /estimates/{leaid}/draws endpoint doesn't exist - only bulk Parquet
shard access is available. Falls back to interval-based rendering.
2026-08-10 16:32:40 -04:00
jared 54814564a9 Hook up raw posterior draws from API to ModelDrawsComparison
Deploy to git-pages / deploy (push) Successful in 11s
- Add fetchDistrictDraws call in ChartPanel for all 4 quadrant models
- Pass modelDraws to ModelDrawsComparison component
- Rewrite Chart 4 with D3 showing:
  - Density histogram from raw draws when available (per-year distribution)
  - Fall back to interval-based rendering otherwise
  - 90% credible intervals from quantile calculation
  - Legend explaining colors (red=observed, teal=one-year, navy=three-year)
2026-08-10 16:27:47 -04:00
jared f65a70fc3e Add D3 and new ridgeline chart matching R aesthetic
Deploy to git-pages / deploy (push) Successful in 19s
- Install d3@7
- Add fetchDistrictDraws API function (for future use with raw draws)
- Create RateDensityRidgeline component using D3.js
  - Density ridges showing posterior distributions per race group
  - Diamond markers for observed rates (matching R design)
  - Dropdown to switch between model specifications
  - Matches Civilytics color palette (navy fill, danger diamonds)
2026-08-10 16:15:33 -04:00
jared c4c0aa63bd Use three-year model (unified_m3_mod) to fetch observed data for all waves
Deploy to git-pages / deploy (push) Successful in 11s
One-year models don't return historical arrest data; three-year models
include observed counts across all CRDC waves (15-16, 17-18, 21-22).
2026-08-10 13:19:39 -04:00
jared 0fa5844a16 Replace logo with Civilytics wordmark from civilyticsR package
Deploy to git-pages / deploy (push) Successful in 11s
2026-08-10 12:54:55 -04:00
jared b135f480c0 Fix chart sizing, model predictions, 90% intervals, and deep link district name
Deploy to git-pages / deploy (push) Successful in 17s
- Increase SVG dimensions in all 6 charts to prevent label truncation
- Fix ModelDrawsComparison: pass quadData so predictions render for three-year models
- Change error bars from 95% to 90% (legend text + SD calculation)
- Widen chart viewport: max-width 70rem vs default 60rem text width
- Fix deep link: fetch district name from API instead of showing 'Loading...'
2026-08-10 12:50:19 -04:00
jared eb396abce7 Restore first three charts (ArrestsOverTime, RateByGroupBar, DistrictVsNational) that were accidentally removed in previous edit
Deploy to git-pages / deploy (push) Successful in 11s
2026-08-10 12:27:04 -04:00
jared 77757348f8 Switch ChartPanel from 3-column grid to single vertical stack
Deploy to git-pages / deploy (push) Successful in 10s
Charts and their titles/captions need more horizontal space for readability. Single column layout ensures each chart renders at full container width without cramped side-by-side placement.
2026-08-10 12:25:31 -04:00
jared 26f72da6db Disable source map generation for cleaner production builds
Deploy to git-pages / deploy (push) Successful in 17s
Source maps cause harmless but noisy 404 errors in browser dev tools when deployed to git-pages, since the hashed filenames change between local and CI builds. Disabling them eliminates these warnings without affecting app functionality.
2026-08-10 12:23:27 -04:00
jared dd69bb6f8f Fix JS syntax error in ObservedRateDensity.jsx — const declarations were embedded inside JSX
Deploy to git-pages / deploy (push) Successful in 11s
The obsWidth, lowerX, upperWidth, and medX variables were declared inline within the SVG <g> element's children, which is invalid JavaScript/JSX. This caused 'ReferenceError: obsWidth is not defined' at runtime when rendering Chart 5.

Fix: Moved all const declarations to the top of the map callback, before the return statement.
2026-08-10 12:21:00 -04:00
jared 116730b685 Fix national_rates.json deployment — copy to public/ directory
Deploy to git-pages / deploy (push) Successful in 19s
- Add missing data/national_rates.json fixture (was only in src/data/, not served by Vite)
- LoadingAnimation.jsx: Use import.meta.env.BASE_URL for national rates path
- Ensures file is available at /crdc-demo/data/national_rates.json on git-pages
2026-08-10 12:17:42 -04:00
jared 2c900b003f Fix static asset paths for /crdc-demo subdirectory deployment
Deploy to git-pages / deploy (push) Successful in 10s
- App.jsx: Use import.meta.env.BASE_URL prefix for civilytics-logo.svg (was hardcoded as '/civilytics-logo.svg' → 404 on pages.civilytics.org/crdc-demo/)
- ChartPanel.jsx: Same fix for /data/national_rates.json path — was causing JSON.parse errors because the 404 HTML response couldn't be parsed
2026-08-10 11:44:26 -04:00
jared 8ba7440a34 Fix routing, data fetching, and rendering issues in CRDC demo app:
Deploy to git-pages / deploy (push) Successful in 10s
- App.jsx: Fix 'search another district' button href to use /crdc-demo/ instead of / (root)
- LoadingAnimation.jsx: Replace broken inline fetch helper with centralized api client from useApi.js — was bypassing envelope unwrapping ({status:'success',data:[...]}), causing all 43 API calls to fail silently
- ChartPanel.jsx: Same fix — replace raw fetch() calls with api.fetchDistrictEstimates(), fix national_rates.json path (remove hardcoded /crdc-demo/ prefix that breaks in dev), parallelize wave/model fetches with Promise.all for faster loading

Root causes fixed:
1. Navigation button linked to site root instead of app subdirectory → users lost their way after selecting a district
2. Dual API client implementations — inline fetch helpers didn't unwrap the JSON envelope structure, causing data parsing failures that made the page appear stuck on 'Organizing data...'
3. Sequential fetches in ChartPanel caused slow loading; parallelized for better UX
2026-08-10 11:42:31 -04:00
jared c475bfe50b Add handoff notes for CORS deployment fix 2026-08-10 11:27:03 -04:00
jared 4ce64c7839 Fix CORS: add proxy support for browser-based API calls on git-pages
Deploy to git-pages / deploy (push) Successful in 20s
- Add VITE_PROXY_URL env var to route requests through a CORS proxy when
  deployed as static site (CRDC API doesn't send Access-Control-Allow-Origin)
- Update useApi.js, LoadingAnimation.jsx, and ChartPanel.jsx to respect proxy
- Add nginx.conf with /api/v1/ reverse proxy + CORS headers for Docker deployment
- Include proxy.php — simple PHP CORS proxy for git-pages hosts that support PHP
- Document CORS workaround in README

Without this fix, browser fetch() calls are blocked by same-origin policy and
the app shows a blank white screen despite serving correct HTML.
2026-08-10 11:22:44 -04:00
jared 194208d817 Fix base URL for subdirectory deployment on pages.civilytics.org
Deploy to git-pages / deploy (push) Successful in 10s
- Set vite.config.mjs base to '/crdc-demo/' so asset paths resolve correctly
  when served from a subdirectory (not root)
- Add favicon.svg in public/ folder
- Update index.html with correct absolute paths including /crdc-demo/ prefix
2026-08-10 11:13:16 -04:00
jared ff4b2ac690 Update README: document Gitea Actions pages workflow
Replaces the GitHub Actions reference with .gitea/workflows/pages.yml,
following the sln-school-comparison deployment pattern. Documents the
git-pages.civilytics.org/crdc-demo/ URL and token configuration.
2026-08-10 11:08:00 -04:00
jared 7cf9777673 Add Gitea Actions workflow for git-pages deployment
Deploy to git-pages / deploy (push) Successful in 10s
Adapts the sln-school-comparison pattern: builds Vite static site in CI,
then copies dist/ to a pages branch served by the git-pages webhook at
pages.civilytics.org/crdc-demo/
2026-08-10 10:57:00 -04:00
jared 52a0c77e31 Initial commit: CRDC Arrests API demo app
Deploy to Git Pages / build-and-deploy (push) Failing after 19s
React + Vite static site demonstrating the CRDC School Arrest Rate API.
Features 6 interactive charts comparing observed arrest data against Bayesian
model estimates across U.S. school districts, with Civilytics visual identity.

- State selector and district search with 'interesting' suggestions (top arrests)
- Animated histogram loading grid showing posterior draw progress
- Charts: time series, rate by group, district vs national, model comparison quadrants, density proxy, exceedance probability
- Static JSON fixture for national rates (no API changes needed)
- Dockerfile + docker-compose.yml for self-hosted deployment
- GitHub Actions workflow and _config.yml for Git Pages
2026-08-10 10:54:51 -04:00
76 changed files with 11535 additions and 33 deletions
+66
View File
@@ -0,0 +1,66 @@
name: Deploy to git-pages
# Builds the Vite static site and pushes it to the `pages` branch, which
# the git-pages webhook receiver at pages.civilytics.org serves.
# URL: https://pages.civilytics.org/crdc-demo/
# Same pattern as Civilytics/sln-school-comparison.
on:
push:
branches: [main]
paths:
- 'src/**'
# Committed fixtures ship with the build: public/data/top_districts.json
# is what the search screen ranks its suggestions from, so regenerating it
# has to be able to trigger a deploy on its own.
- 'public/**'
- 'index.html'
- 'vite.config.mjs'
- 'package.json'
- '.gitea/workflows/pages.yml'
workflow_dispatch:
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
cache: 'npm'
- name: Install dependencies & build static site
run: |
npm ci
# VITE_API_BASE defaults to the production API; override if needed via secrets.
# For git-pages deployment, CORS is handled by setting up a proxy endpoint on the server.
npm run build
- name: Publish dist/ to pages branch
run: |
git config user.name "Gitea Actions"
git config user.email "actions@gitea.civilytics.org"
# Use the GITHUB_TOKEN for authentication (works with Gitea token auth)
TOKEN="${{ secrets.GITHUB_TOKEN }}"
if [ -z "$TOKEN" ]; then
echo "GITHUB_TOKEN not available; skipping deploy."
exit 0
fi
git remote set-url origin "https://token:${TOKEN}@gitea.civilytics.org/Civilytics/crdc-demo.git"
# Create orphan pages branch with only the built files
rm -rf /tmp/pages-branch
git worktree add --orphan -B pages /tmp/pages-branch
cp -r dist/. /tmp/pages-branch/
echo "${GITHUB_SHA}" > /tmp/pages-branch/deploy-sha.txt
cd /tmp/pages-branch
git add --all
git commit -m "Deploy from ${GITHUB_SHA}" || echo "No changes to deploy"
git push origin pages --force
+39
View File
@@ -0,0 +1,39 @@
name: Deploy to Git Pages
on:
push:
branches: [main, master]
workflow_dispatch:
jobs:
build-and-deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
cache: 'npm'
- name: Install dependencies
run: npm ci
- name: Build static site
run: |
# VITE_API_BASE defaults to the production API; override if needed via secrets.
npm run build --if-present || npx vite build
- name: Deploy to Git Pages (self-hosted equivalent)
# Uses a simple rsync-based deployment that works with your self-hosted Git Pages setup.
# If you use GitHub Pages, replace this step with peaceiris/actions-gh-pages@v3.
run: |
echo "Build complete — dist/ ready for upload to your Git Pages host."
ls -la dist/
- name: Upload build artifacts
uses: actions/upload-artifact@v4
with:
name: crdc-demo-dist
path: dist/
+8
View File
@@ -0,0 +1,8 @@
node_modules/
dist/
*.local
*.log
.DS_Store
.env
.vscode/
.idea/
+277
View File
@@ -0,0 +1,277 @@
# Agent Guide — CRDC Demo App
Context, decisions, and guidance for agents working on this codebase.
**`src/` is the source of truth.** If this file and the code disagree, the code
wins and this file is the bug — fix it in the same change.
## Project Overview
A React + Vite static web app demonstrating the
[CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/). Visitors
pick a state, search for a school district, and see what was actually reported
alongside what the Bayesian models estimate. Deployed via Gitea Actions to
`pages.civilytics.org/crdc-demo/`.
The results page is a port of the white paper's Figs 6 and 7
(`wp_fig_group_density` / `wp_fig_group_difference` in
`crdc-arrests/R/paper_figures.R:472-568`). Keeping it recognisably the same
figure is the point — someone who has read the paper should see the paper.
## Architecture Summary
- **Frontend**: React 19 + Vite (static site)
- **Styling**: CSS custom properties mirroring the Civilytics design tokens
(`src/styles/tokens.css`)
- **Charts**: inline SVG that **React owns**. `d3-scale`, `d3-shape`,
`d3-array` and `d3-interpolate` supply scales, path generators and colour
interpolation only. **No d3 selections, no `useEffect` DOM mutation** —
if you find yourself reaching for `d3.select`, the answer is a render.
- **API**: public read-only API called directly from the browser; no backend
- **Deployment**: Gitea Actions (`.gitea/workflows/pages.yml`)
## Key Components and Data Flow
```
App.jsx (router)
→ StateSelector landing screen: state dropdown/grid
→ DistrictSearch live /districts search + suggestions from the
│ committed public/data/top_districts.json fixture
→ LoadingAnimation warms the API, animated histogram grid
→ ChartPanel OWNS all cross-chart state (see below)
├── DistrictSummaryTable observed arrests + enrollment; the checkboxes
│ here are the density panel's group control
├── RateDensityPanel Chart A — posterior density per group,
│ Female over Male, Agresti–Coull rail beneath
├── GroupDifference Chart B — posterior of Δ between two groups
└── ArrestsOverTime observed vs. modelled totals across 3 waves
```
### `ChartPanel` owns the state
Selected model specification, sex pooling, which groups are checked, and the
difference pair all live in `ChartPanel` and are passed down. Charts hold none
of it. That is what keeps the table's checkboxes and the density panel from
drifting apart.
One trap worth knowing: pooled group keys (`'BL'`) and unpooled ones (`'BL_F'`)
are different namespaces. Selections are therefore stored **with the pooling
mode they were made in** and fall back to defaults when the mode changes.
### Data fetching
1. **Wave data** (`ArrestsOverTime`): `unified_m3_mod` for `['15-16', '17-18',
'21-22']`. Three-year models are used because they return observed counts
across all waves; one-year models only cover the most recent one.
2. **Current-wave summary**: fetched for the **selected specification only**.
Enrollment and observed arrests are district facts, not model outputs, and
the modelled shapes now come from real draws — prefetching all four specs'
summaries would be four requests for data three of which are never read.
3. **Posterior draws**: `useDrawDistribution` (see below).
## The draws pipeline — read this before touching `useDrawDistribution.js`
`useDrawDistribution({leaid, state, models, year})` fetches the published
posterior draws from the Hugging Face parquet dataset
(`civilytics/crdc-school-arrest-rates`) and queries them client-side with
`@duckdb/duckdb-wasm` (`src/utils/duckdbClient.js`). There is no server-side
draws endpoint.
**A state's draws are split across multiple parquet parts.** `data_0.parquet`,
`data_1.parquet`, … and the count varies by state: Nevada is one file,
California is **eight** (6.2MB total, of which `data_0` is 37KB and holds 11 of
California's 1,715 districts). The app fetched only `data_0` until 2026-08-12,
which made every CA district except those 11 look absent from the published
data and silently fall back to the approximation — invisible in testing because
Nevada, the district everyone tests with, has exactly one part.
Parts are discovered from the Hugging Face **tree listing API**
(`/api/datasets/{id}/tree/main/parquet/...`), not by probing `data_N` until a
404. A 404 is logged as a console error by the browser's network layer however
cleanly the fetch handles it, and a red error on every page load is
indistinguishable from a real one. Probing (via HEAD) remains the fallback if
the listing API is unavailable. All parts are fetched in parallel, registered
individually, and queried as `read_parquet([...])`.
Four more properties it is easy to break:
- **It returns counts, not rates**, indexed by `draw_id - 1`. Counts are what
make sex pooling and between-group differences possible: both have to sum or
subtract numerators and denominators separately. Callers divide.
- **Indexing by `draw_id`, not push order**, makes DuckDB's row ordering
irrelevant and turns a missing draw into a hole rather than a short array.
- **Incomplete groups are dropped** (`isCompleteDrawSet`). A group present for
300 of 500 draws would otherwise get an interval computed off a biased
subsample that looks identical on screen to a complete one.
- **The reset-before-guard ordering in the effect is load-bearing.** State is
cleared on *every* input change, including ones with nothing to fetch, so a
failed fetch can never leave the previous model's draws on screen under the
new model's label.
`status` is an ANY-model, ANY-group signal. To claim "these are all real draws"
for a specific set of rendered groups, use `hasDrawsForAll`.
### Caption wording: "posterior predictive draws"
`draw_id` is renumbered 1–500 per write batch upstream
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in
different batches, so draw *k* of one group is **not** the same parameter draw
as draw *k* of another. Measured correlation between Black-male and
Hispanic-male `pred` in Clark County was 0.019 even within a batch —
`posterior_predict` observation noise dominates.
The published Fig 7 has the same property, so the app matches the paper. What
neither can claim is a paired-parameter contrast. **Captions must say "posterior
predictive draws" and must never say "paired parameter draws".**
### Fallback path — do not delete
If the draws fetch fails (network, unsupported browser, HF outage),
`RateDensityPanel` falls back **per group** to the analytic approximation in
`src/utils/distributionApprox.js` and shows `<ApproxNote />`. Neither
`distributionApprox.js` nor `ApproxNote.jsx` is dead code.
A pooled group has no fallback shape: adding two groups' interval *bounds*
together is not a pooled interval, so `buildDisplayGroups` sets `modeled: null`
when pooling and the chart omits that group rather than inventing a curve.
The duckdb-wasm engine is ~39MB uncompressed / ~8.86MB gzipped (measured
against the shipped package). It loads via dynamic `import()` only once a
district is selected — never on initial page load — and is browser-cached
thereafter, but it is a real one-time cost.
### Error bar convention: 95%
The API returns **95%** intervals. `validate_interval()` in
`crdc-arrests/api/R/validate.R` defaults to `95L` and this app never passes
`interval=`. `fitSkewedInterval`'s `intervalMass` therefore defaults to `0.95`,
and `ArrestsOverTime`'s legend says "95% interval". (Both said 90% before
2026-08-12; that was a bug, not a convention.)
The observed-data point ranges are a different thing again: a 95%
Agresti–Coull interval computed from observed counts
(`src/utils/agrestiCoull.js`), a direct port of `agresti_coull()` in
`crdc-arrests/R/paper_figures.R:219-237`. Two faithfulness quirks are pinned by
tests and must not be "fixed": the bounds are on the **count** scale, and
`lower` can be **negative** for very small numerators (charts clamp at draw
time, the port does not).
## Tuning decisions with stated rationale — don't re-derive
- **`src/utils/colors.js:9-13`** — the race palette passed the dataviz skill's
CVD validator. Re-run `validate_palette.js` before changing any hex value.
Race is hue; **sex is position, never a second hue**; observed-vs-modelled is
mark type, never a second hue.
- **`src/utils/kde.js`** — `BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`
were tuned for zero-inflated sparse-district posteriors. Both ends matter.
- **`src/utils/pooling.js`** — `POOL_BY_SEX_ARREST_THRESHOLD = 20`, applied to
the district total, not per cell.
- **`src/utils/densityProfile.js`** — `MASS_MAX_DISTINCT = 12`. Below it the
posterior predictive is drawn as discrete mass, because it *is* discrete; a
Gaussian KDE over four achievable values renders as a lumpy smear that reads
as a rendering bug.
- **`src/utils/rateDomain.js`** — `MAX_RATE_DOMAIN = 100` caps the axis so a
four-student cell can't squash every other curve (up from 30, which was exceeded by over 60% of districts). It reports `clipped` so the
chart says so instead of silently cropping.
## Common Pitfalls & Gotchas
### 1. Null safety in chart components
API responses can have empty arrays or missing fields for districts with no
arrests. Use optional chaining and explicit defaults, and never divide by a
denominator you haven't checked:
```javascript
const enroll = row.stu_enroll || 0
const rate = enroll > 0 ? (row.observed_arrests || 0) / enroll * 1000 : 0
```
A rate with no denominator is **not zero and not Infinity — it's undefined**.
`toRates` returns `[]`, `buildDisplayGroups` reports `0` and lets the
enrollment column explain why.
### 2. SVG dimensions and responsiveness
Charts set `width="100%"` with a fixed `viewBox`, wrapped in
`overflowX: 'auto'`. Chart cards are `max-width: 70rem` in `ChartPanel.jsx`
(vs. the ~60rem default text width).
### 3. API endpoint availability
- `/api/v1/estimates/{leaid}` — ✅ summary rows (median, bounds, enrollment,
observed arrests)
- `/api/v1/estimates?state=XX&...` — ✅ but returns rows `ORDER BY LEAID, RACE,
SEX` at 8 per district, capped at `limit=1000`. **Any short read ranks the
lowest-LEAID districts, not the busiest.** Page it with `meta.total`.
- `/api/v1/draws?...` — returns a shard URL + SQL, not draw data. The app goes
to the parquet directly.
### 4. CORS
The API sends no CORS headers. The app auto-detects a proxy via `VITE_PROXY_URL`;
unset, it fetches directly (works same-origin or behind the Docker/nginx proxy).
## The suggestion fixture
`public/data/top_districts.json` holds the top 15 districts per state by
observed arrests, and is generated by a one-off, read-only script:
```bash
node scripts/build-top-districts.mjs # all 51, ~150 requests
node scripts/build-top-districts.mjs --states NV,CA # spot-check
```
It is committed (the `national_rates.json` precedent). Re-run it only when a
new CRDC wave lands. `DistrictSearch` degrades to search-only if it's missing.
## Testing
`npm test` runs `node --test 'src/**/*.test.js'`. The pure utilities are all
covered and **should be written test-first**:
| Module | What its tests pin |
|---|---|
| `agrestiCoull.js` | five cases against real R output, incl. the negative lower bound |
| `pooling.js` | numerator and denominator always drawn from the same groups |
| `densityProfile.js` | the mass/KDE switch, and that KDE delegates to `kde.js` |
| `districtGroups.js` | display-row derivation, defaults, pooled vs unpooled keys |
| `groupDifference.js` | refusal to pair mismatched draw sets |
| `rateDomain.js` | the axis cap, and that it reports clipping |
| `drawGroups.js` | key shapes, and hole detection in draw arrays |
| `kde.js` | bandwidth clamp behaviour |
Components are verified manually — there is no DOM test harness. Useful
districts: **Clark County NV `3200060`** (100 arrests / 148,928 students,
pooling off), **Carson City NV `3200390`** (6 arrests, pooling auto-engages,
discrete mass profile), **Washoe County NV `3200480`** (cross-check the table
against the API), and any California district (large shards, fixture ranking).
## Deployment Checklist
1. `npm run build` succeeds
2. `npm test` passes
3. No console errors after a hard refresh
4. Both charts render for a sample district; Network tab shows **one fetch per
(model, state) shard part** and no repeats when switching specs
5. Test with a **multi-part state** (California), not just Nevada — a
single-part state cannot catch a regression in part discovery
> Caveat on the cache: `App.jsx:112`'s "Search another district" button does
> `window.location.href = '/crdc-demo/'`, a full page reload, which discards the
> module-level `shardCache`. Within one district view — switching specs,
> toggling compare-all — the cache works as intended.
## Git Conventions
- Remote: `https://gitea.civilytics.org/Civilytics/crdc-demo.git`
- Default branch: `main`
- Gitea Actions auto-deploys on push to `main`
- Commit messages explain the *why*: `fix: rank suggestions over the full state
(limit=500 was selecting the lowest 62 LEAIDs)`
## Related Repositories
- **crdc-arrests** — API server (Plumber/R) and the white paper, at
`/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/`
- **civilyticsR** — wordmark and visualization functions
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+27
View File
@@ -0,0 +1,27 @@
# Multi-stage build for the CRDC Arrests API Demo app.
# Stage 1: Build static files with Vite + React
FROM node:22-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci --prefer-offline
COPY . .
ARG VITE_API_BASE=https://crdc-api.civilytics.org/api/v1
ENV VITE_API_BASE=${VITE_API_BASE}
ARG VITE_PROXY_URL=""
ENV VITE_PROXY_URL=${VITE_PROXY_URL}
RUN npm run build
# Stage 2: Serve with nginx (tiny image, ~5MB) + reverse proxy for API CORS bypass
FROM nginx:alpine AS runtime
# Copy static site files
COPY --from=builder /app/dist/ /usr/share/nginx/html/
# Custom nginx config — serves static files AND proxies /api/v1/ to the CRDC API
# with CORS headers, so browser-based requests work without preflight issues.
COPY nginx.conf /etc/nginx/conf.d/default.conf
EXPOSE 80
CMD ["nginx", "-g", "daemon off;"]
+52
View File
@@ -0,0 +1,52 @@
# Handoff: Fix CORS for CRDC Arrests API Demo App
## Current Status
**RESOLVED**: The demo app at `https://pages.civilytics.org/crdc-demo/` is fully functional. All charts render correctly and there are no blocking errors in the console.
The original CORS issue was resolved by deploying updated `plumber.R` with CORS headers to the running API server (`crdc-api.civilytics.org`).
## Historical Context (resolved)
### Original Problem
The demo app loaded HTML correctly but showed a blank white screen because browser-based fetch requests to the CRDC API were blocked by CORS (the API didn't send `Access-Control-Allow-Origin` headers).
### Resolution
The CORS issue was fixed at the API source:
- Updated `crdc-arrests/api/plumber.R` with CORS headers (`Access-Control-Allow-Origin: *`) in the existing `cacheHeaders` filter and OPTIONS preflight handling.
- Deployed to the running API server so it took effect at https://crdc-api.civilytics.org/.
The app now calls the public read-only API directly from the browser without requiring a proxy.
## Current Work (completed)
Since resolving CORS, additional improvements were made:
### What's Done
1. ✅ **CORS resolved**: API server (`crdc-api.civilytics.org`) now sends `Access-Control-Allow-Origin: *` headers — deployed via updated `plumber.R`
2. ✅ **App code**: React + Vite app fully built with 6 charts, Civilytics visual identity (wordmark from civilyticsR package), loading animation, district search
3. ✅ **Git Pages deployment**: Gitea Actions workflow (`.gitea/workflows/pages.yml`) successfully builds `dist/` and deploys to a `pages` branch on `Civilytics/crdc-demo` repo at gitea.civilytics.org
4. ✅ **Chart 5 rewrite**: Replaced interval-proxy ridgeline with proper D3 density ridges (`RateDensityRidgeline.jsx`) using synthetic draws from normal approximation of posterior intervals, smooth `curveBasis` rendering, and diamond markers for observed rates
5. ✅ **Chart 4 enhancement**: Added D3-based quadrant visualization in `ModelDrawsComparison.jsx` showing error bars with proper null safety
### Remaining Proxy Code (optional)
The CORS proxy fallback code (`proxy.php`, nginx reverse proxy config, `VITE_PROXY_URL` support) was added but is **not needed** since the API now sends CORS headers. It remains as a safety net for environments where the API can't be modified.
## What Needs to Be Done Next
### Future Enhancements
- ~~**Raw posterior draws**~~ — Done (2026-08-11). Charts 2 and 3 now fetch real posterior draws client-side via `@duckdb/duckdb-wasm` against the public Hugging Face parquet dataset. See `docs/superpowers/specs/2026-08-11-empirical-draws-wasm-design.md`.
- **Automated testing**: No test suite exists; consider adding basic tests for chart rendering and API error handling.
## Key Files
- **API fix**: `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-arrests/api/plumber.R` (CORS headers added to cacheHeaders filter)
- **App code**: `/home/jared/Nextcloud/Civilytics/Code/Civilytics/crdc-demo/src/` — fully built, deployed to pages branch
- **Proxy fallback**: `proxy.php`, nginx.conf with reverse proxy config
## Verification Steps After Fix
1. Visit https://pages.civilytics.org/crdc-demo/
2. Select a state → search for "Denver" → should see district results
3. Click a district → loading animation runs, then 6 charts appear
4. Check browser dev tools → no CORS errors in console
## Git Status
- `crdc-arrests` repo: plumber.R modified (CORS headers added), not yet committed/pushed to API deployment
- `crdc-demo` repo: All fixes pushed and deployed via Gitea Actions, awaiting CORS resolution at API or server level
+360
View File
@@ -0,0 +1,360 @@
# CRDC Arrests API Demonstration App — Design Proposal
## 1. Overview
A web application hosted at `civilytics.org` that demonstrates the [CRDC School
Arrest Rate API](https://crdc-api.civilytics.org/api/v1/) by letting any visitor
explore school-based arrest estimates for any U.S. school district, with a strong
visual narrative built around Bayesian model comparisons.
The app is a **frontend-only static application** that calls the public read-only
API directly from the browser. No server-side code, no database, no auth — just
HTML/CSS/JS (with optional Docker container for local dev or self-hosting behind
a reverse proxy).
---
## 2. User Flow
```
┌─────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ State │ → │ District Search │ → │ Loading (anim) │ → ╔══════════════╗
│ selector │ │ w/ suggestions │ │ patience msg │ ║ Chart panel ║
└─────────────┘ └──────────────────┘ └─────────────────┘ ║ 6 panels ║
║ see §4 ║
╚══════════════╝
```
### Step 1 — State prompt
- Full-screen landing card with Civilytics branding.
- Dropdown / typeahead for U.S. state (50 states + DC). Two-letter codes map to
the API `state` parameter directly.
- On select → slide to Step 2.
### Step 2 — District search
- Search-as-you-type input, hitting `/api/v1/districts?q=<partial>&state=XX`.
- **Suggested districts** appear as a horizontal carousel of "interesting" picks:
- Top N by total arrest count in the most recent wave (2021-22).
- Fetched once per state via `limit=500` from `/estimates?state=XX&year=21-22`
sorted client-side by `observed_arrests`.
- Each suggestion card shows: district name, enrollment (in small), observed
arrests. Clicking a suggestion jumps straight to loading + charts.
- Typing filters the list live; selecting an item from either source proceeds.
### Step 3 — Loading / patience animation
- A full-screen overlay with an animated visual that conveys "data is being
fetched across multiple endpoints, please be patient."
- The animation should reflect the **multi-model comparison** theme: e.g., a grid
of small histogram-like bars (representing posterior draws) that animate in and
out row by row as each model's endpoint responds.
- Estimated total API calls per district lookup: ~12–15 (one per model × race ×
sex combination, filtered to the selected district). The animation should scale
visually with progress.
### Step 4 — Chart panel (6 charts in a 3×2 grid)
All charts follow Civilytics visual identity: warm paper background (`#FAF7F2`),
civic navy text (`#0E1A2B`), ember accent (`#C25311`). Data-viz palette uses the
supporting colors from `_tokens.scss` (teal, plum, moss, brass).
#### Row 1 — Observed & descriptive
**Chart 1: Arrests over time by CRDC wave (line chart)**
- X-axis: CRDC waves (`2015-16`, `2017-18`, `2021-22`).
- Y-axis: total observed arrest count.
- Line + points; each point labeled with the rate per 1,000 students.
- Source: `/estimates?leaid=XXXXX&year=...` across all three years (or a single
call to `/districts/{leaid}` which returns all demographics for one district).
**Chart 2: Arrest rate by student group — most recent year (bar chart)**
- Bars for each race×sex combination in `AM/BL/HI/WH × F/M`.
- Y-axis: arrests per 1,000 students.
- Color-coded by the supporting palette; legend shows full labels.
**Chart 3: District vs. national — top student group (comparison chart)**
- Identify the district's highest-arrest-rate group.
- Side-by-side bars or a small multiples comparison against the corresponding
**national** rate for that same group.
- National rates fetched via `/states?state=XX&race=&sex=&year=21-22` aggregated,
or more precisely from the national summary (which may need to be computed as
an aggregate across all states).
#### Row 2 — Bayesian model distribution comparisons
All four quadrants show results for the **selected district**, comparing:
- **Column A**: One-year models (`unified_m1_mod`, `unified_m2_mod`) vs.
- **Column B**: Three-year models (`unified_m3_mod` through `unified_m5_mod`).
- Within each column, rows differentiate baseline (no covariate) from covariate
models.
**Chart 4: Predicted draws by year vs. observed (scatter / point-range)**
- For three-year models only: for each of the 3 waves, show the model's predicted
median and 95% interval alongside the observed value as a separate marker.
- Layout: x-axis = wave; y-axis = arrest count; points dodge left (model) vs.
right (observed).
**Chart 5: Observed rate per group against model distribution (density / ridge)**
- For each student group, show the observed rate and overlay the posterior draw
distributions from all four model types as ridgeline or violin plots.
- Mirrors `wp_fig_group_density()` in the white paper.
**Chart 6: Probability district exceeds national rate per group (bar chart)**
- For each race×sex group, compute P(district rate > national rate) using the
posterior draws from a chosen model (e.g., the default `unified_m2_mod`).
- Bars colored by threshold crossing; annotated with exact probability.
---
## 3. Data Sources & API Endpoints Used
| Purpose | Endpoint | Frequency |
|---|---|---|
| State list / validation | Hardcoded enum (`ALLOWED_STATES`) | Once |
| District search | `/api/v1/districts?q=&state=` | On keystroke (debounced) |
| "Interesting" suggestions | `/api/v1/estimates?state=XX&year=21-22` sorted by `observed_arrests desc` | Once per state |
| Single district, all demographics | `/api/v1/distimates/{leaid}` or `/estimates?leaid=` with year/model filters | ~6 calls × 3 years = 18+ |
| National comparison rates | `/api/v1/states?state=&race=&sex=&year=21-22` aggregated across states, OR a dedicated national endpoint if available | Once per group |
| Model metadata | `/api/v1/models` | Once (cache) |
**Key data structures returned by the API:**
The `/estimates/{leaid}` endpoint returns one row per `race × sex × year × model` with:
```json
{
"leaid": "...",
"lea_name": "...",
"state": "TX",
"race": "BL",
"sex": "M",
"year": "21-22",
"model": "unified_m2_mod",
"stu_enroll": 35963,
"observed_arrests": 1,
"rate_median": 0.42,
"rate_lower": 0.18,
"rate_upper": 0.91,
"count_median": 15.3,
"count_lower": 6.7,
"count_upper": 28.9
}
```
The `/draws` endpoint returns a Hugging Face Parquet shard URL + DuckDB SQL for
bulk posterior draw access — useful if we need to compute custom quantities like
P(district > national) without round-tripping through multiple API calls. However,
for a demo app that runs in the browser, relying on 10 HF Parquet shards per model
is impractical (80 GB total). The summary endpoints (`/estimates`) provide enough
aggregated information for all six charts using only `rate_median`, `count_median`,
and interval bounds — **no raw draw access needed**.
---
## 4. Visual Design Language (from reference materials)
### Color palette
```css
:root {
--cv-paper: #FAF7F2; /* warm paper background */
--cv-ink: #0E1A2B; /* primary text — civic navy/black */
--cv-navy-600:#22406A; /* links, accents */
--cv-accent: #C25311; /* ember — alerts, highlights, eyebrows */
/* Data-viz supporting colors (from _tokens.scss) */
--teal-600: #1F6F70;
--plum-600: #6B3A5E;
--moss-600: #4A6B2F;
--brass-600: #B8751C;
}
```
### Typography
- **Display**: Libre Franklin / Inter — headings, stat callouts.
- **Body**: Inter — UI text, labels, captions.
- **Mono**: JetBrains Mono — code snippets, API URLs in footers.
### Chart patterns (from social media posts & white paper)
1. **Ridgeline density plots** for posterior draws (`geom_density_ridges`).
2. **Pointrange / error bar charts** comparing model intervals to observed values.
3. **Faceted small multiples** — always split by `covariate × time` (baseline vs.
covariate; one-year vs. three-year).
4. **Transparent PNG export with watermark logo** in bottom-right corner.
5. **"Stat callout" hero numbers** for key metrics (e.g., total arrests, rate per 1k).
---
## 5. Technology Stack Recommendation
### Recommended: React + Vite + vanilla CSS (static site)
| Layer | Choice | Rationale |
|---|---|---|
| Framework | **React 19** (no framework overhead; Vite dev server) | Component model for charts, built-in state management via hooks |
| Build tool | **Vite** | Fast HMR, native ES modules, trivial static export (`npm run build`) |
| Charting | **`@visx/visx`** or plain SVG/CSS animations | Lightweight, no heavy deps; we control every pixel to match Civilytics style |
| HTTP client | `fetch` with AbortController + exponential backoff | No extra dependency; browser-native |
| Styling | Plain CSS custom properties (no Tailwind) | Zero-runtime, matches the existing SCSS token system exactly |
| Deployment | Static site on any host (GitHub Pages, Vercel, Netlify) or Docker nginx container | Self-hostable; no backend required |
#### Why not Shiny / Quarto?
The API returns JSON and we need a dynamic SPA with loading states, debounced search, and animated transitions. A static React app is the most natural fit. R/Shiny would add unnecessary server-side complexity for what is fundamentally a browser-based data visualization demo. The existing `crdc-arrests` project uses Quarto + ggplot2 for reports; this demo app is a different artifact with different requirements.
#### Why not Svelte or Vue?
React has the largest ecosystem, best tooling (Vite), and most team familiarity. For a 6-chart SPA it's more than sufficient without being overkill.
### Docker option
A minimal `nginx:alpine` container serves the built static files. ~20 MB image. Can be run on your self-hosted fleet (`efron`, `maxwell`, etc.) behind Caddy/Traefik with TLS via Let's Encrypt — consistent with how you host other civilytics.org services.
---
## 6. Build & Deployment Strategy
### Local development
```bash
npm install # one-time: installs React, Vite, dev deps
npm run dev # starts Vite on localhost:5173
# Edit src/App.jsx / src/components/*.jsx — hot reload
```
### Production build (static)
```bash
npm run build # outputs dist/ with index.html + assets
# Upload dist/ to any static host, or:
docker build -t crdc-demo . && docker run -p 8080:80 crdc-demo
```
### Docker deployment (self-hosted on your fleet)
1. Build image locally or via CI: `docker buildx build --platform linux/amd64 -t registry.civilytics.org/crdc-demo:latest .`
2. Push to local registry on `maxwell`.
3. Deploy via docker-compose or a simple systemd service with nginx container + Caddy reverse proxy for TLS termination.
```yaml
# docker-compose.yml (minimal)
services:
app:
image: crdc-demo:latest
ports: ["8080:80"]
restart: unless-stopped
```
### CI/CD (optional, GitHub Actions)
- On push to `main`: run lint + build, publish Docker image to local registry.
- No automated deployment — you control when new versions go live on the fleet.
---
## 7. File Structure
```
crdc-demo/
├── public/ # Static assets (logo, favicon)
│ └── civilytics-logo.svg
├── src/
│ ├── components/ # Reusable UI pieces
│ │ ├── StateSelector.jsx
│ │ ├── DistrictSearch.jsx
│ │ ├── LoadingAnimation.jsx
│ │ └── StatCallout.jsx
│ ├── charts/ # The 6 chart components
│ │ ├── ArrestsOverTime.jsx
│ │ ├── RateByGroupBar.jsx
│ │ ├── DistrictVsNational.jsx
│ │ ├── ModelDrawsComparison.jsx
│ │ ├── ObservedRateDensity.jsx
│ │ └── ExceedanceProbability.jsx
│ ├── hooks/
│ │ ├── useApi.js # fetch wrapper with retry/backoff
│ │ └── useDistrictData.js # orchestrates all 6 charts' data needs
│ ├── styles/
│ │ ├── tokens.css # Civilytics design tokens
│ │ └── main.css
│ ├── App.jsx # Main router: state → search → loading → charts
│ └── main.jsx # React entry point
├── Dockerfile
├── vite.config.js
└── package.json
```
---
## 8. Key Implementation Notes & Risks
### Risk 1 — API response time for multi-model data
Fetching all models × race×sex combinations for a single district requires ~40
individual `/estimates` calls (5 unified + 5 stratified models × 8 groups). Each
API call may take 200–500ms. **Mitigation**: use `Promise.allSettled()` to fire
all requests in parallel; show progress as batches resolve. The loading animation
should reflect this batching visually (e.g., rows of bars filling left-to-right).
### Risk 2 — National rate computation
The `/states` endpoint returns per-state aggregates, not a national total. To get
national arrest rates by student group for Chart 3 and Chart 6:
- **Option A**: Fetch all states (`limit=100`, iterate through ~50 pages) and sum.
Too slow for a browser app.
- **Option B**: Add a `/api/v1/national` endpoint to the API (server-side aggregate).
Requires modifying `crdc-arrests/api/` — quick R/SQL change but needs your approval.
- **Recommended short-term**: Cache national rates as a static JSON file generated
during data release and committed alongside this app. The white paper already
computes these values; we can extract them into a small fixture.
### Risk 3 — Chart complexity (Chart 5 & 6)
Charts 5 (density/ridge) and 6 (exceedance probability) require either raw posterior
draws or sufficient summary statistics to reconstruct distributions. The `/estimates`
endpoint provides `count_median`, `count_lower`, `count_upper` but not the full draw
distribution. **Mitigation**: Use interval bounds + median as a proxy for ridge plots
(showing just the 95% HPD region), and compute exceedance probability using a normal
approximation to the posterior (mean=median, sd derived from interval width). This is
a reasonable approximation for demonstration purposes but should be clearly labeled.
### Risk 4 — Mobile responsiveness
The social media figures are desktop-first PNGs. The web app must work on mobile:
- Use CSS Grid with `auto-fit` for the chart panel (1 column on mobile, 3 on desktop).
- Make the loading animation responsive.
- Ensure touch targets in district search are ≥44px.
---
## 9. Open Questions for You
Before implementation begins, I need your input on these decisions:
### Q1 — National rates source
Chart 3 and Chart 6 compare a district's arrest rate to the **national average**
for each student group. The API doesn't have a `/national` endpoint. How should we handle this?
- **(A)** Cache national rates as a static JSON fixture (generated from the white paper / release data). [Recommended — simplest, no API changes]
- **(B)** Add a `/api/v1/national` aggregate endpoint to `crdc-arrests/api/` and have the demo call it live.
- **(C)** Approximate by fetching the top 5 largest states' rates weighted by enrollment (fast but imprecise).
### Q2 — Model selection for distribution charts
Charts 4–6 show results from multiple model types (one-year, three-year, baseline, covariate). Should users be able to toggle which models are displayed, or should we always show all four quadrants as specified in your requirements?
- **(A)** Always show all four quadrants (1yr no-cov, 1yr cov, 3yr no-cov, 3yr cov) — matches the white paper's `wp_fig_district_intervals()` layout. [Recommended]
- **(B)** Let users toggle between unified vs. stratified models via a dropdown.
- **(C)** Default to showing only the recommended model (`unified_m2_mod`) with an "advanced" expando for all four.
### Q3 — Loading animation style
The patience animation should convey that data is being fetched across multiple endpoints/models. What visual approach do you prefer?
- **(A)** Animated grid of histogram bars (one per model×group) that fill sequentially as API calls resolve, with a counter showing "X of ~40 datasets loaded". [Recommended — matches the Bayesian posterior theme]
- **(B)** A simple spinner + text message ("Fetching 12 model results from the CRDC API…").
- **(C)** An abstract animation (e.g., particles converging) that loops while loading, with no per-request feedback.
### Q4 — Deployment target
Where should this be hosted? Your infrastructure is self-hosted (`efron`, `maxwell`, etc.), but you also mentioned civilytics.org. Which deployment approach do you want me to implement first?
- **(A)** Docker container ready for your fleet (nginx + static files) — I'll write the Dockerfile and docker-compose.yml, deploy on a dev host. [Recommended]
- **(B)** Static site optimized for GitHub Pages / Vercel with CI/CD via GitHub Actions.
- **(C)** Both A and B (Docker for self-hosting, plus GH Pages config).
---
## 10. Next Steps
Once you answer the four questions above, I'll begin implementation:
1. **Week 1**: Scaffold project (Vite + React), implement design tokens & typography, build state selector + district search with "interesting" suggestions.
2. **Week 2**: Implement loading animation; fetch all model data in parallel; build Charts 1–3 (observed/descriptive).
3. **Week 3**: Build Charts 4–6 (model comparison distributions); wire up national rate fixture if Q1=A.
4. **Week 4**: Polish transitions, mobile responsiveness, Dockerfile + deployment configs; write README with usage/deployment docs.
The app will be fully self-contained — no external dependencies beyond the public CRDC API and optionally a cached JSON fixture for national rates. All code follows Civilytics visual identity as documented in `theme/_tokens.scss` and demonstrated in `social_media_posts.md`.
+162
View File
@@ -0,0 +1,162 @@
# CRDC Arrests API Demo
A demonstration web application for the [CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/), showing Bayesian model comparisons of school-based arrest rates across U.S. districts. Built with React + Vite, deployable as a static site or Docker container.
## Quick Start (Development)
```bash
# Install dependencies
npm ci
# Start dev server on localhost:5173
npm run dev
# Build for production
npm run build # outputs to dist/
npm run preview # serve built files locally
```
## What It Does
Visitors select a U.S. state, search for a school district (with suggestions of the districts reporting the most arrests), and see what was actually reported alongside what the Bayesian models estimate. The results page ports the white paper's Figs 6 and 7 (`wp_fig_group_density` / `wp_fig_group_difference`):
1. **Reported arrests and enrollment** — a summary table, one row per student group plus a district total: students, observed arrests, and rate per 1,000. Its checkboxes double as the legend and the group control for the density panel. In sparse districts (fewer than 20 arrests district-wide) Female and Male are pooled within each race, with a banner explaining the rule and a switch to override it.
2. **Arrest rate probability density** — each selected group's posterior predictive distribution as a filled area, Female over Male sharing one axis, direct-labelled at the peak. Beneath each panel, a rail of 95% Agresti–Coull point ranges for the observed rate. A segmented control switches between the four Bayesian specifications; an opt-in toggle compares all four at once. Draws that take only a handful of distinct values are drawn as discrete probability mass rather than smoothed — in a small district the posterior predictive genuinely *is* discrete.
3. **Model estimated differences** — the posterior of Δ = rate(A) − rate(B) per 1,000, computed at each draw, filled with a diverging ramp centred at zero, with a dashed rule at no-difference and a plain-language `Pr(Δ > 0)` readout.
4. **Arrests over time** — observed counts by CRDC wave against the three-year model's median and 95% interval, inline SVG.
Distributions come from the published 500-draw posteriors, fetched client-side via duckdb-wasm from the public Hugging Face parquet dataset, and fall back per group to an analytic approximation (with a visible note) when those draws can't be fetched.
## Architecture
### Tech Stack
- **React 19** + **Vite** (static site generation, no backend required)
- Plain CSS custom properties for styling (matches Civilytics design tokens exactly)
- Inline SVG rendering that **React owns** — no charting library. `d3-scale`, `d3-shape`, `d3-array` and `d3-interpolate` supply scales, path generators and colour interpolation only; no d3 selections and no `useEffect` DOM mutation
- Embeds **DuckDB-Wasm** (`@duckdb/duckdb-wasm`, ~39MB uncompressed / ~8.8MB gzipped, loaded on demand only after a district is selected) to query real posterior draws client-side from a public Hugging Face Parquet dataset
- Calls the public read-only API directly from the browser
### File Structure
```
crdc-demo/
├── index.html # Entry point
├── vite.config.mjs # Vite build config
├── public/ # Static assets (wordmark, favicon, fixtures)
│ ├── civilytics-wordmark.svg # Civilytics wordmark from civilyticsR package
│ └── data/
│ ├── national_rates.json # National rates fixture for comparisons
│ └── top_districts.json # Top 15 districts per state by observed arrests
├── scripts/
│ └── build-top-districts.mjs # One-off generator for top_districts.json
├── src/
│ ├── main.jsx # React entry
│ ├── App.jsx # Main router (state → search → loading → charts)
│ ├── hooks/
│ │ ├── useApi.js # API client with retry/backoff + endpoint wrappers
│ │ └── useDrawDistribution.js # Posterior draw counts by draw_id (duckdb-wasm + HF parquet)
│ ├── components/
│ │ ├── StateSelector.jsx # Landing screen — state dropdown/grid
│ │ ├── DistrictSearch.jsx # Search + suggestions from the committed fixture
│ │ ├── LoadingAnimation.jsx # Animated histogram grid during data fetch
│ │ ├── ChartLegend.jsx # Shared legend row
│ │ ├── ApproxNote.jsx # "shape estimated from interval bounds" caption
│ │ ├── DistrictSummaryTable.jsx # Observed arrests table — also the chart's legend/control
│ │ └── ChartPanel.jsx # Owns cross-chart state + data fetching
│ ├── charts/
│ │ ├── ArrestsOverTime.jsx # Observed vs. modelled counts by wave (SVG)
│ │ ├── RateDensityPanel.jsx # Chart A — posterior density per group + AC rail
│ │ └── GroupDifference.jsx # Chart B — posterior of Δ between two groups
│ ├── utils/
│ │ ├── duckdbClient.js # Lazy duckdb-wasm bundle loader (dynamic import)
│ │ ├── kde.js # Empirical density from real draws (+ .test.js)
│ │ ├── densityProfile.js # Discrete-mass vs. KDE profile choice (+ .test.js)
│ │ ├── agrestiCoull.js # Frequentist interval, ported from R (+ .test.js)
│ │ ├── pooling.js # Sex pooling for sparse districts (+ .test.js)
│ │ ├── districtGroups.js # Display-row derivation and defaults (+ .test.js)
│ │ ├── groupDifference.js # Per-draw Δ and its summary (+ .test.js)
│ │ ├── rateDomain.js # Shared x-axis domain and clip flag (+ .test.js)
│ │ ├── drawGroups.js # Draw-map key format + coverage check (+ .test.js)
│ │ └── distributionApprox.js # Analytic fallback when draws are unavailable
│ └── styles/tokens.css # Civilytics design tokens (colors, fonts, spacing)
├── Dockerfile # Multi-stage build → nginx static server
├── docker-compose.yml # Local dev / self-hosted deployment
└── .gitea/workflows/pages.yml # CI/CD — builds and deploys to pages branch
```
### API Endpoints Used
| Endpoint | Purpose | Frequency |
|---|---|---|
| `/api/v1/models` | List available Bayesian model specs | Once (cached) |
| `/api/v1/districts?q=&state=` | District name/geo lookup → LEAID | On keystroke |
| `/api/v1/estimates/{leaid}?model=X&year=Y` | Estimates for one district/model/year/group | 3 waves + 1 per selected spec |
| `/api/v1/estimates?state=XX&year=Y` | Not called at runtime — rows come back `ORDER BY LEAID` at 8 per district, so any short read ranks the lowest-LEAID districts. Paged with `meta.total` by `scripts/build-top-districts.mjs` | Build-time only |
| `/api/v1/draws?...` | Locate raw-posterior Parquet shard | Not called from app — the app fetches shards directly from Hugging Face via duckdb-wasm; see `src/hooks/useDrawDistribution.js` |
| `/data/top_districts.json` | Suggested districts per state (committed fixture) | Once per session |
| `/data/national_rates.json` | Static national rates fixture (committed) | Once per session |
## Deployment
### Option A: Git Pages via Gitea Actions (recommended)
The app is a static site deployed automatically via Gitea Actions. On every push to `main`, the workflow in
`.gitea/workflows/pages.yml` builds the Vite app and copies `dist/` to a `pages` branch served by the
git-pages webhook receiver.
**URL**: `https://pages.civilytics.org/crdc-demo/`
> **CORS note**: The CRDC API (`crdc-api.civilytics.org`) does not send CORS headers. When deployed as a static site,
> browser-based fetch requests will be blocked by the same-origin policy. To fix this, deploy `proxy.php` to your git-pages host
> and set the `VITE_PROXY_URL` environment variable (e.g., `/crdc-demo/proxy.php`). The app automatically detects
> proxy availability — when unset, it attempts direct API calls (works via Docker/nginx or same-origin setups).
The workflow follows the same pattern as `Civilytics/sln-school-comparison`:
1. Triggers on push to `main` (only when source files change)
2. Builds with Node 22 + Vite (`npm ci && npm run build`)
3. Copies `dist/` contents to an orphan `pages` branch
4. Forces push — the git-pages webhook picks it up automatically
> **Note**: Ensure `${{ secrets.GITHUB_TOKEN }}` is configured in repo settings on gitea.civilytics.org.
### Option A.5: CORS Proxy for Git Pages
If you can't deploy PHP, create a simple reverse proxy in nginx or use an edge function that:
1. Accepts `?target=/api/v1/...` as a query parameter
2. Forwards the request to `https://crdc-api.civilytics.org/api/v1/...
3. Returns the response with `Access-Control-Allow-Origin: *`
4. Set `VITE_PROXY_URL` at build time to point at this endpoint.
### Option B: Docker (self-hosted fleet) — includes CORS proxy nginx config:
```bash
docker compose up -d # builds + serves on localhost:8080
# Or build manually for your registry:
docker buildx build --platform linux/amd64 -t registry.civilytics.org/crdc-demo:latest .
docker push registry.civilytics.org/crdc-demo:latest
```
The container is nginx Alpine (~5MB) plus the built site, which is dominated by the ~39MB DuckDB-Wasm engine. `nginx.conf` serves that engine gzipped (~8.8MB on the wire) with immutable cache headers, so it is fetched once per browser. A `/healthz` endpoint supports container orchestration health checks.
### Environment Variables
| Variable | Default | Description |
|---|---|---|
| `VITE_API_BASE` | `https://crdc-api.civilytics.org/api/v1` | Override to point at a staging API (set in Dockerfile build or docker-compose) |
## Visual Design
This app follows the Civilytics visual identity as defined in:
- `theme/_tokens.scss` — colors, typography, spacing tokens
- `social_media_posts.md` — chart patterns from published social media figures
- `white_paper.qmd` / `R/paper_figures.R` — white paper figure builders
Key design decisions:
- **Color palette**: Warm paper background (`#FAF7F2`), civic navy text (`#0E1A2B`), ember accent (`#C25311`)
- **Data viz colors**: Teal, plum, moss, brass from the supporting palette
- **Fonts**: Libre Franklin/Inter for display and body, JetBrains Mono for code
- **Chart patterns**: Pointrange comparisons (model vs. observed), ridgeline-style density bars, faceted small multiples
## Citation & Attribution
The CRDC School Arrest Rate API: Knowles, J.E., & Miller, H. (2025). *CRDC School Arrest Rate API v1*. Civilytics Consulting. https://crdc-api.civilytics.org/api/v1/
Data source: US Department of Education, Office for Civil Rights. Civil Rights Data Collection (2021-22).
This research was supported by a grant from the American Educational Research Association which receives funds for its "AERA Grants Program" from the National Science Foundation under NSF award NSF-DRL #1749275. Opinions reflect those of the author and do not necessarily reflect those AERA or NSF.
+7
View File
@@ -0,0 +1,7 @@
# Git Pages configuration — for self-hosted static site deployment.
# This file is informational; the actual build + deploy happens via .gitea/workflows/pages.yml.
# The workflow builds the Vite app, then copies dist/ to a `pages` branch served by git-pages.
title: CRDC Arrests API Demo
description: "Explore school-based arrest rates by U.S. district with Bayesian model comparisons."
author: Civilytics Consulting LLC
File diff suppressed because one or more lines are too long
@@ -1 +0,0 @@
var e=`/crdc-demo/assets/duckdb-browser-mvp.worker-C9hF7LGh.js`;export{e as default};
File diff suppressed because one or more lines are too long
-1
View File
@@ -1 +0,0 @@
var e=`/crdc-demo/assets/duckdb-mvp-BP0pRkMH.wasm`;export{e as default};
Binary file not shown.
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
-1
View File
@@ -1 +0,0 @@
654b42ca71f3654156e9b7a5e669dfa3c2fff41f
+12
View File
@@ -0,0 +1,12 @@
# Docker Compose for local development or self-hosted deployment on your fleet.
# Usage: docker compose up -d
services:
crdc-demo:
build: .
ports:
- "8080:80" # Map host port 8080 → container port 80 (nginx)
restart: unless-stopped
environment:
# Override to point at a staging API if needed; defaults to production.
VITE_API_BASE: https://crdc-api.civilytics.org/api/v1
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,219 @@
# CRDC demo — rebuild the visuals around the white-paper figures
## Context
The demo app at `pages.civilytics.org/crdc-demo/` works, but the charts don't show off the
modelling. The navigation (state → district) is good and stays. The three existing charts get
cut to one, replaced by a summary table and two charts ported from the white paper:
`wp_fig_group_density` (Fig 6) and `wp_fig_group_difference` (Fig 7) in
`crdc-arrests/R/paper_figures.R:472-568`.
Two real defects surfaced while scoping:
1. **The "suggested districts" list is ranked over a truncated set.** `DistrictSearch.jsx:22`
calls `/estimates?state=XX&year=21-22&limit=500`, but that endpoint returns rows
`ORDER BY LEAID` (`crdc-arrests/api/R/handlers_estimates.R:46`) at 8 rows per district. So
the app ranks the ~62 lowest-LEAID districts in the state. California has 11,488 rows.
2. **The analytic fallback assumes 90% bounds but the API returns 95%.**
`distributionApprox.js` defaults `intervalMass = 0.90`; the API default is `interval=95`
(`crdc-arrests/api/R/handlers_estimates.R`). The new charts compute intervals from real
draws, so this only affects the fallback path — fix the default while we're in there.
### What the data supports (verified, not assumed)
- The HF parquet carries `LEAID, RACE, SEX, pred, draw_id, subgroup_id, batch_num`. `draw_id`
is dense 1–500 for every group. The current query at `useDrawDistribution.js:93` just
doesn't select it — **that one column is what unlocks both pooling and differences.**
- **Sex pooling is the project's own house method.** `build_state_summary()` in
`crdc-arrests/R/summarize_draws.R:152-182` pools across LEAs by summing `pred` and
`stu_enroll` within each draw, then summarizing across draws. Pooling M+F within a district
is the identical operation on a different axis. Honest effort estimate: **~3–4 hours**, most
of it UI and labelling, not statistics.
- **Caveat to word carefully:** `draw_id` is renumbered 1–500 per write batch
(`crdc-arrests/R/postprocess.R:106-133`), and a district's groups land in different batches,
so cross-group draw pairing is effectively independent. Measured correlation between
Black-male and Hispanic-male `pred` in Clark County was 0.019 even *within* a batch —
observation noise from `posterior_predict` dominates. Published Fig 7 has the same property,
so the app matches the paper. Captions should say "posterior predictive draws", never
"paired parameter draws".
- Enrollment covers only AM/BL/HI/WH (verified against the API). Clark County sums to 148,928
against a district enrollment near 304,000. The table must label this.
## Decisions taken
| Question | Decision |
|---|---|
| Model specs | One selected spec by default; opt-in "compare all four" expands to 4 ridge rows |
| Existing charts | Keep `ArrestsOverTime`; delete `RateByGroupBar` and `RateDensityRidgeline` |
| Pooling trigger | Whole-district: pool when total observed arrests across the 8 cells < 20 |
| Rendering | React owns the DOM; add d3 submodules for scales/paths/interpolation |
---
## Work
### 1. Draws pipeline — expose `draw_id`, return counts not rates
**`src/hooks/useDrawDistribution.js`** — the one structural change everything else rests on.
- Query becomes `SELECT RACE, SEX, draw_id, pred FROM read_parquet(...) WHERE LEAID = ?`.
- Return **raw counts indexed by draw**, not rates: `countsByGroup[key][draw_id - 1] = pred`.
Indexing by `draw_id` rather than push-order means row ordering from DuckDB is irrelevant.
- Accept `models: string[]` instead of a single `model`, returning `{status, byModel, nDraws}`.
A fixed-length array avoids conditional hooks when "compare all four" is on. The existing
module-level `shardCache` already keys on `(model, year, state)`, so four models is four
cache entries with no other change.
- Reject a group whose count array has holes (fewer entries than `nDraws`) — a partial group
must fall back, not silently render a short draw set.
- Keep the reset-before-guard ordering at `useDrawDistribution.js:79-85`; it exists to stop one
model's draws being shown under another model's label.
**`src/utils/pooling.js`** (new, + test) — pure functions, no React:
- `poolBySex(countsByGroup, enrollByGroup)` → sums counts within each draw index across
`SEX ∈ {F,M}` and sums enrollment, keyed by race alone.
- `toRates(counts, enroll)` → per-1,000 array.
- `shouldPoolBySex(rows)` → total `observed_arrests` across rows < `POOL_BY_SEX_ARREST_THRESHOLD`
(20, a named constant with the rationale in a comment).
**`src/utils/drawGroups.js`** — extend `groupKey` to handle a pooled key (race only) and update
`hasDrawsForAll` for the new count-array shape.
### 2. Frequentist interval
**`src/utils/agrestiCoull.js`** (new, + test) — direct port of
`crdc-arrests/R/paper_figures.R:219-237`, including the zero-numerator branch
(`ci_upper = -log(1 - level)`, the rule of three; `ci_lower = 0`). Note the R function returns
`c(upper, lower, sd, se, phat)` — **upper first**. Return a named object here instead. Needs a
`qnorm`/probit; `distributionApprox.js` already has one — reuse it rather than adding a second.
Tests should pin at least one case against R output (e.g. `agresti_coull(15, 499, 0.95)`).
### 3. Density profile — handle discrete posteriors honestly
**`src/utils/densityProfile.js`** (new, + test).
In sparse districts the posterior predictive is a discrete count distribution. Carson City NV
(`3200390`) has a 53-student AI/AN female cell where one arrest is 18.9 per 1,000 — the draws
take four distinct values and a Gaussian KDE renders them as a lumpy smear that reads as a
rendering bug.
- `densityProfile(counts, enroll, domain)` returns `{kind: 'kde'|'mass', points}`.
- `kind: 'mass'` when the draws take ≤ 12 distinct values: probability mass at each achievable
rate, drawn as a filled staircase so it visually rhymes with the smooth areas beside it.
- Otherwise delegate to the existing `kdeCurve` in `src/utils/kde.js` — its bandwidth clamp
(`BANDWIDTH_FLOOR_DIVISOR` / `BANDWIDTH_CEILING_DIVISOR`) was tuned for exactly these
zero-inflated posteriors and should not be touched.
### 4. Summary table (top of results)
**`src/components/DistrictSummaryTable.jsx`** (new). One row per student group plus a total:
| Student group | Students | Observed arrests | Rate per 1,000 |
- Sorted by observed arrests descending. Zero-arrest rows de-emphasized, not hidden.
- Each row carries the checkbox that drives chart A — the table *is* the legend and the control.
Default checked = `observed_arrests > 0`; if no group has any, check the two largest by
enrollment and say so.
- Footnote: students counted are those in the four modeled race groups (AI/AN, Black, Hispanic,
White), not total district enrollment.
- When pooling is active, rows collapse to four races and a banner states the rule in one
sentence, with a switch to force it off.
### 5. Chart A — "Arrest rate probability density"
**`src/charts/RateDensityPanel.jsx`** (new). Replaces `RateDensityRidgeline.jsx`.
- Two stacked sub-panels, Female over Male, sharing one x-axis (per 1,000). Collapses to a
single panel when pooled. This preserves the palette contract documented at
`src/utils/colors.js:9-13`: race is hue, sex is position — never a second hue.
- Within a sub-panel, selected groups overlap as filled areas (fill ~0.4 opacity, 2px stroke in
the race color), direct-labelled at each peak so there's no legend hunting.
- Below each sub-panel's baseline, a thin rail stacks one Agresti–Coull point-range per selected
group in the matching color — the R figure's `position_nudge` idea, but un-overplotted.
- Segmented control for the four unified quadrant specs. A "Compare all four specifications"
switch expands to four ridge rows (matching Fig 1's structure) and triggers four shard
fetches — cheap for NV (~100KB each), ~6.3MB each for CA, so it stays opt-in with a spinner.
- x-domain: max of the density supports and the frequentist upper bounds, with the existing cap
logic from `RateDensityRidgeline.jsx:101` and `niceTicks`.
- Caption states 500 posterior predictive draws and a 95% Agresti–Coull observed interval.
### 6. Chart B — "Model Estimated Differences"
**`src/charts/GroupDifference.jsx`** (new).
- Two group pickers; defaults are the two groups with the most observed arrests (pooled groups
when pooling is on). Δ = rate(A) − rate(B) per 1,000, computed per draw index.
- Single density, filled with an SVG `linearGradient` mapped across x. Use a **diverging ramp
centered at zero** — navy for Δ<0, paper at 0, ember for Δ>0 — rather than the paper's YlOrRd:
the quantity is signed, and diverging-at-zero is the honest encoding. On-brand via
`tokens.css`.
- Dashed vertical rule at 0 in `--cv-danger`, matching the paper's red line.
- Large readout `Pr(Δ > 0)` with a plain-language sentence beneath ("In 94.4% of posterior
draws, the Black male arrest rate exceeds the White male rate"), plus median Δ and an 80%/95%
interval as a point-range.
- Degrade explicitly when fewer than two groups have usable draws — say why, don't render empty.
### 7. Fix the district suggestions
**`scripts/build-top-districts.mjs`** (new) — pages `/estimates?state=XX&year=21-22&limit=1000`
using `meta.total` (confirmed present in the envelope) across all 51 states, aggregates observed
arrests per LEAID, and writes `public/data/top_districts.json` with the top 15 per state
(leaid, name, arrests, enrollment, rate). Roughly 140 requests as a one-off; the output is
~50KB and gets committed, following the `public/data/national_rates.json` precedent.
**`src/components/DistrictSearch.jsx`** — read the fixture instead of calling
`fetchStateDistricts` at runtime. The search screen loses a multi-second fetch and the ranking
becomes correct. Keep live name search on `/districts` unchanged. Drop the hardcoded
"Try Derby (KS), Paterson (NJ)…" hint at `DistrictSearch.jsx:164` — the real list supersedes it.
### 8. Wiring, deletions, docs
- **`src/components/ChartPanel.jsx`** — owns pooling state, selected groups, selected spec, and
the difference pair; passes them down. Keep `ArrestsOverTime`. Drop `QUADRANT_MODELS`
prefetch of all four models' *summaries* if only the selected one is needed.
- **Delete**: `src/charts/RateByGroupBar.jsx`, `src/charts/RateDensityRidgeline.jsx`.
- **Keep**: `distributionApprox.js` and `ApproxNote.jsx` — still the fallback when draws can't
be fetched (`AGENTS.md:91-101`). Fix its `intervalMass` default to 0.95 to match the API.
- **`package.json`** — add `d3-scale`, `d3-shape`, `d3-array`, `d3-interpolate` as real
`dependencies` (the existing deps are all miscategorised under `devDependencies`; leave that
alone unless it's breaking the build).
- **Docs**: `AGENTS.md` still describes 6 charts, D3 selections, and `DistrictVsNational` /
`ModelDrawsComparison` / `ExceedanceProbability` — none of which exist. Rewrite the
architecture, data-flow, and interval sections. Update `README.md`'s chart list.
---
## Verification
1. `npm run build` clean; `npm test` (`node --test 'src/**/*.test.js'`) passes, including new
tests for `agrestiCoull`, `pooling`, `densityProfile`, and the extended `drawGroups`.
2. `npm run dev`, then walk these districts:
- **Clark County NV (`3200060`, 100 arrests / 148,928 students)** — pooling stays off, all
four races render, differences chart defaults to the top two groups.
- **Carson City NV (`3200390`, 6 arrests / 4,073 students)** — pooling auto-engages (6 < 20),
banner appears, table collapses to four races. The AI/AN cell should render as a discrete
mass profile, not a smear. Verified against the draws: pooling narrows AI/AN's 90% interval
from 37.7 to 27.0 per 1,000, and Hispanic male's from 4.9 to 2.5.
- **Washoe County NV (`3200480`)** — cross-check the summary table's observed counts and
rates against `/api/v1/estimates/3200480?model=unified_m4_mod&year=21-22`.
- **A California district** — confirm the suggestion fixture ranks correctly (this is the
case the current code gets wrong), and that "compare all four" warns/spins before pulling
~25MB of shards.
3. Toggle every group off, then on; switch specs; flip pooling manually — no stale draws from a
previous model should ever appear under a new label.
4. Compare chart A against `wp_fig_group_density` output for Clark County: same curve shapes,
same point-range positions.
5. Browser console clean on hard refresh; check the Network tab shows one shard fetch per
(model, state) and no repeats when navigating between districts.
## Effort
Roughly **2–3 focused days** end to end: ~1 day for the draws/pooling/util layer with tests,
~1 day for the two charts, ~half a day for the table, the suggestion fixture, and docs. At ~10
hours a week that's about two calendar weeks.
The R Shiny alternative would be slower, not faster — it trades a zero-server static site for a
container, an R runtime, and server-side access to either the 91GB draws DuckDB or the 51-state
parquet tree, and turns every toggle into a round-trip re-render. The React app already fetches
real draws client-side and already carries the design tokens.
@@ -0,0 +1,187 @@
# Empirical draw distributions via DuckDB-Wasm — Design Spec
**Date:** 2026-08-11
**Status:** Draft for review
---
## 1. Purpose & context
Two of this app's three charts currently show a *modeled* distribution shape that is
not the real posterior — `src/utils/distributionApprox.js` fits a two-piece-normal
curve to each group's `(median, lower, upper)` summary stats returned by the
`/estimates` API, because the API's raw posterior draws are only available in bulk
as Hive-partitioned Parquet on Hugging Face
(`civilytics/crdc-school-arrest-rates`), meant for DuckDB/bulk consumption, not
browser fetches (see `AGENTS.md` §"API Endpoint Availability" and
`HANDOFF.md` §"Future Enhancements").
This spec replaces the approximation with the **real** empirical draws (500 per
group), fetched client-side using `@duckdb/duckdb-wasm` to query the actual Parquet
shard for the district's state directly from Hugging Face — no new backend
endpoint, no change to the existing summary API calls.
**Verified feasibility (2026-08-11):**
- The HF dataset is public, non-gated. Each `(model_id, YEAR, LEA_STATE)` partition
is a single file (`data_0.parquet`).
- File sizes range from ~130KB (DC) to ~6.4MB (CA, the largest state) — confirmed
by resolving the `resolve/main/...` redirect to the actual CDN blob.
- The redirect target sends `access-control-allow-origin: *` and
`accept-ranges: bytes` — browser `fetch()` works directly, no proxy needed.
- Schema (from `crdc-arrests/R/postprocess.R` + `R/export_parquet.R`):
`LEAID, LEA_STATE, YEAR, RACE, SEX, model_id, subgroup_id, draw_id, pred`, sorted
within each shard by `(LEAID, RACE, SEX)`. `LEA_STATE`/`YEAR`/`model_id` are
Hive-partition columns (encoded in the path, not repeated in every row).
`stu_enroll` is **not** in the draws table — it's already available in this app
from the existing `/estimates` summary call.
---
## 2. Architecture
```
district selected (leaid, state)
│
▼
resolve HF parquet URL for (model_id, YEAR=21-22, LEA_STATE=state)
e.g. https://huggingface.co/datasets/civilytics/crdc-school-arrest-rates/
resolve/main/parquet/model_id=unified_m4_mod/YEAR=21-22/LEA_STATE=CO/data_0.parquet
│
▼
fetch() the shard (native fetch, follows the HF→CDN redirect automatically)
│
▼
duckdb-wasm: registerFileBuffer + query
SELECT RACE, SEX, pred FROM shard WHERE LEAID = '<leaid>'
│
▼
join `pred` (posterior count draws) against stu_enroll already in app state
(from the existing /estimates summary call) → rate-per-1000 draws per group
│
▼
KDE per race×sex group → smooth density curve, same shape the charts draw today
```
Given verified shard sizes (≤6.4MB), the design fetches the **whole shard** with a
plain `fetch()` and queries it in-memory via duckdb-wasm, rather than relying on
fine-grained HTTP range / row-group pruning. This is simpler and more robust than
depending on duckdb-wasm's HTTP virtual filesystem correctly handling the HF→CDN
redirect chain under partial-range requests — an unverified behavior — for a
saving that wouldn't matter at these file sizes.
`@duckdb/duckdb-wasm` is MIT-licensed and runs entirely client-side in a Web
Worker. It introduces no new server dependency and no new hosted service beyond
the Hugging Face dataset the `crdc-arrests` project's `/draws` endpoint already
points to (per `2026-05-30-draws-api-design.md`, decision #5) — this spec doesn't
introduce that dependency, it makes the demo app actually use data that was
already published there for exactly this purpose.
**Deployment risk:** duckdb-wasm's threaded ("eh") bundle requires
`Cross-Origin-Opener-Policy` / `Cross-Origin-Embedder-Policy` response headers
(for `SharedArrayBuffer`), which the git-pages static host does not send today.
This design uses the **single-threaded ("mvp") bundle** instead — at these file
sizes threading has no meaningful benefit, and it avoids needing new headers on
both the git-pages and Docker/nginx deploy paths.
---
## 3. File-level changes
### New files
- **`src/utils/duckdbClient.js`** — lazy-initialized singleton. Dynamic-imports
`@duckdb/duckdb-wasm`, selects the MVP (non-threaded) bundle, starts the worker
once. Dynamic `import()` keeps the ~3–5MB wasm payload out of the main bundle;
it only loads when a chart actually needs draws.
- **`src/utils/kde.js`** — Gaussian KDE over an array of numbers (Silverman
bandwidth). Takes over the role `distributionApprox.js`'s `densityCurve` plays
today, fed real empirical draws instead of a parametric fit.
- **`src/hooks/useDrawDistribution.js`** — given
`{ leaid, state, model, year, groups }` (groups = race/sex + `stu_enroll` already
in app state), resolves the HF URL, fetches, registers the buffer with
duckdb-wasm, runs the query, joins enrollment, and returns
`{ status: 'loading' | 'ready' | 'error', drawsByGroup }`. Owns an in-memory
`Map` cache keyed by `model+state+year` so re-selecting a model in Chart 3's
dropdown, or viewing another district in the same state, reuses the shard
already fetched.
### Modified files
- **`RateDensityRidgeline.jsx`** — on model-dropdown change, calls
`useDrawDistribution` for the selected model; replaces
`fitSkewedInterval`/`densityCurve` with the hook's real draws → `kde.js`. Shows
an inline spinner in the ridge area while that model's shard is loading (Charts
1–2 aren't blocked).
- **`RateByGroupBar.jsx`** — fetches draws for `unified_m3_mod` (the one model
this chart uses) alongside its existing data fetch. `q1`/`q3` become exact
empirical quantiles from the 500 real draws — removes this chart's use of
`fitSkewedInterval`'s fitted quantile function entirely.
- **`ChartPanel.jsx`** — passes `district.leaid`, `state`, and each group's
`stu_enroll` (already fetched) down to the two charts above.
- **`ApproxNote.jsx`** — becomes conditional: renders the "estimated shape" note
only when a chart is in fallback mode; charts backed by real draws show no note
(or a neutral "500 posterior draws" caption).
- **`package.json` / `vite.config.mjs`** — add `@duckdb/duckdb-wasm`; wasm/worker
assets are pulled in via Vite's native `?url` imports, which already respect the
`/crdc-demo/` `base` path — no bundler plugin needed.
### Unchanged
- **`ArrestsOverTime.jsx`** — already uses real summary stats (point-range from
`/estimates`), no approximation involved; out of scope.
- **`useApi.js`** — still the source for medians, intervals, and enrollment.
- **`distributionApprox.js`** — kept as the fallback path (see §4).
---
## 4. Error handling & caching
**Fallback:** `distributionApprox.js` is retained. `useDrawDistribution` catches
fetch/wasm/query failures and returns `status: 'error'`; both chart components
branch on that to render the current analytic-approximation path with
`<ApproxNote />` visible. A Hugging Face outage, a network failure, or an
unsupported browser degrades to today's behavior rather than breaking the chart.
**Caching:** in-memory only (a `Map` inside the hook), scoped to the browser
session. No IndexedDB/persistent cache in this iteration — a demo session
typically covers one or two districts, and shards are cheap enough to refetch on
reload.
---
## 5. Testing
This project has no automated test suite (per `AGENTS.md`); this follows the
existing manual-verification convention:
1. `npm run dev`; walk a small state (DC or WY, ~130–200KB shard) and a large one
(CA or TX, several MB) through the full district-search flow.
2. Confirm both charts render from real draws; confirm the Network tab shows the
expected parquet fetch(es) and sizes.
3. Simulate failure (block the `huggingface.co` / CDN domain in devtools) and
confirm both charts fall back to the analytic approximation with the note
visible, rather than breaking.
4. Confirm the model dropdown in Chart 3 re-fetches on first selection and is
instant on re-selection (cache hit).
---
## 6. Open risk to de-risk first
Before wiring up the full UI, spike: does duckdb-wasm's MVP bundle load and query
correctly when deployed under the `/crdc-demo/` subpath on git-pages, and does
nginx/git-pages serve `.wasm` with a usable content type? Everything else in this
design is standard Vite asset handling already exercised elsewhere in the app, but
this specific combination (wasm worker + subpath base + static host) hasn't been
verified end-to-end and should be checked with a throwaway spike rather than
assumed.
---
## 7. Explicitly out of scope
- `ArrestsOverTime.jsx` (Chart 1) — no approximation to replace.
- A new server-side `/draws`-streaming API endpoint — explicitly rejected in favor
of client-side wasm access, per the brainstorming decision that led to this spec.
- Persistent (IndexedDB) caching of fetched shards.
- Fine-grained HTTP range / row-group-level partial reads — shard sizes are small
enough that whole-file fetch is simpler and sufficiently fast.
- Extending empirical draws to national/exceedance-probability views — those
charts aren't part of the current 3-chart demo.
+1 -2
View File
@@ -5,11 +5,10 @@
<link rel="icon" type="image/svg+xml" href="/crdc-demo/favicon.svg" /> <link rel="icon" type="image/svg+xml" href="/crdc-demo/favicon.svg" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" /> <meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>CRDC Arrests API Demo — Civilytics</title> <title>CRDC Arrests API Demo — Civilytics</title>
<script type="module" crossorigin src="/crdc-demo/assets/index-naTE57ov.js"></script>
<link rel="stylesheet" crossorigin href="/crdc-demo/assets/index-DtZIgSE6.css">
</head> </head>
<body> <body>
<div id="root"></div> <div id="root"></div>
<!-- Vite injects the correct script tag with base-relative path --> <!-- Vite injects the correct script tag with base-relative path -->
<script type="module" src="/src/main.jsx"></script>
</body> </body>
</html> </html>
+77
View File
@@ -0,0 +1,77 @@
# nginx config for CRDC Arrests Demo — static SPA + API reverse proxy with CORS.
# When served via Docker, this proxies /api/v1/ requests to the CRDC API server-side,
# adding Access-Control-Allow-Origin: * so browser-based fetch works without issues.
server {
listen 80;
server_name _;
# Proxy all /api/v1/ requests to the CRDC Arrest Rate API with CORS headers
location /api/v1/ {
proxy_pass https://crdc-api.civilytics.org/api/v1/;
proxy_set_header Host crdc-api.civilytics.org;
proxy_ssl_verify off;
# Add CORS headers so browser-based requests work
add_header Access-Control-Allow-Origin "*" always;
add_header Access-Control-Allow-Methods "GET, OPTIONS" always;
add_header Access-Control-Allow-Headers "*" always;
if ($request_method = 'OPTIONS') {
return 204;
}
}
# Also support a simple proxy endpoint for git-pages-style deployments
location /crdc-demo/proxy/ {
# Extract the target path from query string and forward to CRDC API
proxy_pass https://crdc-api.civilytics.org$arg_target;
proxy_set_header Host crdc-api.civilytics.org;
proxy_ssl_verify off;
add_header Access-Control-Allow-Origin "*" always;
add_header Access-Control-Allow-Methods "GET, OPTIONS" always;
}
# Serve static files from the Vite build output
root /usr/share/nginx/html/crdc-demo/; # Adjust if base is different
index index.html;
# Compress text assets and — most importantly — the duckdb-wasm engine,
# which is ~39MB uncompressed and ~8.8MB gzipped. Without this it ships
# uncompressed on every cold load. Dynamic gzip rather than gzip_static
# because the Vite build emits no pre-compressed .gz files.
gzip on;
gzip_vary on;
gzip_min_length 1024;
gzip_proxied any;
gzip_comp_level 6;
gzip_types
text/plain
text/css
application/javascript
text/javascript
application/json
image/svg+xml
application/wasm;
location / {
try_files $uri $uri/ /index.html;
}
# Civilytics design tokens: immutable cache headers for static assets.
# `wasm` belongs here too — the duckdb engine is content-hashed by Vite and
# is by far the largest asset in the build, so it must not be re-fetched.
location ~* \.(js|css|wasm|png|jpg|jpeg|gif|svg|woff2|ttf)$ {
expires 1y;
add_header Cache-Control "public, max-age=31536000, immutable";
try_files $uri =404;
}
# Health check endpoint for container orchestration
location /healthz {
access_log off;
return 200 "ok";
add_header Content-Type text/plain;
}
}
+2704
View File
File diff suppressed because it is too large Load Diff
+37
View File
@@ -0,0 +1,37 @@
{
"name": "crdc-arrests-demo",
"version": "0.1.0",
"description": "CRDC School Arrest Rate API demonstration app — shows Bayesian model comparisons for school-based arrest rates across U.S. districts.",
"scripts": {
"dev": "vite",
"build": "vite build",
"preview": "vite preview",
"test": "node --test 'src/**/*.test.js'",
"lint": "eslint src/ --ext .js,.jsx,.ts,.tsx",
"format": "prettier --write \"src/**/*.{js,jsx,css}\""
},
"keywords": [
"crdc",
"school-arrests",
"api-demo",
"civilytics"
],
"author": "Civilytics Consulting LLC",
"license": "MIT",
"devDependencies": {
"@duckdb/duckdb-wasm": "^1.32.0",
"@vitejs/plugin-react-swc": "^4.3.3",
"eslint": "^8.57.1",
"prettier": "^3.9.6",
"react": "^19.2.8",
"react-dom": "^19.2.8",
"vite": "^8.2.1"
},
"type": "module",
"dependencies": {
"d3-array": "^3.2.4",
"d3-interpolate": "^3.0.1",
"d3-scale": "^4.0.2",
"d3-shape": "^3.2.0"
}
}
+57
View File
@@ -0,0 +1,57 @@
<?php
/**
* Simple CORS proxy for CRDC Arrest Rate API.
*
* Deploy this file to your git-pages host alongside the static site (e.g., at /crdc-demo/proxy.php).
* Then set VITE_PROXY_URL=/crdc-demo/proxy.php in your build environment.
*
* Usage: GET /proxy.php?target=/api/v1/districts?q=Denver&state=CO
*/
// Allow cross-origin requests from any origin
header('Access-Control-Allow-Origin: *');
header('Access-Control-Allow-Methods: GET, OPTIONS');
header('Access-Control-Allow-Headers: *');
header('Content-Type: application/json');
if ($_SERVER['REQUEST_METHOD'] === 'OPTIONS') {
http_response_code(204);
exit;
}
// Extract and validate the target path from query string
$target = isset($_GET['target']) ? $_GET['target'] : '';
if (empty($target)) {
http_response_code(400);
echo json_encode(['error' => 'Missing "target" parameter']);
exit;
}
// Construct the full API URL — only allow requests to crdc-api.civilytics.org for security
$apiBase = 'https://crdc-api.civilytics.org';
$url = $apiBase . $target;
// Validate that we're not proxying arbitrary URLs (prevent SSRF)
if (!str_starts_with($url, $apiBase . '/api/v1/')) {
http_response_code(403);
echo json_encode(['error' => 'Invalid target — must be under /api/v1/']);
exit;
}
// Fetch the API response and return it with CORS headers
$ch = curl_init($url);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_TIMEOUT => 30,
CURLOPT_SSL_VERIFYPEER => false,
]);
$response = curl_exec($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
// Forward the response with appropriate status code and CORS headers
http_response_code($httpCode);
header('Access-Control-Allow-Origin: *');
echo $response;

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 3.0 KiB

Before

Width:  |  Height:  |  Size: 5.9 KiB

After

Width:  |  Height:  |  Size: 5.9 KiB

View File

Before

Width:  |  Height:  |  Size: 128 B

After

Width:  |  Height:  |  Size: 128 B

+392
View File
@@ -0,0 +1,392 @@
/**
* Analysis: distribution of maximum x-axis values for the rate density chart.
*
* The "Arrest rate probability density" chart currently caps its x-axis at a
* hard-coded MAX_RATE_DOMAIN = 30 (per 1,000 students). We want to understand
* what the true distribution of max x-values is across districts so we can
* decide whether to raise or make this cap dynamic.
*
* What feeds computeRateDomain():
* 1. Per-group rate arrays from real posterior draws → 0.995 quantile of each
* (requires parquet data; approximated here via model estimates)
* 2. Agresti-Coull upper bound rates = (ac.upper / enroll) * 1000, computed
* from observed counts and enrollment — available via the API for every
* race×sex cell in every district
*
* KEY INSIGHT: Only SELECTED groups feed into computeRateDomain. Per
* `defaultSelectedKeys` in districtGroups.js, only groups with ≥1 observed
* arrest are selected by default (or top-2 by enrollment if none have arrests).
* Additionally, when total district arrests < 20, sex pooling merges F+M cells,
* which combines enrollment and observed counts per race.
*/
const BASE = 'https://crdc-api.civilytics.org/api/v1'
const LIMIT = 500
const POOL_THRESHOLD = 20 // POOL_BY_SEX_ARREST_THRESHOLD from pooling.js
/** Fetch JSON from the API, retrying on transient failures. */
async function apiFetch(path) {
const url = `${BASE}${path}`
let lastErr
for (let attempt = 0; attempt <= 3; attempt++) {
try {
const res = await fetch(url, { signal: AbortSignal.timeout(30000) })
if (!res.ok) throw new Error(`HTTP ${res.status}`)
return await res.json()
} catch (err) {
lastErr = err
if (attempt < 3) {
const wait = 500 * Math.pow(2, attempt) + Math.random() * 100
console.error(` retry ${attempt + 1}/3 after ${Math.round(wait)}ms:`, err.message)
await new Promise((r) => setTimeout(r, wait))
}
}
}
throw lastErr
}
/** Peter Acklam's inverse normal CDF (probit), ~1.15e-9 relative error. */
function probit(p) {
const a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02,
1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00]
const b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02,
6.680131188771972e+01, -1.328068155288572e+01]
const c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00,
-2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00]
const d = [7.784695709041462e-03, 3.224671290700398e-01, 2.445134137142996e+00,
3.754408661907416e+00]
const pLow = 0.02425, pHigh = 1 - pLow
if (p < pLow) {
const q = Math.sqrt(-2 * Math.log(p))
return (((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) /
((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)
}
if (p <= pHigh) {
const q = p - 0.5, r = q * q
return (((((a[0]*r+a[1])*r+a[2])*r+a[3])*r+a[4])*r+a[5]) * q /
(((((b[0]*r+b[1])*r+b[2])*r+b[3])*r+b[4])*r+1)
}
const q = Math.sqrt(-2 * Math.log(1 - p))
return -(((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) /
((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)
}
/** Agresti-Coull upper bound (count scale), port of agrestiCoull.js. */
function acUpperBound(numerator, denominator, confidenceLevel = 0.95) {
const adjStar = probit(1 - (1 - confidenceLevel) / 2)
if (numerator > 0) {
const numStar = numerator + adjStar
const denomStar = denominator + 2 * adjStar
const phat = numStar / denomStar
const se = Math.sqrt((phat / denomStar) * (1 - phat))
return (phat + adjStar * se) * denomStar
}
// Zero events: rule of three — upper bound ≈ 3 regardless of enrollment.
return -Math.log(1 - confidenceLevel)
}
/** Estimate the 0.995 quantile of posterior predictive draw rates. */
function estimateDrawQuantile(row, targetP = 0.995) {
const enroll = row.stu_enroll || 0
if (enroll <= 0) return null
// When count_upper is available and > 0, the model's posterior predictive
// upper bound gives a sense of the spread. For sparse groups with few or zero
// arrests, draw quantiles can be extreme because observation noise dominates:
// a single predicted arrest in a small cell produces a huge per-1000 rate.
const countUpper = row.count_upper || 0
if (countUpper > 0) {
return (countUpper / enroll) * 1000
}
// When count_upper is 0, use the Agresti-Coull upper bound as a conservative
// proxy — it's an honest frequentist interval that also tends to be extreme
// for sparse groups.
const ac = acUpperBound(row.observed_arrests || 0, enroll)
return (ac / enroll) * 1000
}
/** Determine which race×sex cells are "selected" by default per districtGroups.js logic. */
function determineSelectedCells(cells, pooled) {
const usable = cells.filter(
(r) => ['WH', 'BL', 'HI', 'AM'].includes(r.race) && ['F', 'M'].includes(r.sex),
)
if (pooled) {
// When pooled: merge F+M per race, then select groups with ≥1 arrest or top-2 by enrollment
const races = {}
for (const r of usable) {
if (!races[r.race]) races[r.race] = { race: r.race, enroll: 0, observed: 0 }
races[r.race].enroll += r.stu_enroll || 0
races[r.race].observed += r.observed_arrests || 0
}
const merged = Object.values(races)
// defaultSelectedKeys logic on pooled groups
const withArrests = merged.filter((r) => r.observed > 0)
if (withArrests.length > 0) {
return withArrests.map((r) => ({ race: r.race, sex: null, enroll: r.enroll, observed: r.observed }))
}
return [...merged]
.sort((a, b) => b.enroll - a.enroll)
.slice(0, 2)
.map((r) => ({ race: r.race, sex: null, enroll: r.enroll, observed: r.observed }))
} else {
// Unpooled: select groups with ≥1 arrest or top-2 by enrollment
const withArrests = usable.filter((r) => (r.observed_arrests || 0) > 0)
if (withArrests.length > 0) return withArrests
return [...usable]
.sort((a, b) => (b.stu_enroll || 0) - (a.stu_enroll || 0))
.slice(0, 2)
}
}
/** Compute max x-value candidate from selected cells. */
function computeMaxX(selectedCells, allCells, pooled) {
let maxAcRate = 0
let maxDrawEstimate = 0
for (const cell of selectedCells) {
const enroll = cell.enroll || 0
if (enroll <= 0) continue
const observed = cell.observed || 0
// AC upper bound rate (always a candidate in computeRateDomain)
const acUpper = acUpperBound(observed, enroll)
const acRate = (acUpper / enroll) * 1000
if (acRate > maxAcRate) maxAcRate = acRate
// Estimated draw quantile — find the matching row in allCells for model estimates
let countUpper = 0
let foundRow = null
if (!pooled && cell.sex) {
foundRow = allCells.find((r) => r.race === cell.race && r.sex === cell.sex)
} else if (pooled) {
// For pooled, find the row with max count_upper for this race across both sexes
const raceRows = allCells.filter((r) => r.race === cell.race && ['F', 'M'].includes(r.sex))
for (const rr of raceRows) {
if ((rr.count_upper || 0) > countUpper) countUpper = rr.count_upper || 0
}
}
if (!pooled && foundRow) {
countUpper = foundRow.count_upper || 0
}
const enroll_ = cell.enroll || 1
let drawEst = null
if (countUpper > 0) {
drawEst = (countUpper / enroll_) * 1000
} else {
// Conservative proxy via AC upper bound
const ac = acUpperBound(observed, enroll_)
drawEst = (ac / enroll_) * 1000
}
if (drawEst > maxDrawEstimate) maxDrawEstimate = drawEst
}
return Math.max(maxAcRate, maxDrawEstimate)
}
async function main() {
console.log('=== Rate Domain Analysis ===\n')
// Step 1: Get all states from the API
const statesResp = await apiFetch('/states?limit=500')
const stateSet = new Set(statesResp.data.map((r) => r.state))
const states = [...stateSet].sort()
console.log(`Found ${states.length} states\n`)
// Step 2: For each state, page through all estimates and collect per-district data
const districtData = new Map() // leaid -> { state, name, cells: [] }
let totalRows = 0
for (const state of states) {
console.log(`Fetching ${state}...`)
const firstResp = await apiFetch(`/estimates?state=${state}&limit=${LIMIT}`)
const total = firstResp.meta.total
const nPages = Math.ceil(total / LIMIT)
let rows = [...firstResp.data]
for (let page = 1; page < nPages; page++) {
process.stderr.write(` ${state} page ${page + 1}/${nPages}\r`)
const resp = await apiFetch(`/estimates?state=${state}&limit=${LIMIT}&page=${page}`)
rows.push(...resp.data)
}
totalRows += rows.length
for (const row of rows) {
const key = `${row.state}|${row.leaid}`
if (!districtData.has(key)) {
districtData.set(key, { state: row.state, leaid: row.leaid, name: row.lea_name, cells: [] })
}
districtData.get(key).cells.push(row)
}
}
console.log(`\nTotal rows fetched: ${totalRows}`)
console.log(`Total districts: ${districtData.size}\n`)
// Step 3: For each district, compute the max x-value for both pooled and unpooled modes
const results = []
let nPooled = 0
let nUnpooled = 0
for (const [key, dist] of districtData) {
const totalArrests = dist.cells.reduce((sum, r) => sum + (r.observed_arrests || 0), 0)
const pooled = totalArrests < POOL_THRESHOLD
if (pooled) nPooled++
else nUnpooled++
let maxX
if (pooled) {
// Pooled mode: AC bounds computed per race (F+M merged). Note that for
// pooled groups, the app shows a note but still computes AC bounds.
const selected = determineSelectedCells(dist.cells, true)
maxX = computeMaxX(selected, dist.cells, true)
} else {
const selected = determineSelectedCells(dist.cells, false)
maxX = computeMaxX(selected, dist.cells, false)
}
let totalEnroll = 0
for (const cell of dist.cells) {
if ((cell.stu_enroll || 0) > 0) totalEnroll += cell.stu_enroll
}
results.push({
key, state: dist.state, leaid: dist.leaid, name: dist.name,
totalEnroll, pooled, maxX: maxX * 1.15, // HEADROOM factor from rateDomain.js
})
}
console.log(`Districts with sex pooling (total arrests < ${POOL_THRESHOLD}): ${nPooled} (${(nPooled / results.length * 100).toFixed(1)}%)`)
console.log(`Districts without pooling: ${nUnpooled} (${(nUnpooled / results.length * 100).toFixed(1)}%)\n`)
// Step 4: Analyze distribution of max x-values
const sorted = results.sort((a, b) => a.maxX - b.maxX)
const n = sorted.length
console.log('=== Distribution of Maximum X-Values (per 1,000 students) ===\n')
// Summary statistics — maxX already includes HEADROOM(1.15) factor
const percentiles = [5, 10, 25, 50, 75, 90, 95, 99, 99.9]
console.log('Percentiles of max x-value (includes HEADROOM=1.15):')
for (const p of percentiles) {
const idx = Math.floor((p / 100) * (n - 1))
console.log(` ${p.toFixed(1)}th: ${sorted[idx].maxX.toFixed(2)} per 1,000`)
}
console.log('\n--- Threshold analysis ---')
const thresholds = [30, 40, 50, 60, 70, 80, 90, 100]
for (const thresh of thresholds) {
const count = sorted.filter((r) => r.maxX > thresh).length
const pct = (count / n) * 100
console.log(` Exceeds ${thresh}: ${count} districts (${pct.toFixed(2)}%)`)
}
// Step 5: Break down by pooling status
console.log('\n--- By pooling status ---')
const pooledResults = sorted.filter((r) => r.pooled)
const unpooledResults = sorted.filter((r) => !r.pooled)
for (const thresh of [30, 50, 100]) {
const pClipped = pooledResults.filter((r) => r.maxX > thresh).length
const uClipped = unpooledResults.filter((r) => r.maxX > thresh).length
console.log(` Cap=${thresh}: pooled ${pClipped}/${pooledResults.length} (${(pClipped / pooledResults.length * 100).toFixed(1)}%), ` +
`unpooled ${uClipped}/${unpooledResults.length} (${(uClipped / unpooledResults.length * 100).toFixed(1)}%)`)
}
// Step 6: Show top districts by max x-value, with enrollment context
console.log('\n--- Top 25 districts by max x-value ---')
const top = sorted.slice(-25).reverse()
for (const r of top) {
const pooledStr = r.pooled ? ' [pooled]' : ''
const clippedAt30 = r.maxX > 30 ? ' *** CLIPPED at 30' : ''
console.log(` ${r.state} | LEAID ${r.leaid} | enroll=${r.totalEnroll.toLocaleString()}${pooledStr} | ` +
`max_x=${r.maxX.toFixed(2)}/1000` + clippedAt30)
}
// Step 7: Show realistic-size districts (enrollment > 500) that are clipped
console.log('\n--- Clipped districts with enrollment > 500 ---')
const realClipped = sorted.filter((r) => r.maxX > 30 && r.totalEnroll >= 500).sort((a, b) => b.maxX - a.maxX).slice(0, 15)
for (const r of realClipped) {
const pooledStr = r.pooled ? ' [pooled]' : ''
console.log(` ${r.state} | ${r.name} | enroll=${r.totalEnroll.toLocaleString()}${pooledStr} | ` +
`max_x=${r.maxX.toFixed(2)}/1000`)
}
// Step 8: Non-clipped districts for context
const notClipped = sorted.filter((r) => r.maxX <= 30).sort((a, b) => a.maxX - b.maxX)
console.log(`\n--- Non-clipped districts (max_x ≤ 30): ${notClipped.length} (${(notClipped.length / n * 100).toFixed(2)}%) ---`)
if (notClipped.length > 0) {
const midIdx = Math.floor(notClipped.length / 2)
console.log(' Sample non-clipped districts:')
for (let i = Math.max(0, midIdx - 3); i < Math.min(notClipped.length, midIdx + 4); i++) {
const r = notClipped[i]
console.log(` ${r.state} | ${r.name} | enroll=${r.totalEnroll.toLocaleString()} | max_x=${r.maxX.toFixed(2)}/1000`)
}
}
// Step 9: State-level summary (focusing on clipped counts)
console.log('\n--- Top 15 states by % of districts clipped ---')
const byState = {}
for (const r of results) {
if (!byState[r.state]) byState[r.state] = []
byState[r.state].push(r.maxX)
}
const stateStats = Object.entries(byState).map(([state, vals]) => ({
state, n: vals.length, median: percentile(vals, 50), p95: percentile(vals, 95),
max: Math.max(...vals), clipped: vals.filter((v) => v > 30).length,
})).sort((a, b) => (b.clipped / b.n) - (a.clipped / a.n))
for (const s of stateStats.slice(0, 15)) {
console.log(` ${s.state}: n=${s.n}, median=${s.median.toFixed(1)}, p95=${s.p95.toFixed(1)}, ` +
`max=${s.max.toFixed(1)}, clipped>30: ${s.clipped} (${(s.clipped / s.n * 100).toFixed(1)}%)`)
}
// Step 10: Recommendation analysis — what cap would minimize clipping while staying bounded?
console.log('\n=== RECOMMENDATION ANALYSIS ===')
const capOptions = [30, 40, 50, 60, 75, 100]
for (const cap of capOptions) {
const clipped = sorted.filter((r) => r.maxX > cap).length
console.log(` Cap=${cap}: ${clipped} districts clipped (${(clipped / n * 100).toFixed(2)}%)`)
}
// Step 11: Key findings summary
console.log('\n--- Key Findings ---')
const pctExceed30 = (sorted.filter((r) => r.maxX > 30).length / n) * 100
const pctExceed50 = (sorted.filter((r) => r.maxX > 50).length / n) * 100
console.log(`- ${pctExceed30.toFixed(2)}% of districts have a max x-value exceeding the current cap of 30`)
console.log(`- ${pctExceed50.toFixed(2)}% exceed 50 per 1,000`)
const p99 = sorted[Math.floor(0.99 * (n - 1))].maxX
const max = sorted[n - 1].maxX
console.log(`- 99th percentile: ${p99.toFixed(2)} per 1,000`)
console.log(`- Maximum observed: ${max.toFixed(2)} per 1,000`)
// Analyze the nature of clipped districts — are they sparse or not?
const extreme = sorted.filter((r) => r.maxX > 30).sort((a, b) => a.totalEnroll - b.totalEnroll)
console.log(`\n--- Clipped district enrollment distribution ---`)
for (const p of [10, 25, 50, 75, 90]) {
const idx = Math.floor((p / 100) * (extreme.length - 1))
console.log(` ${p}th percentile enrollment: ${extreme[idx].totalEnroll.toLocaleString()}`)
}
// How many clipped districts have "normal" school sizes (>1000 students)?
const normalClipped = extreme.filter((r) => r.totalEnroll >= 1000).length
console.log(`\n- ${normalClipped} of ${extreme.length} clipped districts have enrollment ≥ 1,000 (${(normalClipped / extreme.length * 100).toFixed(1)}%)`)
// Analyze what's driving the extremes — AC bounds vs draw estimates
console.log('\n--- What drives extreme values? ---')
const acDriven = sorted.filter((r) => r.maxX > 30 && !r.pooled).length
console.log(`- Unpooled districts clipped: ${acDriven} (${(acDriven / n * 100).toFixed(2)}% of all)`)
function percentile(arr, p) {
const s = [...arr].sort((a, b) => a - b)
return s[Math.floor((p / 100) * (s.length - 1))]
}
}
main().catch(console.error)
+204
View File
@@ -0,0 +1,204 @@
#!/usr/bin/env node
/**
* Builds `public/data/top_districts.json` — the "suggested districts" list the
* search screen shows before you type anything.
*
* Why this exists: the app used to build that list at runtime from
* `/estimates?state=XX&year=21-22&limit=500`. That endpoint returns rows
* `ORDER BY LEAID, RACE, SEX` at eight rows per district, so a 500-row cap is
* the ~62 *lowest-LEAID* districts in the state, not the busiest ones —
* California alone has 11,488 rows. The list was therefore ranked over a
* truncated and essentially arbitrary slice of each state.
*
* This script pages the whole state using `meta.total` from the response
* envelope, aggregates observed arrests per district, and commits the answer as
* a fixture (the `public/data/national_rates.json` precedent). The search screen
* then loses a multi-second fetch and gets a correct ranking.
*
* Read-only against the public API. Roughly 150 requests as a one-off; re-run it
* only when a new CRDC wave lands.
*
* node scripts/build-top-districts.mjs
* node scripts/build-top-districts.mjs --states NV,CA # spot-check a few
*/
import { writeFile, mkdir } from 'node:fs/promises'
import { dirname, resolve } from 'node:path'
import { fileURLToPath } from 'node:url'
const BASE_URL = process.env.CRDC_API_BASE || 'https://crdc-api.civilytics.org/api/v1'
const YEAR = '21-22'
// Pinned rather than left to the API default so a change to that default can't
// silently alter the fixture. Enrollment and observed arrests are the same in
// every specification; only the modelled columns differ, and we read none.
const MODEL = 'unified_m2_mod'
const PAGE_SIZE = 1000 // the API's LIMIT_CAP
const TOP_N = 15
const CONCURRENCY = 3
const MAX_RETRIES = 4
const ALL_STATES = [
'AL', 'AK', 'AZ', 'AR', 'CA', 'CO', 'CT', 'DE', 'DC', 'FL', 'GA', 'HI',
'ID', 'IL', 'IN', 'IA', 'KS', 'KY', 'LA', 'ME', 'MD', 'MA', 'MI', 'MN',
'MS', 'MO', 'MT', 'NE', 'NV', 'NH', 'NJ', 'NM', 'NY', 'NC', 'ND', 'OH',
'OK', 'OR', 'PA', 'RI', 'SC', 'SD', 'TN', 'TX', 'UT', 'VT', 'VA', 'WA',
'WV', 'WI', 'WY',
]
const OUT_PATH = resolve(
dirname(fileURLToPath(import.meta.url)),
'..',
'public',
'data',
'top_districts.json',
)
function parseStates() {
const flag = process.argv.indexOf('--states')
if (flag === -1) return ALL_STATES
const requested = (process.argv[flag + 1] || '').split(',').map((s) => s.trim().toUpperCase())
const unknown = requested.filter((s) => !ALL_STATES.includes(s))
if (unknown.length) throw new Error(`Unknown state code(s): ${unknown.join(', ')}`)
return requested
}
const sleep = (ms) => new Promise((r) => setTimeout(r, ms))
/** GET one page, returning the full envelope (we need `meta.total`). */
async function fetchPage(state, page) {
const params = new URLSearchParams({
state,
year: YEAR,
model: MODEL,
limit: String(PAGE_SIZE),
page: String(page),
})
const url = `${BASE_URL}/estimates?${params}`
let lastError
for (let attempt = 0; attempt <= MAX_RETRIES; attempt++) {
try {
const res = await fetch(url, { signal: AbortSignal.timeout(60000) })
if (!res.ok) throw new Error(`HTTP ${res.status} ${res.statusText}`)
const envelope = await res.json()
if (envelope.status !== 'success') throw new Error(envelope.error || 'Unknown API error')
if (!envelope.meta || typeof envelope.meta.total !== 'number') {
throw new Error('Response envelope is missing meta.total — cannot page safely')
}
return envelope
} catch (err) {
lastError = err
if (attempt === MAX_RETRIES) break
await sleep(500 * 2 ** attempt)
}
}
throw new Error(`${state} page ${page}: ${lastError.message}`)
}
async function collectState(state) {
const first = await fetchPage(state, 0)
const total = first.meta.total
const rows = [...first.data]
const pages = Math.ceil(total / PAGE_SIZE)
for (let page = 1; page < pages; page++) {
const envelope = await fetchPage(state, page)
rows.push(...envelope.data)
}
if (rows.length !== total) {
// Loud rather than silent: a short read here would quietly produce a
// truncated ranking, which is the exact bug this script exists to fix.
throw new Error(`${state}: expected ${total} rows, collected ${rows.length}`)
}
const byLeaid = new Map()
for (const row of rows) {
const leaid = row.leaid
if (!leaid) continue
const prev = byLeaid.get(leaid) || { leaid, name: row.lea_name || leaid, arrests: 0, enrollment: 0 }
byLeaid.set(leaid, {
...prev,
name: prev.name || row.lea_name || leaid,
arrests: prev.arrests + (row.observed_arrests || 0),
enrollment: prev.enrollment + (row.stu_enroll || 0),
})
}
const ranked = [...byLeaid.values()]
.filter((d) => d.arrests > 0)
.sort((a, b) => b.arrests - a.arrests || a.leaid.localeCompare(b.leaid))
.slice(0, TOP_N)
.map((d) => ({
leaid: d.leaid,
name: d.name,
arrests: d.arrests,
enrollment: d.enrollment,
rate: d.enrollment > 0 ? Math.round((d.arrests / d.enrollment) * 1000 * 100) / 100 : 0,
}))
return { state, districts: ranked, districtsSeen: byLeaid.size, rows: total, pages }
}
/** Small fixed-size worker pool — polite to a single public API host. */
async function mapWithConcurrency(items, limit, worker) {
const results = new Array(items.length)
let next = 0
const runners = Array.from({ length: Math.min(limit, items.length) }, async () => {
while (next < items.length) {
const i = next++
results[i] = await worker(items[i], i)
}
})
await Promise.all(runners)
return results
}
async function main() {
const states = parseStates()
console.log(`Fetching ${states.length} state(s) from ${BASE_URL} (year ${YEAR}, model ${MODEL})…`)
let done = 0
let requests = 0
const collected = await mapWithConcurrency(states, CONCURRENCY, async (state) => {
const result = await collectState(state)
requests += result.pages
done += 1
console.log(
` [${String(done).padStart(2)}/${states.length}] ${state}: ` +
`${result.rows} rows / ${result.pages} page(s), ` +
`${result.districtsSeen} districts, top ${result.districts.length} kept`,
)
return result
})
const byState = {}
for (const { state, districts } of collected.sort((a, b) => a.state.localeCompare(b.state))) {
byState[state] = districts
}
const payload = {
metadata: {
source: 'CRDC School Arrest Rate API (Knowles & Miller 2025)',
endpoint: `${BASE_URL}/estimates`,
year: YEAR,
model: MODEL,
description:
`Top ${TOP_N} school districts per state by total observed arrests in ${YEAR}, ` +
'summed across the eight modelled race×sex groups. Generated by ' +
'scripts/build-top-districts.mjs over the complete paged result set for each state.',
generated_states: states.length,
generated_requests: requests,
},
states: byState,
}
await mkdir(dirname(OUT_PATH), { recursive: true })
await writeFile(OUT_PATH, `${JSON.stringify(payload, null, 2)}\n`, 'utf8')
console.log(`\nWrote ${OUT_PATH} (${requests} requests, ${states.length} states).`)
}
main().catch((err) => {
console.error('\nbuild-top-districts failed:', err.message)
process.exit(1)
})
+122
View File
@@ -0,0 +1,122 @@
import { useState, useEffect } from 'react'
import StateSelector from './components/StateSelector.jsx'
import DistrictSearch from './components/DistrictSearch.jsx'
import LoadingAnimation from './components/LoadingAnimation.jsx'
import ChartPanel from './components/ChartPanel.jsx'
import Footer from './components/Footer.jsx'
import { fetchDistrictEstimates } from './hooks/useApi.js'
// Two-letter USPS state code. The deep-link value flows into API request URLs
// and into the Hugging Face parquet shard path duckdb-wasm reads, so it gets
// validated here at the boundary rather than propagated verbatim.
const STATE_CODE_PATTERN = /^[A-Z]{2}$/
/**
* CRDC Arrests API Demo App — main router.
* Flow: state → district search (with interesting suggestions) → loading animation → charts
*/
export default function App() {
const [step, setStep] = useState('state') // 'state' | 'search' | 'loading' | 'results'
const [selectedState, setSelectedState] = useState(null)
const [district, setDistrict] = useState(null)
// On first load: check URL for ?leaid= and state to allow deep-linking
useEffect(() => {
const params = new URLSearchParams(window.location.search)
const leaid = params.get('leaid')
const rawState = params.get('state')
const stateParam = rawState ? rawState.trim().toUpperCase() : null
if (rawState && !STATE_CODE_PATTERN.test(stateParam)) {
// Malformed deep link — drop it and start at the state selector rather
// than passing an arbitrary string into fetch URLs and shard paths.
console.warn('Ignoring deep link: `state` is not a two-letter state code.')
return
}
if (leaid && stateParam) {
// Deep link (including browser back/forward landing on this URL): resolve
// the district name via a real estimates row, keyed by LEAID. The previous
// version called searchDistricts(leaid, ...), a name/text search — passing
// an LEAID as search text never matches, so it always fell back to "Unknown
// District".
setSelectedState(stateParam)
fetchDistrictEstimates(leaid, { model: 'unified_m2_mod', year: '21-22' }).then(rows => {
const row = rows?.[0]
setDistrict(row
? { leaid: row.leaid, lea_name: row.lea_name, state: row.state }
: { leaid, lea_name: 'Unknown District', state: stateParam })
setStep('loading')
}).catch(() => {
// Fallback on error
setDistrict({ leaid, lea_name: 'Unknown District', state: stateParam })
setStep('loading')
})
}
}, [])
const handleStateChange = (stateCode) => {
setSelectedState(stateCode)
setStep('search')
}
const handleDistrictSelect = (dist) => {
// Update URL for shareability
const params = new URLSearchParams({ leaid: dist.leaid, state: selectedState })
window.history.replaceState(null, '', `?${params}`)
setDistrict(dist)
setStep('loading')
}
return (
<div className="cv-app">
{/* Header */}
<header className="cv-header cv-wrap">
<a href="/crdc-demo" aria-label="Civilytics" style={{ display: 'block' }}>
<img
src={import.meta.env.BASE_URL + 'civilytics-wordmark.svg'}
alt="Civilytics — social science for the public good"
style={{ height: '2.5rem', width: 'auto' }}
/>
</a>
</header>
{/* Main content */}
<main style={{ maxWidth: '60rem', margin: '0 auto', padding: '0 var(--space-3)' }}>
{step === 'state' && (
<StateSelector onNext={handleStateChange} />
)}
{step === 'search' && selectedState && (
<DistrictSearch
state={selectedState}
onSelect={handleDistrictSelect}
onBack={() => setStep('state')}
/>
)}
{step === 'loading' && district && (
<LoadingAnimation district={district} state={selectedState} />
)}
{/* ChartPanel listens for data-ready event from LoadingAnimation */}
{step === 'results' && district && (
<ChartPanel district={district} state={selectedState} />
)}
{(step === 'loading' || step === 'results') && (
<button
className="btn-outline"
style={{ margin: 'var(--space-3) 0', display: 'block' }}
onClick={() => { window.location.href = '/crdc-demo/' }}
>
← Search another district
</button>
)}
</main>
<Footer />
</div>
)
}
+227
View File
@@ -0,0 +1,227 @@
import ChartLegend from '../components/ChartLegend.jsx'
import { OBSERVED_MARK_COLOR, MODELED_AGGREGATE_COLOR } from '../utils/colors.js'
import { niceTicks } from '../utils/niceTicks.js'
import { MODEL_QUADRANT_LABEL } from '../utils/labels.js'
/**
* Chart 1: Observed arrests by CRDC wave (line, diamond markers — the same
* "diamond = observed" convention as every other chart), overlaid with one
* model's predicted total (point-range) per wave, aligned to the same x
* position as its observed point rather than dodged aside, so the two are
* read as a pair.
*/
/**
* Says which interval is actually on screen, and offers the better one when it
* hasn't been fetched. The distinction is not cosmetic: summing each group's
* own bounds answers "what if every group hit its extreme at once", which is a
* far wider claim than "how many arrests does the model think there were".
*/
function TotalIntervalNote({ totals, usingExact, onRequest }) {
const base = { fontSize: '0.75rem', margin: 'var(--space-1) 0 0' }
if (usingExact) {
return (
<p style={{ ...base, color: 'var(--cv-ink-3)' }}>
The modeled band is the 95% interval of the district <em>total</em>, taken from{' '}
{(totals?.byYear && Object.values(totals.byYear).find(Boolean)?.nDraws?.toLocaleString()) || '500'}{' '}
posterior predictive draws — summed across student groups within each draw, then
summarized across draws.
</p>
)
}
if (totals?.status === 'loading') {
return <p style={{ ...base, color: 'var(--cv-ink-3)' }}>Computing the total from posterior draws…</p>
}
const mb = totals?.bytes ? (totals.bytes / 1048576).toFixed(1) : null
return (
<p style={{ ...base, color: 'var(--cv-ink-3)' }}>
The modeled band sums each student group&rsquo;s own 95% bounds, which is wider than the
interval of the total — it is the case where every group lands at its extreme in the same
draw.{' '}
{totals?.status === 'error' ? (
<>The exact total could not be computed from the draws for this district.</>
) : (
onRequest && (
<button type="button" onClick={onRequest} style={linkButton}>
Compute the exact interval from the draws{mb ? ` (${mb} MB)` : ''}
</button>
)
)}
</p>
)
}
const linkButton = {
background: 'none',
border: 'none',
padding: 0,
font: 'inherit',
color: 'var(--cv-navy-600)',
textDecoration: 'underline',
cursor: 'pointer',
}
const WAVE_LABELS = { '15-16': '2015–16', '17-18': '2017–18', '21-22': '2021–22' }
const HEADROOM = 26 // px reserved at the top of the plot so a point's rate label has room to sit above it
// Matches RateDensityPanel/GroupDifference. These SVGs are rendered at
// width="100%" inside the same card, so the viewBox width sets the scale factor
// for everything in it: at 360 this chart's 0.7rem text came out roughly twice
// the size of the other charts' 0.62rem, which is what made it look like a
// different chart from a different app.
const WIDTH = 760
const HEIGHT = 300
const MARGIN = { top: 20, right: 24, bottom: 52, left: 56 }
// Keeps the first and last wave off the plot edges. Without it the end markers
// sit flush against the y-axis and the right border, so half of each diamond's
// glyph is visually clipped and its label has nowhere to go.
const X_PAD = 64
export default function ArrestsOverTime({ data, modelId, totals, onRequestExactTotals }) {
// Prefer the interval computed from the draws (summed within each draw) over
// the sum of each group's own bounds, which is not the interval of the total
// and comes out systematically too wide.
const exact = totals?.status === 'ready' ? totals.byYear : null
const series = data.map((d) => {
const iv = exact?.[d.year]
return iv
? { ...d, modeledMedian: iv.median, modeledLower: iv.lower, modeledUpper: iv.upper, exact: true }
: { ...d, exact: false }
})
const usingExact = series.some((d) => d.exact)
const maxArrests = Math.max(
...series.map((d) => Math.max(d.arrests, d.modeledUpper ?? 0)),
1
)
const { ticks, niceMax } = niceTicks(maxArrests)
const modelLabel = MODEL_QUADRANT_LABEL[modelId] || modelId
const width = WIDTH
const height = HEIGHT
const margin = MARGIN
const innerWidth = width - margin.left - margin.right
const innerHeight = height - margin.top - margin.bottom
const plotLeft = margin.left + X_PAD
const plotRight = margin.left + innerWidth - X_PAD
const xScale = (i) => plotLeft + (i / Math.max(series.length - 1, 1)) * (plotRight - plotLeft)
const yScale = (val) => HEADROOM + (innerHeight - HEADROOM) * (1 - val / niceMax)
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={{ fontSize: '0.85rem', marginBottom: 0, color: 'var(--cv-ink-2)' }}>
Arrests over time — observed vs. modeled
</h3>
{modelLabel && (
<p style={{ fontSize: '0.78rem', fontStyle: 'italic', color: 'var(--cv-ink-3)', margin: '0.15rem 0 0' }}>
Modeled intervals use the {modelLabel.toLowerCase()} model.
</p>
)}
<svg width="100%" viewBox={`0 0 ${width} ${height}`}
style={{ maxWidth: '100%', minWidth: '320px', marginTop: 'var(--space-1)' }}
role="img" aria-label="Total arrests by CRDC wave, observed against the modeled total">
{ticks.map((val) => {
const y = margin.top + yScale(val)
return (
<g key={`y-${val}`}>
<line x1={margin.left} y1={y} x2={margin.left + innerWidth} y2={y}
stroke="var(--cv-rule)" strokeWidth={1} />
<text x={margin.left - 8} y={y + 3} textAnchor="end"
fontSize="0.62rem" fill="var(--cv-ink-3)">{val.toLocaleString()}</text>
</g>
)
})}
{/* Baseline, so the wave labels read as sitting on an axis */}
<line x1={margin.left} y1={margin.top + yScale(0)} x2={margin.left + innerWidth}
y2={margin.top + yScale(0)} stroke="var(--cv-rule-strong)" strokeWidth={1} />
<text x={14} y={margin.top + innerHeight / 2} textAnchor="middle"
fontSize="0.64rem" fill="var(--cv-ink-3)" transform={`rotate(-90 14 ${margin.top + innerHeight / 2})`}>
Total arrests
</text>
{series.map((d, i) => (
<text key={d.year} x={xScale(i)} y={margin.top + yScale(0) + 18}
textAnchor="middle" fontSize="0.62rem" fill="var(--cv-ink-2)">
{WAVE_LABELS[d.year] || d.label}
</text>
))}
<text x={margin.left + innerWidth / 2} y={height - 6} textAnchor="middle"
fontSize="0.64rem" fill="var(--cv-ink-3)">CRDC wave</text>
{/* Modeled point-range per wave — drawn first, directly under the
observed marks it pairs with, and kept visually subordinate (thinner,
no halo) so the observed series reads as the primary line. */}
{series.map((d, i) => {
if (d.modeledMedian == null) return null
const cx = xScale(i)
const cyMedian = margin.top + yScale(d.modeledMedian)
const cyLower = margin.top + yScale(d.modeledLower ?? d.modeledMedian)
const cyUpper = margin.top + yScale(d.modeledUpper ?? d.modeledMedian)
return (
<g key={`modeled-${d.year}`}>
<line x1={cx} y1={cyLower} x2={cx} y2={cyUpper}
stroke={MODELED_AGGREGATE_COLOR} strokeWidth={1.5} strokeLinecap="round" opacity={0.75} />
<circle cx={cx} cy={cyMedian} r={3.5} fill={MODELED_AGGREGATE_COLOR}
stroke="var(--cv-paper)" strokeWidth={1.25} />
</g>
)
})}
{/* Observed line */}
{series.length > 1 && (
<polyline
points={series.map((d, i) => `${xScale(i)},${margin.top + yScale(d.arrests)}`).join(' ')}
fill="none" stroke={OBSERVED_MARK_COLOR} strokeWidth={2}
strokeLinejoin="round" strokeLinecap="round"
/>
)}
{/* Observed diamonds + rate-per-1k labels. X_PAD keeps the end markers
clear of the plot edges, so every label can be centred over its own
point instead of being anchored outward to avoid an overflow. */}
{series.map((d, i) => {
const cx = xScale(i)
const cy = margin.top + yScale(d.arrests)
const ratePerK = d.enroll > 0 ? (d.arrests / (d.enroll / 1000)).toFixed(2) : '0.00'
return (
<g key={d.year}>
<rect x={cx - 4.5} y={cy - 4.5} width={9} height={9}
fill={OBSERVED_MARK_COLOR} stroke="var(--cv-paper)" strokeWidth={1.5}
transform={`rotate(45 ${cx} ${cy})`} />
<text x={cx} y={cy - 13} textAnchor="middle" fontSize="0.66rem" fontWeight={700}
fill="var(--cv-ink)" stroke="var(--cv-paper)" strokeWidth={3.5}
strokeLinejoin="round" paintOrder="stroke">
{ratePerK}
<tspan fontSize="0.58rem" fontWeight={500} fill="var(--cv-ink-3)"> per 1,000</tspan>
</text>
</g>
)
})}
</svg>
<ChartLegend items={[
{ shape: 'diamond', color: OBSERVED_MARK_COLOR, label: 'Observed' },
// 95%, not 90%: the API's count_lower/count_upper are a 95% interval
// (validate_interval() defaults to 95), and the draws-based total is
// computed at the same mass, so the two are directly comparable.
{ shape: 'dot', color: MODELED_AGGREGATE_COLOR, label: 'Modeled (median + 95% interval)' },
]} />
<TotalIntervalNote totals={totals} usingExact={usingExact} onRequest={onRequestExactTotals} />
<p style={{ fontSize: '0.75rem', color: 'var(--cv-ink-3)', marginTop: 'var(--space-1)' }}>
Rate per 1,000 students labeled above each observed point. Data from CRDC waves 2015–16
through 2021–22.
</p>
</div>
)
}
+394
View File
@@ -0,0 +1,394 @@
import { useMemo } from 'react'
import { area, curveMonotoneX, curveStep } from 'd3-shape'
import { scaleLinear } from 'd3-scale'
import { interpolateRgb } from 'd3-interpolate'
import { rateProfile } from '../utils/densityProfile.js'
import { differenceRates, differenceSummary } from '../utils/groupDifference.js'
import { displayDraws } from '../utils/districtGroups.js'
import { niceTicks } from '../utils/niceTicks.js'
/**
* Chart B — "Model estimated differences".
*
* Port of the white paper's Fig 7 (`wp_fig_group_difference`): the posterior
* distribution of Δ = rate(A) − rate(B), per 1,000 students, computed at each
* draw index.
*
* Two deliberate departures from the R figure:
*
* - The fill is a **diverging** ramp centred at zero (navy below, paper at
* zero, ember above) rather than the paper's sequential YlOrRd. Δ is a signed
* quantity; a sequential ramp encodes "more" where the data means "which
* direction", and would make a large negative difference read as a small one.
* - The readout is spelled out in a sentence. `Pr(Δ > 0) = 94.4%` is the number
* a reader is most likely to misread as "94.4% more arrests".
*/
const NEGATIVE_COLOR = '#22406A' // --cv-navy-600
const ZERO_COLOR = '#F2EDE4' // --cv-paper-2
const POSITIVE_COLOR = '#C25311' // --cv-accent
const GRADIENT_STOPS = 24
const WIDTH = 760
const HEIGHT = 250
const MARGIN = { top: 18, right: 22, bottom: 46, left: 22 }
export default function GroupDifference({
groups,
enrollByGroup,
byModel,
selectedModel,
status,
pooled,
pair,
onPairChange,
}) {
const { counts, enroll } = useMemo(
() => displayDraws(byModel?.[selectedModel]?.counts, enrollByGroup, pooled),
[byModel, selectedModel, enrollByGroup, pooled],
)
// Only groups with a usable draw set can be differenced at all — offering the
// others in the picker would produce an empty chart with no explanation.
const comparable = useMemo(
() => groups.filter((g) => counts?.[g.key]?.length > 0 && enroll?.[g.key] > 0),
[groups, counts, enroll],
)
const [keyA, keyB] = pair || []
const groupA = comparable.find((g) => g.key === keyA)
const groupB = comparable.find((g) => g.key === keyB)
const deltas = useMemo(
() =>
groupA && groupB
? differenceRates(counts[groupA.key], enroll[groupA.key], counts[groupB.key], enroll[groupB.key])
: [],
[groupA, groupB, counts, enroll],
)
const summary = useMemo(() => differenceSummary(deltas), [deltas])
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<div style={headerStyle}>
<div>
<h3 style={cardTitle}>Model estimated differences</h3>
<p style={subtitleStyle}>
How much higher is one group&rsquo;s modelled arrest rate than another&rsquo;s, and how
sure is the model of the direction?
</p>
</div>
{comparable.length >= 2 && (
<PairPickers
comparable={comparable}
keyA={keyA}
keyB={keyB}
onPairChange={onPairChange}
/>
)}
</div>
{status === 'loading' ? (
<p style={emptyStyle}>Loading posterior draws…</p>
) : comparable.length < 2 ? (
<Degraded
reason={
comparable.length === 1
? `Only ${comparable[0].label} has a usable set of posterior draws in this district, so there is no second group to compare it against.`
: 'No student group in this district has a usable set of posterior draws, so no difference can be computed. This is usually because the district is absent from the published draw shard for this model.'
}
/>
) : !groupA || !groupB || groupA.key === groupB.key ? (
<Degraded reason="Pick two different student groups to compare." />
) : !summary ? (
<Degraded
reason={`${groupA.label} and ${groupB.label} do not have matching draw sets in this model specification, so their difference cannot be computed draw by draw.`}
/>
) : (
<>
<Readout summary={summary} groupA={groupA} groupB={groupB} />
<DifferencePlot deltas={deltas} summary={summary} groupA={groupA} groupB={groupB} />
<p style={captionStyle}>
Δ is computed at each of {summary.n.toLocaleString()} posterior predictive draws as{' '}
{groupA.label} minus {groupB.label}, per 1,000 students. The dashed line marks no
difference. Bars beneath the curve are the 80% and 95% intervals around the median Δ.
</p>
</>
)}
</div>
)
}
function Readout({ summary, groupA, groupB }) {
const pct = summary.prGreater * 100
const higher = summary.median >= 0 ? groupA : groupB
const lower = summary.median >= 0 ? groupB : groupA
const share = summary.median >= 0 ? pct : 100 - pct
return (
<div style={readoutStyle}>
<div>
<span style={readoutNumber}>{formatPercent(pct)}</span>
<span style={readoutLabel}>Pr(Δ &gt; 0)</span>
</div>
<p style={{ margin: 0, fontSize: '0.88rem', maxWidth: '38rem' }}>
In {formatPercent(share)} of posterior predictive draws, the {higher.sentenceLabel} arrest
rate exceeds the {lower.sentenceLabel} rate. The median difference is{' '}
<strong>{formatDelta(summary.median)}</strong> per 1,000 students.
</p>
</div>
)
}
function DifferencePlot({ deltas, summary, groupA, groupB }) {
const innerWidth = WIDTH - MARGIN.left - MARGIN.right
const baselineY = HEIGHT - MARGIN.bottom
const { ticks, min, max } = useMemo(() => symmetricDomain(summary), [summary])
const x = scaleLinear().domain([min, max]).range([MARGIN.left, MARGIN.left + innerWidth])
const profile = useMemo(() => rateProfile(deltas, { min, max, n: 80 }), [deltas, min, max])
if (!profile) return null
const y = scaleLinear().domain([0, profile.maxY || 1]).range([baselineY, MARGIN.top])
const points =
profile.kind === 'mass' ? padMassPoints(profile.points, profile.step, min, max) : profile.points
const areaGen = area()
.x((p) => x(clamp(p.x, min, max)))
.y0(baselineY)
.y1((p) => y(p.y))
.curve(profile.kind === 'mass' ? curveStep : curveMonotoneX)
const gradientId = `delta-gradient-${groupA.key}-${groupB.key}`
return (
<div style={{ overflowX: 'auto' }}>
<svg
width="100%"
viewBox={`0 0 ${WIDTH} ${HEIGHT}`}
style={{ maxWidth: '100%', minWidth: '320px' }}
role="img"
aria-label={`Posterior distribution of the difference in arrest rate between ${groupA.label} and ${groupB.label}`}
>
<defs>
<linearGradient id={gradientId} x1="0" y1="0" x2="1" y2="0">
{divergingStops(min, max).map((s) => (
<stop key={s.offset} offset={`${s.offset * 100}%`} stopColor={s.color} />
))}
</linearGradient>
</defs>
{ticks.map((t) => (
<line key={t} x1={x(t)} y1={MARGIN.top} x2={x(t)} y2={baselineY} stroke="var(--cv-rule)" strokeWidth={1} />
))}
<path d={areaGen(points)} fill={`url(#${gradientId})`} opacity={0.85} />
<path
d={areaGen.lineY1()(points)}
fill="none"
stroke="var(--cv-ink-2)"
strokeWidth={1.5}
strokeLinejoin="round"
/>
{/* No difference — the paper's red vertical line. */}
<line
x1={x(0)}
y1={MARGIN.top - 4}
x2={x(0)}
y2={baselineY + 6}
stroke="var(--cv-danger)"
strokeWidth={2}
strokeDasharray="5 4"
/>
<text x={x(0)} y={MARGIN.top - 7} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-danger)">
no difference
</text>
<line x1={MARGIN.left} y1={baselineY} x2={MARGIN.left + innerWidth} y2={baselineY} stroke="var(--cv-rule-strong)" strokeWidth={1} />
{/* Median with 80% (thick) and 95% (thin) intervals. */}
<g>
<line x1={x(clamp(summary.lower95, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper95, min, max))} y2={baselineY + 13} stroke="var(--cv-ink-2)" strokeWidth={1.5} strokeLinecap="round" />
<line x1={x(clamp(summary.lower80, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper80, min, max))} y2={baselineY + 13} stroke="var(--cv-ink)" strokeWidth={4} strokeLinecap="round" />
<circle cx={x(clamp(summary.median, min, max))} cy={baselineY + 13} r={3.5} fill="var(--cv-paper)" stroke="var(--cv-ink)" strokeWidth={2} />
</g>
{ticks.map((t) => (
<text key={t} x={x(t)} y={baselineY + 32} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-ink-3)">
{formatTick(t)}
</text>
))}
<text x={MARGIN.left} y={HEIGHT - 4} textAnchor="start" fontSize="0.62rem" fill="var(--cv-ink-3)">
← {groupB.shortLabel} higher
</text>
<text x={MARGIN.left + innerWidth} y={HEIGHT - 4} textAnchor="end" fontSize="0.62rem" fill="var(--cv-ink-3)">
{groupA.shortLabel} higher →
</text>
</svg>
</div>
)
}
/**
* A domain centred on zero. A signed quantity drawn on an off-centre axis makes
* the eye read the *position* of the curve as the size of the difference, so
* zero sits in the middle even when every draw falls on one side of it.
*/
function symmetricDomain(summary) {
const extent = Math.max(
Math.abs(summary.lower95),
Math.abs(summary.upper95),
Math.abs(summary.median),
1e-6,
)
const { niceMax } = niceTicks(extent * 1.25, 4)
const step = niceMax / 4
const ticks = []
for (let i = -4; i <= 4; i++) ticks.push(Math.round(step * i * 1e6) / 1e6)
return { ticks, min: -niceMax, max: niceMax }
}
/** Diverging ramp, with the paper-coloured midpoint pinned to Δ = 0. */
function divergingStops(min, max) {
const zeroOffset = (0 - min) / (max - min)
const toNegative = interpolateRgb(NEGATIVE_COLOR, ZERO_COLOR)
const toPositive = interpolateRgb(ZERO_COLOR, POSITIVE_COLOR)
const stops = []
for (let i = 0; i <= GRADIENT_STOPS; i++) {
const offset = i / GRADIENT_STOPS
const color =
offset <= zeroOffset
? toNegative(zeroOffset > 0 ? offset / zeroOffset : 1)
: toPositive(zeroOffset < 1 ? (offset - zeroOffset) / (1 - zeroOffset) : 0)
stops.push({ offset, color })
}
return stops
}
function padMassPoints(points, step, min, max) {
const half = Math.max(step, 1e-6) / 2
return [
{ x: Math.max(points[0].x - half, min), y: 0 },
...points,
{ x: Math.min(points[points.length - 1].x + half, max), y: 0 },
]
}
function PairPickers({ comparable, keyA, keyB, onPairChange }) {
return (
<div style={{ display: 'flex', alignItems: 'center', gap: '0.4rem', flexWrap: 'wrap' }}>
<GroupSelect
label="Compare"
value={keyA}
options={comparable}
onChange={(v) => onPairChange([v, keyB])}
/>
<span style={{ fontSize: '0.8rem', color: 'var(--cv-ink-3)' }}>against</span>
<GroupSelect
label="against"
value={keyB}
options={comparable}
onChange={(v) => onPairChange([keyA, v])}
/>
</div>
)
}
function GroupSelect({ label, value, options, onChange }) {
return (
<select
aria-label={label}
value={value || ''}
onChange={(e) => onChange(e.target.value)}
style={{ padding: '0.25rem 0.4rem', fontFamily: 'var(--font-sans)', fontSize: '0.78rem' }}
>
{options.map((g) => (
<option key={g.key} value={g.key}>
{g.label}
</option>
))}
</select>
)
}
function Degraded({ reason }) {
return (
<div style={degradedStyle}>
<p style={{ margin: 0, fontSize: '0.85rem' }}>{reason}</p>
</div>
)
}
function clamp(v, min, max) {
return Math.min(Math.max(v, min), max)
}
function formatPercent(pct) {
return `${pct.toLocaleString(undefined, { minimumFractionDigits: 1, maximumFractionDigits: 1 })}%`
}
function formatDelta(v) {
const sign = v > 0 ? '+' : ''
return `${sign}${v.toLocaleString(undefined, { minimumFractionDigits: 2, maximumFractionDigits: 2 })}`
}
function formatTick(v) {
if (v === 0) return '0'
const abs = Math.abs(v)
return v.toLocaleString(undefined, {
minimumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
maximumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
})
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const headerStyle = {
display: 'flex',
justifyContent: 'space-between',
alignItems: 'flex-start',
flexWrap: 'wrap',
gap: 'var(--space-2)',
}
const subtitleStyle = { fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: 0, maxWidth: '32rem' }
const emptyStyle = { color: 'var(--cv-ink-3)', fontSize: '0.85rem', padding: 'var(--space-2) 0' }
const captionStyle = { fontSize: '0.74rem', color: 'var(--cv-ink-3)', margin: '0.4rem 0 0', lineHeight: 1.5 }
const readoutStyle = {
display: 'flex',
alignItems: 'center',
gap: 'var(--space-3)',
flexWrap: 'wrap',
padding: 'var(--space-2) 0',
borderTop: '3px double var(--cv-ink)',
borderBottom: '1px solid var(--cv-rule)',
margin: 'var(--space-2) 0',
}
const readoutNumber = {
fontFamily: "'Source Serif 4', Georgia, serif",
fontWeight: 700,
fontSize: '2.75rem',
lineHeight: 1,
letterSpacing: '-0.02em',
color: 'var(--cv-ink)',
fontVariantNumeric: 'tabular-nums',
display: 'block',
}
const readoutLabel = {
fontSize: '0.72rem',
fontWeight: 600,
textTransform: 'uppercase',
letterSpacing: '0.08em',
color: 'var(--cv-ink-3)',
display: 'block',
marginTop: '0.3rem',
}
const degradedStyle = {
background: 'var(--cv-paper-2)',
border: '1px solid var(--cv-rule)',
borderLeft: '3px solid var(--cv-ink-4)',
borderRadius: 'var(--radius-md)',
padding: 'var(--space-2)',
marginTop: 'var(--space-2)',
color: 'var(--cv-ink-2)',
}
+580
View File
@@ -0,0 +1,580 @@
import { useMemo } from 'react'
import { area, curveMonotoneX, curveStep } from 'd3-shape'
import { scaleLinear } from 'd3-scale'
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
import { raceColor } from '../utils/colors.js'
import { agrestiCoull } from '../utils/agrestiCoull.js'
import { densityProfile } from '../utils/densityProfile.js'
import { displayDraws } from '../utils/districtGroups.js'
import { computeRateDomain } from '../utils/rateDomain.js'
import { layoutPeakLabels } from '../utils/labelLayout.js'
import { toRates } from '../utils/pooling.js'
import { densityCurve, fitSkewedInterval } from '../utils/distributionApprox.js'
import ApproxNote from '../components/ApproxNote.jsx'
/**
* Chart A — "Arrest rate probability density".
*
* Port of the white paper's Fig 6 (`wp_fig_group_density` in
* crdc-arrests/R/paper_figures.R): each selected student group's posterior
* predictive arrest rate as a filled density, with that group's frequentist
* Agresti–Coull interval on a rail beneath, so model and observation can be
* read against each other.
*
* Layout follows the palette contract in utils/colors.js: **race is hue, sex is
* position**. Female and Male are two stacked sub-panels sharing one x-axis,
* never two hues; when sex pooling is on there is a single panel and the sex
* dimension disappears entirely rather than being recoloured.
*
* React owns the DOM here — d3 supplies scales and path generators only. No
* selections, no imperative mutation.
*/
const AGRESTI_COULL_LEVEL = 0.95
const FILL_OPACITY = 0.4
const STROKE_WIDTH = 2
const CHART_WIDTH = 760
const MARGIN = { top: 22, right: 18, bottom: 34, left: 16 }
const ROW_HEIGHT = 118
const COMPARE_ROW_HEIGHT = 74
const RAIL_HEIGHT = 26
// Direct labels live in a reserved band above each row rather than floating at
// each curve's apex. Every density is normalized to its own peak, so all the
// apexes sit at the same height — placing labels there stacks them on one line
// and they overprint into gibberish as soon as two groups have similar rates.
const LABEL_FONT_PX = 10.4 // 0.65rem
const COMPACT_LABEL_FONT_PX = 9.3 // 0.58rem
// Comfortably above the rendered line box (a 10.4px label measures ~12.1px tall
// with descenders), so adjacent lanes clear each other instead of just touching.
const LABEL_LINE_HEIGHT = 14
const COMPACT_LABEL_LINE_HEIGHT = 12
const LABEL_GAP = 8
const SPEC_LABEL_HEIGHT = 13
export default function RateDensityPanel({
groups,
selectedKeys,
enrollByGroup,
byModel,
status,
nDraws,
selectedModel,
onSelectModel,
compareAll,
onCompareAllChange,
pooled,
}) {
const activeModels = useMemo(
() => (compareAll ? MODEL_QUADRANTS.map((q) => q.model) : [selectedModel]),
[compareAll, selectedModel],
)
const selected = useMemo(
() => groups.filter((g) => selectedKeys.includes(g.key)),
[groups, selectedKeys],
)
// Agresti–Coull is computed from observed counts, so it is identical across
// model specifications — the same rail is drawn on every model row, which is
// exactly what makes the rows comparable.
const acByKey = useMemo(() => {
const out = {}
for (const g of selected) {
if (!(g.enroll > 0)) continue
const ac = agrestiCoull(g.observed, g.enroll, AGRESTI_COULL_LEVEL)
out[g.key] = {
// The R original can return a negative lower bound for a single event;
// a negative arrest count is not drawable, so clamp here rather than in
// the port (see utils/agrestiCoull.js).
lower: Math.max(0, (ac.lower / g.enroll) * 1000),
upper: (ac.upper / g.enroll) * 1000,
point: (g.observed / g.enroll) * 1000,
}
}
return out
}, [selected])
// Draws for every active model, moved into display (pooled or unpooled) space.
const drawsByModel = useMemo(() => {
const out = {}
for (const model of activeModels) {
out[model] = displayDraws(byModel?.[model]?.counts, enrollByGroup, pooled)
}
return out
}, [activeModels, byModel, enrollByGroup, pooled])
const domain = useMemo(() => {
const rateArrays = []
for (const model of activeModels) {
const { counts, enroll } = drawsByModel[model] || {}
for (const g of selected) rateArrays.push(toRates(counts?.[g.key], enroll?.[g.key]))
}
return computeRateDomain(rateArrays, Object.values(acByKey).map((a) => a.upper))
}, [activeModels, drawsByModel, selected, acByKey])
// One entry per (model, group): the shape to draw and where it came from.
const rowsByModel = useMemo(() => {
const out = {}
for (const model of activeModels) {
const { counts, enroll } = drawsByModel[model] || {}
out[model] = selected.map((g) => ({
group: g,
...profileFor(g, counts?.[g.key], enroll?.[g.key], domain.niceMax),
}))
}
return out
}, [activeModels, drawsByModel, selected, domain.niceMax])
const anyApproximated = Object.values(rowsByModel)
.flat()
.some((r) => r.source !== 'draws')
const loading = status === 'loading'
const sexPanels = pooled
? [{ sex: null, label: null }]
: [{ sex: 'F', label: 'Female students' }, { sex: 'M', label: 'Male students' }]
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<div style={headerStyle}>
<div>
<h3 style={cardTitle}>Arrest rate probability density</h3>
<p style={subtitleStyle}>
Each curve is one student group&rsquo;s modelled arrest rate. The bar beneath each panel
is what was actually reported, with its {Math.round(AGRESTI_COULL_LEVEL * 100)}%
Agresti&ndash;Coull interval.
</p>
</div>
<SpecControls
selectedModel={selectedModel}
onSelectModel={onSelectModel}
compareAll={compareAll}
onCompareAllChange={onCompareAllChange}
loading={loading}
/>
</div>
{/* Only once the fetch has settled. While it is in flight every group is
nominally "approximated" (there are no draws yet), so showing the note
then would claim a fallback that hasn't happened — next to a spinner
saying the real draws are still coming. */}
{!loading && anyApproximated && <ApproxNote />}
{selected.length === 0 ? (
<p style={emptyStyle}>
No student groups are selected. Check a group in the table above to show its distribution.
</p>
) : loading ? (
<p style={emptyStyle}>
{compareAll
? 'Loading posterior draws for all four specifications…'
: 'Loading posterior draws…'}
</p>
) : (
sexPanels.map((panel) => (
<SexPanel
key={panel.sex || 'pooled'}
label={panel.label}
sex={panel.sex}
domain={domain}
rowsByModel={rowsByModel}
activeModels={activeModels}
acByKey={acByKey}
compareAll={compareAll}
/>
))
)}
{domain.clipped && (
<p style={noteStyle}>
The axis stops at {domain.niceMax} per 1,000; one or more groups extend beyond it. Those
curves and intervals are cut off at the right edge, marked &rsaquo;.
</p>
)}
<p style={captionStyle}>
Densities are drawn from {nDraws > 0 ? nDraws.toLocaleString() : '500'} posterior predictive
draws per group. Point ranges are the observed rate with its{' '}
{Math.round(AGRESTI_COULL_LEVEL * 100)}% Agresti&ndash;Coull interval. Source: CRDC School
Arrest Rate API (Knowles &amp; Miller 2025), 2021&ndash;22 Civil Rights Data Collection.
</p>
</div>
)
}
/**
* Picks the shape for one group: real draws when we have them, the fitted
* summary-interval approximation when we don't, and nothing at all for a pooled
* group with no draws — there is no honest pooled shape to fit (see
* `buildDisplayGroups`).
*/
function profileFor(group, counts, enroll, niceMax) {
const domain = { min: 0, max: niceMax, n: 60 }
const fromDraws = densityProfile(counts, enroll, domain)
if (fromDraws) return { profile: fromDraws, source: 'draws' }
if (!group.modeled || !(group.modeled.upper > 0)) return { profile: null, source: 'none' }
const fit = fitSkewedInterval({ ...group.modeled, intervalMass: 0.95 })
const points = densityCurve(fit, domain)
return {
profile: { kind: 'kde', points, maxY: Math.max(...points.map((p) => p.y)), step: 0 },
source: 'approx',
}
}
/** x of a profile's tallest point, clamped into the plotted domain. */
function peakX(profile, x, niceMax) {
const peak = profile.points.reduce((best, p) => (p.y > best.y ? p : best), profile.points[0])
return x(Math.min(peak.x, niceMax))
}
function SexPanel({ label, sex, domain, rowsByModel, activeModels, acByKey, compareAll }) {
const rows = activeModels.map((model) => ({
model,
label: MODEL_QUADRANTS.find((q) => q.model === model)?.label || model,
entries: (rowsByModel[model] || []).filter((r) => (sex ? r.group.sex === sex : true)),
}))
const hasAnything = rows.some((r) => r.entries.some((e) => e.profile))
const rowHeight = compareAll ? COMPARE_ROW_HEIGHT : ROW_HEIGHT
const innerWidth = CHART_WIDTH - MARGIN.left - MARGIN.right
const fontPx = compareAll ? COMPACT_LABEL_FONT_PX : LABEL_FONT_PX
const lineHeight = compareAll ? COMPACT_LABEL_LINE_HEIGHT : LABEL_LINE_HEIGHT
const x = scaleLinear().domain([0, domain.niceMax]).range([MARGIN.left, MARGIN.left + innerWidth])
// Lay the labels out before sizing the SVG: how many lanes they need decides
// how much room each row has to reserve above its curves.
let cursor = MARGIN.top
const placedRows = rows.map((row) => {
const drawable = row.entries.filter((e) => e.profile)
const { labels, lanes } = layoutPeakLabels(
drawable.map((e) => ({
key: e.group.key,
x: peakX(e.profile, x, domain.niceMax),
text: e.group.shortLabel,
color: raceColor(e.group.race),
})),
{ min: MARGIN.left, max: MARGIN.left + innerWidth, fontPx, gap: LABEL_GAP },
)
const specLabel = compareAll ? SPEC_LABEL_HEIGHT : 0
const band = specLabel + (lanes > 0 ? lanes * lineHeight + 3 : 0)
const top = cursor
cursor += band + rowHeight + RAIL_HEIGHT
return { row, drawable, labels, band, specLabel, top, lineHeight, fontPx }
})
const plotBottom = cursor
const height = plotBottom + MARGIN.bottom
return (
<figure style={{ margin: '0 0 var(--space-2)' }}>
{label && <figcaption style={panelLabelStyle}>{label}</figcaption>}
{!hasAnything ? (
<p style={emptyStyle}>No modelled distribution is available for these groups.</p>
) : (
<div style={{ overflowX: 'auto' }}>
<svg
width="100%"
viewBox={`0 0 ${CHART_WIDTH} ${height}`}
style={{ maxWidth: '100%', minWidth: '320px' }}
role="img"
aria-label={`Modelled arrest rate distributions${label ? ` for ${label.toLowerCase()}` : ''}`}
>
{domain.ticks.map((t) => (
<line
key={t}
x1={x(t)}
y1={MARGIN.top}
x2={x(t)}
y2={plotBottom}
stroke="var(--cv-rule)"
strokeWidth={1}
/>
))}
{placedRows.map((placed) => (
<ModelRow
key={placed.row.model}
placed={placed}
compareAll={compareAll}
x={x}
rowHeight={rowHeight}
acByKey={acByKey}
niceMax={domain.niceMax}
/>
))}
{domain.ticks.map((t) => (
<text
key={t}
x={x(t)}
y={plotBottom + 15}
textAnchor="middle"
fontSize="0.62rem"
fill="var(--cv-ink-3)"
>
{t}
</text>
))}
<text
x={MARGIN.left + innerWidth / 2}
y={height - 4}
textAnchor="middle"
fontSize="0.64rem"
fill="var(--cv-ink-3)"
>
Arrests per 1,000 students
</text>
</svg>
</div>
)}
</figure>
)
}
function ModelRow({ placed, compareAll, x, rowHeight, acByKey, niceMax }) {
const { row, drawable, labels, band, specLabel, top, lineHeight, fontPx } = placed
const baselineY = top + band + rowHeight
const peakHeight = rowHeight * 0.86
return (
<g>
{compareAll && (
<text x={MARGIN.left + 2} y={top + 10} fontSize="0.62rem" fontWeight={700} fill="var(--cv-ink-2)">
{row.label}
</text>
)}
{/* Direct labels, packed into lanes so overlapping densities stay legible.
Colour is what ties each one to its curve — the same hue the table's
swatch uses — so no leader lines are needed. */}
{labels.map((label) => (
<text
key={label.key}
x={label.x}
y={top + specLabel + label.lane * lineHeight + fontPx}
textAnchor={label.anchor}
fontSize={`${fontPx}px`}
fontWeight={700}
fill={label.color}
stroke="var(--cv-paper)"
strokeWidth={3}
strokeLinejoin="round"
paintOrder="stroke"
>
{label.text}
</text>
))}
{drawable.map((entry) => (
<GroupArea
key={entry.group.key}
entry={entry}
x={x}
baselineY={baselineY}
// Each profile is normalized to its own peak. The two profile kinds
// carry different y units (probability mass vs. density), so a shared
// maximum would squash whichever kind happened to peak lower — and
// what the reader is comparing here is location and spread, not peak
// height.
peakHeight={peakHeight}
niceMax={niceMax}
/>
))}
<line
x1={MARGIN.left}
y1={baselineY}
x2={x(niceMax)}
y2={baselineY}
stroke="var(--cv-rule-strong)"
strokeWidth={1}
/>
{row.entries.map((entry, i) => {
const ac = acByKey[entry.group.key]
if (!ac) return null
return (
<PointRange
key={entry.group.key}
ac={ac}
color={raceColor(entry.group.race)}
x={x}
y={baselineY + 8 + (i % 3) * 5}
niceMax={niceMax}
/>
)
})}
</g>
)
}
function GroupArea({ entry, x, baselineY, peakHeight, niceMax }) {
const { profile, group } = entry
const color = raceColor(group.race)
const y = scaleLinear().domain([0, profile.maxY || 1]).range([baselineY, baselineY - peakHeight])
const points =
profile.kind === 'mass'
? padMassPoints(profile.points, profile.step, niceMax)
: profile.points
const areaGen = area()
.x((p) => x(Math.min(p.x, niceMax)))
.y0(baselineY)
.y1((p) => y(p.y))
// A staircase for discrete mass, a smooth curve for a KDE. curveStep keeps
// the mass profile honest — it says "these values and no others" — while
// still visually rhyming with the filled areas beside it.
.curve(profile.kind === 'mass' ? curveStep : curveMonotoneX)
return (
<g>
<path d={areaGen(points)} fill={color} opacity={FILL_OPACITY} />
<path
d={areaGen.lineY1()(points)}
fill="none"
stroke={color}
strokeWidth={STROKE_WIDTH}
strokeLinejoin="round"
/>
</g>
)
}
/**
* A staircase needs a floor to close against on both sides: without the zero
* pads, curveStep leaves the first and last bars open and the fill bleeds to
* the panel edge.
*/
function padMassPoints(points, step, niceMax) {
const half = Math.max(step, 1e-6) / 2
const first = points[0]
const last = points[points.length - 1]
return [
{ x: Math.max(0, first.x - half), y: 0 },
...points,
{ x: Math.min(last.x + half, niceMax), y: 0 },
]
}
function PointRange({ ac, color, x, y, niceMax }) {
const lo = Math.min(ac.lower, niceMax)
const hi = Math.min(ac.upper, niceMax)
const point = Math.min(ac.point, niceMax)
const clipped = ac.upper > niceMax
return (
<g>
<line x1={x(lo)} y1={y} x2={x(hi)} y2={y} stroke={color} strokeWidth={2} strokeLinecap="round" />
<circle cx={x(point)} cy={y} r={3} fill={color} stroke="var(--cv-paper)" strokeWidth={1} />
{clipped && (
<text x={x(niceMax) + 3} y={y + 3} fontSize="0.6rem" fill={color}>
&rsaquo;
</text>
)}
</g>
)
}
function SpecControls({ selectedModel, onSelectModel, compareAll, onCompareAllChange, loading }) {
return (
<div style={{ display: 'flex', flexDirection: 'column', gap: '0.4rem', alignItems: 'flex-end' }}>
<div role="group" aria-label="Model specification" style={segmentedStyle}>
{MODEL_QUADRANTS.map((q) => {
const active = !compareAll && q.model === selectedModel
return (
<button
key={q.model}
type="button"
onClick={() => onSelectModel(q.model)}
aria-pressed={active}
disabled={compareAll}
style={{
...segmentStyle,
background: active ? 'var(--cv-navy-600)' : 'transparent',
color: active ? '#fff' : 'var(--cv-ink-2)',
opacity: compareAll ? 0.5 : 1,
cursor: compareAll ? 'not-allowed' : 'pointer',
}}
>
{q.label}
</button>
)
})}
</div>
<label style={switchLabel}>
<input
type="checkbox"
checked={compareAll}
onChange={(e) => onCompareAllChange(e.target.checked)}
/>
<span>Compare all four specifications</span>
{loading && compareAll && <Spinner />}
</label>
</div>
)
}
function Spinner() {
return (
<span
aria-label="Loading"
style={{
display: 'inline-block',
width: '11px',
height: '11px',
border: '2px solid var(--cv-rule-strong)',
borderTopColor: 'var(--cv-accent)',
borderRadius: '50%',
animation: 'cv-spin 700ms linear infinite',
}}
/>
)
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const headerStyle = {
display: 'flex',
justifyContent: 'space-between',
alignItems: 'flex-start',
flexWrap: 'wrap',
gap: 'var(--space-2)',
}
const subtitleStyle = { fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: 0, maxWidth: '34rem' }
const panelLabelStyle = {
fontSize: '0.75rem',
fontWeight: 700,
textTransform: 'uppercase',
letterSpacing: '0.06em',
color: 'var(--cv-ink-3)',
marginBottom: '0.1rem',
}
const emptyStyle = { color: 'var(--cv-ink-3)', fontSize: '0.85rem', padding: 'var(--space-2) 0' }
const noteStyle = { fontSize: '0.76rem', fontStyle: 'italic', color: 'var(--cv-ink-3)', margin: '0 0 0.4rem' }
const captionStyle = { fontSize: '0.74rem', color: 'var(--cv-ink-3)', margin: '0.4rem 0 0', lineHeight: 1.5 }
const segmentedStyle = {
display: 'inline-flex',
border: '1px solid var(--cv-rule)',
borderRadius: 'var(--radius-md)',
overflow: 'hidden',
}
const segmentStyle = {
padding: '0.25rem 0.55rem',
border: 'none',
borderRight: '1px solid var(--cv-rule)',
fontFamily: 'var(--font-sans)',
fontSize: '0.72rem',
fontWeight: 600,
}
const switchLabel = {
display: 'inline-flex',
alignItems: 'center',
gap: '0.4rem',
fontSize: '0.76rem',
color: 'var(--cv-ink-2)',
cursor: 'pointer',
}
+10
View File
@@ -0,0 +1,10 @@
import { DISTRIBUTION_APPROX_NOTE } from '../utils/distributionApprox.js'
/** Consistent caption for any chart that renders an approximated distribution shape. */
export default function ApproxNote() {
return (
<p style={{ fontSize: '0.78rem', fontStyle: 'italic', color: 'var(--cv-ink-3)', margin: '0.25rem 0 0' }}>
{DISTRIBUTION_APPROX_NOTE}
</p>
)
}
+43
View File
@@ -0,0 +1,43 @@
/**
* Plain-HTML chart legend — replaces every chart's hand-rolled SVG legend.
* Both of the app's stray-`}` CSS bugs lived inside hand-written SVG legend
* `fill="..."` strings; keeping legends out of SVG removes that bug class.
*
* items: [{ color, label, shape: 'swatch' | 'diamond' | 'line' | 'dot' }]
*/
export default function ChartLegend({ items }) {
return (
<div style={{ display: 'flex', flexWrap: 'wrap', gap: '1.25rem', marginTop: 'var(--space-2)' }}>
{items.map((item) => (
<div key={item.label} style={{ display: 'flex', alignItems: 'center', gap: '0.4rem' }}>
<LegendMark shape={item.shape} color={item.color} />
<span style={{ fontSize: '0.8rem', color: 'var(--cv-ink-2)' }}>{item.label}</span>
</div>
))}
</div>
)
}
function LegendMark({ shape, color }) {
if (shape === 'diamond') {
return (
<span
style={{
display: 'inline-block',
width: '9px',
height: '9px',
background: color,
transform: 'rotate(45deg)',
flexShrink: 0,
}}
/>
)
}
if (shape === 'line') {
return <span style={{ display: 'inline-block', width: '16px', height: '3px', background: color, borderRadius: '2px', flexShrink: 0 }} />
}
if (shape === 'dot') {
return <span style={{ display: 'inline-block', width: '10px', height: '10px', background: color, borderRadius: '50%', flexShrink: 0 }} />
}
return <span style={{ display: 'inline-block', width: '11px', height: '11px', background: color, borderRadius: '3px', flexShrink: 0 }} />
}
+298
View File
@@ -0,0 +1,298 @@
import { useCallback, useEffect, useMemo, useState } from 'react'
import * as api from '../hooks/useApi.js'
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
import { useDistrictTotalDraws, useDrawDistribution } from '../hooks/useDrawDistribution.js'
import {
buildDisplayGroups,
defaultDiffPair,
defaultSelectedKeys,
enrollByGroupKey,
} from '../utils/districtGroups.js'
import { shouldPoolBySex } from '../utils/pooling.js'
import ArrestsOverTime from '../charts/ArrestsOverTime.jsx'
import GroupDifference from '../charts/GroupDifference.jsx'
import RateDensityPanel from '../charts/RateDensityPanel.jsx'
import DistrictSummaryTable from './DistrictSummaryTable.jsx'
const ALL_WAVES = ['15-16', '17-18', '21-22']
const WAVE_MODEL = 'unified_m3_mod' // three-year, no covariate — powers the time series
const DEFAULT_SPEC = 'unified_m4_mod' // three-year + referral rate
const CURRENT_WAVE = '21-22'
const QUADRANT_MODELS = MODEL_QUADRANTS.map((q) => q.model)
// Below this, the draws for all three waves are fetched without asking. Sized
// from the real shards: Nevada's three waves are 0.31MB, California's 13.9MB
// and Texas's 14.3MB, so 3MB cleanly separates "free" from "worth a click".
const AUTO_TOTAL_DRAW_BYTES = 3 * 1024 * 1024
/**
* The results page: what was observed (summary table), what the model says the
* rate is (density panel), and what it says about the gap between two groups
* (difference chart) — plus the district's arrest history.
*
* This component owns every piece of cross-chart state: the selected model
* specification, whether sexes are pooled, which groups are checked, and which
* pair is being differenced. The charts are given values and callbacks and hold
* none of it, so the table's checkboxes and the density panel can't drift apart.
*/
export default function ChartPanel({ district, state }) {
const [waveData, setWaveData] = useState(null)
const [summaryByModel, setSummaryByModel] = useState({})
const [loading, setLoading] = useState(true)
const [selectedModel, setSelectedModel] = useState(DEFAULT_SPEC)
const [compareAll, setCompareAll] = useState(false)
const [poolOverride, setPoolOverride] = useState(null) // null = follow the rule
// Selections are stored with the pooling mode they were made in: pooled keys
// ('BL') and unpooled keys ('BL_F') are different namespaces, so a selection
// made in one mode must not be reapplied in the other.
const [selection, setSelection] = useState(null)
const [pairSelection, setPairSelection] = useState(null)
const [exactTotals, setExactTotals] = useState(false)
// ——— Fetch: the three waves for the time series ———
useEffect(() => {
let cancelled = false
setLoading(true)
setWaveData(null)
setSummaryByModel({})
setPoolOverride(null)
setSelection(null)
setPairSelection(null)
setCompareAll(false)
setExactTotals(false)
async function run() {
const waves = {}
await Promise.all(
ALL_WAVES.map(async (year) => {
try {
waves[year] = await api.fetchDistrictEstimates(district.leaid, { model: WAVE_MODEL, year })
} catch (err) {
console.error(`ChartPanel: wave ${year} fetch failed:`, err)
}
}),
)
if (cancelled) return
setWaveData(waves)
setLoading(false)
}
run()
return () => {
cancelled = true
}
}, [district.leaid])
// ——— Fetch: the current wave's summary for whichever spec is selected ———
// Only the selected specification is fetched. Enrollment and observed arrest
// counts are district facts, not model outputs, and the modelled distributions
// now come from real draws rather than from these rows — so prefetching all
// four specs' summaries would be four requests for data three of which are
// never read.
useEffect(() => {
let cancelled = false
if (summaryByModel[selectedModel]) return
async function run() {
try {
const rows = await api.fetchDistrictEstimates(district.leaid, {
model: selectedModel,
year: CURRENT_WAVE,
})
if (!cancelled) setSummaryByModel((prev) => ({ ...prev, [selectedModel]: rows }))
} catch (err) {
console.error(`ChartPanel: summary fetch failed for ${selectedModel}:`, err)
if (!cancelled) setSummaryByModel((prev) => ({ ...prev, [selectedModel]: [] }))
}
}
run()
return () => {
cancelled = true
}
}, [district.leaid, selectedModel, summaryByModel])
const rows = useMemo(() => summaryByModel[selectedModel] || [], [summaryByModel, selectedModel])
const autoPooled = useMemo(() => shouldPoolBySex(rows), [rows])
const pooled = poolOverride ?? autoPooled
const groups = useMemo(() => buildDisplayGroups(rows, pooled), [rows, pooled])
const enrollByGroup = useMemo(() => enrollByGroupKey(rows), [rows])
const defaultKeys = useMemo(() => defaultSelectedKeys(groups), [groups])
const selectedKeys = selection?.pooled === pooled ? selection.keys : defaultKeys
const defaultPair = useMemo(() => defaultDiffPair(groups), [groups])
const pair = pairSelection?.pooled === pooled ? pairSelection.pair : defaultPair
const selectionIsEnrollmentFallback =
groups.length > 0 && groups.every((g) => g.observed === 0) && selection === null
const handleToggleGroup = useCallback(
(key) => {
const next = selectedKeys.includes(key)
? selectedKeys.filter((k) => k !== key)
: [...selectedKeys, key]
setSelection({ pooled, keys: next })
},
[selectedKeys, pooled],
)
const handlePooledChange = useCallback((next) => {
setPoolOverride(next)
// Both selections live in the other namespace now — drop them and let the
// defaults recompute for the new group set.
setSelection(null)
setPairSelection(null)
}, [])
const handlePairChange = useCallback((next) => setPairSelection({ pooled, pair: next }), [pooled])
const drawModels = useMemo(
() => (compareAll ? QUADRANT_MODELS : [selectedModel]),
[compareAll, selectedModel],
)
const { status, byModel, nDraws } = useDrawDistribution({
leaid: district.leaid,
state,
models: drawModels,
year: CURRENT_WAVE,
})
// The time series' band, computed from the draws rather than by summing each
// group's own bounds. Fetched automatically only where it is cheap — three
// waves is 0.3MB for Nevada but 13.9MB for California — otherwise the chart
// offers it as an explicit choice with the size shown.
const totals = useDistrictTotalDraws({
leaid: district.leaid,
state,
model: WAVE_MODEL,
years: ALL_WAVES,
enabled: exactTotals,
})
useEffect(() => {
if (totals.bytes > 0 && totals.bytes <= AUTO_TOTAL_DRAW_BYTES) setExactTotals(true)
}, [totals.bytes])
if (loading || !waveData) return <LoadingCharts />
const timeSeriesData = ALL_WAVES.map((year) => {
const yearRows = waveData[year] || []
return {
year,
label: `20${year.replace('-', '-')}`,
arrests: yearRows.reduce((sum, r) => sum + (r.observed_arrests || 0), 0),
enroll: yearRows.reduce((sum, r) => sum + (r.stu_enroll || 0), 0),
modeledMedian: yearRows.reduce((sum, r) => sum + (r.count_median || 0), 0),
modeledLower: yearRows.reduce((sum, r) => sum + (r.count_lower || 0), 0),
modeledUpper: yearRows.reduce((sum, r) => sum + (r.count_upper || 0), 0),
}
})
return (
<div style={{ padding: 'var(--space-3) 0 var(--space-7)' }}>
<div style={{ marginBottom: 'var(--space-4)' }}>
<span className="eyebrow">District estimates — Bayesian model comparison</span>
<h2 style={{ marginTop: 'var(--space-1)', marginBottom: 0 }}>
{district.lea_name} ({state}) — School-based arrest rates, 2021–22 CRDC
</h2>
</div>
<div
style={{
display: 'flex',
flexDirection: 'column',
gap: 'var(--space-5)',
maxWidth: '70rem',
marginLeft: 'auto',
marginRight: 'auto',
}}
>
<DistrictSummaryTable
groups={groups}
selectedKeys={selectedKeys}
onToggleGroup={handleToggleGroup}
pooled={pooled}
autoPooled={autoPooled}
onPooledChange={handlePooledChange}
selectionIsEnrollmentFallback={selectionIsEnrollmentFallback}
/>
<RateDensityPanel
groups={groups}
selectedKeys={selectedKeys}
enrollByGroup={enrollByGroup}
byModel={byModel}
status={status}
nDraws={nDraws}
selectedModel={selectedModel}
onSelectModel={setSelectedModel}
compareAll={compareAll}
onCompareAllChange={setCompareAll}
pooled={pooled}
/>
<GroupDifference
groups={groups}
enrollByGroup={enrollByGroup}
byModel={byModel}
selectedModel={selectedModel}
status={status}
pooled={pooled}
pair={pair}
onPairChange={handlePairChange}
/>
<ArrestsOverTime
data={timeSeriesData}
modelId={WAVE_MODEL}
totals={totals}
onRequestExactTotals={() => setExactTotals(true)}
/>
</div>
<div
style={{
marginTop: 'var(--space-6)',
padding: 'var(--space-3) 0',
borderTop: '1px solid var(--cv-rule)',
}}
>
<span className="eyebrow">Methodology</span>
<p style={{ fontSize: '0.85rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-1)' }}>
Estimates are from the CRDC School Arrest Rate API (Knowles &amp; Miller 2025). Data shown
spans three waves of the Civil Rights Data Collection (2015–16, 2017–18, 2021–22) and lets
you explore four Bayesian model specifications: one-year vs. three-year models with and
without referral-rate covariates. All rates are per 1,000 students. Modelled distributions
are drawn from the published posterior predictive draws, fetched in the browser.
</p>
<p style={{ fontSize: '0.85rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-2)' }}>
<strong>Student groups not shown.</strong> The CRDC also reports arrests for Asian,
Hawaiian/Pacific Islander, and multiracial students. The underlying research does not
model these groups due to extreme sparsity in the outcome of interest and computational
limits. Fitting every reported group would also have meant fitting 70
stratified models rather than 40. Their absence here reflects that scope, not a finding
about those students, and no estimate for them should be inferred from this page. See the{' '}
<a href="https://crdc-api.civilytics.org/api/v1/" target="_blank" rel="noopener noreferrer">
white paper
</a>{' '}
for the full rationale. Arrests that cannot be disaggregated by both race and sex —
including Section 504 students — are excluded upstream, as are the small number of arrests in
schools not enrolling grade 7 or above.
</p>
</div>
</div>
)
}
function LoadingCharts() {
return (
<div style={{ padding: 'var(--space-5) 0 var(--space-7)', textAlign: 'center' }}>
<span className="eyebrow">Preparing charts</span>
<p style={{ color: 'var(--cv-ink-3)' }}>Organizing data across all model specifications…</p>
</div>
)
}
+196
View File
@@ -0,0 +1,196 @@
import { useState, useEffect } from 'react'
import { searchDistricts } from '../hooks/useApi.js'
const TOP_DISTRICTS_URL = `${import.meta.env.BASE_URL}data/top_districts.json`
const SUGGESTION_COUNT = 8
// Module-level cache: the fixture covers every state, so it is fetched at most
// once per page load no matter how many states the visitor browses through.
let topDistrictsPromise = null
function loadTopDistricts() {
if (!topDistrictsPromise) {
topDistrictsPromise = fetch(TOP_DISTRICTS_URL)
.then((res) => {
if (!res.ok) throw new Error(`HTTP ${res.status} loading ${TOP_DISTRICTS_URL}`)
return res.json()
})
.catch((err) => {
// Don't poison the cache with a rejected promise — let a later visit retry.
topDistrictsPromise = null
throw err
})
}
return topDistrictsPromise
}
/**
* District search screen with:
* 1. "Interesting" suggestions — the districts with the most arrests in the
* selected state, read from the committed `public/data/top_districts.json`
* fixture (built by `scripts/build-top-districts.mjs`).
* 2. Live-search as you type → /api/v1/districts?q=...&state=XX
*
* The suggestions used to be computed at runtime from
* `/estimates?state=XX&year=21-22&limit=500`. That endpoint returns rows
* `ORDER BY LEAID, RACE, SEX` at eight rows per district, so the cap selected
* the ~62 lowest-LEAID districts in the state rather than the busiest ones —
* California has 11,488 rows. The fixture is ranked over the complete result
* set, and it also removes a multi-second fetch from this screen.
*/
export default function DistrictSearch({ state, onSelect, onBack }) {
const [query, setQuery] = useState('')
const [searchResults, setSearchResults] = useState([])
const [suggestions, setSuggestions] = useState([]) // "interesting" districts sorted by arrests desc
const [loadingSugg, setLoadingSugg] = useState(true)
const [loadingSearch, setLoadingSearch] = useState(false)
// ——— Read the suggestion fixture for this state ———
useEffect(() => {
let cancelled = false
setLoadingSugg(true)
loadTopDistricts()
.then((fixture) => {
if (cancelled) return
const forState = fixture?.states?.[state] || []
setSuggestions(
forState.slice(0, SUGGESTION_COUNT).map((d) => ({
leaid: d.leaid,
lea_name: d.name,
state,
observed_arrests: d.arrests,
})),
)
})
.catch((err) => {
console.error('Failed to load suggested districts:', err)
if (!cancelled) setSuggestions([]) // degrade gracefully — just show search box
})
.finally(() => {
if (!cancelled) setLoadingSugg(false)
})
return () => { cancelled = true }
}, [state])
// ——— Live search as user types (debounced by React's natural event batching) ———
useEffect(() => {
if (!query.trim()) { setSearchResults([]); return }
const timer = setTimeout(async () => {
setLoadingSearch(true)
try {
const data = await searchDistricts(query, state)
setSearchResults(data || [])
} catch (err) {
console.error('District search failed:', err)
setSearchResults([])
} finally {
setLoadingSearch(false)
}
}, 300)
return () => clearTimeout(timer)
}, [query, state])
// ——— Render a district card (used for both suggestions and search results) ———
const renderDistrictCard = (dist) => (
<button
key={dist.leaid}
className="cv-card"
style={{ textAlign: 'left', padding: 'var(--space-2)' }}
onClick={() => onSelect(dist)}
>
<div style={{ display: 'flex', justifyContent: 'space-between', alignItems: 'baseline' }}>
<span style={{ fontWeight: 600, color: 'var(--cv-navy-700)', fontSize: '0.95rem' }}>
{dist.lea_name || dist.name}
</span>
{dist.observed_arrests !== undefined && (
<span className="stat-label accent" style={{ marginTop: 0 }}>
{dist.observed_arrests.toLocaleString()} arrests in 2021–22
</span>
)}
</div>
</button>
)
return (
<div style={{ padding: 'var(--space-5) 0 var(--space-7)' }}>
{/* Back button */}
<button className="btn-outline" onClick={onBack} style={{ marginBottom: 'var(--space-3)' }}>
← Change state
</button>
<h2 style={{ marginTop: 0, marginBottom: 'var(--space-1)' }}>
{state === 'DC' ? "District of Columbia" : `Find a district in ${getStateName(state)}`}
</h2>
<p style={{ color: 'var(--cv-ink-2)', marginBottom: 'var(--space-3)' }}>
Search by district name, or try one of the suggested districts with the highest arrest counts.
</p>
{/* "Interesting" suggestions */}
{!query && (
<div style={{ marginTop: 'var(--space-4)' }}>
<span className="eyebrow">Suggested districts — most arrests in 2021–22</span>
{loadingSugg ? (
<p style={{ color: 'var(--cv-ink-3)', padding: 'var(--space-2) 0' }}>Loading suggestions…</p>
) : suggestions.length > 0 ? (
<div style={{ display: 'grid', gridTemplateColumns: 'repeat(auto-fill, minmax(280px, 1fr))', gap: 'var(--space-1)' }}>
{suggestions.map(renderDistrictCard)}
</div>
) : (
<p style={{ color: 'var(--cv-ink-3)', padding: 'var(--space-2) 0' }}>
No districts with reported arrests found in this state. Try searching by name below.
</p>
)}
</div>
)}
{/* Live search */}
<div style={{ marginTop: query ? 0 : 'var(--space-5)' }}>
<label htmlFor="district-search" className="stat-label">Search district name:</label>
<input
id="district-search"
type="text"
placeholder="Type a district name… (e.g., Madison, Derby, Paterson)"
value={query}
onChange={(e) => setQuery(e.target.value)}
style={{
width: '100%', padding: 'var(--space-2)', fontSize: '1rem',
border: '1px solid var(--cv-rule)', borderRadius: 'var(--radius-md)',
fontFamily: 'var(--font-sans)'
}}
/>
{query && (
<div style={{ marginTop: 'var(--space-2)' }}>
{loadingSearch ? (
<p style={{ color: 'var(--cv-ink-3)' }}>Searching…</p>
) : searchResults.length > 0 ? (
<div style={{ display: 'grid', gap: 'var(--space-1)' }}>
{searchResults.slice(0, 20).map(renderDistrictCard)}
</div>
) : (
<p style={{ color: 'var(--cv-ink-3)' }}>No districts found matching "{query}".</p>
)}
{/* Back to suggestions */}
{searchResults.length > 0 && (
<button className="btn-outline" onClick={() => setQuery('')} style={{ marginTop: 'var(--space-2)' }}>
← Show suggested districts again
</button>
)}
</div>
)}
</div>
</div>
)
}
function getStateName(code) {
const names = {
AL:'Alabama',AK:'Alaska',AZ:'Arizona',AR:'Arkansas',CA:'California',CO:'Colorado',CT:'Connecticut',DE:'Delaware',DC:'District of Columbia',FL:'Florida',GA:'Georgia',HI:'Hawaii',ID:'Idaho',IL:'Illinois',IN:'Indiana',IA:'Iowa',KS:'Kansas',KY:'Kentucky',LA:'Louisiana',ME:'Maine',MD:'Maryland',MA:'Massachusetts',MI:'Michigan',MN:'Minnesota',MS:'Mississippi',MO:'Missouri',MT:'Montana',NE:'Nebraska',NV:'Nevada',NH:'New Hampshire',NJ:'New Jersey',NM:'New Mexico',NY:'New York',NC:'North Carolina',ND:'North Dakota',OH:'Ohio',OK:'Oklahoma',OR:'Oregon',PA:'Pennsylvania',RI:'Rhode Island',SC:'South Carolina',SD:'South Dakota',TN:'Tennessee',TX:'Texas',UT:'Utah',VT:'Vermont',VA:'Virginia',WA:'Washington',WV:'West Virginia',WI:'Wisconsin',WY:'Wyoming'
}
return names[code] || code
}
+234
View File
@@ -0,0 +1,234 @@
import { raceColor } from '../utils/colors.js'
import { POOL_BY_SEX_ARREST_THRESHOLD } from '../utils/pooling.js'
/**
* The results page's first block: what was actually observed, before any
* modelling. One row per student group plus a district total.
*
* The table doubles as the legend and the control for the density panel — each
* row carries the checkbox that adds or removes that group's curve, and the
* colour swatch is the same hue the curve is drawn in. That's deliberate:
* a separate legend plus a separate group picker would make the reader hold
* three mappings in their head instead of one.
*
* @param {{
* groups: Array<{key:string, race:string, sex:string|null, label:string,
* enroll:number, observed:number, rate:number}>,
* selectedKeys: string[],
* onToggleGroup: (key: string) => void,
* pooled: boolean,
* autoPooled: boolean,
* onPooledChange: (pooled: boolean) => void,
* selectionIsEnrollmentFallback: boolean,
* }} props
*/
export default function DistrictSummaryTable({
groups,
selectedKeys,
onToggleGroup,
pooled,
autoPooled,
onPooledChange,
selectionIsEnrollmentFallback,
}) {
if (!groups?.length) {
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={cardTitle}>Reported arrests and enrollment</h3>
<p style={{ color: 'var(--cv-ink-3)', margin: 0 }}>
No student-group data is available for this district in 2021–22.
</p>
</div>
)
}
const selected = new Set(selectedKeys)
const totalEnroll = groups.reduce((sum, g) => sum + g.enroll, 0)
const totalObserved = groups.reduce((sum, g) => sum + g.observed, 0)
const totalRate = totalEnroll > 0 ? (totalObserved / totalEnroll) * 1000 : 0
return (
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
<h3 style={cardTitle}>Reported arrests and enrollment, 2021–22</h3>
<p style={{ fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: '0 0 var(--space-2)' }}>
Check a group to show its modelled arrest-rate distribution in the chart below.
</p>
{pooled && (
<PoolingBanner autoPooled={autoPooled} totalObserved={totalObserved} onPooledChange={onPooledChange} />
)}
{selectionIsEnrollmentFallback && (
<p style={noteStyle}>
No group in this district reported an arrest in 2021–22, so the two largest groups by
enrollment are shown by default.
</p>
)}
<div style={{ overflowX: 'auto' }}>
<table style={tableStyle}>
<caption style={{ captionSide: 'bottom', textAlign: 'left', paddingTop: 'var(--space-1)', fontSize: '0.75rem', color: 'var(--cv-ink-3)' }}>
Students counted are those in the four modelled race groups (American Indian / Alaska
Native, Black, Hispanic, White) — not the district&rsquo;s total enrollment, which also
includes groups this model does not estimate.
</caption>
<thead>
<tr>
<th scope="col" style={{ ...thStyle, textAlign: 'left' }}>Student group</th>
<th scope="col" style={thStyle}>Students</th>
<th scope="col" style={thStyle}>Observed arrests</th>
<th scope="col" style={thStyle}>Rate per 1,000</th>
</tr>
</thead>
<tbody>
{groups.map((g) => (
<GroupRow
key={g.key}
group={g}
checked={selected.has(g.key)}
onToggle={() => onToggleGroup(g.key)}
/>
))}
</tbody>
<tfoot>
<tr>
<th scope="row" style={{ ...tdStyle, textAlign: 'left', fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
All modelled groups
</th>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{totalEnroll.toLocaleString()}
</td>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{totalObserved.toLocaleString()}
</td>
<td style={{ ...numStyle, fontWeight: 700, borderTop: '2px solid var(--cv-ink)' }}>
{formatRate(totalRate)}
</td>
</tr>
</tfoot>
</table>
</div>
{!pooled && (
<label style={{ ...toggleLabel, marginTop: 'var(--space-2)' }}>
<input type="checkbox" checked={false} onChange={() => onPooledChange(true)} />
<span>Combine Female and Male within each race</span>
</label>
)}
</div>
)
}
function GroupRow({ group, checked, onToggle }) {
const zero = group.observed === 0
const inputId = `group-toggle-${group.key}`
return (
<tr style={{ opacity: zero ? 0.62 : 1 }}>
<td style={{ ...tdStyle, textAlign: 'left' }}>
<label htmlFor={inputId} style={{ display: 'flex', alignItems: 'center', gap: '0.5rem', cursor: 'pointer' }}>
<input id={inputId} type="checkbox" checked={checked} onChange={onToggle} />
<span
aria-hidden="true"
style={{
width: '11px',
height: '11px',
borderRadius: '3px',
flexShrink: 0,
// Hollow when deselected: the swatch tracks whether that curve is
// on screen, so the table stays a truthful legend.
background: checked ? raceColor(group.race) : 'transparent',
border: `2px solid ${raceColor(group.race)}`,
}}
/>
<span>{group.label}</span>
</label>
</td>
<td style={numStyle}>{group.enroll.toLocaleString()}</td>
<td style={numStyle}>{group.observed.toLocaleString()}</td>
<td style={numStyle}>{formatRate(group.rate)}</td>
</tr>
)
}
function PoolingBanner({ autoPooled, totalObserved, onPooledChange }) {
return (
<div style={bannerStyle}>
<p style={{ margin: 0, fontSize: '0.82rem' }}>
{autoPooled ? (
<>
This district reported {totalObserved.toLocaleString()} arrests in total — fewer than{' '}
{POOL_BY_SEX_ARREST_THRESHOLD} — so Female and Male students are combined within each
race to give each estimate more data to stand on.
</>
) : (
<>Female and Male students are combined within each race.</>
)}
</p>
<label style={toggleLabel}>
<input type="checkbox" checked onChange={() => onPooledChange(false)} />
<span>Combined — uncheck to show Female and Male separately</span>
</label>
</div>
)
}
function formatRate(rate) {
if (!(rate > 0)) return '0.0'
return rate.toLocaleString(undefined, { minimumFractionDigits: 1, maximumFractionDigits: 1 })
}
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
const tableStyle = {
width: '100%',
borderCollapse: 'collapse',
fontSize: '0.85rem',
fontVariantNumeric: 'tabular-nums',
}
const thStyle = {
textAlign: 'right',
padding: '0.35rem 0.5rem',
borderBottom: '1px solid var(--cv-rule-strong)',
fontSize: '0.72rem',
fontWeight: 600,
textTransform: 'uppercase',
letterSpacing: '0.06em',
color: 'var(--cv-ink-3)',
whiteSpace: 'nowrap',
}
const tdStyle = {
padding: '0.35rem 0.5rem',
borderBottom: '1px solid var(--cv-rule)',
}
const numStyle = { ...tdStyle, textAlign: 'right', whiteSpace: 'nowrap' }
const bannerStyle = {
background: 'var(--cv-paper-2)',
border: '1px solid var(--cv-rule)',
borderLeft: '3px solid var(--cv-accent)',
borderRadius: 'var(--radius-md)',
padding: 'var(--space-1) var(--space-2)',
marginBottom: 'var(--space-2)',
display: 'flex',
flexDirection: 'column',
gap: '0.4rem',
}
const toggleLabel = {
display: 'inline-flex',
alignItems: 'center',
gap: '0.45rem',
fontSize: '0.78rem',
color: 'var(--cv-ink-2)',
cursor: 'pointer',
}
const noteStyle = {
fontSize: '0.78rem',
fontStyle: 'italic',
color: 'var(--cv-ink-3)',
margin: '0 0 var(--space-2)',
}
+24
View File
@@ -0,0 +1,24 @@
/**
* Footer — matches the Civilytics visual style from theme/civilytics.scss.
*/
export default function Footer() {
return (
<footer className="cv-footer cv-wrap" style={{ marginTop: 'var(--space-5)' }}>
<span>© Civilytics — social science for the public good.</span>
<div style={{ display: 'flex', gap: 'var(--space-2)', flexWrap: 'wrap' }}>
<a href="https://civilytics.com">civilytics.com</a>
<a href="https://github.com/civilytics/crdc-arrests" target="_blank" rel="noopener noreferrer">Source code</a>
<a href="https://crdc-api.civilytics.org/api/v1/" target="_blank" rel="noopener noreferrer">API docs</a>
</div>
{/* Methodology footnote */}
<div style={{ width: '100%', marginTop: 'var(--space-2)', padding: 'var(--space-1) 0', borderTop: '1px solid var(--cv-rule)' }}>
<span className="stat-label">About this demo</span>
<p style={{ fontSize: '0.8rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-0.5)' }}>
This application demonstrates the CRDC School Arrest Rate API, which provides Bayesian small-area estimates of school-based arrest rates from the US Department of Education Civil Rights Data Collection. Estimates use hierarchical models that improve precision for rare events — enabling meaningful comparisons across districts, student groups, and time. Cite: Knowles & Miller 2025.
</p>
</div>
</footer>
)
}
+209
View File
@@ -0,0 +1,209 @@
import { useState, useEffect } from 'react'
import ChartPanel from './ChartPanel.jsx'
import * as api from '../hooks/useApi.js'
import { STUDENT_GROUPS, MODEL_QUADRANTS } from '../hooks/useApi.js'
import { shortGroupLabel } from '../utils/colors.js'
// CRDC waves to fetch (for Chart 1 — time series)
const WAVES = ['21-22', '17-18', '15-16']
/**
* Loading screen with animated histogram grid.
* Fetches all required data in parallel, shows progress as bars fill.
* When complete → renders ChartPanel directly (no transition needed).
*/
export default function LoadingAnimation({ district, state }) {
const [loadedCount, setLoadedCount] = useState(0)
const [totalCalls, setTotalCalls] = useState(0)
const [error, setError] = useState(null)
// Build the full list of API calls needed for all 3 charts
useEffect(() => {
let cancelled = false
async function loadData() {
try {
// ——— Chart 1: Arrests over time (3 waves × default model) ———
const wavePromises = WAVES.map((year) =>
fetchDistrictEstimatesBatch(district.leaid, year)
.then(() => !cancelled && setLoadedCount(c => c + 1))
.catch(() => { /* individual failure doesn't block */ })
)
// ——— Chart 2: Rate by group (8 groups × default model, most recent wave) ———
const groupPromises = STUDENT_GROUPS.map((sg) =>
fetchDistrictEstimatesBatch(district.leaid, '21-22', sg.race, sg.sex)
.then(() => !cancelled && setLoadedCount(c => c + 1))
.catch(() => {})
)
// ——— Chart 3: posterior density, all 4 quadrant models so the dropdown can switch (8 groups × 4 = 32 calls) ———
const modelPromises = MODEL_QUADRANTS.flatMap((quad) =>
STUDENT_GROUPS.map((sg) =>
fetchDistrictEstimatesBatch(district.leaid, '21-22', sg.race, sg.sex, quad.model)
.then(() => !cancelled && setLoadedCount(c => c + 1))
.catch(() => {})
)
)
// Set total before starting (for progress bar)
const allPromises = [...wavePromises, ...groupPromises, ...modelPromises]
if (!cancelled) setTotalCalls(allPromises.length)
await Promise.all(allPromises)
if (!cancelled) {
// All data loaded — ChartPanel renders in place of this component
// We use a render prop pattern: return <ChartPanel /> when done
setLoadedCount(allPromises.length)
}
} catch (err) {
if (!cancelled) setError(err.message || 'Failed to load data')
}
}
loadData()
return () => { cancelled = true }
}, [district.leaid])
// ——— Render the histogram grid animation or ChartPanel when done ———
const isComplete = loadedCount >= totalCalls && totalCalls > 0
if (isComplete) {
return <ChartPanel district={district} state={state} />
}
if (error) {
return renderError(error, district)
}
// Build the grid of "bars" — one per API call needed.
// Total = WAVES (3) + STUDENT_GROUPS (8) + MODEL_QUADRANTS×STUDENT_GROUPS (4×8=32)
const TOTAL_BARS = WAVES.length + STUDENT_GROUPS.length + (MODEL_QUADRANTS.length * STUDENT_GROUPS.length)
// Build bar metadata for rendering
const allBars = Array.from({ length: TOTAL_BARS }, (_, i) => {
if (i < WAVES.length) return { subLabel: 'Time series', group: 'waves' }
if (i < WAVES.length + STUDENT_GROUPS.length) {
const sg = STUDENT_GROUPS[i - WAVES.length]
return { subLabel: shortGroupLabel(sg.race, sg.sex), group: 'groups' }
}
const quadIdx = Math.floor((i - WAVES.length - STUDENT_GROUPS.length) / STUDENT_GROUPS.length)
const sgIdx = (i - WAVES.length - STUDENT_GROUPS.length) % STUDENT_GROUPS.length
const sg = STUDENT_GROUPS[sgIdx]
return { subLabel: shortGroupLabel(sg.race, sg.sex), group: 'models', modelIdx: quadIdx }
})
return (
<div style={{ padding: 'var(--space-5) 0 var(--space-7)' }}>
{/* Header */}
<span className="eyebrow">Fetching data from the CRDC Arrest Rate API</span>
<h2 style={{ marginTop: 'var(--space-2)', marginBottom: 'var(--space-3)' }}>
Building estimates for {district.lea_name} in {state === 'DC' ? "District of Columbia" : getStateName(state)}
</h2>
{/* Progress counter */}
<div style={{ display: 'flex', justifyContent: 'space-between', marginBottom: 'var(--space-3)' }}>
<span style={{ color: 'var(--cv-navy-600)', fontWeight: 600, fontSize: '1.1rem' }}>
{loadedCount} of {totalCalls || allBars.length} datasets loaded
</span>
<span className="stat-label">
{Math.round((loadedCount / (totalCalls || allBars.length)) * 100)}% complete
</span>
</div>
{/* Animated histogram grid — mirrors the Bayesian posterior draw theme */}
<div style={{
display: 'grid',
gridTemplateColumns: `repeat(${3}, 1fr)`,
gap: 'var(--space-2)',
marginTop: 'var(--space-4)'
}}>
{/* Column headers */}
{['Time series (Chart 1)', 'By student group (Chart 2)', 'Posterior density (Chart 3)'].map((h, i) => (
<div key={i} style={{ textAlign: 'center', paddingBottom: 'var(--space-1)' }}>
<span className="stat-label" style={{ display: 'block' }}>{h}</span>
</div>
))}
{/* Grid of animated bars — each represents one API call */}
{allBars.map((bar, i) => {
const filled = i < loadedCount
// Random height for "data viz" aesthetic (deterministic via seed = index)
const heightSeed = ((i * 37) % 100) + 20 // 20–120px range
return (
<div key={i} style={{ display: 'flex', flexDirection: 'column', alignItems: 'center' }}>
{/* The animated bar */}
<div style={{
width: '100%', height: `${heightSeed}px`,
background: filled
? `var(--cv-navy-${600 - (i % 2) * 100})`
: 'var(--paper-3)',
border: '1px solid var(--cv-rule)',
borderRadius: '4px 4px 0 0',
transition: 'height 300ms ease, background 300ms ease',
position: 'relative',
overflow: 'hidden'
}}>
{filled && (
<div style={{
position: 'absolute', bottom: 0, left: 0, right: 0, height: '100%',
background: `linear-gradient(to top, var(--teal-600) 0%, transparent ${70 + (i % 30)}%)`,
opacity: 0.4, transition: 'opacity 500ms ease'
}} />
)}
</div>
{/* Label below */}
<span style={{ fontSize: '0.65rem', color: filled ? 'var(--cv-ink)' : 'var(--cv-ink-4)', marginTop: '0.25rem', textAlign: 'center' }}>
{bar.subLabel}
</span>
</div>
)
})}
</div>
{/* Context message */}
<p style={{ fontSize: '0.9rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-4)' }}>
Fetching {totalCalls || allBars.length} datasets across 3 CRDC waves, 8 student groups, and 4 Bayesian model specifications...
</p>
{/* Model legend */}
<div style={{ display: 'flex', gap: 'var(--space-2)', marginTop: 'var(--space-3)', flexWrap: 'wrap' }}>
{MODEL_QUADRANTS.map((q, i) => (
<span key={i} className="stat-label" style={{ display: 'inline-flex', alignItems: 'center', gap: '0.4rem' }}>
<span style={{ width: 12, height: 12, background: `var(--cv-navy-${600 - i * 150})`, borderRadius: 2 }} />
{q.label}
</span>
))}
</div>
</div>
)
}
// ——— API helper using the centralized client from useApi.js ———
// This ensures consistent error handling, envelope unwrapping, and CORS proxy support.
async function fetchDistrictEstimatesBatch(leaid, year, race = null, sex = null, model = 'unified_m2_mod') {
return api.fetchDistrictEstimates(leaid, { year, model, race, sex })
}
function getStateName(code) {
const names = {
AL:'Alabama',AK:'Alaska',AZ:'Arizona',AR:'Arkansas',CA:'California',CO:'Colorado',CT:'Connecticut',DE:'Delaware',DC:'District of Columbia',FL:'Florida',GA:'Georgia',HI:'Hawaii',ID:'Idaho',IL:'Illinois',IN:'Indiana',IA:'Iowa',KS:'Kansas',KY:'Kentucky',LA:'Louisiana',ME:'Maine',MD:'Maryland',MA:'Massachusetts',MI:'Michigan',MN:'Minnesota',MS:'Mississippi',MO:'Missouri',MT:'Montana',NE:'Nebraska',NV:'Nevada',NH:'New Hampshire',NJ:'New Jersey',NM:'New Mexico',NY:'New York',NC:'North Carolina',ND:'North Dakota',OH:'Ohio',OK:'Oklahoma',OR:'Oregon',PA:'Pennsylvania',RI:'Rhode Island',SC:'South Carolina',SD:'South Dakota',TN:'Tennessee',TX:'Texas',UT:'Utah',VT:'Vermont',VA:'Virginia',WA:'Washington',WV:'West Virginia',WI:'Wisconsin',WY:'Wyoming'
}
return names[code] || code
}
function renderError(error, district) {
return (
<div style={{ padding: 'var(--space-5) 0 var(--space-7)', textAlign: 'center' }}>
<span className="stat-label" style={{ color: 'var(--cv-danger)' }}>ERROR</span>
<h3 style={{ marginTop: 'var(--space-2)' }}>Failed to load data for {district.lea_name}</h3>
<p style={{ color: 'var(--cv-ink-2)', maxWidth: '40rem', margin: '0 auto var(--space-3)' }}>
{error}. This may be a temporary API issue or the district may not have available estimates.
</p>
<button className="btn-outline" onClick={() => window.location.href = '/'}>
← Try another district
</button>
</div>
)
}
+137
View File
@@ -0,0 +1,137 @@
import { useState } from 'react'
// All 50 states + DC, matching the API's ALLOWED_STATES
const STATES = [
'AL','AK','AZ','AR','CA','CO','CT','DE','DC','FL','GA','HI',
'ID','IL','IN','IA','KS','KY','LA','ME','MD','MA','MI','MN','MS','MO','MT',
'NE','NV','NH','NJ','NM','NY','NC','ND','OH','OK','OR','PA','RI','SC','SD',
'TN','TX','UT','VT','VA','WA','WV','WI','WY'
]
const STATE_NAMES = {
AL: 'Alabama', AK: 'Alaska', AZ: 'Arizona', AR: 'Arkansas', CA: 'California',
CO: 'Colorado', CT: 'Connecticut', DE: 'Delaware', DC: 'District of Columbia', FL: 'Florida',
GA: 'Georgia', HI: 'Hawaii', ID: 'Idaho', IL: 'Illinois', IN: 'Indiana', IA: 'Iowa',
KS: 'Kansas', KY: 'Kentucky', LA: 'Louisiana', ME: 'Maine', MD: 'Maryland', MA: 'Massachusetts',
MI: 'Michigan', MN: 'Minnesota', MS: 'Mississippi', MO: 'Missouri', MT: 'Montana',
NE: 'Nebraska', NV: 'Nevada', NH: 'New Hampshire', NJ: 'New Jersey', NM: 'New Mexico',
NY: 'New York', NC: 'North Carolina', ND: 'North Dakota', OH: 'Ohio', OK: 'Oklahoma',
OR: 'Oregon', PA: 'Pennsylvania', RI: 'Rhode Island', SC: 'South Carolina', SD: 'South Dakota',
TN: 'Tennessee', TX: 'Texas', UT: 'Utah', VT: 'Vermont', VA: 'Virginia',
WA: 'Washington', WV: 'West Virginia', WI: 'Wisconsin', WY: 'Wyoming'
}
/**
* Landing screen — prompts user to select a state.
* Matches the Civilytics visual style with eyebrow label, stat callout, and clean card layout.
*/
export default function StateSelector({ onNext }) {
const [query, setQuery] = useState('')
// Filter states by code or name as user types
const filtered = STATES.filter(
(code) =>
code.toLowerCase().includes(query.toLowerCase()) ||
STATE_NAMES[code].toLowerCase().includes(query.toLowerCase())
)
return (
<div style={{ padding: 'var(--space-5) 0 var(--space-7)' }}>
{/* Eyebrow */}
<span className="eyebrow">US Department of Education Civil Rights Data Collection</span>
{/* Hero */}
<h1 style={{ marginTop: 'var(--space-2)', marginBottom: 'var(--space-3)' }}>
Explore school-based arrest rates by district
</h1>
<p style={{ fontSize: '1.1rem', lineHeight: 1.7, maxWidth: '48rem' }}>
The Civil Rights Data Collection (CRDC) is the only source of data on school-related
arrests for all U.S. public schools and districts. Enter a state to see estimates for
any district — with Bayesian model comparisons that reveal how confident we can be in
each rate, even when arrests are rare events.
</p>
{/* Stat callout — key context */}
<div className="stat-callout" style={{ marginTop: 'var(--space-4)' }}>
<div>
<span className="stat-num">34,846</span>
<span className="stat-label">students arrested in 2021–22</span>
</div>
<div>
<span className="stat-num">0.72</span>
<span className="stat-label accent">arrests per 1,000 students (national rate)</span>
</div>
<div>
<span className="stat-num">11.6%</span>
<span className="stat-label">of districts reported &gt;0 arrests</span>
</div>
</div>
{/* Search input */}
<div style={{ marginTop: 'var(--space-5)' }}>
<label htmlFor="state-search" style={{ display: 'block', marginBottom: 'var(--space-2)', fontWeight: 600 }}>
Select a state to begin:
</label>
<input
id="state-search"
type="text"
placeholder="Type a state name or code (e.g., Texas, TX)…"
value={query}
onChange={(e) => setQuery(e.target.value)}
style={{
width: '100%', padding: 'var(--space-2)', fontSize: '1rem',
border: '1px solid var(--cv-rule)', borderRadius: 'var(--radius-md)',
fontFamily: 'var(--font-sans)'
}}
autoFocus
/>
{/* Filtered grid of states */}
{filtered.length > 0 && (
<div style={{
display: 'grid',
gridTemplateColumns: 'repeat(auto-fill, minmax(120px, 1fr))',
gap: 'var(--space-1)',
marginTop: 'var(--space-3)'
}}>
{filtered.map((code) => (
<button
key={code}
className="cv-card"
style={{ textAlign: 'center', padding: 'var(--space-2)' }}
onClick={() => onNext(code)}
>
<span style={{ fontSize: '1.25rem', fontWeight: 700, color: 'var(--cv-navy-600)' }}>
{code}
</span>
<span className="stat-label" style={{ marginTop: '0.25rem' }}>
{STATE_NAMES[code]}
</span>
</button>
))}
</div>
)}
{filtered.length === 0 && (
<p style={{ color: 'var(--cv-ink-3)', marginTop: 'var(--space-2)' }}>
No states match "{query}". Try a full state name or two-letter code.
</p>
)}
</div>
{/* Context footer */}
<div style={{ marginTop: 'var(--space-5)', paddingBottom: 'var(--space-3)', borderBottom: '1px solid var(--cv-rule)' }}>
<span className="eyebrow">How it works</span>
<p style={{ fontSize: '0.9rem', color: 'var(--cv-ink-2)', marginTop: 'var(--space-1)' }}>
After selecting a state, you'll search for a school district and see arrest estimates from 10 Bayesian models —
including one-year vs. three-year specifications and baseline vs. referral-rate covariate models.
The app fetches data directly from the{' '}
<a href="https://crdc-api.civilytics.org/api/v1/" target="_blank" rel="noopener noreferrer">
CRDC Arrest Rate API
</a>. No account or key required.
</p>
</div>
</div>
)
}
+157
View File
@@ -0,0 +1,157 @@
/**
* Civilytics CRDC Arrest Rate API client.
* Wraps fetch() with retry, backoff, and typed error handling.
* Base URL: https://crdc-api.civilytics.org/api/v1
*/
// The git-pages static host serves from pages.civilytics.org/crdc-demo/ but the CRDC API
// at crdc-api.civilytics.org does not send CORS headers. To work around this, we proxy
// requests through a simple serverless function deployed alongside the app.
// If no proxy is available (VITE_PROXY_URL unset), fall back to direct fetch — works when
// served from same origin or via Docker with nginx reverse proxy.
const BASE_URL = import.meta.env.VITE_API_BASE || 'https://crdc-api.civilytics.org/api/v1'
const PROXY_URL = import.meta.env.VITE_PROXY_URL // e.g., https://pages.civilytics.org/crdc-demo/proxy/
// Student group labels matching the API enum (race=AM|BL|HI|WH; sex=F|M)
export const STUDENT_GROUPS = [
{ race: 'WH', sex: 'F', label: 'White Female' },
{ race: 'WH', sex: 'M', label: 'White Male' },
{ race: 'BL', sex: 'F', label: 'Black Female' },
{ race: 'BL', sex: 'M', label: 'Black Male' },
{ race: 'HI', sex: 'F', label: 'Hispanic Female' },
{ race: 'HI', sex: 'M', label: 'Hispanic Male' },
{ race: 'AM', sex: 'F', label: 'American Indian\nFemale' },
{ race: 'AM', sex: 'M', label: 'American Indian\nMale' },
]
// CRDC wave labels (year=15-16|17-18|21-22)
export const CRDC_WAVES = [
{ value: '15-16', label: '2015–16' },
{ value: '17-18', label: '2017–18' },
{ value: '21-22', label: '2021–22' },
]
// Model IDs returned by /models endpoint
export const MODEL_IDS = [
'unified_m1_mod', 'unified_m2_mod', 'unified_m3_mod', 'unified_m4_mod', 'unified_m5_mod',
'stratified_m1_mod', 'stratified_m2_mod', 'stratified_m3_mod', 'stratified_m4_mod', 'stratified_m5_mod',
]
// The "four quadrants" for distribution charts (Charts 4–6)
export const MODEL_QUADRANTS = [
{ model: 'unified_m1_mod', label: 'One-year, no covariate' },
{ model: 'unified_m2_mod', label: 'One-year + referral rate' },
{ model: 'unified_m3_mod', label: 'Three-year, no covariate' },
{ model: 'unified_m4_mod', label: 'Three-year + referral rate' },
]
// Default/recommended model (from validate.R)
export const DEFAULT_MODEL = 'unified_m2_mod'
/** Exponential backoff fetch wrapper */
async function apiFetch(path, { retries = 3, delay = 500 } = {}) {
// If a proxy URL is configured, route through it to bypass CORS restrictions.
// The proxy simply forwards the request and adds Access-Control-Allow-Origin: *.
const url = PROXY_URL ? `${PROXY_URL}?target=${encodeURIComponent(path)}` : `${BASE_URL}${path}`
let lastError
for (let attempt = 0; attempt <= retries; attempt++) {
try {
const res = await fetch(url, { signal: AbortSignal.timeout(30000) })
if (!res.ok) throw new Error(`HTTP ${res.status}: ${res.statusText}`)
/** @type {{status:string,data:any,error?:string,meta:object}} */
const envelope = PROXY_URL ? await res.json() : (await res.json())
// When using a proxy that returns the raw API response, unwrap it.
if (PROXY_URL && envelope.status) return envelope.data
if (!PROXY_URL && envelope.status === 'success') return envelope.data
throw new Error(envelope.error || 'Unknown API error')
} catch (err) {
lastError = err
if (attempt < retries && !(err instanceof DOMException)) {
const wait = delay * Math.pow(2, attempt) + Math.random() * 100
await new Promise((r) => setTimeout(r, wait))
} else {
throw lastError
}
}
}
throw lastError
}
// ——— API endpoint wrappers (return the `data` array from the envelope) ———
/** GET /models — static list of available Bayesian model specifications */
export async function fetchModels() {
return apiFetch('/models')
}
/** GET /districts?q=<partial>&state=XX — name/geo lookup → LEAID */
export async function searchDistricts(query, state) {
const params = new URLSearchParams({ q: query, state })
return apiFetch(`/districts?${params}`)
}
/** GET /estimates/{leaid}?model=X&year=Y — all demographics for one district/year/model */
export async function fetchDistrictEstimates(leaid, options = {}) {
const params = new URLSearchParams()
if (options.model) params.set('model', options.model)
if (options.year) params.set('year', options.year)
if (options.race) params.set('race', options.race)
if (options.sex) params.set('sex', options.sex)
const qs = params.toString()
return apiFetch(`/estimates/${leaid}${qs ? `?${qs}` : ''}`)
}
/** GET /states/{state}?race=X&sex=Y&year=Z — state-level aggregate for one group */
export async function fetchStateEstimates(state, options = {}) {
const params = new URLSearchParams()
if (options.race) params.set('race', options.race)
if (options.sex) params.set('sex', options.sex)
if (options.year) params.set('year', options.year)
const qs = params.toString()
return apiFetch(`/states/${state}${qs ? `?${qs}` : ''}`)
}
/**
* GET /estimates?state=XX&year=Y — every estimate row in a state.
*
* Not used by the running app. The rows come back `ORDER BY LEAID, RACE, SEX`
* at eight per district, so any `limit` short of the state's full row count
* selects the lowest-LEAID districts rather than a meaningful sample —
* `DistrictSearch` reads the pre-ranked `public/data/top_districts.json`
* fixture instead. Kept for scripts and ad-hoc use; page it with `meta.total`
* as `scripts/build-top-districts.mjs` does.
*/
export async function fetchStateDistricts(state, year = '21-22', limit = 500) {
return apiFetch(`/estimates?state=${state}&year=${year}&limit=${limit}`)
}
/** GET /draws?... — returns HF parquet shard URL + DuckDB SQL (bulk only, not used in browser app) */
export async function fetchDrawLocation(state, race, sex, year, model) {
const params = new URLSearchParams({ state })
if (race) params.set('race', race)
if (sex) params.set('sex', sex)
if (year) params.set('year', year)
if (model) params.set('model', model)
return apiFetch(`/draws?${params}`)
}
/** GET /estimates/{leaid}/draws?model=X&year=Y — fetch posterior draws for a district */
export async function fetchDistrictDraws(leaid, options = {}) {
const params = new URLSearchParams()
if (options.model) params.set('model', options.model)
if (options.year) params.set('year', options.year)
if (options.race) params.set('race', options.race)
if (options.sex) params.set('sex', options.sex)
const qs = params.toString()
return apiFetch(`/estimates/${leaid}/draws${qs ? `?${qs}` : ''}`)
}
+397
View File
@@ -0,0 +1,397 @@
import { useEffect, useMemo, useState } from 'react'
import { getDb } from '../utils/duckdbClient.js'
import { groupKey, isCompleteDrawSet } from '../utils/drawGroups.js'
import { totalInterval, totalPerDraw } from '../utils/districtTotal.js'
const HF_DATASET = 'civilytics/crdc-school-arrest-rates'
const HF_BASE = `https://huggingface.co/datasets/${HF_DATASET}/resolve/main/parquet`
const HF_TREE = `https://huggingface.co/api/datasets/${HF_DATASET}/tree/main/parquet`
// Module-level cache: the registered duckdb-wasm file buffers for one
// (model, year, state) shard, shared across every component instance and
// district navigated to in this browser session. See Global Constraints —
// in-memory only, no persistence across page loads. Four models is simply four
// cache entries; nothing else about this cache changes when comparing specs.
const shardCache = new Map()
// Runaway guard on a malformed directory listing, not an expected limit.
const MAX_SHARD_PARTS = 64
function shardKey(model, year, state) {
return `${model}__${year}__${state}`
}
function shardDir(model, year, state) {
return `model_id=${model}/YEAR=${year}/LEA_STATE=${state}`
}
// Directory listings are ~1KB and carry each part's byte size, so the UI can
// tell the reader what a fetch will cost before committing to it. Cached
// separately from the buffers: listing a shard is cheap, downloading it is not.
const listingCache = new Map()
/**
* @returns {Promise<Array<{part: number, size: number}>>} parquet parts, in order
*/
function listShardEntries(model, year, state) {
const dir = shardDir(model, year, state)
if (!listingCache.has(dir)) {
listingCache.set(
dir,
(async () => {
try {
const res = await fetch(`${HF_TREE}/${dir}`)
if (!res.ok) throw new Error(`tree listing HTTP ${res.status}`)
const entries = await res.json()
return entries
.filter((e) => e?.type === 'file' && /\/data_\d+\.parquet$/.test(e.path || ''))
.map((e) => ({ part: Number(e.path.match(/data_(\d+)\.parquet$/)[1]), size: e.size || 0 }))
.sort((a, b) => a.part - b.part)
} catch (err) {
listingCache.delete(dir)
throw err
}
})(),
)
}
return listingCache.get(dir)
}
/**
* Total bytes of the draw shards for one model across several years — what a
* "compute this from the draws" action will actually download.
*
* @returns {Promise<number>} bytes, or 0 if the size can't be determined
*/
export async function measureShardBytes(model, years, state) {
try {
const sizes = await Promise.all(
years.map(async (year) =>
(await listShardEntries(model, year, state)).reduce((sum, p) => sum + p.size, 0),
),
)
return sizes.reduce((sum, n) => sum + n, 0)
} catch {
return 0
}
}
/**
* Lists the parquet parts published for one (model, year, state).
*
* A state's draws are split across `data_0.parquet`, `data_1.parquet`, … and
* the part count varies by state: Nevada is one file, California is eight
* (6.2MB in total, of which data_0 is 37KB). Reading only data_0 covers 11 of
* California's 1,715 districts and makes every other CA district look absent
* from the published data, silently falling back to the approximation.
*
* The directory listing is used rather than probing `data_N` until a 404
* because a 404 is logged as a console error by the browser's network layer no
* matter how cleanly the fetch handles it — and a red error on every load is
* indistinguishable from a real one. Probing remains the fallback if the
* listing API is unavailable or changes shape.
*/
async function listShardParts(model, year, state) {
const dir = shardDir(model, year, state)
try {
const parts = await listShardEntries(model, year, state)
if (parts.length > 0) return parts.map((p) => p.part)
throw new Error('tree listing contained no parquet parts')
} catch (err) {
console.warn(`useDrawDistribution: falling back to sequential part probing for ${dir}:`, err.message)
const parts = []
for (let part = 0; part < MAX_SHARD_PARTS; part++) {
const res = await fetch(`${HF_BASE}/${dir}/data_${part}.parquet`, { method: 'HEAD' })
if (!res.ok) break
parts.push(part)
}
return parts
}
}
function ensureShardRegistered(db, model, year, state) {
const key = shardKey(model, year, state)
if (!shardCache.has(key)) {
shardCache.set(
key,
(async () => {
try {
const dir = shardDir(model, year, state)
const parts = await listShardParts(model, year, state)
if (parts.length === 0) {
throw new Error(`No draw shard published for ${state}/${year}/${model}`)
}
// Parts are independent files, so fetch them together rather than
// walking them one at a time — California is eight round trips.
const fileNames = await Promise.all(
parts.map(async (part) => {
const res = await fetch(`${HF_BASE}/${dir}/data_${part}.parquet`)
if (!res.ok) throw new Error(`Failed to fetch draw shard part ${part}: HTTP ${res.status}`)
const buffer = new Uint8Array(await res.arrayBuffer())
const fileName = `${key}__${part}.parquet`
await db.registerFileBuffer(fileName, buffer)
return fileName
}),
)
return fileNames
} catch (err) {
// Don't let a transient failure (network blip, HF outage) poison the
// cache forever — remove the rejected entry so the next caller for
// this shard gets a fresh attempt instead of the same dead promise.
shardCache.delete(key)
throw err
}
})(),
)
}
return shardCache.get(key)
}
/**
* Reads one model's shard and returns that district's predicted counts keyed by
* group and indexed by draw.
*
* Indexing by `draw_id - 1` rather than by push order means DuckDB's row
* ordering is irrelevant, and it makes a missing draw a hole instead of a
* silently shorter array — which `isCompleteDrawSet` then rejects.
*/
async function fetchModelCounts(db, { leaid, state, model, year }) {
const fileNames = await ensureShardRegistered(db, model, year, state)
let conn
try {
conn = await db.connect()
// read_parquet over the full part list — a district lives in exactly one
// part, and which one is not predictable from its LEAID.
const fileList = fileNames.map((f) => `'${f}'`).join(', ')
const stmt = await conn.prepare(
`SELECT RACE, SEX, draw_id, pred FROM read_parquet([${fileList}]) WHERE LEAID = ?`,
)
const table = await stmt.query(leaid)
await stmt.close()
const rows = table.toArray().map((r) => r.toJSON())
const raw = {}
let nDraws = 0
for (const row of rows) {
const drawId = Number(row.draw_id)
const pred = Number(row.pred)
if (!Number.isFinite(drawId) || drawId < 1 || !Number.isFinite(pred)) continue
const key = groupKey(row.RACE, row.SEX)
;(raw[key] ??= [])[drawId - 1] = pred
if (drawId > nDraws) nDraws = drawId
}
// A group present for only part of the draw set must fall back, not render
// a short draw set: its density and interval would be computed off a
// biased subsample and look identical on screen to a complete one.
const counts = {}
let dropped = 0
for (const [key, arr] of Object.entries(raw)) {
if (isCompleteDrawSet(arr, nDraws)) counts[key] = arr
else {
dropped += 1
console.warn('useDrawDistribution: dropping incomplete draw set for', { leaid, model, key, got: arr.length, want: nDraws })
}
}
// `dropped` matters to any consumer that aggregates *across* groups (the
// district total): a missing group there is an undercount, not a gap.
return Object.keys(counts).length > 0 ? { counts, nDraws, dropped } : null
} finally {
if (conn) await conn.close()
}
}
/**
* Fetches real posterior predictive draws for one district from the Hugging
* Face parquet dataset via duckdb-wasm, for one or more model specifications at
* once, and returns **predicted counts indexed by draw** — not rates.
*
* Counts rather than rates is what unlocks the rest of the app: sex pooling has
* to sum numerators and denominators separately (see `utils/pooling.js`), and
* a between-group difference has to be taken at a common draw index. Callers
* divide by their own enrollment, which is why this hook no longer takes a
* `groups` argument at all.
*
* A caveat worth carrying into any caption: `draw_id` is renumbered 1–500 per
* write batch upstream, and a district's groups land in different batches, so
* draws are **not** paired parameter draws across groups — the pairing is
* effectively independent, dominated by `posterior_predict` observation noise.
* Published Fig 7 has the same property. Say "posterior predictive draws".
*
* `status` is an ANY-model, ANY-group signal: 'ready' means at least one model
* returned at least one complete group. A caller claiming "these are all real
* draws" must check its own rendered groups with `hasDrawsForAll`.
*
* @param {{leaid: string, state: string, models: string[], year: string}} params
* @returns {{status: 'loading'|'ready'|'error',
* byModel: Record<string, {counts: Record<string, number[]>, nDraws: number} | null> | null,
* nDraws: number}} `nDraws` is the largest draw count among models that came
* back; per-model counts live in `byModel[model].nDraws`.
*/
export function useDrawDistribution({ leaid, state, models, year }) {
const [status, setStatus] = useState('loading')
const [byModel, setByModel] = useState(null)
// `models` is typically a fresh array literal every render; derive a stable
// primitive so the effect only re-runs when its actual content changes.
const modelsSignature = (models || []).filter(Boolean).join(',')
useEffect(() => {
let cancelled = false
// Reset BEFORE the input guard below, not after. Clearing any previous
// model/district's draws has to happen on every input change, including the
// ones that have nothing to fetch. Without this ordering, a chart that
// varies its model selection across renders and lands on a model whose
// fetch fails would keep reporting 'ready' and keep handing back the
// *previous* model's real draws — keyed by the same group strings — under
// the newly selected model's label, silently mixing two models' data.
setStatus('loading')
setByModel(null)
const modelList = modelsSignature ? modelsSignature.split(',') : []
// Nothing to fetch: stay in 'loading' with no draws, which every consumer
// already treats as "fall back to the approximation". No cleanup needed —
// nothing async was started.
if (!leaid || !state || !year || modelList.length === 0) return
async function run() {
try {
const db = await getDb()
// Per-model try/catch: one shard 404ing (or one model missing for this
// state) must not blank out the models that did load.
const results = await Promise.all(
modelList.map(async (model) => {
try {
return [model, await fetchModelCounts(db, { leaid, state, model, year })]
} catch (err) {
console.error(`useDrawDistribution: model ${model} failed:`, err)
return [model, null]
}
}),
)
if (cancelled) return
const next = Object.fromEntries(results)
const anyReady = Object.values(next).some((v) => v !== null)
if (anyReady) {
setByModel(next)
setStatus('ready')
} else {
console.warn('useDrawDistribution: no usable draws for', { leaid, state, year, models: modelList })
setByModel(null)
setStatus('error')
}
} catch (err) {
// Only reached when the shared duckdb engine itself fails to load.
console.error('useDrawDistribution failed:', err)
if (!cancelled) {
// Belt-and-suspenders alongside the setByModel(null) at the top of
// this effect: a failed load must never leave a *previous* model's
// real draws in place under the newly-selected model's label.
setByModel(null)
setStatus('error')
}
}
}
run()
return () => {
cancelled = true
}
}, [leaid, state, year, modelsSignature])
const nDraws = useMemo(
() => Math.max(0, ...Object.values(byModel || {}).map((v) => v?.nDraws || 0)),
[byModel],
)
return { status, byModel, nDraws }
}
/**
* The district's total arrest count per wave, with a 95% interval computed from
* the draws — summed within each draw, then summarized across draws.
*
* Gated behind `enabled` because the cost is wildly uneven: three waves of
* Nevada is 0.3MB, but California is 13.9MB and Texas 14.3MB. `bytes` is
* probed from the directory listings up front (~1KB per year) so the caller can
* either fetch automatically when it's cheap or show the reader the price
* first.
*
* @param {{leaid: string, state: string, model: string, years: string[], enabled: boolean}} params
* @returns {{status: 'idle'|'loading'|'ready'|'error',
* byYear: Record<string, {lower:number, median:number, upper:number, nDraws:number}> | null,
* bytes: number}}
*/
export function useDistrictTotalDraws({ leaid, state, model, years, enabled }) {
const [status, setStatus] = useState('idle')
const [byYear, setByYear] = useState(null)
const [bytes, setBytes] = useState(0)
const yearsSignature = (years || []).join(',')
// Probe sizes regardless of `enabled` — this is what lets the UI decide.
useEffect(() => {
let cancelled = false
setBytes(0)
if (!state || !model || !yearsSignature) return
measureShardBytes(model, yearsSignature.split(','), state).then((n) => {
if (!cancelled) setBytes(n)
})
return () => {
cancelled = true
}
}, [state, model, yearsSignature])
useEffect(() => {
let cancelled = false
setStatus(enabled ? 'loading' : 'idle')
setByYear(null)
if (!enabled || !leaid || !state || !model || !yearsSignature) return
async function run() {
try {
const db = await getDb()
const yearList = yearsSignature.split(',')
const results = await Promise.all(
yearList.map(async (year) => {
try {
const model_ = await fetchModelCounts(db, { leaid, state, model, year })
if (!model_ || model_.dropped > 0) return [year, null]
const interval = totalInterval(totalPerDraw(model_.counts, model_.nDraws))
return [year, interval]
} catch (err) {
console.error(`useDistrictTotalDraws: ${year} failed:`, err)
return [year, null]
}
}),
)
if (cancelled) return
const next = Object.fromEntries(results)
if (Object.values(next).some((v) => v !== null)) {
setByYear(next)
setStatus('ready')
} else {
setByYear(null)
setStatus('error')
}
} catch (err) {
console.error('useDistrictTotalDraws failed:', err)
if (!cancelled) {
setByYear(null)
setStatus('error')
}
}
}
run()
return () => {
cancelled = true
}
}, [leaid, state, model, yearsSignature, enabled])
return { status, byYear, bytes }
}
+10
View File
@@ -0,0 +1,10 @@
import React from 'react'
import ReactDOM from 'react-dom/client'
import App from './App.jsx'
import './styles/tokens.css'
ReactDOM.createRoot(document.getElementById('root')).render(
<React.StrictMode>
<App />
</React.StrictMode>,
)
+278
View File
@@ -0,0 +1,278 @@
/* =============================================================
Civilytics Design Tokens — single source of truth.
Mirrored from crdc-arrests/theme/_tokens.scss so this demo app
stays in sync with the white paper / social media visual style.
Neutrals: warm paper → civic ink
Brand: Civic Navy (primary) + Ember (accent)
VIZ: Supporting palette for data series
============================================================ */
:root {
/* — Paper (backgrounds) — */
--cv-paper: #FAF7F2;
--cv-paper-2: #F2EDE4;
--cv-paper-3: #E6DFD1;
/* — Rule / borders — */
--cv-rule: #D6CEBD;
--cv-rule-strong: #B8AE97;
/* — Ink (text) — */
--cv-ink: #0E1A2B;
--cv-ink-2: #2B3A52;
--cv-ink-3: #5A6A82;
--cv-ink-4: #8C97AB;
/* — Civic Navy (primary brand) — */
--cv-navy-50: #EEF3FA;
--cv-navy-100: #DDE6F2;
--cv-navy-200: #B3C6E0;
--cv-navy-300: #7A9BCA;
--cv-navy-400: #4A74B0;
--cv-navy-500: #2E5590;
--cv-navy-600: #22406A; /* links, primary buttons */
--cv-navy-700: #1A2E4A; /* hover state */
/* — Ember (accent) — */
--cv-accent: #C25311; /* main accent color */
--cv-accent-text: #A04400;
--cv-accent-dark: #923D00;
--cv-accent-hover: #7A3600;
/* — Supporting data-viz palette (from _tokens.scss) — */
--teal-600: #1F6F70;
--plum-600: #6B3A5E;
--moss-600: #4A6B2F;
--brass-600: #B8751C;
/* — Semantic status colors — */
--cv-success: #4A6B2F; /* moss */
--cv-warning: #9A5F18;
--cv-danger: #A6271D;
/* — Font stacks (from _tokens.scss) — */
--font-display: 'Libre Franklin', 'Franklin Gothic', 'Inter', system-ui, sans-serif;
--font-sans: 'Inter', -apple-system, BlinkMacSystemFont, 'Segoe UI', Helvetica, Arial, sans-serif;
--font-mono: 'JetBrains Mono', 'SF Mono', Menlo, Consolas, monospace;
/* — Spacing scale (from _tokens.scss) — */
--space-1: 0.5rem;
--space-2: 1rem;
--space-3: 1.5rem;
--space-4: 2rem;
--space-5: 3rem;
--space-6: 4rem;
--space-7: 6rem;
/* — Border radius (from _tokens.scss) — */
--radius-sm: 4px;
--radius-md: 6px;
--radius-lg: 12px;
/* — Race categorical palette (validated, see src/utils/colors.js) — */
--race-wh: #3D6FC4;
--race-bl: #C98A2A;
--race-hi: #9C3F86;
--race-am: #3D9A6B;
/* — Chart-specific colors — aliased onto the tokens charts actually use */
--chart-modeled: var(--cv-navy-600);
--chart-frequentist: var(--cv-ink-4);
--chart-observed: var(--cv-ink);
/* — Model quadrant colors — */
--model-one-year-baseline: var(--cv-navy-600);
--model-three-year-baseline: var(--teal-600);
}
/* =============================================================
Base reset + typography (from crdc-arrests/theme/civilytics.scss)
============================================================ */
* { box-sizing: border-box; }
html { font-size: 100%; -webkit-font-smoothing: antialiased; text-rendering: optimizeLegibility; }
body {
margin: 0;
background: var(--cv-paper);
color: var(--cv-ink);
font-family: var(--font-sans);
line-height: 1.65;
}
h1, h2, h3, h4 {
font-family: var(--font-display);
font-weight: 800;
letter-spacing: -0.025em;
line-height: 1.08;
color: var(--cv-ink);
text-wrap: pretty;
margin: 0 0 .4em;
}
h1 { font-size: clamp(2.25rem, 4.2vw, 3.75rem); font-weight: 900; letter-spacing: -0.035em; line-height: 1.0; }
h2 { font-size: clamp(1.75rem, 2.8vw, 2.375rem); }
h3 { font-size: 1.5rem; font-weight: 700; letter-spacing: -0.022em; color: var(--cv-ink-2); }
p, li { line-height: 1.7; text-wrap: pretty; }
a {
color: var(--cv-navy-600);
text-decoration: underline;
text-decoration-thickness: 1px;
text-underline-offset: 2px;
transition: color 120ms ease;
}
a:hover { color: var(--cv-navy-700); text-decoration-thickness: 2px; }
code, .mono { font-family: var(--font-mono); font-feature-settings: "tnum" 1, "zero" 1; }
:focus-visible { outline: 2px solid var(--cv-accent); outline-offset: 2px; }
/* =============================================================
Layout helpers (from civilytics.scss)
============================================================ */
.cv-wrap { max-width: 60rem; margin: 0 auto; padding: 0 var(--space-3); }
.cv-header {
display: flex; align-items: center; justify-content: space-between;
padding: var(--space-3) 0; border-bottom: 1px solid var(--cv-rule);
}
.cv-header .cv-logo { height: 28px; display: block; }
/* Eyebrow / section label (from civilytics.scss + extras.css) */
.eyebrow, .cv-eyebrow {
font-family: var(--font-sans);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.1em;
color: var(--cv-accent-text);
margin: 0 0 0.5rem;
display: inline-flex;
align-items: center;
gap: 0.5rem;
}
.eyebrow::before, .cv-eyebrow::before {
content: "";
display: inline-block;
width: 24px; height: 2px;
background: var(--cv-accent);
}
/* Stat callout — hero numbers (from extras.css) */
.stat-callout {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(180px, 1fr));
gap: 2rem;
margin: 2rem 0 2.5rem;
padding: 1.5rem 0;
border-top: 3px double var(--cv-ink);
border-bottom: 1px solid var(--cv-rule);
}
.stat-callout .stat-num {
font-family: 'Source Serif 4', Georgia, serif;
font-weight: 700;
font-size: 3rem;
line-height: 1;
letter-spacing: -0.02em;
color: var(--cv-ink);
font-variant-numeric: tabular-nums;
display: block;
}
.stat-callout .stat-label {
font-family: var(--font-sans);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--cv-ink-3);
margin-top: 0.4rem;
}
.stat-callout .stat-label.accent { color: var(--cv-accent); }
/* Card component (used for district suggestions, model cards) */
.cv-card {
display: flex;
flex-direction: column;
gap: 0.25rem;
background: #fff;
border: 1px solid var(--cv-rule);
border-top: 3px solid var(--cv-accent);
border-radius: var(--radius-lg);
padding: var(--space-2) var(--space-3);
text-decoration: none;
transition: border-color 120ms ease, box-shadow 120ms ease;
}
.cv-card:hover {
border-color: var(--cv-ink-3);
box-shadow: 0 2px 8px rgba(0,0,0,0.04);
}
/* Button (primary) */
.btn-primary {
display: inline-flex;
align-items: center;
justify-content: center;
gap: 0.5rem;
padding: 0.625rem 1.25rem;
background: var(--cv-navy-600);
color: #fff;
border: none;
border-radius: var(--radius-md);
font-family: var(--font-sans);
font-size: 0.9rem;
font-weight: 600;
cursor: pointer;
transition: background 120ms ease, transform 20ms ease;
}
.btn-primary:hover { background: var(--cv-navy-700); }
.btn-primary:active { transform: scale(0.98); }
.btn-primary:disabled { opacity: 0.5; cursor: not-allowed; }
/* Button (secondary / outline) */
.btn-outline {
display: inline-flex;
align-items: center;
justify-content: center;
gap: 0.5rem;
padding: 0.5rem 1rem;
background: transparent;
color: var(--cv-navy-600);
border: 1px solid var(--cv-rule);
border-radius: var(--radius-md);
font-family: var(--font-sans);
font-size: 0.85rem;
font-weight: 500;
cursor: pointer;
transition: background 120ms ease, color 120ms ease;
}
.btn-outline:hover {
background: var(--cv-paper-2);
color: var(--cv-navy-700);
}
/* Footer */
.cv-footer {
border-top: 1px solid var(--cv-rule);
padding: var(--space-4) 0;
color: var(--cv-ink-3);
font-size: 0.9rem;
display: flex;
flex-wrap: wrap;
gap: var(--space-2);
justify-content: space-between;
align-items: center;
}
/* Responsive */
@media (max-width: 680px) { .cv-footer { flex-direction: column; text-align: center; } }
/* Spinner used while opt-in multi-shard draw fetches are in flight */
@keyframes cv-spin { to { transform: rotate(360deg); } }
@media (prefers-reduced-motion: reduce) {
@keyframes cv-spin { to { transform: none; } }
}
+66
View File
@@ -0,0 +1,66 @@
/**
* Agresti–Coull approximate interval for a rare-event count.
*
* Direct port of `agresti_coull()` in crdc-arrests/R/paper_figures.R:219-237 —
* the frequentist point range drawn beside the posterior densities in the white
* paper's Figs 6 and 7. Keeping this a faithful port is the whole point: the
* app's error bars have to be the *same* interval the paper published, so the
* two can be compared directly.
*
* Two consequences of that faithfulness, both deliberate:
*
* 1. The bounds are on the **count** scale, not the proportion scale — the R
* function multiplies back up by the adjusted denominator. Divide by
* enrollment yourself to plot per-1,000.
* 2. `lower` can be **negative** for very small numerators (e.g. 1 arrest in 53
* students gives -0.32). Clamp for display at the call site; do not clamp
* here, or this stops matching the published figures.
*
* The R original returns an unnamed vector in the order
* `c(ci_upper, ci_lower, sd, phat_se, phat)` — upper *first*, which is easy to
* transcribe backwards. This returns a named object instead.
*/
import { probit } from './distributionApprox.js'
/**
* @param {number} numerator - observed events (arrests)
* @param {number} denominator - trials (students enrolled)
* @param {number} [confidenceLevel=0.95] - e.g. 0.95 for a 95% interval
* @returns {{upper: number, lower: number, sd: number, se: number, phat: number}}
* `upper`/`lower`/`sd` are counts. In the zero-numerator branch the R
* original also reports `phat` as a count (the interval midpoint) rather than
* a proportion; that quirk is preserved.
*/
export function agrestiCoull(numerator, denominator, confidenceLevel = 0.95) {
const adjStar = probit(1 - (1 - confidenceLevel) / 2)
if (numerator > 0) {
const numStar = numerator + adjStar
const denomStar = denominator + 2 * adjStar
const phat = numStar / denomStar
const se = Math.sqrt((phat / denomStar) * (1 - phat))
return {
upper: (phat + adjStar * se) * denomStar,
lower: (phat - adjStar * se) * denomStar,
sd: se * denomStar,
se,
phat,
}
}
// Zero events: the rule of three. The R original writes the bound as
// denominator * (-log(1 - level) / denominator), where the denominator
// cancels — so the upper bound is -log(1 - level) ≈ 3 at 95% regardless of
// how many students were enrolled. Kept in the cancelled form so the value
// is identical rather than merely close.
const upper = -Math.log(1 - confidenceLevel)
const midpoint = (upper + 0) / 2
return {
upper,
lower: 0,
sd: midpoint / confidenceLevel,
se: 0,
phat: midpoint,
}
}
+109
View File
@@ -0,0 +1,109 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { agrestiCoull } from './agrestiCoull.js'
/**
* Reference values produced by the R original, `agresti_coull()` in
* crdc-arrests/R/paper_figures.R:219-237, printed at 15 significant digits:
*
* agresti_coull(15, 499, 0.95) c(24.89431440404100, 9.02561356503911,
* 4.04821235598516, 0.00804941727469993,
* 0.03372299056239180)
* agresti_coull(0, 53, 0.95) c(2.99573227355399, 0, 1.57670119660736,
* 0, 1.49786613677699)
* agresti_coull(1, 53, 0.95) c(6.24314599198645, -0.32321802290634,
* 1.67512364173205, 0.02942947578293,
* 0.05200224403214)
* agresti_coull(100, 148928, 0.95) c(121.743969120453, 82.1759588486273,
* 10.0940656522091, 6.77763749845643e-05,
* 6.84607866694077e-04)
* agresti_coull(3, 200, 0.90) c(8.14909680812749, 1.14061044577545,
* 2.13042858267619, 0.01047976610058,
* 0.02284844466400)
*
* R's qnorm is exact to double precision; this port uses Acklam's rational
* probit (relative error < 1.15e-9), so equality is asserted to 1e-7 relative.
*/
const REL_TOL = 1e-7
function assertClose(actual, expected, label) {
const scale = Math.max(Math.abs(expected), 1e-9)
assert.ok(
Math.abs(actual - expected) / scale < REL_TOL,
`${label}: expected ${expected}, got ${actual}`,
)
}
function assertMatchesR(result, [upper, lower, sd, se, phat], label) {
assertClose(result.upper, upper, `${label} upper`)
assertClose(result.lower, lower, `${label} lower`)
assertClose(result.sd, sd, `${label} sd`)
assertClose(result.se, se, `${label} se`)
assertClose(result.phat, phat, `${label} phat`)
}
test('agrestiCoull: matches R for a typical rare-event cell', () => {
assertMatchesR(
agrestiCoull(15, 499),
[24.894314404041, 9.02561356503911, 4.04821235598516, 0.00804941727469993, 0.0337229905623918],
'ac(15, 499, 0.95)',
)
})
test('agrestiCoull: matches R for a large district cell', () => {
assertMatchesR(
agrestiCoull(100, 148928),
[121.743969120453, 82.1759588486273, 10.0940656522091, 6.77763749845643e-5, 6.84607866694077e-4],
'ac(100, 148928, 0.95)',
)
})
test('agrestiCoull: matches R at a non-default confidence level', () => {
assertMatchesR(
agrestiCoull(3, 200, 0.9),
[8.14909680812749, 1.14061044577545, 2.13042858267619, 0.0104797661005796, 0.0228484446639996],
'ac(3, 200, 0.90)',
)
})
test('agrestiCoull: zero numerator uses the rule of three', () => {
// -log(1 - 0.95) = 2.9957…, the classic "rule of three" upper bound for zero
// events. The R original writes it as denominator * (-log(1-cl)/denominator),
// which cancels — the bound does not depend on the denominator at all.
const r = agrestiCoull(0, 53)
assertMatchesR(r, [2.99573227355399, 0, 1.57670119660736, 0, 1.49786613677699], 'ac(0, 53, 0.95)')
assert.equal(r.lower, 0)
})
test('agrestiCoull: the zero-numerator upper bound ignores the denominator', () => {
assert.equal(agrestiCoull(0, 53).upper, agrestiCoull(0, 500000).upper)
})
test('agrestiCoull: keeps R\'s negative lower bound for a single event', () => {
// Faithful to the R original: with numerator = 1 the lower bound goes below
// zero. Callers clamp for display — do NOT clamp here, or this port silently
// stops matching the published figures.
const r = agrestiCoull(1, 53)
assertMatchesR(
r,
[6.24314599198645, -0.323218022906343, 1.67512364173205, 0.029429475782928, 0.0520022440321421],
'ac(1, 53, 0.95)',
)
assert.ok(r.lower < 0, 'lower bound should be negative for numerator = 1, n = 53')
})
test('agrestiCoull: bounds are on the count scale, not the proportion scale', () => {
const r = agrestiCoull(15, 499)
assert.ok(r.upper > 15 && r.lower < 15, 'the interval should bracket the observed count')
})
test('agrestiCoull: defaults to 95%', () => {
assert.deepEqual(agrestiCoull(15, 499), agrestiCoull(15, 499, 0.95))
})
test('agrestiCoull: a wider confidence level gives a wider interval', () => {
const narrow = agrestiCoull(15, 499, 0.8)
const wide = agrestiCoull(15, 499, 0.99)
assert.ok(wide.upper > narrow.upper)
assert.ok(wide.lower < narrow.lower)
})
+56
View File
@@ -0,0 +1,56 @@
/**
* Fixed categorical race palette — validated with the dataviz skill's
* validate_palette.js (passes lightness band, chroma floor, CVD separation
* ΔE 12.8, normal-vision floor ΔE 26.4; gold's sub-3:1 contrast WARN is
* mitigated by always-visible direct labels/legend in every chart that uses
* it). Order is fixed (WH, BL, HI, AM) and must not be re-derived from array
* position — re-run the validator before changing any of these hex values:
* node scripts/validate_palette.js "#3D6FC4,#C98A2A,#9C3F86,#3D9A6B" --mode light
*
* Sex (F/M) is deliberately NOT a second hue — it's encoded by row/facet
* position in every chart. Observed-vs-modeled is deliberately NOT a second
* hue either — it's encoded by mark type (diamond vs. filled bar/density).
*/
export const RACE_COLORS = { WH: '#3D6FC4', BL: '#C98A2A', HI: '#9C3F86', AM: '#3D9A6B' }
export const RACE_LABELS = { WH: 'White', BL: 'Black', HI: 'Hispanic', AM: 'American Indian / Alaska Native' }
// Compact form for axis ticks and other space-constrained labels, where the
// full RACE_LABELS text would overflow its container and get clipped by the
// SVG viewBox.
export const SHORT_RACE_LABEL = { WH: 'White', BL: 'Black', HI: 'Hispanic', AM: 'AI/AN' }
// Mark-type colors for the observed-vs-modeled convention. Used where a
// chart has no race facet of its own (aggregate totals).
export const OBSERVED_MARK_COLOR = 'var(--cv-ink)'
export const MODELED_AGGREGATE_COLOR = 'var(--cv-navy-600)'
// Fallback for an unrecognized race key — kept neutral rather than defaulting to a real race hue.
export const REFERENCE_GRAY = 'var(--cv-ink-4)'
export function raceColor(race) {
return RACE_COLORS[race] || REFERENCE_GRAY
}
// Both label helpers take an optional `sex`: omitting it names a sex-pooled
// group (race alone), mirroring `groupKey(race)` in utils/drawGroups.js. Don't
// let a missing sex fall through to "Male".
export function groupLabel(race, sex) {
const race_ = RACE_LABELS[race] || race
if (!sex) return race_
return `${race_} ${sex === 'F' ? 'Female' : 'Male'}`
}
export function shortGroupLabel(race, sex) {
const race_ = SHORT_RACE_LABEL[race] || race
return sex ? `${race_} ${sex}` : race_
}
// For labels that appear mid-sentence ("…the Black male arrest rate exceeds the
// White male rate"). The race is a proper noun and keeps its capital; only the
// sex word is lowercased. Don't reach for .toLowerCase() on groupLabel() — it
// turns "Black Male" into "black male".
export function sentenceGroupLabel(race, sex) {
const race_ = RACE_LABELS[race] || race
if (!sex) return race_
return `${race_} ${sex === 'F' ? 'female' : 'male'}`
}
+114
View File
@@ -0,0 +1,114 @@
/**
* Chooses an honest visual profile for one group's posterior predictive draws.
*
* In a sparse district the posterior predictive is a *discrete count*
* distribution, not a smooth one. Carson City NV (3200390) has a 53-student
* AI/AN female cell where a single arrest is 18.9 per 1,000: the draws take
* four distinct values, and a Gaussian KDE renders those four spikes as a lumpy
* smear that reads as a rendering bug rather than as a finding about the data.
*
* So: few distinct values → draw the actual probability mass at each achievable
* rate. Many → the smooth KDE, delegated to `kde.js` unchanged (its bandwidth
* clamp was tuned for exactly these zero-inflated posteriors — see
* BANDWIDTH_FLOOR_DIVISOR there; do not re-derive it here).
*/
import { kdeCurve } from './kde.js'
import { toRates } from './pooling.js'
/**
* At or below this many distinct predicted counts, the draws are shown as
* discrete mass rather than smoothed.
*
* 12 is comfortably above the 3–5 distinct values a genuinely sparse cell
* produces and comfortably below the ~40+ a district with real arrest volume
* produces, so the switch happens well away from either regime rather than
* flickering at the boundary.
*/
export const MASS_MAX_DISTINCT = 12
/**
* @param {number[] | null | undefined} counts - predicted counts, indexed by draw
* @param {number | null | undefined} enroll - students in the group
* @param {{min?: number, max?: number, n?: number}} [domain] - the x-range the
* profile will be drawn over, in rate per 1,000. Passed straight to
* `kdeCurve`, whose bandwidth clamp is relative to this width.
* @returns {{kind: 'kde'|'mass', points: Array<{x: number, y: number}>,
* maxY: number, step: number} | null}
* `null` when there is nothing to draw (no draws, or no denominator).
* For `kind: 'mass'`, `y` is a probability and `step` is the spacing between
* achievable rates — the bar width for a filled staircase. For `kind: 'kde'`,
* `y` is a density and `step` is 0.
* `maxY` is the profile's own peak: because the two kinds carry different y
* units, a chart overlaying several groups must normalize each profile by its
* own `maxY` rather than by a shared maximum.
*/
export function densityProfile(counts, enroll, domain = {}) {
const rates = toRates(counts, enroll)
if (!rates.length) return null
// Counts are integers, so consecutive achievable rates are exactly one
// student-rate apart — pass that explicitly rather than inferring it from the
// gaps actually observed, which overstates the bar width when (say) only
// counts 0 and 3 appear.
return rateProfile(rates, domain, 1000 / enroll)
}
/**
* The same choice made directly on a set of rates, for quantities that are
* already differences rather than count/denominator pairs (see
* `utils/groupDifference.js`). A difference of two discrete posteriors is
* itself discrete, and deserves the same honesty.
*
* @param {number[] | null | undefined} rates - values on the plotted scale
* @param {{min?: number, max?: number, n?: number}} [domain]
* @param {number} [explicitStep] - known spacing between achievable values;
* inferred from the smallest observed gap when omitted.
* @returns {{kind: 'kde'|'mass', points: Array<{x: number, y: number}>,
* maxY: number, step: number} | null}
*/
export function rateProfile(rates, domain = {}, explicitStep = 0) {
if (!rates?.length) return null
const { min = 0, max, n = 60 } = domain
// Round before tallying so two draws that differ only in floating-point noise
// count as one achievable value rather than two.
const tally = new Map()
for (const r of rates) {
const key = Math.round(r * 1e9) / 1e9
tally.set(key, (tally.get(key) || 0) + 1)
}
if (tally.size <= MASS_MAX_DISTINCT) {
const total = rates.length
const points = [...tally.entries()]
.map(([x, freq]) => ({ x, y: freq / total }))
.sort((a, b) => a.x - b.x)
return {
kind: 'mass',
points,
maxY: Math.max(...points.map((p) => p.y)),
step: explicitStep > 0 ? explicitStep : inferStep(points, max, min),
}
}
const points = kdeCurve(rates, { min, max, n })
return {
kind: 'kde',
points,
maxY: Math.max(...points.map((p) => p.y)),
step: 0,
}
}
/** Smallest gap between achievable values; a visible default for a single point. */
function inferStep(points, max, min) {
let smallest = Infinity
for (let i = 1; i < points.length; i++) {
const gap = points[i].x - points[i - 1].x
if (gap > 0 && gap < smallest) smallest = gap
}
if (Number.isFinite(smallest)) return smallest
const width = Number.isFinite(max) && max > min ? max - min : 1
return width / 40
}
+135
View File
@@ -0,0 +1,135 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { MASS_MAX_DISTINCT, densityProfile, rateProfile } from './densityProfile.js'
import { kdeCurve } from './kde.js'
import { toRates } from './pooling.js'
/** counts whose distinct-value count is exactly `k` (values 0…k-1, padded). */
function countsWithDistinct(k, total = 500) {
const out = []
for (let i = 0; i < total; i++) out.push(i % k)
return out
}
test('densityProfile: discrete posterior returns a mass profile', () => {
// 500 draws over 4 achievable counts is the Carson City AI/AN case: a KDE
// renders it as a lumpy smear that reads as a rendering bug.
const profile = densityProfile(countsWithDistinct(4), 53, { max: 100 })
assert.equal(profile.kind, 'mass')
})
test('densityProfile: mass points are probabilities at achievable rates', () => {
const profile = densityProfile([0, 0, 0, 1], 500, { max: 10 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [
{ x: 0, y: 0.75 },
{ x: 2, y: 0.25 },
])
})
test('densityProfile: mass probabilities sum to 1', () => {
const profile = densityProfile([0, 0, 1, 2, 2, 5], 1000, { max: 10 })
const total = profile.points.reduce((sum, p) => sum + p.y, 0)
assert.ok(Math.abs(total - 1) < 1e-12, `probabilities summed to ${total}`)
})
test('densityProfile: mass points are sorted by rate ascending', () => {
const profile = densityProfile([5, 0, 3, 1], 1000, { max: 10 })
const xs = profile.points.map((p) => p.x)
assert.deepEqual(xs, [...xs].sort((a, b) => a - b))
})
test('densityProfile: mass carries the achievable-rate step for staircase width', () => {
// Counts are integers, so achievable rates are spaced 1000/enroll apart.
// The chart needs that width to draw a bar rather than a hairline.
const profile = densityProfile([0, 1], 250, { max: 10 })
assert.equal(profile.step, 4)
})
test('densityProfile: a fully degenerate draw set is a single mass point', () => {
// 500 identical zeros is common in small districts. This is the case a
// Gaussian KDE turns into a delta spike.
const profile = densityProfile(new Array(500).fill(0), 4073, { max: 20 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [{ x: 0, y: 1 }])
})
test('densityProfile: switches to KDE above the distinct-value cutoff', () => {
const atCutoff = densityProfile(countsWithDistinct(MASS_MAX_DISTINCT), 1000, { max: 50 })
const aboveCutoff = densityProfile(countsWithDistinct(MASS_MAX_DISTINCT + 1), 1000, { max: 50 })
assert.equal(atCutoff.kind, 'mass')
assert.equal(aboveCutoff.kind, 'kde')
})
test('densityProfile: KDE branch delegates to kdeCurve over the given domain', () => {
// The bandwidth clamp in kde.js was tuned for exactly these zero-inflated
// posteriors — this must delegate, not re-derive.
const counts = countsWithDistinct(40)
const enroll = 1000
const profile = densityProfile(counts, enroll, { min: 0, max: 50, n: 60 })
const expected = kdeCurve(toRates(counts, enroll), { min: 0, max: 50, n: 60 })
assert.equal(profile.kind, 'kde')
assert.deepEqual(profile.points, expected)
})
test('densityProfile: reports the profile peak for per-group normalization', () => {
const profile = densityProfile([0, 0, 0, 1], 500, { max: 10 })
assert.equal(profile.maxY, 0.75)
const kde = densityProfile(countsWithDistinct(40), 1000, { max: 50 })
assert.equal(kde.maxY, Math.max(...kde.points.map((p) => p.y)))
})
test('densityProfile: null for unusable input rather than an empty curve', () => {
assert.equal(densityProfile([], 500, { max: 10 }), null)
assert.equal(densityProfile(undefined, 500, { max: 10 }), null)
assert.equal(densityProfile([0, 1], 0, { max: 10 }), null)
assert.equal(densityProfile([0, 1], undefined, { max: 10 }), null)
})
test('densityProfile: does not mutate its input counts', () => {
const counts = [3, 1, 2]
densityProfile(counts, 1000, { max: 10 })
assert.deepEqual(counts, [3, 1, 2])
})
// ——— rateProfile (used for already-differenced quantities) ———
test('rateProfile: mass profile straight from rate values', () => {
const profile = rateProfile([-1, -1, 0, 2], { min: -5, max: 5 })
assert.equal(profile.kind, 'mass')
assert.deepEqual(profile.points, [
{ x: -1, y: 0.5 },
{ x: 0, y: 0.25 },
{ x: 2, y: 0.25 },
])
})
test('rateProfile: infers the staircase step from the smallest observed gap', () => {
assert.equal(rateProfile([0, 0.5, 2], { min: 0, max: 5 }).step, 0.5)
})
test('rateProfile: falls back to a visible step for a single achievable value', () => {
const profile = rateProfile([3, 3, 3], { min: 0, max: 40 })
assert.ok(profile.step > 0, 'a single mass point still needs a drawable width')
})
test('rateProfile: an explicit step wins over the inferred one', () => {
assert.equal(rateProfile([0, 3], { min: 0, max: 5 }, 0.25).step, 0.25)
})
test('rateProfile: merges values differing only by floating-point noise', () => {
const profile = rateProfile([0.1 + 0.2, 0.3, 0.3], { min: 0, max: 1 })
assert.equal(profile.points.length, 1)
assert.equal(profile.points[0].y, 1)
})
test('rateProfile: switches to KDE above the distinct-value cutoff', () => {
const many = Array.from({ length: 300 }, (_, i) => (i % (MASS_MAX_DISTINCT + 1)) * 0.7)
assert.equal(rateProfile(many, { min: 0, max: 10 }).kind, 'kde')
})
test('rateProfile: null for an empty or missing set', () => {
assert.equal(rateProfile([], { max: 5 }), null)
assert.equal(rateProfile(undefined, { max: 5 }), null)
})
+136
View File
@@ -0,0 +1,136 @@
/**
* Approximates a distribution shape from summary statistics only
* (median + interval bounds) when raw posterior draws aren't available
* client-side. This is NOT the true posterior — always pair its use with
* DISTRIBUTION_APPROX_NOTE (rendered via <ApproxNote />).
*
* Arrest rates/counts are right-skewed, not normal, so a single symmetric
* normal (the old approach) systematically misrepresents the shape. This
* fits two normal halves — one on each side of the median, each sized to
* its own interval bound — and splices their CDFs at the median. Because
* each half's CDF independently spans [0, 0.5] or [0.5, 1], the join is
* exactly the median by construction (cdf(median) === 0.5 always), unlike
* the classical two-piece-normal parameterization, which biases the median
* away from the split point whenever the two sigmas differ. Do not "fix"
* this toward that textbook formula — losing median-exactness is the bug,
* not a missing feature.
*/
const SQRT_2PI = Math.sqrt(2 * Math.PI)
const MIN_ABS_SIGMA = 1e-3
function standardNormalPdf(z) {
return Math.exp(-0.5 * z * z) / SQRT_2PI
}
// Abramowitz & Stegun 7.1.26 approximation of erf, ~1.5e-7 max error.
function erf(x) {
const sign = x < 0 ? -1 : 1
const ax = Math.abs(x)
const a1 = 0.254829592
const a2 = -0.284496736
const a3 = 1.421413741
const a4 = -1.453152027
const a5 = 1.061405429
const p = 0.3275911
const t = 1 / (1 + p * ax)
const y = 1 - (((((a5 * t + a4) * t) + a3) * t + a2) * t + a1) * t * Math.exp(-ax * ax)
return sign * y
}
function standardNormalCdf(z) {
return 0.5 * (1 + erf(z / Math.SQRT2))
}
/**
* Peter Acklam's rational approximation of the inverse standard normal CDF
* (probit / R's `qnorm`), relative error < 1.15e-9. Supports arbitrary
* intervalMass.
*
* Exported so `agrestiCoull.js` can reuse it — the app should carry exactly one
* probit implementation.
*
* @param {number} p - probability in (0, 1)
* @returns {number}
*/
export function probit(p) {
const a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02, 1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00]
const b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02, 6.680131188771972e+01, -1.328068155288572e+01]
const c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00, -2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00]
const d = [7.784695709041462e-03, 3.224671290700398e-01, 2.445134137142996e+00, 3.754408661907416e+00]
const pLow = 0.02425
const pHigh = 1 - pLow
if (p < pLow) {
const q = Math.sqrt(-2 * Math.log(p))
return (((((c[0] * q + c[1]) * q + c[2]) * q + c[3]) * q + c[4]) * q + c[5]) /
((((d[0] * q + d[1]) * q + d[2]) * q + d[3]) * q + 1)
}
if (p <= pHigh) {
const q = p - 0.5
const r = q * q
return (((((a[0] * r + a[1]) * r + a[2]) * r + a[3]) * r + a[4]) * r + a[5]) * q /
(((((b[0] * r + b[1]) * r + b[2]) * r + b[3]) * r + b[4]) * r + 1)
}
const q = Math.sqrt(-2 * Math.log(1 - p))
return -(((((c[0] * q + c[1]) * q + c[2]) * q + c[3]) * q + c[4]) * q + c[5]) /
((((d[0] * q + d[1]) * q + d[2]) * q + d[3]) * q + 1)
}
/**
* @param {{median:number, lower:number, upper:number, intervalMass?:number, floorAtZero?:boolean}} p
* intervalMass: fraction of probability covered by [lower, upper]. The API's
* rate/count bounds are a **95%** interval — `validate_interval()` in
* crdc-arrests/api/R/validate.R defaults to 95 and the app never passes
* `interval=` — so the default is 0.95. Fitting 95% bounds as if they were
* 90% understates sigma by ~16% and draws a distribution narrower than the
* model's own.
* @returns {{ median:number, sigmaLeft:number, sigmaRight:number, pdf:(x:number)=>number, cdf:(x:number)=>number }}
*/
export function fitSkewedInterval({ median, lower, upper, intervalMass = 0.95, floorAtZero = true }) {
const z = probit((1 + intervalMass) / 2)
let sigmaLeft = (median - lower) / z
let sigmaRight = (upper - median) / z
if (!(sigmaLeft > 0)) sigmaLeft = Math.max(sigmaRight * 0.15, MIN_ABS_SIGMA)
if (!(sigmaRight > 0)) sigmaRight = Math.max(sigmaLeft * 0.15, MIN_ABS_SIGMA)
sigmaLeft = Math.max(sigmaLeft, MIN_ABS_SIGMA)
sigmaRight = Math.max(sigmaRight, MIN_ABS_SIGMA)
function pdf(x) {
if (floorAtZero && x < 0) return 0
const sigma = x < median ? sigmaLeft : sigmaRight
return standardNormalPdf((x - median) / sigma) / sigma
}
function cdf(x) {
if (floorAtZero && x < 0) return 0
const sigma = x <= median ? sigmaLeft : sigmaRight
return standardNormalCdf((x - median) / sigma)
}
/** Inverse cdf — exact round-trip of fitSkewedInterval's own construction (quantile(fit, 0.05) === lower, etc). */
function quantile(p) {
const z = probit(p)
return median + (p <= 0.5 ? sigmaLeft : sigmaRight) * z
}
return { median, sigmaLeft, sigmaRight, pdf, cdf, quantile }
}
/** n evenly spaced {x,y} points of the fitted pdf, for drawing a smooth curve. */
export function densityCurve(fit, { min = 0, max, n = 60 } = {}) {
const hi = max ?? fit.median + fit.sigmaRight * 3.5
const lo = Math.min(min, fit.median - fit.sigmaLeft * 0.1)
const step = (hi - lo) / (n - 1)
const points = []
for (let i = 0; i < n; i++) {
const x = lo + step * i
points.push({ x, y: fit.pdf(x) })
}
return points
}
export const DISTRIBUTION_APPROX_NOTE =
'Distribution shape estimated from interval bounds — not raw posterior draws.'
+128
View File
@@ -0,0 +1,128 @@
/**
* Turns the API's per-race×sex estimate rows into the row set the results page
* actually displays, in one place, so the summary table, the density panel and
* the difference chart can never disagree about which groups exist, what they
* are called, or which are selected by default.
*
* "Display space" is either eight race×sex groups or — when sex pooling is on —
* four race groups. Keys come from `groupKey`, so the same map lookups work in
* both modes.
*/
import { groupLabel, sentenceGroupLabel, shortGroupLabel } from './colors.js'
import { groupKey } from './drawGroups.js'
import { poolBySex } from './pooling.js'
/** The four modeled race groups, in the fixed palette order (see colors.js). */
export const RACE_ORDER = ['WH', 'BL', 'HI', 'AM']
const SEXES = ['F', 'M']
/**
* @param {Array<object>} rows - estimate rows for one district/model/year
* @param {boolean} pooled - collapse Female + Male into one row per race
* @returns {Array<{key: string, race: string, sex: string|null, label: string,
* shortLabel: string, enroll: number, observed: number, rate: number}>}
* sorted by observed arrests descending.
*/
export function buildDisplayGroups(rows, pooled) {
const usable = (rows || []).filter(
(r) => RACE_ORDER.includes(r.race) && SEXES.includes(r.sex),
)
const acc = new Map()
for (const r of usable) {
const sex = pooled ? null : r.sex
const key = groupKey(r.race, sex)
const prev = acc.get(key) || { key, race: r.race, sex, enroll: 0, observed: 0 }
acc.set(key, {
...prev,
enroll: prev.enroll + (r.stu_enroll || 0),
observed: prev.observed + (r.observed_arrests || 0),
// The model's own summary interval, carried per 1,000 so a chart can fall
// back to the fitted approximation when real draws can't be fetched.
// Null when pooled: adding two groups' interval *bounds* together is not
// a pooled interval, and there is no honest way to fake one without the
// draws. A pooled group with no draws simply has no modelled shape.
modeled: pooled
? null
: {
median: (r.rate_median || 0) * 1000,
lower: (r.rate_lower || 0) * 1000,
upper: (r.rate_upper || 0) * 1000,
},
})
}
return [...acc.values()]
.map((g) => ({
...g,
label: groupLabel(g.race, g.sex),
shortLabel: shortGroupLabel(g.race, g.sex),
sentenceLabel: sentenceGroupLabel(g.race, g.sex),
// A rate with no denominator is not a large rate — it is no rate. Report
// 0 and let the table's enrollment column show why.
rate: g.enroll > 0 ? (g.observed / g.enroll) * 1000 : 0,
}))
.sort((a, b) => b.observed - a.observed || b.enroll - a.enroll || a.key.localeCompare(b.key))
}
/**
* Groups checked on first render: every group with at least one observed
* arrest. When a district reports none at all, falls back to the two largest by
* enrollment so the chart still shows something explainable rather than an
* empty panel (the caller says so in the UI).
*
* @param {ReturnType<typeof buildDisplayGroups>} groups
* @returns {string[]} group keys
*/
export function defaultSelectedKeys(groups) {
if (!groups?.length) return []
const withArrests = groups.filter((g) => g.observed > 0)
if (withArrests.length > 0) return withArrests.map((g) => g.key)
return [...groups]
.sort((a, b) => b.enroll - a.enroll)
.slice(0, 2)
.map((g) => g.key)
}
/**
* The difference chart's opening pair: the two groups with the most observed
* arrests. Null when there aren't two groups to compare.
*
* @param {ReturnType<typeof buildDisplayGroups>} groups
* @returns {[string, string] | null}
*/
export function defaultDiffPair(groups) {
if (!groups || groups.length < 2) return null
return [groups[0].key, groups[1].key]
}
/**
* Enrollment for every race×sex cell, keyed for `poolBySex`/`displayDraws`.
* Always unpooled — pooling sums these itself.
*
* @param {Array<object>} rows
* @returns {Record<string, number>}
*/
export function enrollByGroupKey(rows) {
const out = {}
for (const r of rows || []) {
if (!RACE_ORDER.includes(r.race) || !SEXES.includes(r.sex)) continue
out[groupKey(r.race, r.sex)] = r.stu_enroll || 0
}
return out
}
/**
* Moves one model's raw race×sex draw counts into display space.
*
* @param {Record<string, number[]> | null | undefined} counts - keyed race×sex
* @param {Record<string, number>} enrollByGroup - keyed race×sex
* @param {boolean} pooled
* @returns {{counts: Record<string, number[]>, enroll: Record<string, number>}}
*/
export function displayDraws(counts, enrollByGroup, pooled) {
if (!counts) return { counts: {}, enroll: {} }
if (pooled) return poolBySex(counts, enrollByGroup)
return { counts, enroll: enrollByGroup || {} }
}
+169
View File
@@ -0,0 +1,169 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
buildDisplayGroups,
defaultDiffPair,
defaultSelectedKeys,
displayDraws,
enrollByGroupKey,
} from './districtGroups.js'
const row = (race, sex, stu_enroll, observed_arrests) => ({ race, sex, stu_enroll, observed_arrests })
const CLARK = [
row('WH', 'F', 30000, 8),
row('WH', 'M', 31000, 20),
row('BL', 'F', 12000, 18),
row('BL', 'M', 12500, 40),
row('HI', 'F', 30000, 5),
row('HI', 'M', 31000, 8),
row('AM', 'F', 200, 0),
row('AM', 'M', 228, 1),
]
// ——— buildDisplayGroups ———
test('buildDisplayGroups: one row per race×sex when not pooled', () => {
const groups = buildDisplayGroups(CLARK, false)
assert.equal(groups.length, 8)
assert.deepEqual(
groups.map((g) => g.key).sort(),
['AM_F', 'AM_M', 'BL_F', 'BL_M', 'HI_F', 'HI_M', 'WH_F', 'WH_M'],
)
})
test('buildDisplayGroups: collapses to four race rows when pooled', () => {
const groups = buildDisplayGroups(CLARK, true)
assert.equal(groups.length, 4)
const bl = groups.find((g) => g.key === 'BL')
assert.equal(bl.enroll, 24500)
assert.equal(bl.observed, 58)
assert.equal(bl.sex, null)
})
test('buildDisplayGroups: sorted by observed arrests descending', () => {
const observed = buildDisplayGroups(CLARK, false).map((g) => g.observed)
assert.deepEqual(observed, [...observed].sort((a, b) => b - a))
})
test('buildDisplayGroups: rate is per 1,000 students', () => {
const bl = buildDisplayGroups(CLARK, false).find((g) => g.key === 'BL_M')
assert.ok(Math.abs(bl.rate - (40 / 12500) * 1000) < 1e-12)
})
test('buildDisplayGroups: rate is 0 rather than Infinity with no enrollment', () => {
const groups = buildDisplayGroups([row('BL', 'F', 0, 3)], false)
assert.equal(groups[0].rate, 0)
})
test('buildDisplayGroups: labels a pooled group by race alone', () => {
const bl = buildDisplayGroups(CLARK, true).find((g) => g.key === 'BL')
assert.equal(bl.label, 'Black')
const blf = buildDisplayGroups(CLARK, false).find((g) => g.key === 'BL_F')
assert.equal(blf.label, 'Black Female')
})
test('buildDisplayGroups: the sentence label keeps the race capitalized', () => {
// Guards the readout in GroupDifference: lowercasing the whole label turned
// "Black Male" into "black male" mid-sentence.
const groups = buildDisplayGroups(CLARK, false)
assert.equal(groups.find((g) => g.key === 'BL_M').sentenceLabel, 'Black male')
assert.equal(groups.find((g) => g.key === 'WH_F').sentenceLabel, 'White female')
assert.equal(buildDisplayGroups(CLARK, true).find((g) => g.key === 'HI').sentenceLabel, 'Hispanic')
})
test('buildDisplayGroups: carries the modeled interval per 1,000 when not pooled', () => {
const groups = buildDisplayGroups(
[{ race: 'BL', sex: 'M', stu_enroll: 1000, observed_arrests: 4, rate_median: 0.004, rate_lower: 0.001, rate_upper: 0.009 }],
false,
)
assert.deepEqual(groups[0].modeled, { median: 4, lower: 1, upper: 9 })
})
test('buildDisplayGroups: pooled groups carry no modeled interval', () => {
// Summing two groups' interval bounds is not a pooled interval — there is no
// honest fallback shape without the draws, so don't invent one.
const groups = buildDisplayGroups(CLARK, true)
assert.ok(groups.every((g) => g.modeled === null))
})
test('buildDisplayGroups: ignores races and sexes outside the modeled set', () => {
const groups = buildDisplayGroups([...CLARK, row('AS', 'F', 900, 4), row('WH', 'X', 5, 5)], false)
assert.equal(groups.length, 8)
assert.ok(!groups.some((g) => g.race === 'AS'))
})
test('buildDisplayGroups: missing counts default to zero', () => {
const groups = buildDisplayGroups([{ race: 'BL', sex: 'F' }], false)
assert.equal(groups[0].observed, 0)
assert.equal(groups[0].enroll, 0)
})
test('buildDisplayGroups: empty input gives an empty list', () => {
assert.deepEqual(buildDisplayGroups([], false), [])
assert.deepEqual(buildDisplayGroups(undefined, true), [])
})
// ——— defaultSelectedKeys ———
test('defaultSelectedKeys: every group with at least one observed arrest', () => {
const groups = buildDisplayGroups(CLARK, false)
const keys = defaultSelectedKeys(groups)
assert.ok(!keys.includes('AM_F'))
assert.ok(keys.includes('AM_M'))
assert.equal(keys.length, 7)
})
test('defaultSelectedKeys: falls back to the two largest by enrollment', () => {
// A district with no arrests anywhere still has to show something, or the
// chart renders empty with no explanation.
const groups = buildDisplayGroups(
[row('WH', 'F', 900, 0), row('WH', 'M', 1000, 0), row('AM', 'F', 20, 0)],
false,
)
assert.deepEqual(defaultSelectedKeys(groups).sort(), ['WH_F', 'WH_M'])
})
test('defaultSelectedKeys: empty for no groups', () => {
assert.deepEqual(defaultSelectedKeys([]), [])
})
// ——— defaultDiffPair ———
test('defaultDiffPair: the two groups with the most observed arrests', () => {
assert.deepEqual(defaultDiffPair(buildDisplayGroups(CLARK, false)), ['BL_M', 'WH_M'])
})
test('defaultDiffPair: null when fewer than two groups exist', () => {
assert.equal(defaultDiffPair(buildDisplayGroups([row('BL', 'F', 10, 1)], false)), null)
assert.equal(defaultDiffPair([]), null)
})
// ——— enrollByGroupKey ———
test('enrollByGroupKey: maps race×sex keys to enrollment', () => {
const map = enrollByGroupKey(CLARK)
assert.equal(map.BL_M, 12500)
assert.equal(Object.keys(map).length, 8)
})
// ——— displayDraws ———
test('displayDraws: passes raw counts straight through when not pooled', () => {
const counts = { BL_F: [1, 2], BL_M: [3, 4] }
const out = displayDraws(counts, { BL_F: 100, BL_M: 200 }, false)
assert.deepEqual(out.counts, counts)
assert.deepEqual(out.enroll, { BL_F: 100, BL_M: 200 })
})
test('displayDraws: pools by sex when pooling is on', () => {
const out = displayDraws({ BL_F: [1, 2], BL_M: [3, 4] }, { BL_F: 100, BL_M: 200 }, true)
assert.deepEqual(out.counts, { BL: [4, 6] })
assert.deepEqual(out.enroll, { BL: 300 })
})
test('displayDraws: empty structures for missing counts', () => {
const out = displayDraws(null, { BL_F: 100 }, false)
assert.deepEqual(out.counts, {})
assert.deepEqual(out.enroll, {})
})
+104
View File
@@ -0,0 +1,104 @@
/**
* The district-wide arrest total, taken from the posterior predictive draws.
*
* Why this exists: the time-series chart used to draw its band by summing each
* student group's own `count_lower`/`count_upper`. That is not the interval of
* the total. The sum of per-group 97.5th percentiles is the value you would see
* if *every* group simultaneously landed at its own extreme in the same draw,
* which is far less likely than any one group doing so — so the band came out
* systematically too wide. Summing within each draw and taking quantiles of the
* resulting totals answers the actual question: how many arrests does the model
* think this district had?
*
* This is `build_state_summary()`'s method from crdc-arrests/R/summarize_draws.R
* applied on a different axis — sum inside the draw, summarize across draws —
* the same operation `poolBySex` performs for sex pooling.
*
* One caveat to carry into any caption. Upstream, `draw_id` is renumbered per
* write batch, so draw *k* of one group is not the same posterior sample as
* draw *k* of another; measured cross-group correlation is ≈ 0.02. Summing at a
* draw index therefore convolves what are effectively independent draws. That
* is a reasonable model here — these are *posterior predictive* draws, whose
* observation noise is independent across groups by construction and dominates
* the parameter-level covariance — but it does mean the result should be read
* as a predictive total, not as a contrast that preserves parameter
* correlation. It is still much closer to the truth than summing bounds.
*/
import { quantile } from './kde.js'
import { isCompleteDrawSet } from './drawGroups.js'
/** Matches the API's default interval, so the two are directly comparable. */
export const TOTAL_INTERVAL_MASS = 0.95
/**
* District total at each draw index: the sum across every student group.
*
* Returns null unless *every* group has a complete draw set. A group silently
* missing from the sum would undercount the total at every draw and shift the
* whole interval down, which is indistinguishable on screen from a real result.
*
* @param {Record<string, number[]> | null | undefined} countsByGroup
* @param {number} nDraws
* @returns {number[] | null}
*/
export function totalPerDraw(countsByGroup, nDraws) {
const groups = Object.values(countsByGroup || {})
if (!groups.length || !(nDraws > 0)) return null
if (!groups.every((counts) => isCompleteDrawSet(counts, nDraws))) return null
const totals = new Array(nDraws).fill(0)
for (const counts of groups) {
for (let i = 0; i < nDraws; i++) totals[i] += counts[i]
}
return totals
}
/**
* Narrowest interval containing `mass` of the values — the highest-density
* interval, matching what the API stores.
*
* The API's `count_lower`/`count_upper` are HPD bounds (`hpd_bounds_sql` in
* crdc-arrests/R/summarize_draws.R), not equal-tailed quantiles, and the white
* paper's figures use the same. Computing the total's interval the same way
* keeps one definition of "95% interval" on the page: the difference is only
* two or three arrests on a Clark County band of ~50, but the chart's fallback
* mode shows summed HPD bounds, and mixing conventions between the two modes
* would be a distinction with no explanation.
*
* Contiguous by construction, as the R original is. For a strongly bimodal
* posterior that is a simplification, but a count total summed across eight
* groups is unimodal in practice.
*
* @param {number[]} values
* @param {number} mass - e.g. 0.95
* @returns {[number, number]}
*/
export function hpdBounds(values, mass) {
const sorted = [...values].sort((a, b) => a - b)
const n = sorted.length
const span = Math.ceil(mass * n) - 1
if (span <= 0) return [sorted[0], sorted[0]]
if (span >= n - 1) return [sorted[0], sorted[n - 1]]
let best = [sorted[0], sorted[span]]
let bestWidth = sorted[span] - sorted[0]
for (let i = 1; i + span < n; i++) {
const width = sorted[i + span] - sorted[i]
if (width < bestWidth) {
bestWidth = width
best = [sorted[i], sorted[i + span]]
}
}
return best
}
/**
* @param {number[] | null | undefined} totals - district total per draw
* @returns {{lower: number, median: number, upper: number, nDraws: number} | null}
*/
export function totalInterval(totals) {
if (!totals?.length) return null
const [lower, upper] = hpdBounds(totals, TOTAL_INTERVAL_MASS)
return { lower, median: quantile(totals, 0.5), upper, nDraws: totals.length }
}
+139
View File
@@ -0,0 +1,139 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { TOTAL_INTERVAL_MASS, hpdBounds, totalInterval, totalPerDraw } from './districtTotal.js'
// ——— totalPerDraw ———
test('totalPerDraw: sums every group at the same draw index', () => {
const counts = { BL_F: [1, 2, 3], BL_M: [10, 20, 30], WH_F: [100, 200, 300] }
assert.deepEqual(totalPerDraw(counts, 3), [111, 222, 333])
})
test('totalPerDraw: a single group is its own total', () => {
assert.deepEqual(totalPerDraw({ BL_M: [4, 5] }, 2), [4, 5])
})
test('totalPerDraw: null when any group is short of nDraws', () => {
// Summing a 2-draw group into a 3-draw total would silently undercount the
// last draw, biasing the whole interval downward.
assert.equal(totalPerDraw({ BL_F: [1, 2, 3], BL_M: [1, 2] }, 3), null)
})
test('totalPerDraw: null when a group has a hole', () => {
const holey = [1, 2, 3]
delete holey[1]
assert.equal(totalPerDraw({ BL_F: holey }, 3), null)
})
test('totalPerDraw: null for empty or missing input', () => {
assert.equal(totalPerDraw({}, 500), null)
assert.equal(totalPerDraw(null, 500), null)
assert.equal(totalPerDraw({ BL_F: [1] }, 0), null)
})
test('totalPerDraw: does not mutate its input', () => {
const counts = { BL_F: [1, 2], BL_M: [3, 4] }
totalPerDraw(counts, 2)
assert.deepEqual(counts, { BL_F: [1, 2], BL_M: [3, 4] })
})
// ——— totalInterval ———
test('totalInterval: median and 95% bounds from the draw totals', () => {
// Symmetric and unimodal, so the HPD should sit close to the equal-tailed
// 25–975 without being required to equal it.
const totals = []
for (let i = 0; i < 1000; i++) {
totals.push(500 + 120 * (Math.sin(i * 1.7) + Math.sin(i * 0.31) + Math.sin(i * 2.9)) / 3)
}
const iv = totalInterval(totals)
assert.equal(iv.nDraws, 1000)
assert.ok(iv.lower < iv.median && iv.median < iv.upper)
assert.ok(Math.abs(iv.median - 500) < 25, `median ${iv.median}`)
})
test('totalInterval: bounds are the narrowest window covering 95% of draws', () => {
const totals = Array.from({ length: 1000 }, (_, i) => i)
const iv = totalInterval(totals)
const inside = totals.filter((t) => t >= iv.lower && t <= iv.upper).length
assert.ok(inside >= 950, `only ${inside} of 1000 draws inside`)
// A uniform has no denser region, so the narrowest window is ~95% of the range.
assert.ok(iv.upper - iv.lower <= 951, `width ${iv.upper - iv.lower}`)
})
test('hpdBounds: never wider than the equal-tailed interval', () => {
// Right-skewed, which is where the two definitions diverge most.
const skewed = Array.from({ length: 2000 }, (_, i) => Math.round(((i * 7919) % 1000) ** 1.6 / 1000))
const [lo, hi] = hpdBounds(skewed, 0.95)
const sorted = [...skewed].sort((a, b) => a - b)
const etLo = sorted[Math.floor(0.025 * (sorted.length - 1))]
const etHi = sorted[Math.ceil(0.975 * (sorted.length - 1))]
assert.ok(hi - lo <= etHi - etLo, `hpd ${hi - lo} vs equal-tailed ${etHi - etLo}`)
})
test('hpdBounds: stays in the dense region when it holds enough mass', () => {
// 970 draws at 0-9 and 30 stragglers out at 500. The cluster alone covers 97%,
// so the narrowest 95% window fits inside it and the far tail is excluded.
const values = [
...Array.from({ length: 970 }, (_, i) => i % 10),
...Array.from({ length: 30 }, () => 500),
]
const [lo, hi] = hpdBounds(values, 0.95)
assert.equal(lo, 0)
assert.ok(hi < 500, `upper bound ${hi} should exclude the far cluster`)
})
test('hpdBounds: still reaches the tail when the cluster is too small', () => {
// The mirror case, and the honest one: 90% in the cluster cannot cover a 95%
// interval, so the bound must extend outward rather than under-covering.
const values = [
...Array.from({ length: 900 }, (_, i) => i % 10),
...Array.from({ length: 100 }, () => 500),
]
const [, hi] = hpdBounds(values, 0.95)
assert.equal(hi, 500)
})
test('hpdBounds: degenerate input collapses to a point', () => {
assert.deepEqual(hpdBounds(new Array(100).fill(7), 0.95), [7, 7])
})
test('hpdBounds: does not mutate its input', () => {
const values = [5, 1, 3]
hpdBounds(values, 0.95)
assert.deepEqual(values, [5, 1, 3])
})
test('totalInterval: uses a 95% mass, matching the API convention', () => {
assert.equal(TOTAL_INTERVAL_MASS, 0.95)
})
test('totalInterval: lower <= median <= upper', () => {
const totals = Array.from({ length: 500 }, (_, i) => (i * 7919) % 331)
const iv = totalInterval(totals)
assert.ok(iv.lower <= iv.median && iv.median <= iv.upper)
})
test('totalInterval: a degenerate total collapses to a point', () => {
const iv = totalInterval(new Array(500).fill(12))
assert.deepEqual([iv.lower, iv.median, iv.upper], [12, 12, 12])
})
test('totalInterval: null for empty or missing totals', () => {
assert.equal(totalInterval([]), null)
assert.equal(totalInterval(null), null)
})
test('totalInterval: is narrower than summing the groups own bounds', () => {
// The property that motivates this whole path. Two independent groups, each
// roughly uniform on 0..100: summing each group's 97.5th percentile gives
// ~200, but the 97.5th percentile of the *sum* is well below that, because
// both groups landing at their extreme in the same draw is rare.
const n = 4000
const a = Array.from({ length: n }, (_, i) => (i * 37) % 101)
const b = Array.from({ length: n }, (_, i) => (i * 61) % 101)
const iv = totalInterval(totalPerDraw({ A: a, B: b }, n))
const summedBounds = { lower: 0, upper: 100 + 100 }
assert.ok(iv.upper < summedBounds.upper, `sum-of-quantiles ${summedBounds.upper} vs quantile-of-sum ${iv.upper}`)
assert.ok(iv.upper - iv.lower < summedBounds.upper - summedBounds.lower)
})
+72
View File
@@ -0,0 +1,72 @@
/**
* Shared key format and coverage checks for the per-group posterior draws
* returned by `useDrawDistribution`. Both the hook (which builds the map) and
* the charts (which read it, and decide whether to claim "real draws") go
* through here so the key format lives in exactly one place.
*
* Draw arrays are *counts indexed by draw* — `counts[draw_id - 1]` — not rates
* and not push-ordered. That makes row order out of DuckDB irrelevant, and it
* makes a missing draw a hole rather than a silently shorter array, which is
* why `isCompleteDrawSet` checks every index instead of trusting `length`.
*/
/**
* Canonical key for one group in a draws map.
*
* With a sex, this is one race×sex cell (`'BL_F'`). Without one — the pooled
* mode sparse districts fall back to — it is the race alone (`'BL'`). The two
* shapes can never collide, so a single map type serves both modes.
*
* @param {string} race
* @param {string} [sex] - omit (or pass null/'') for a sex-pooled group
* @returns {string}
*/
export function groupKey(race, sex) {
return sex ? `${race}_${sex}` : `${race}`
}
/**
* True when `counts` holds a finite value at every draw index 0…nDraws-1.
*
* A partial group must fall back rather than render a short draw set: a group
* present for only 300 of 500 draws would otherwise get a density and an
* interval computed off a biased subsample, indistinguishable on screen from
* a complete one.
*
* @param {number[] | null | undefined} counts
* @param {number} nDraws
* @returns {boolean}
*/
export function isCompleteDrawSet(counts, nDraws) {
if (!Array.isArray(counts) || !(nDraws > 0) || counts.length !== nDraws) return false
for (let i = 0; i < nDraws; i++) {
if (!Number.isFinite(counts[i])) return false
}
return true
}
/**
* True only when *every* group in `groups` has a complete draw array.
*
* `useDrawDistribution`'s `status` is an any-group signal: it reports 'ready'
* as soon as one group has real draws. Charts fall back per-group, so the
* chart-wide "these are real posterior draws" note/caption must be gated on
* complete coverage instead — a group can be missing because its (LEAID, RACE,
* SEX) isn't in the parquet shard.
*
* Returns false for an empty group list (nothing rendered means nothing to
* claim) and for a null map (loading, or the fetch failed outright).
*
* @param {Record<string, number[]> | null | undefined} drawsByGroup
* @param {Array<{race: string, sex?: string}>} groups - the groups a chart is
* actually rendering, not everything the API returned. Omit `sex` for pooled
* groups.
* @returns {boolean}
*/
export function hasDrawsForAll(drawsByGroup, groups) {
if (!drawsByGroup || !groups?.length) return false
return groups.every((g) => {
const counts = drawsByGroup[groupKey(g.race, g.sex)]
return isCompleteDrawSet(counts, counts?.length ?? 0)
})
}
+100
View File
@@ -0,0 +1,100 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { groupKey, hasDrawsForAll, isCompleteDrawSet } from './drawGroups.js'
test('groupKey: joins race and sex with an underscore', () => {
assert.equal(groupKey('BL', 'F'), 'BL_F')
})
test('groupKey: returns the race alone for a pooled group', () => {
// Pooled rows are keyed by race only, so the same helper builds both key
// shapes and nothing downstream has to know which mode it is in.
assert.equal(groupKey('BL'), 'BL')
assert.equal(groupKey('BL', null), 'BL')
assert.equal(groupKey('BL', ''), 'BL')
})
test('groupKey: a pooled key never collides with an unpooled one', () => {
assert.notEqual(groupKey('BL'), groupKey('BL', 'F'))
assert.notEqual(groupKey('BL'), groupKey('BL', 'M'))
})
test('hasDrawsForAll: true when every rendered group has draws', () => {
const map = { WH_F: [1, 2], BL_F: [3] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), true)
})
test('hasDrawsForAll: false when one rendered group is missing', () => {
// The any-group 'ready' status would still be true here — this is exactly the
// case where the chart must keep showing the approximation note.
const map = { WH_F: [1, 2] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), false)
})
test('hasDrawsForAll: false when a rendered group has an empty draw array', () => {
const map = { WH_F: [1, 2], BL_F: [] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), false)
})
test('hasDrawsForAll: false for a null map (loading or failed fetch)', () => {
assert.equal(hasDrawsForAll(null, [{ race: 'WH', sex: 'F' }]), false)
assert.equal(hasDrawsForAll(undefined, [{ race: 'WH', sex: 'F' }]), false)
})
test('hasDrawsForAll: false for an empty or missing group list', () => {
assert.equal(hasDrawsForAll({ WH_F: [1] }, []), false)
assert.equal(hasDrawsForAll({ WH_F: [1] }, undefined), false)
})
test('hasDrawsForAll: ignores groups the chart is not rendering', () => {
// Extra keys in the map (e.g. a race outside RACE_ORDER) must not block the
// claim for the groups actually on screen.
const map = { WH_F: [1], BL_F: [2], AS_F: [3] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH', sex: 'F' }, { race: 'BL', sex: 'F' }]), true)
})
test('hasDrawsForAll: works on pooled (race-only) groups', () => {
const map = { WH: [1], BL: [2] }
assert.equal(hasDrawsForAll(map, [{ race: 'WH' }, { race: 'BL' }]), true)
assert.equal(hasDrawsForAll(map, [{ race: 'WH' }, { race: 'HI' }]), false)
})
test('hasDrawsForAll: false when a count array has a hole', () => {
// Counts are indexed by draw_id, so a group missing draw 2 leaves a hole
// rather than a short array. length alone would call this complete.
const holey = [1, 2, 3]
delete holey[1]
assert.equal(hasDrawsForAll({ BL_F: holey }, [{ race: 'BL', sex: 'F' }]), false)
})
// ——— isCompleteDrawSet ———
test('isCompleteDrawSet: true for a dense array of the expected length', () => {
assert.equal(isCompleteDrawSet([0, 1, 2], 3), true)
})
test('isCompleteDrawSet: false when the array is shorter than nDraws', () => {
assert.equal(isCompleteDrawSet([0, 1], 3), false)
})
test('isCompleteDrawSet: false when the array is longer than nDraws', () => {
assert.equal(isCompleteDrawSet([0, 1, 2, 3], 3), false)
})
test('isCompleteDrawSet: false when a draw index was never filled', () => {
const holey = new Array(3)
holey[0] = 1
holey[2] = 3
assert.equal(isCompleteDrawSet(holey, 3), false)
})
test('isCompleteDrawSet: false for a non-finite entry', () => {
assert.equal(isCompleteDrawSet([1, NaN, 3], 3), false)
assert.equal(isCompleteDrawSet([1, Infinity, 3], 3), false)
})
test('isCompleteDrawSet: false for a missing array or a zero-draw expectation', () => {
assert.equal(isCompleteDrawSet(null, 3), false)
assert.equal(isCompleteDrawSet(undefined, 3), false)
assert.equal(isCompleteDrawSet([], 0), false)
})
+26
View File
@@ -0,0 +1,26 @@
/**
* Lazy-initialized singleton AsyncDuckDB instance running the single-threaded
* MVP wasm bundle only — never eh/coi, which need Cross-Origin-Opener-Policy /
* Cross-Origin-Embedder-Policy response headers this static host doesn't send.
* See docs/superpowers/specs/2026-08-11-empirical-draws-wasm-design.md.
*/
let dbPromise = null
/** @returns {Promise<import('@duckdb/duckdb-wasm').AsyncDuckDB>} */
export function getDb() {
if (!dbPromise) dbPromise = initDb()
return dbPromise
}
async function initDb() {
const duckdb = await import('@duckdb/duckdb-wasm')
const mvpWorkerUrl = (await import('@duckdb/duckdb-wasm/dist/duckdb-browser-mvp.worker.js?url')).default
const mvpWasmUrl = (await import('@duckdb/duckdb-wasm/dist/duckdb-mvp.wasm?url')).default
const worker = new Worker(mvpWorkerUrl)
const logger = new duckdb.ConsoleLogger(duckdb.LogLevel.WARNING)
const db = new duckdb.AsyncDuckDB(logger, worker)
await db.instantiate(mvpWasmUrl, null)
return db
}
+57
View File
@@ -0,0 +1,57 @@
/**
* The difference in modelled arrest rate between two student groups, taken
* draw by draw — the quantity behind the white paper's Fig 7
* (`wp_fig_group_difference`).
*
* Read the wording carefully before writing a caption from these numbers.
* Upstream, `draw_id` is renumbered 1–500 per write batch, and a district's
* groups land in different batches, so draw *k* of group A and draw *k* of
* group B are not the same parameter draw. Measured correlation between two
* groups' `pred` in Clark County was ≈ 0.02 even within a batch — observation
* noise from `posterior_predict` dominates. The published figure has exactly
* the same property, so the app matches the paper; what neither can claim is a
* paired-parameter contrast. Always say **posterior predictive draws**, never
* "paired parameter draws".
*/
import { quantile } from './kde.js'
/**
* Per-draw difference in rate per 1,000: group A minus group B.
*
* Returns an empty array rather than a truncated one when the two draw sets
* disagree in length — pairing draw 3 of one group with draw 7 of another would
* fabricate a difference distribution out of unrelated draws.
*
* @param {number[] | null | undefined} countsA - predicted counts, indexed by draw
* @param {number | null | undefined} enrollA
* @param {number[] | null | undefined} countsB
* @param {number | null | undefined} enrollB
* @returns {number[]}
*/
export function differenceRates(countsA, enrollA, countsB, enrollB) {
if (!countsA?.length || !countsB?.length) return []
if (countsA.length !== countsB.length) return []
if (!(enrollA > 0) || !(enrollB > 0)) return []
return countsA.map((a, i) => (a / enrollA) * 1000 - (countsB[i] / enrollB) * 1000)
}
/**
* @param {number[] | null | undefined} deltas
* @returns {{n: number, prGreater: number, median: number, lower80: number,
* upper80: number, lower95: number, upper95: number} | null}
* `prGreater` counts draws **strictly** above zero, so a group whose every
* draw ties reads as 0%, not 100%.
*/
export function differenceSummary(deltas) {
if (!deltas?.length) return null
return {
n: deltas.length,
prGreater: deltas.filter((d) => d > 0).length / deltas.length,
median: quantile(deltas, 0.5),
lower80: quantile(deltas, 0.1),
upper80: quantile(deltas, 0.9),
lower95: quantile(deltas, 0.025),
upper95: quantile(deltas, 0.975),
}
}
+77
View File
@@ -0,0 +1,77 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { differenceRates, differenceSummary } from './groupDifference.js'
// ——— differenceRates ———
test('differenceRates: subtracts rates at the same draw index', () => {
// A: 2 and 4 arrests in 1,000 students → 2 and 4 per 1,000
// B: 1 and 1 arrests in 500 students → 2 and 2 per 1,000
const deltas = differenceRates([2, 4], 1000, [1, 1], 500)
assert.deepEqual(deltas, [0, 2])
})
test('differenceRates: differences can be negative', () => {
assert.deepEqual(differenceRates([0], 1000, [5], 1000), [-5])
})
test('differenceRates: empty when the draw sets are different lengths', () => {
// Pairing draw 3 of one group with draw 7 of another would fabricate a
// difference distribution out of unrelated draws.
assert.deepEqual(differenceRates([1, 2, 3], 1000, [1, 2], 1000), [])
})
test('differenceRates: empty when either denominator is missing', () => {
assert.deepEqual(differenceRates([1], 0, [1], 1000), [])
assert.deepEqual(differenceRates([1], 1000, [1], undefined), [])
})
test('differenceRates: empty when either draw set is missing', () => {
assert.deepEqual(differenceRates(null, 1000, [1], 1000), [])
assert.deepEqual(differenceRates([1], 1000, [], 1000), [])
})
test('differenceRates: does not mutate its inputs', () => {
const a = [1, 2]
const b = [3, 4]
differenceRates(a, 1000, b, 1000)
assert.deepEqual(a, [1, 2])
assert.deepEqual(b, [3, 4])
})
// ——— differenceSummary ———
test('differenceSummary: Pr(delta > 0) is the share of draws strictly above zero', () => {
const summary = differenceSummary([-1, 0, 1, 2])
assert.equal(summary.prGreater, 0.5)
})
test('differenceSummary: a draw of exactly zero does not count as greater', () => {
assert.equal(differenceSummary([0, 0, 0, 0]).prGreater, 0)
})
test('differenceSummary: reports the median and both interval widths', () => {
const deltas = Array.from({ length: 101 }, (_, i) => i) // 0…100
const summary = differenceSummary(deltas)
assert.equal(summary.median, 50)
assert.equal(summary.lower80, 10)
assert.equal(summary.upper80, 90)
assert.equal(summary.lower95, 2.5)
assert.equal(summary.upper95, 97.5)
})
test('differenceSummary: the 95% interval contains the 80% interval', () => {
const deltas = Array.from({ length: 500 }, (_, i) => Math.sin(i) * 4)
const s = differenceSummary(deltas)
assert.ok(s.lower95 <= s.lower80)
assert.ok(s.upper95 >= s.upper80)
})
test('differenceSummary: reports the draw count it summarized', () => {
assert.equal(differenceSummary([1, 2, 3]).n, 3)
})
test('differenceSummary: null for an empty or missing set', () => {
assert.equal(differenceSummary([]), null)
assert.equal(differenceSummary(undefined), null)
})
+110
View File
@@ -0,0 +1,110 @@
/**
* Empirical density utilities for posterior draw arrays — Gaussian KDE with
* Silverman's rule-of-thumb bandwidth, plus a linear-interpolated quantile.
* Used in place of distributionApprox.js's analytic fitSkewedInterval/
* densityCurve approximation whenever real posterior draws are available.
*/
const SQRT_2PI = Math.sqrt(2 * Math.PI)
// Absolute last-resort floor, used only when no plotting domain is known.
const MIN_BANDWIDTH = 1e-3
// Bandwidth is clamped relative to the plotting domain, not to absolute units,
// because these draws are rate-per-1,000 values whose scale varies by orders of
// magnitude between districts. domainWidth/50 is slightly wider than one render
// step at the charts' n = 60 (step = domainWidth/59), so a degenerate draw set
// (e.g. 500 identical zeros, common for small districts) resolves as a narrow
// bump instead of a delta spike that flattens every other ridge sharing the
// column's maxPdf. It is also loose enough not to bind on an ordinary posterior:
// a spread wider than ~8% of the domain keeps its own Silverman bandwidth.
// domainWidth/6 stops a handful of extreme draws from inflating sd until the
// curve is a flat line.
const BANDWIDTH_FLOOR_DIVISOR = 50
const BANDWIDTH_CEILING_DIVISOR = 6
/**
* Linear-interpolated quantile (R type-7). Does not mutate `draws`.
* @param {number[]} draws
* @param {number} p - probability in [0, 1]
* @returns {number}
*/
export function quantile(draws, p) {
const sorted = [...draws].sort((a, b) => a - b)
const idx = p * (sorted.length - 1)
const lo = Math.floor(idx)
const hi = Math.ceil(idx)
if (lo === hi) return sorted[lo]
const frac = idx - lo
return sorted[lo] * (1 - frac) + sorted[hi] * frac
}
function standardDeviation(draws) {
const n = draws.length
const mean = draws.reduce((sum, d) => sum + d, 0) / n
const variance = draws.reduce((sum, d) => sum + (d - mean) ** 2, 0) / (n - 1)
return Math.sqrt(variance)
}
/**
* Silverman's rule-of-thumb bandwidth (robust variant using the smallest
* *positive* spread estimate among sd and IQR/1.34), clamped to a fraction of
* the plotting domain.
*
* Both ends of the clamp matter for the zero-inflated count posteriors small
* districts produce. Without the floor, an all-identical draw set (sd = IQR = 0)
* collapses to a delta-function spike. Without the ceiling, a group whose IQR is
* 0 (the normal case when most draws are 0) falls back to raw sd, which a
* handful of extreme draws inflates until the ridge is a featureless flat line.
*
* @param {number[]} draws
* @param {number} [domainWidth] - width of the x-range the curve will be drawn
* over. Omit only when no domain is known; the clamp then degrades to the
* absolute MIN_BANDWIDTH floor.
* @returns {number}
*/
export function silvermanBandwidth(draws, domainWidth = 0) {
const width = domainWidth > 0 ? domainWidth : 0
const floor = width ? width / BANDWIDTH_FLOOR_DIVISOR : MIN_BANDWIDTH
const ceiling = width ? width / BANDWIDTH_CEILING_DIVISOR : Infinity
const n = draws.length
if (n < 2) return floor
const sd = standardDeviation(draws)
const iqr = quantile(draws, 0.75) - quantile(draws, 0.25)
// Degrade gracefully: keep the robust rule when IQR is informative, use sd
// when it isn't, and let the floor handle a fully degenerate draw set —
// rather than treating a zero spread as "no estimate available".
const candidates = [sd, iqr / 1.34].filter((v) => v > 0)
const spread = candidates.length ? Math.min(...candidates) : 0
const raw = 0.9 * spread * Math.pow(n, -0.2)
return Math.min(Math.max(raw, floor), ceiling)
}
/**
* n evenly spaced {x, y} points of a Gaussian KDE over `draws` — same shape
* contract as distributionApprox.js's densityCurve, so chart code can switch
* between the two without changing its rendering path.
* @param {number[]} draws
* @param {{min?: number, max?: number, n?: number}} [options]
* @returns {Array<{x: number, y: number}>}
*/
export function kdeCurve(draws, { min = 0, max, n = 60 } = {}) {
const hi = max ?? Math.max(...draws) * 1.1
// Guarded against a degenerate/inverted domain so the bandwidth clamp can
// never be handed a negative width.
const domainWidth = Math.max(hi - min, 0)
const h = silvermanBandwidth(draws, domainWidth)
const step = (hi - min) / (n - 1)
const points = []
for (let i = 0; i < n; i++) {
const x = min + step * i
let sum = 0
for (const d of draws) {
const z = (x - d) / h
sum += Math.exp(-0.5 * z * z) / SQRT_2PI
}
points.push({ x, y: sum / (draws.length * h) })
}
return points
}
+107
View File
@@ -0,0 +1,107 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { quantile, silvermanBandwidth, kdeCurve } from './kde.js'
test('quantile: median of an odd-length array', () => {
assert.equal(quantile([3, 1, 2], 0.5), 2)
})
test('quantile: linear interpolation between two ranks', () => {
// sorted: [10, 20, 30, 40] — p=0.25 -> index 0.75 -> interpolate 10..20
assert.equal(quantile([40, 10, 30, 20], 0.25), 17.5)
})
test('quantile: does not mutate its input array', () => {
const input = [5, 3, 4, 1, 2]
quantile(input, 0.5)
assert.deepEqual(input, [5, 3, 4, 1, 2])
})
test('silvermanBandwidth: positive, finite floor for identical draws', () => {
const h = silvermanBandwidth([7, 7, 7, 7, 7], 10)
assert.ok(h > 0 && Number.isFinite(h))
assert.equal(h, 10 / 50, 'degenerate spread falls back to the domain-relative floor')
})
test('silvermanBandwidth: positive, finite floor for a single draw', () => {
const h = silvermanBandwidth([7], 10)
assert.ok(h > 0 && Number.isFinite(h))
assert.equal(h, 10 / 50)
})
test('silvermanBandwidth: positive, finite floor when no domain is supplied', () => {
assert.ok(silvermanBandwidth([7, 7, 7, 7, 7]) > 0)
assert.ok(silvermanBandwidth([7]) > 0)
})
test('silvermanBandwidth: clamped to [domainWidth/50, domainWidth/6]', () => {
const domainWidth = 30
// sd is inflated by 50 extreme draws; without the ceiling this is ~15.6.
const outlierDraws = [...Array(450).fill(0), ...Array(50).fill(200)]
assert.equal(silvermanBandwidth(outlierDraws, domainWidth), domainWidth / 6)
// Fully degenerate: sd = IQR = 0.
assert.equal(silvermanBandwidth(Array(500).fill(0), domainWidth), domainWidth / 50)
})
test('silvermanBandwidth: leaves an ordinary spread untouched by the clamp', () => {
const draws = Array.from({ length: 500 }, (_, i) => 3 + Math.sin(i) * 1.2 + (i % 7) * 0.15)
const h = silvermanBandwidth(draws, 10)
assert.ok(h > 10 / 50 && h < 10 / 6, `expected an unclamped bandwidth, got ${h}`)
})
test('kdeCurve: all-identical draws do not produce a delta-function spike', () => {
// 500 identical zeros is the common case for a small district's rare-event
// count posterior. The old absolute 1e-3 floor gave maxY ~399 here (peak
// ~3989x a uniform density over the same domain), which flattened every other
// ridge sharing the column's maxPdf to sub-pixel height.
const domainWidth = 10
const curve = kdeCurve(Array(500).fill(0), { min: 0, max: domainWidth, n: 60 })
const maxY = Math.max(...curve.map((p) => p.y))
assert.ok(Number.isFinite(maxY) && maxY > 0)
assert.ok(
maxY * domainWidth < 25,
`peak density should stay within ~25x a uniform density over the domain, got ${maxY * domainWidth}x`,
)
})
test('kdeCurve: zero-inflated draws keep their shape (not over-smoothed to flat)', () => {
const domainWidth = 10
const draws = [...Array(450).fill(0), ...Array(50).fill(1)]
const curve = kdeCurve(draws, { min: 0, max: domainWidth, n: 60 })
const ys = curve.map((p) => p.y)
const maxY = Math.max(...ys)
const meanY = ys.reduce((sum, y) => sum + y, 0) / ys.length
assert.ok(maxY / meanY > 3, `expected a peaked curve, got peak/mean ${maxY / meanY}`)
})
test('kdeCurve: outlier-inflated sd does not flatten the curve', () => {
// IQR is 0 here (most draws are 0), so the rule falls back to sd — which these
// 50 extreme draws inflate to ~60. Unclamped that gives h ~15.6 on a domain of
// 30, i.e. peak/mean ~1.6: a near-flat line claiming maximal uncertainty.
const domainWidth = 30
const draws = [...Array(450).fill(0), ...Array(50).fill(200)]
const curve = kdeCurve(draws, { min: 0, max: domainWidth, n: 60 })
const ys = curve.map((p) => p.y)
const maxY = Math.max(...ys)
const meanY = ys.reduce((sum, y) => sum + y, 0) / ys.length
assert.ok(maxY / meanY > 3, `expected a peaked curve, got peak/mean ${maxY / meanY}`)
})
test('kdeCurve: returns n points spanning [min, max]', () => {
const draws = [1, 2, 2, 3, 4, 5, 5, 5, 6, 8]
const curve = kdeCurve(draws, { min: 0, max: 10, n: 60 })
assert.equal(curve.length, 60)
assert.equal(curve[0].x, 0)
assert.ok(Math.abs(curve[curve.length - 1].x - 10) < 1e-9)
})
test('kdeCurve: density integrates to ~1 over a wide domain (trapezoidal check)', () => {
const draws = [1, 2, 2, 3, 4, 5, 5, 5, 6, 8]
const curve = kdeCurve(draws, { min: -20, max: 30, n: 2000 })
let area = 0
for (let i = 1; i < curve.length; i++) {
const dx = curve[i].x - curve[i - 1].x
area += (dx * (curve[i].y + curve[i - 1].y)) / 2
}
assert.ok(Math.abs(area - 1) < 0.01, `expected area ~1, got ${area}`)
})
+81
View File
@@ -0,0 +1,81 @@
/**
* Collision-free placement for direct labels on a shared axis.
*
* The density panel labels each curve at its own peak rather than shipping a
* legend, which only works if the labels don't collide — and in this app they
* collide constantly, because the interesting districts are exactly the ones
* where several groups have similar rates. Worse, each density is normalized to
* its own peak height, so every curve peaks at the *same* y and a naive
* placement puts every label on one line: three overlapping groups rendered as
* "HiWhite:F: Black F".
*
* So labels are packed into horizontal lanes: first-fit by x, dropping to a new
* lane only when the previous one is occupied at that position. Groups that are
* far apart still share a lane, which keeps the common case compact.
*
* Pure geometry — no DOM, no measurement. SVG text can't be measured before
* render, so widths are estimated from character count; the estimate is
* deliberately generous so labels err toward extra separation rather than
* overlap.
*/
// Mean advance width of a glyph as a fraction of font size, for the app's sans
// stack at the weights these labels use. Slightly over the true average so the
// packing errs toward separation.
const CHAR_WIDTH_RATIO = 0.62
/**
* @param {string} text
* @param {number} fontPx
* @returns {number} approximate rendered width in px
*/
export function estimateTextWidth(text, fontPx) {
return (text || '').length * fontPx * CHAR_WIDTH_RATIO
}
/**
* @param {Array<{key: string, x: number, text: string}>} items - one per label,
* `x` being the point it wants to sit above (a curve's peak).
* @param {{min: number, max: number, fontPx: number, gap?: number}} options -
* `min`/`max` are the plot's horizontal bounds; labels are kept inside them.
* @returns {{labels: Array<{key: string, text: string, x: number,
* anchor: 'start'|'middle'|'end', width: number, left: number, right: number,
* lane: number}>, lanes: number}}
* `labels` is in input order; `lanes` is how many rows the caller must
* reserve above the plot.
*/
export function layoutPeakLabels(items, { min, max, fontPx, gap = 6 } = {}) {
if (!items?.length) return { labels: [], lanes: 0 }
const placed = items.map((item) => {
const width = estimateTextWidth(item.text, fontPx)
const half = width / 2
// Anchor outward near the edges so a centred label can't overflow the plot.
let anchor = 'middle'
if (item.x - half < min) anchor = 'start'
else if (item.x + half > max) anchor = 'end'
// Clamp the anchor point itself, so a peak clipped to the axis edge still
// yields a label fully inside the frame.
let x = item.x
if (anchor === 'start') x = Math.max(min, Math.min(x, max - width))
else if (anchor === 'end') x = Math.min(max, Math.max(x, min + width))
else x = Math.min(Math.max(x, min + half), max - half)
const left = anchor === 'start' ? x : anchor === 'end' ? x - width : x - half
return { ...item, width, anchor, x, left, right: left + width, lane: 0 }
})
// First-fit by left edge. Sorting only decides lane order; the returned array
// keeps the caller's original order.
const laneRightEdge = []
for (const label of [...placed].sort((a, b) => a.left - b.left)) {
let lane = 0
while (lane < laneRightEdge.length && laneRightEdge[lane] + gap > label.left) lane++
laneRightEdge[lane] = label.right
label.lane = lane
}
return { labels: placed, lanes: laneRightEdge.length }
}
+115
View File
@@ -0,0 +1,115 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { estimateTextWidth, layoutPeakLabels } from './labelLayout.js'
const BOUNDS = { min: 0, max: 400, fontPx: 10 }
test('estimateTextWidth: grows with text length and font size', () => {
assert.ok(estimateTextWidth('AB', 10) > estimateTextWidth('A', 10))
assert.ok(estimateTextWidth('ABC', 20) > estimateTextWidth('ABC', 10))
assert.ok(estimateTextWidth('', 10) >= 0)
})
test('layoutPeakLabels: well-separated labels all sit on lane 0', () => {
const out = layoutPeakLabels(
[{ key: 'a', x: 20, text: 'A' }, { key: 'b', x: 200, text: 'B' }, { key: 'c', x: 380, text: 'C' }],
BOUNDS,
)
assert.deepEqual(out.labels.map((l) => l.lane), [0, 0, 0])
assert.equal(out.lanes, 1)
})
test('layoutPeakLabels: overlapping labels are pushed to separate lanes', () => {
// This is the failing case from the live site: three densities peaking at
// nearly the same rate rendered as "HiWhite:F: Black F".
const out = layoutPeakLabels(
[
{ key: 'hi', x: 100, text: 'Hispanic F' },
{ key: 'wh', x: 104, text: 'White F' },
{ key: 'bl', x: 108, text: 'Black F' },
],
BOUNDS,
)
const lanes = out.labels.map((l) => l.lane).sort()
assert.deepEqual(lanes, [0, 1, 2])
assert.equal(out.lanes, 3)
})
test('layoutPeakLabels: identical positions never share a lane', () => {
const out = layoutPeakLabels(
[
{ key: 'a', x: 200, text: 'Hispanic F' },
{ key: 'b', x: 200, text: 'Hispanic M' },
],
BOUNDS,
)
assert.notEqual(out.labels[0].lane, out.labels[1].lane)
})
test('layoutPeakLabels: a lane is reused once there is horizontal room', () => {
const out = layoutPeakLabels(
[
{ key: 'a', x: 20, text: 'A' },
{ key: 'b', x: 24, text: 'B' },
{ key: 'c', x: 380, text: 'C' },
],
BOUNDS,
)
const byKey = Object.fromEntries(out.labels.map((l) => [l.key, l.lane]))
assert.equal(byKey.a, 0)
assert.equal(byKey.b, 1)
// 'c' is far away, so it drops back to the first lane rather than stacking.
assert.equal(byKey.c, 0)
})
test('layoutPeakLabels: anchors outward at the edges so text stays in frame', () => {
const out = layoutPeakLabels(
[{ key: 'l', x: 0, text: 'Hispanic F' }, { key: 'r', x: 400, text: 'Hispanic F' }],
BOUNDS,
)
const byKey = Object.fromEntries(out.labels.map((l) => [l.key, l]))
assert.equal(byKey.l.anchor, 'start')
assert.equal(byKey.r.anchor, 'end')
})
test('layoutPeakLabels: a mid-plot label stays centred on its peak', () => {
const out = layoutPeakLabels([{ key: 'm', x: 200, text: 'Black F' }], BOUNDS)
assert.equal(out.labels[0].anchor, 'middle')
assert.equal(out.labels[0].x, 200)
})
test('layoutPeakLabels: no label extends outside the plot bounds', () => {
const out = layoutPeakLabels(
[
{ key: 'l', x: -50, text: 'American Indian / Alaska Native' },
{ key: 'r', x: 900, text: 'American Indian / Alaska Native' },
],
BOUNDS,
)
for (const l of out.labels) {
assert.ok(l.left >= BOUNDS.min - 0.01, `${l.key} left ${l.left} < ${BOUNDS.min}`)
assert.ok(l.right <= BOUNDS.max + 0.01, `${l.key} right ${l.right} > ${BOUNDS.max}`)
}
})
test('layoutPeakLabels: preserves input order in the output', () => {
// Lane assignment sorts internally; callers still key off their own order.
const out = layoutPeakLabels(
[{ key: 'z', x: 300, text: 'Z' }, { key: 'a', x: 10, text: 'A' }],
BOUNDS,
)
assert.deepEqual(out.labels.map((l) => l.key), ['z', 'a'])
})
test('layoutPeakLabels: empty input yields zero lanes', () => {
const out = layoutPeakLabels([], BOUNDS)
assert.deepEqual(out.labels, [])
assert.equal(out.lanes, 0)
})
test('layoutPeakLabels: wider gap forces more lanes', () => {
const items = [{ key: 'a', x: 100, text: 'A' }, { key: 'b', x: 130, text: 'B' }]
const tight = layoutPeakLabels(items, { ...BOUNDS, gap: 0 })
const loose = layoutPeakLabels(items, { ...BOUNDS, gap: 40 })
assert.ok(loose.lanes > tight.lanes)
})
+6
View File
@@ -0,0 +1,6 @@
import { MODEL_QUADRANTS } from '../hooks/useApi.js'
/** model id -> display label, derived from the single source of truth in useApi.js. */
export const MODEL_QUADRANT_LABEL = Object.fromEntries(
MODEL_QUADRANTS.map((q) => [q.model, q.label])
)
+40
View File
@@ -0,0 +1,40 @@
/**
* "Nice numbers" axis tick generator (Heckbert 1990) — produces clean,
* human-friendly tick values (0 / 50 / 100 / 150 rather than 0 / 37 / 74 /
* 111, an artifact of slicing frac*max into equal fractions) for any linear
* axis, from just the data's max value.
*/
function niceNumber(range, round) {
const exponent = Math.floor(Math.log10(range))
const fraction = range / 10 ** exponent
let niceFraction
if (round) {
if (fraction < 1.5) niceFraction = 1
else if (fraction < 3) niceFraction = 2
else if (fraction < 7) niceFraction = 5
else niceFraction = 10
} else {
if (fraction <= 1) niceFraction = 1
else if (fraction <= 2) niceFraction = 2
else if (fraction <= 5) niceFraction = 5
else niceFraction = 10
}
return niceFraction * 10 ** exponent
}
/**
* @param {number} maxValue - the largest value the axis must cover
* @param {number} tickCount - approximate desired number of ticks
* @returns {{ ticks: number[], niceMax: number }}
*/
export function niceTicks(maxValue, tickCount = 5) {
if (!(maxValue > 0)) return { ticks: [0], niceMax: 1 }
const niceRange = niceNumber(maxValue, false)
const niceStep = niceNumber(niceRange / (tickCount - 1), true)
const niceMax = Math.ceil(maxValue / niceStep) * niceStep
const ticks = []
for (let v = 0; v <= niceMax + niceStep / 2; v += niceStep) {
ticks.push(Math.round(v * 1e6) / 1e6)
}
return { ticks, niceMax }
}
+115
View File
@@ -0,0 +1,115 @@
/**
* Sex pooling for sparse districts.
*
* This is the project's own house method applied on a different axis:
* `build_state_summary()` in `crdc-arrests/R/summarize_draws.R` pools across
* LEAs by summing `pred` and `stu_enroll` *within each draw* and only then
* summarizing across draws. Pooling Female + Male inside one district is the
* identical operation — sum the numerators draw-by-draw, sum the denominators,
* divide once. Summing rates instead would weight a 53-student cell the same as
* a 3,000-student one.
*
* Pure functions; no React, no fetching.
*/
import { groupKey } from './drawGroups.js'
/**
* Whole-district trigger for sex pooling: fewer than this many observed arrests
* across all eight race×sex cells.
*
* 20 is chosen so that the average cell carries at least ~2 arrests before the
* app is willing to show eight separate posteriors. Below it, most cells are
* zero-count and their posteriors are dominated by the prior and by
* observation noise, so eight ridges read as eight findings when they are
* really one. The threshold is on the *district* total rather than per cell so
* the table's row set doesn't change shape group by group.
*/
export const POOL_BY_SEX_ARREST_THRESHOLD = 20
/**
* True when a district is sparse enough that its eight race×sex cells should
* collapse to four race rows.
*
* Returns false for an empty row set: no rows is not "a sparse district", it is
* no district data at all, and firing the pooling banner there would explain a
* rule that never applied.
*
* @param {Array<{observed_arrests?: number|null}>} rows - estimate rows for one
* district/model/year, one per race×sex cell.
* @returns {boolean}
*/
export function shouldPoolBySex(rows) {
if (!rows?.length) return false
const total = rows.reduce((sum, r) => sum + (r.observed_arrests || 0), 0)
return total < POOL_BY_SEX_ARREST_THRESHOLD
}
/**
* Per-1,000 rate array from an array of predicted counts.
*
* Returns an empty array when the denominator is missing or non-positive: a
* rate with no denominator is undefined, not zero, and every consumer already
* treats an empty draw array as "nothing to draw here".
*
* @param {number[] | null | undefined} counts - predicted counts, indexed by draw
* @param {number | null | undefined} enroll - students in the group
* @returns {number[]}
*/
export function toRates(counts, enroll) {
if (!counts?.length || !(enroll > 0)) return []
return counts.map((c) => (c / enroll) * 1000)
}
/**
* Pools Female + Male within each race, summing predicted counts at the same
* draw index and summing enrollment.
*
* Two refusals keep the arithmetic honest:
*
* - A race whose sexes have different draw-array lengths is dropped entirely.
* Element-wise addition would pair unrelated draws and truncate to the
* shorter array, producing a narrower pooled interval than the data supports.
* - Numerator and denominator are built from the same set of sexes. Adding both
* sexes' enrollment to one sex's counts would roughly halve the rate — an
* invented improvement, not a pooled estimate — and adding a sex's counts
* without its enrollment inflates it the same way.
*
* @param {Record<string, number[]>} countsByGroup - keyed by `groupKey(race, sex)`
* @param {Record<string, number>} enrollByGroup - keyed by `groupKey(race, sex)`
* @returns {{counts: Record<string, number[]>, enroll: Record<string, number>}}
* both keyed by race alone (see `groupKey(race)`).
*/
export function poolBySex(countsByGroup, enrollByGroup) {
const counts = {}
const enroll = {}
if (!countsByGroup) return { counts, enroll }
const byRace = {}
for (const key of Object.keys(countsByGroup)) {
const [race] = key.split('_')
;(byRace[race] ??= []).push(key)
}
for (const [race, keys] of Object.entries(byRace)) {
// A group must bring both halves of the fraction. Counts with no matching
// enrollment would land in the numerator while contributing nothing to the
// denominator, inflating the pooled rate.
const present = keys.filter((k) => countsByGroup[k]?.length > 0 && enrollByGroup?.[k] > 0)
if (present.length === 0) continue
const nDraws = countsByGroup[present[0]].length
if (present.some((k) => countsByGroup[k].length !== nDraws)) continue
const summed = new Array(nDraws).fill(0)
for (const k of present) {
const arr = countsByGroup[k]
for (let i = 0; i < nDraws; i++) summed[i] += arr[i]
}
counts[groupKey(race)] = summed
enroll[groupKey(race)] = present.reduce((sum, k) => sum + (enrollByGroup?.[k] || 0), 0)
}
return { counts, enroll }
}
+110
View File
@@ -0,0 +1,110 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
POOL_BY_SEX_ARREST_THRESHOLD,
poolBySex,
shouldPoolBySex,
toRates,
} from './pooling.js'
// ——— toRates ———
test('toRates: converts counts to a per-1,000 rate array', () => {
assert.deepEqual(toRates([0, 1, 2], 500), [0, 2, 4])
})
test('toRates: does not mutate its input', () => {
const counts = [1, 2]
toRates(counts, 1000)
assert.deepEqual(counts, [1, 2])
})
test('toRates: returns an empty array for non-positive enrollment', () => {
// A rate with no denominator is not zero, it is undefined — return nothing
// rather than a column of Infinity that a chart would happily plot.
assert.deepEqual(toRates([1, 2], 0), [])
assert.deepEqual(toRates([1, 2], -5), [])
assert.deepEqual(toRates([1, 2], undefined), [])
})
test('toRates: returns an empty array when counts are missing', () => {
assert.deepEqual(toRates(undefined, 100), [])
assert.deepEqual(toRates([], 100), [])
})
// ——— shouldPoolBySex ———
test('shouldPoolBySex: true when total observed arrests is under the threshold', () => {
const rows = [{ observed_arrests: 4 }, { observed_arrests: 2 }]
assert.equal(shouldPoolBySex(rows), true)
})
test('shouldPoolBySex: false at exactly the threshold', () => {
const rows = [{ observed_arrests: POOL_BY_SEX_ARREST_THRESHOLD }]
assert.equal(shouldPoolBySex(rows), false)
})
test('shouldPoolBySex: treats missing observed_arrests as zero', () => {
assert.equal(shouldPoolBySex([{ observed_arrests: null }, {}]), true)
})
test('shouldPoolBySex: false for an empty or missing row set', () => {
// No rows is not "a sparse district" — it is no district data at all. Firing
// the pooling banner there would explain a rule that never applied.
assert.equal(shouldPoolBySex([]), false)
assert.equal(shouldPoolBySex(undefined), false)
})
// ——— poolBySex ———
test('poolBySex: sums counts draw-by-draw and enrollment across F and M', () => {
const counts = { BL_F: [1, 2, 3], BL_M: [10, 20, 30] }
const enroll = { BL_F: 100, BL_M: 300 }
const pooled = poolBySex(counts, enroll)
assert.deepEqual(pooled.counts, { BL: [11, 22, 33] })
assert.deepEqual(pooled.enroll, { BL: 400 })
})
test('poolBySex: keys the result by race alone', () => {
const pooled = poolBySex(
{ WH_F: [1], WH_M: [1], HI_F: [2], HI_M: [2] },
{ WH_F: 10, WH_M: 10, HI_F: 20, HI_M: 20 },
)
assert.deepEqual(Object.keys(pooled.counts).sort(), ['HI', 'WH'])
})
test('poolBySex: pools a race present for only one sex', () => {
const pooled = poolBySex({ AM_F: [1, 1] }, { AM_F: 53, AM_M: 61 })
// Only F has draws, so only F's enrollment may go in the denominator —
// summing both sexes' enrollment against one sex's counts would halve the rate.
assert.deepEqual(pooled.counts, { AM: [1, 1] })
assert.deepEqual(pooled.enroll, { AM: 53 })
})
test('poolBySex: drops a race whose sexes have mismatched draw counts', () => {
// Adding a 500-draw array to a 400-draw array element-wise would silently
// pair unrelated draws and truncate; refuse instead.
const pooled = poolBySex({ BL_F: [1, 2, 3], BL_M: [1, 2] }, { BL_F: 10, BL_M: 10 })
assert.deepEqual(pooled.counts, {})
assert.deepEqual(pooled.enroll, {})
})
test('poolBySex: ignores a group with no enrollment entry', () => {
const pooled = poolBySex({ BL_F: [1], BL_M: [2] }, { BL_F: 100 })
assert.deepEqual(pooled.counts, { BL: [1] })
assert.deepEqual(pooled.enroll, { BL: 100 })
})
test('poolBySex: returns empty structures for empty input', () => {
const pooled = poolBySex({}, {})
assert.deepEqual(pooled.counts, {})
assert.deepEqual(pooled.enroll, {})
})
test('poolBySex: does not mutate its inputs', () => {
const counts = { BL_F: [1], BL_M: [2] }
const enroll = { BL_F: 10, BL_M: 20 }
poolBySex(counts, enroll)
assert.deepEqual(counts, { BL_F: [1], BL_M: [2] })
assert.deepEqual(enroll, { BL_F: 10, BL_M: 20 })
})
+67
View File
@@ -0,0 +1,67 @@
/**
* Shared x-domain for the arrest-rate charts, in rate per 1,000.
*
* Every panel and every model row has to sit on the same axis or the reader
* will compare curves that aren't comparable, so the domain is computed once
* from everything that will be drawn.
*/
import { quantile } from './kde.js'
import { niceTicks } from './niceTicks.js'
/**
* Hard ceiling on the axis.
*
* A four-student cell can produce posterior draws in the hundreds per 1,000;
* letting those set the axis squashes every other group into the leftmost few
* pixels. `computeRateDomain` reports `clipped` whenever the cap actually
* binds, so the chart can say so rather than cropping data off the edge in
* silence.
*
* Set to 100 per 1,000 (one arrest per ten students). The previous value of
* 30 was exceeded by over 60% of districts — mostly sparse groups with small
* enrollment cells whose Agresti–Coull upper bounds genuinely extend that far.
* A cap of 100 still guards against the most degenerate cases (e.g. a single
* predicted arrest in a four-student cell yielding ~250 per 1,000) while
* letting realistic data drive the axis for the vast majority of districts.
*/
export const MAX_RATE_DOMAIN = 100
/**
* Ignore the very top of each draw set when sizing the axis. The posterior
* predictive has a long right tail by construction; one draw in 500 should not
* decide the axis for all of them.
*/
const DOMAIN_QUANTILE = 0.995
const HEADROOM = 1.15
/**
* Computes a shared x-domain for the rate density and over-time charts.
*
* The domain covers everything that will actually be drawn: the 0.995
* quantile of each group's posterior-predictive draw rates, plus every
* Agresti–Coull upper bound from selected groups. `niceMax` is a clean
* "nice number" ceiling (via `niceTicks`) that may exceed the raw data max by
* up to one tick step, and is capped at `MAX_RATE_DOMAIN`.
*
* @param {Array<number[]>} rateArrays - per-group draw rates, per 1,000
* @param {number[]} acUpperRates - Agresti–Coull upper bounds, per 1,000
* @returns {{ticks: number[], niceMax: number, clipped: boolean}}
* `clipped` is true when something that will be drawn extends past `niceMax`.
*/
export function computeRateDomain(rateArrays, acUpperRates) {
const candidates = []
for (const rates of rateArrays || []) {
if (rates?.length) candidates.push(quantile(rates, DOMAIN_QUANTILE))
}
for (const upper of acUpperRates || []) {
if (Number.isFinite(upper)) candidates.push(upper)
}
const dataMax = candidates.length ? Math.max(...candidates) : 0
const rawMax = Math.max(dataMax, 1) * HEADROOM
const { ticks, niceMax } = niceTicks(Math.min(rawMax, MAX_RATE_DOMAIN), 5)
return { ticks, niceMax, clipped: dataMax > niceMax }
}
+73
View File
@@ -0,0 +1,73 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { MAX_RATE_DOMAIN, computeRateDomain } from './rateDomain.js'
test('computeRateDomain: covers the draws and the frequentist bounds', () => {
const domain = computeRateDomain([[0, 1, 2, 3, 4]], [6])
assert.ok(domain.niceMax >= 6, `niceMax ${domain.niceMax} should cover the AC upper bound`)
assert.equal(domain.clipped, false)
})
test('computeRateDomain: ticks start at zero and end at niceMax', () => {
const domain = computeRateDomain([[0, 1, 2]], [4])
assert.equal(domain.ticks[0], 0)
assert.equal(domain.ticks[domain.ticks.length - 1], domain.niceMax)
})
test('computeRateDomain: one extreme draw does not blow out the axis', () => {
// A single 400-per-1,000 draw in a 53-student cell would otherwise squash
// every other group into the leftmost pixel.
const bulk = new Array(500).fill(2)
bulk[0] = 400
const domain = computeRateDomain([bulk], [3])
assert.ok(domain.niceMax < 10, `niceMax was ${domain.niceMax}`)
})
test('computeRateDomain: caps at MAX_RATE_DOMAIN=100 and reports the clip', () => {
// The cap keeps a tiny-denominator group from flattening every other curve,
// but the caller has to be able to say so on screen rather than silently
// cropping data off the right edge.
const wide = new Array(500).fill(0).map((_, i) => (i < 490 ? 200 : 0))
const domain = computeRateDomain([wide], [300])
assert.equal(domain.niceMax, MAX_RATE_DOMAIN)
assert.equal(domain.clipped, true)
})
test('computeRateDomain: realistic sparse district data is not clipped', () => {
// Carson City NV (LEAID 3200390): pooled BL cell with 34 students and 0
// observed arrests → AC upper bound ≈ 88 per 1,000. This should NOT be
// clipped under the new cap of 100.
const blRates = new Array(500).fill(0)
const domain = computeRateDomain([blRates], [88])
assert.ok(domain.niceMax >= 88, `niceMax ${domain.niceMax} should cover AC bound of 88`)
assert.equal(domain.clipped, false)
})
test('computeRateDomain: never returns a zero-width domain', () => {
const domain = computeRateDomain([[0, 0, 0]], [0])
assert.ok(domain.niceMax > 0)
assert.equal(domain.clipped, false)
})
test('computeRateDomain: handles no data at all', () => {
const domain = computeRateDomain([], [])
assert.ok(domain.niceMax > 0)
assert.ok(domain.ticks.length > 1)
})
test('computeRateDomain: ignores empty rate arrays', () => {
const domain = computeRateDomain([[], [1, 2, 3]], [])
assert.ok(domain.niceMax > 0 && domain.niceMax < MAX_RATE_DOMAIN)
})
test('computeRateDomain: degenerate tiny-denominator cell is clipped', () => {
// A four-student cell with a single predicted arrest yields ~250 per
// 1,000 — this should be capped at MAX_RATE_DOMAIN=100 and flagged as
// clipped so the chart can warn the reader.
const tinyRates = new Array(500).fill(0)
for (let i = 0; i < 250; i++) tinyRates[i] = 250
const domain = computeRateDomain([tinyRates], [300])
assert.equal(domain.niceMax, MAX_RATE_DOMAIN)
assert.equal(domain.clipped, true)
})
+17
View File
@@ -0,0 +1,17 @@
import { defineConfig } from 'vite'
import react from '@vitejs/plugin-react-swc'
// The app is served at /crdc-demo/ on pages.civilytics.org.
// Set base to the subdirectory so asset paths resolve correctly.
export default defineConfig({
plugins: [react()],
server: { port: 5173 },
build: { outDir: 'dist' },
base: '/crdc-demo/', // Required for subdirectory deployment on git-pages
optimizeDeps: {
// duckdb-wasm ships its own worker + wasm binaries resolved via `?url`
// imports; esbuild's dev-server pre-bundling can rewrite those import
// paths and break worker instantiation. Exclude it from pre-bundling.
exclude: ['@duckdb/duckdb-wasm'],
},
})