Ports Figs 6 and 7 from crdc-arrests/R/paper_figures.R (wp_fig_group_density and
wp_fig_group_difference) so someone who has read the paper sees the paper. The
three old charts become a summary table and two charts; ArrestsOverTime stays.
Draws pipeline. useDrawDistribution now selects draw_id and returns predicted
counts indexed by draw rather than pre-divided rates. That one column is what
unlocks the rest: pooling has to sum numerators and denominators separately, and
a between-group difference has to be taken at a common draw index. Indexing by
draw_id rather than push order makes DuckDB's row ordering irrelevant and turns
a missing draw into a hole, which isCompleteDrawSet then rejects — a group
present for 300 of 500 draws would otherwise get an interval computed off a
biased subsample that looks identical on screen to a complete one. It also takes
models[] instead of a single model, so "compare all four" needs no conditional
hooks. The reset-before-guard ordering is preserved.
Also fixes a pre-existing bug in that pipeline: a state's draws are split across
data_0.parquet, data_1.parquet, … and the part count varies by state. The app
only ever fetched data_0. Nevada has one part, so this was invisible in every
Nevada test; California has eight, totalling 6.2MB, of which data_0 is 37KB and
holds 11 of California's 1,715 districts. Every other CA district looked absent
from the published data and silently fell back to the approximation. Parts are
now discovered from the Hugging Face tree listing API and fetched in parallel —
listing rather than probing data_N until a 404, because the browser logs a 404
as a console error however cleanly the fetch handles it, and a red error on
every load is indistinguishable from a real one. HEAD probing is the fallback.
New pure utils, all written against tests first:
agrestiCoull faithful port incl. the zero-numerator rule of three and the
negative lower bound at (1, 53); pinned to five R outputs
pooling sex pooling for sparse districts; numerator and denominator
are always drawn from the same set of groups
densityProfile discrete probability mass below 12 distinct values, KDE above;
KDE delegates to kde.js, whose bandwidth clamp is untouched
districtGroups display-row derivation, defaults, pooled vs unpooled keys
groupDifference per-draw delta; refuses to pair mismatched draw sets
rateDomain shared x-axis, with a clip flag so the cap is never silent
Chart A keeps the palette contract: race is hue, sex is position. Female and
Male are stacked panels sharing one axis, collapsing to one panel when pooled —
never a second hue. Chart B uses a diverging ramp centred at zero rather than
the paper's sequential YlOrRd, because delta is signed and a sequential ramp
encodes "more" where the data means "which direction".
Captions say "posterior predictive draws", never "paired parameter draws":
draw_id is renumbered per write batch upstream and a district's groups land in
different batches, so cross-group pairing is effectively independent (measured
cor ~= 0.02). The published figure has the same property; what neither can claim
is a paired-parameter contrast.
Sparse districts (under 20 arrests district-wide) pool Female and Male within
each race, with a banner stating the rule and a switch to override it. A pooled
group carries no modelled interval — summing two groups' interval bounds is not
a pooled interval, and there is no honest way to fake one without the draws.
React still owns the DOM. d3-scale/shape/array/interpolate supply scales, path
generators and colour interpolation; no selections, no useEffect DOM mutation.
Deletes RateByGroupBar and RateDensityRidgeline. Keeps distributionApprox.js and
ApproxNote.jsx — still the per-group fallback when draws can't be fetched.
112 tests pass; npm run build clean.
395 lines
14 KiB
React
395 lines
14 KiB
React
import { useMemo } from 'react'
|
||
import { area, curveMonotoneX, curveStep } from 'd3-shape'
|
||
import { scaleLinear } from 'd3-scale'
|
||
import { interpolateRgb } from 'd3-interpolate'
|
||
import { rateProfile } from '../utils/densityProfile.js'
|
||
import { differenceRates, differenceSummary } from '../utils/groupDifference.js'
|
||
import { displayDraws } from '../utils/districtGroups.js'
|
||
import { niceTicks } from '../utils/niceTicks.js'
|
||
|
||
/**
|
||
* Chart B — "Model estimated differences".
|
||
*
|
||
* Port of the white paper's Fig 7 (`wp_fig_group_difference`): the posterior
|
||
* distribution of Δ = rate(A) − rate(B), per 1,000 students, computed at each
|
||
* draw index.
|
||
*
|
||
* Two deliberate departures from the R figure:
|
||
*
|
||
* - The fill is a **diverging** ramp centred at zero (navy below, paper at
|
||
* zero, ember above) rather than the paper's sequential YlOrRd. Δ is a signed
|
||
* quantity; a sequential ramp encodes "more" where the data means "which
|
||
* direction", and would make a large negative difference read as a small one.
|
||
* - The readout is spelled out in a sentence. `Pr(Δ > 0) = 94.4%` is the number
|
||
* a reader is most likely to misread as "94.4% more arrests".
|
||
*/
|
||
|
||
const NEGATIVE_COLOR = '#22406A' // --cv-navy-600
|
||
const ZERO_COLOR = '#F2EDE4' // --cv-paper-2
|
||
const POSITIVE_COLOR = '#C25311' // --cv-accent
|
||
const GRADIENT_STOPS = 24
|
||
|
||
const WIDTH = 760
|
||
const HEIGHT = 250
|
||
const MARGIN = { top: 18, right: 22, bottom: 46, left: 22 }
|
||
|
||
export default function GroupDifference({
|
||
groups,
|
||
enrollByGroup,
|
||
byModel,
|
||
selectedModel,
|
||
status,
|
||
pooled,
|
||
pair,
|
||
onPairChange,
|
||
}) {
|
||
const { counts, enroll } = useMemo(
|
||
() => displayDraws(byModel?.[selectedModel]?.counts, enrollByGroup, pooled),
|
||
[byModel, selectedModel, enrollByGroup, pooled],
|
||
)
|
||
|
||
// Only groups with a usable draw set can be differenced at all — offering the
|
||
// others in the picker would produce an empty chart with no explanation.
|
||
const comparable = useMemo(
|
||
() => groups.filter((g) => counts?.[g.key]?.length > 0 && enroll?.[g.key] > 0),
|
||
[groups, counts, enroll],
|
||
)
|
||
|
||
const [keyA, keyB] = pair || []
|
||
const groupA = comparable.find((g) => g.key === keyA)
|
||
const groupB = comparable.find((g) => g.key === keyB)
|
||
|
||
const deltas = useMemo(
|
||
() =>
|
||
groupA && groupB
|
||
? differenceRates(counts[groupA.key], enroll[groupA.key], counts[groupB.key], enroll[groupB.key])
|
||
: [],
|
||
[groupA, groupB, counts, enroll],
|
||
)
|
||
const summary = useMemo(() => differenceSummary(deltas), [deltas])
|
||
|
||
return (
|
||
<div className="cv-card" style={{ padding: 'var(--space-2)' }}>
|
||
<div style={headerStyle}>
|
||
<div>
|
||
<h3 style={cardTitle}>Model estimated differences</h3>
|
||
<p style={subtitleStyle}>
|
||
How much higher is one group’s modelled arrest rate than another’s, and how
|
||
sure is the model of the direction?
|
||
</p>
|
||
</div>
|
||
{comparable.length >= 2 && (
|
||
<PairPickers
|
||
comparable={comparable}
|
||
keyA={keyA}
|
||
keyB={keyB}
|
||
onPairChange={onPairChange}
|
||
/>
|
||
)}
|
||
</div>
|
||
|
||
{status === 'loading' ? (
|
||
<p style={emptyStyle}>Loading posterior draws…</p>
|
||
) : comparable.length < 2 ? (
|
||
<Degraded
|
||
reason={
|
||
comparable.length === 1
|
||
? `Only ${comparable[0].label} has a usable set of posterior draws in this district, so there is no second group to compare it against.`
|
||
: 'No student group in this district has a usable set of posterior draws, so no difference can be computed. This is usually because the district is absent from the published draw shard for this model.'
|
||
}
|
||
/>
|
||
) : !groupA || !groupB || groupA.key === groupB.key ? (
|
||
<Degraded reason="Pick two different student groups to compare." />
|
||
) : !summary ? (
|
||
<Degraded
|
||
reason={`${groupA.label} and ${groupB.label} do not have matching draw sets in this model specification, so their difference cannot be computed draw by draw.`}
|
||
/>
|
||
) : (
|
||
<>
|
||
<Readout summary={summary} groupA={groupA} groupB={groupB} />
|
||
<DifferencePlot deltas={deltas} summary={summary} groupA={groupA} groupB={groupB} />
|
||
<p style={captionStyle}>
|
||
Δ is computed at each of {summary.n.toLocaleString()} posterior predictive draws as{' '}
|
||
{groupA.label} minus {groupB.label}, per 1,000 students. The dashed line marks no
|
||
difference. Bars beneath the curve are the 80% and 95% intervals around the median Δ.
|
||
</p>
|
||
</>
|
||
)}
|
||
</div>
|
||
)
|
||
}
|
||
|
||
function Readout({ summary, groupA, groupB }) {
|
||
const pct = summary.prGreater * 100
|
||
const higher = summary.median >= 0 ? groupA : groupB
|
||
const lower = summary.median >= 0 ? groupB : groupA
|
||
const share = summary.median >= 0 ? pct : 100 - pct
|
||
return (
|
||
<div style={readoutStyle}>
|
||
<div>
|
||
<span style={readoutNumber}>{formatPercent(pct)}</span>
|
||
<span style={readoutLabel}>Pr(Δ > 0)</span>
|
||
</div>
|
||
<p style={{ margin: 0, fontSize: '0.88rem', maxWidth: '38rem' }}>
|
||
In {formatPercent(share)} of posterior predictive draws, the {higher.sentenceLabel} arrest
|
||
rate exceeds the {lower.sentenceLabel} rate. The median difference is{' '}
|
||
<strong>{formatDelta(summary.median)}</strong> per 1,000 students.
|
||
</p>
|
||
</div>
|
||
)
|
||
}
|
||
|
||
function DifferencePlot({ deltas, summary, groupA, groupB }) {
|
||
const innerWidth = WIDTH - MARGIN.left - MARGIN.right
|
||
const baselineY = HEIGHT - MARGIN.bottom
|
||
|
||
const { ticks, min, max } = useMemo(() => symmetricDomain(summary), [summary])
|
||
const x = scaleLinear().domain([min, max]).range([MARGIN.left, MARGIN.left + innerWidth])
|
||
|
||
const profile = useMemo(() => rateProfile(deltas, { min, max, n: 80 }), [deltas, min, max])
|
||
if (!profile) return null
|
||
|
||
const y = scaleLinear().domain([0, profile.maxY || 1]).range([baselineY, MARGIN.top])
|
||
const points =
|
||
profile.kind === 'mass' ? padMassPoints(profile.points, profile.step, min, max) : profile.points
|
||
|
||
const areaGen = area()
|
||
.x((p) => x(clamp(p.x, min, max)))
|
||
.y0(baselineY)
|
||
.y1((p) => y(p.y))
|
||
.curve(profile.kind === 'mass' ? curveStep : curveMonotoneX)
|
||
|
||
const gradientId = `delta-gradient-${groupA.key}-${groupB.key}`
|
||
|
||
return (
|
||
<div style={{ overflowX: 'auto' }}>
|
||
<svg
|
||
width="100%"
|
||
viewBox={`0 0 ${WIDTH} ${HEIGHT}`}
|
||
style={{ maxWidth: '100%', minWidth: '320px' }}
|
||
role="img"
|
||
aria-label={`Posterior distribution of the difference in arrest rate between ${groupA.label} and ${groupB.label}`}
|
||
>
|
||
<defs>
|
||
<linearGradient id={gradientId} x1="0" y1="0" x2="1" y2="0">
|
||
{divergingStops(min, max).map((s) => (
|
||
<stop key={s.offset} offset={`${s.offset * 100}%`} stopColor={s.color} />
|
||
))}
|
||
</linearGradient>
|
||
</defs>
|
||
|
||
{ticks.map((t) => (
|
||
<line key={t} x1={x(t)} y1={MARGIN.top} x2={x(t)} y2={baselineY} stroke="var(--cv-rule)" strokeWidth={1} />
|
||
))}
|
||
|
||
<path d={areaGen(points)} fill={`url(#${gradientId})`} opacity={0.85} />
|
||
<path
|
||
d={areaGen.lineY1()(points)}
|
||
fill="none"
|
||
stroke="var(--cv-ink-2)"
|
||
strokeWidth={1.5}
|
||
strokeLinejoin="round"
|
||
/>
|
||
|
||
{/* No difference — the paper's red vertical line. */}
|
||
<line
|
||
x1={x(0)}
|
||
y1={MARGIN.top - 4}
|
||
x2={x(0)}
|
||
y2={baselineY + 6}
|
||
stroke="var(--cv-danger)"
|
||
strokeWidth={2}
|
||
strokeDasharray="5 4"
|
||
/>
|
||
<text x={x(0)} y={MARGIN.top - 7} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-danger)">
|
||
no difference
|
||
</text>
|
||
|
||
<line x1={MARGIN.left} y1={baselineY} x2={MARGIN.left + innerWidth} y2={baselineY} stroke="var(--cv-rule-strong)" strokeWidth={1} />
|
||
|
||
{/* Median with 80% (thick) and 95% (thin) intervals. */}
|
||
<g>
|
||
<line x1={x(clamp(summary.lower95, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper95, min, max))} y2={baselineY + 13} stroke="var(--cv-ink-2)" strokeWidth={1.5} strokeLinecap="round" />
|
||
<line x1={x(clamp(summary.lower80, min, max))} y1={baselineY + 13} x2={x(clamp(summary.upper80, min, max))} y2={baselineY + 13} stroke="var(--cv-ink)" strokeWidth={4} strokeLinecap="round" />
|
||
<circle cx={x(clamp(summary.median, min, max))} cy={baselineY + 13} r={3.5} fill="var(--cv-paper)" stroke="var(--cv-ink)" strokeWidth={2} />
|
||
</g>
|
||
|
||
{ticks.map((t) => (
|
||
<text key={t} x={x(t)} y={baselineY + 32} textAnchor="middle" fontSize="0.62rem" fill="var(--cv-ink-3)">
|
||
{formatTick(t)}
|
||
</text>
|
||
))}
|
||
|
||
<text x={MARGIN.left} y={HEIGHT - 4} textAnchor="start" fontSize="0.62rem" fill="var(--cv-ink-3)">
|
||
← {groupB.shortLabel} higher
|
||
</text>
|
||
<text x={MARGIN.left + innerWidth} y={HEIGHT - 4} textAnchor="end" fontSize="0.62rem" fill="var(--cv-ink-3)">
|
||
{groupA.shortLabel} higher →
|
||
</text>
|
||
</svg>
|
||
</div>
|
||
)
|
||
}
|
||
|
||
/**
|
||
* A domain centred on zero. A signed quantity drawn on an off-centre axis makes
|
||
* the eye read the *position* of the curve as the size of the difference, so
|
||
* zero sits in the middle even when every draw falls on one side of it.
|
||
*/
|
||
function symmetricDomain(summary) {
|
||
const extent = Math.max(
|
||
Math.abs(summary.lower95),
|
||
Math.abs(summary.upper95),
|
||
Math.abs(summary.median),
|
||
1e-6,
|
||
)
|
||
const { niceMax } = niceTicks(extent * 1.25, 4)
|
||
const step = niceMax / 4
|
||
const ticks = []
|
||
for (let i = -4; i <= 4; i++) ticks.push(Math.round(step * i * 1e6) / 1e6)
|
||
return { ticks, min: -niceMax, max: niceMax }
|
||
}
|
||
|
||
/** Diverging ramp, with the paper-coloured midpoint pinned to Δ = 0. */
|
||
function divergingStops(min, max) {
|
||
const zeroOffset = (0 - min) / (max - min)
|
||
const toNegative = interpolateRgb(NEGATIVE_COLOR, ZERO_COLOR)
|
||
const toPositive = interpolateRgb(ZERO_COLOR, POSITIVE_COLOR)
|
||
const stops = []
|
||
for (let i = 0; i <= GRADIENT_STOPS; i++) {
|
||
const offset = i / GRADIENT_STOPS
|
||
const color =
|
||
offset <= zeroOffset
|
||
? toNegative(zeroOffset > 0 ? offset / zeroOffset : 1)
|
||
: toPositive(zeroOffset < 1 ? (offset - zeroOffset) / (1 - zeroOffset) : 0)
|
||
stops.push({ offset, color })
|
||
}
|
||
return stops
|
||
}
|
||
|
||
function padMassPoints(points, step, min, max) {
|
||
const half = Math.max(step, 1e-6) / 2
|
||
return [
|
||
{ x: Math.max(points[0].x - half, min), y: 0 },
|
||
...points,
|
||
{ x: Math.min(points[points.length - 1].x + half, max), y: 0 },
|
||
]
|
||
}
|
||
|
||
function PairPickers({ comparable, keyA, keyB, onPairChange }) {
|
||
return (
|
||
<div style={{ display: 'flex', alignItems: 'center', gap: '0.4rem', flexWrap: 'wrap' }}>
|
||
<GroupSelect
|
||
label="Compare"
|
||
value={keyA}
|
||
options={comparable}
|
||
onChange={(v) => onPairChange([v, keyB])}
|
||
/>
|
||
<span style={{ fontSize: '0.8rem', color: 'var(--cv-ink-3)' }}>against</span>
|
||
<GroupSelect
|
||
label="against"
|
||
value={keyB}
|
||
options={comparable}
|
||
onChange={(v) => onPairChange([keyA, v])}
|
||
/>
|
||
</div>
|
||
)
|
||
}
|
||
|
||
function GroupSelect({ label, value, options, onChange }) {
|
||
return (
|
||
<select
|
||
aria-label={label}
|
||
value={value || ''}
|
||
onChange={(e) => onChange(e.target.value)}
|
||
style={{ padding: '0.25rem 0.4rem', fontFamily: 'var(--font-sans)', fontSize: '0.78rem' }}
|
||
>
|
||
{options.map((g) => (
|
||
<option key={g.key} value={g.key}>
|
||
{g.label}
|
||
</option>
|
||
))}
|
||
</select>
|
||
)
|
||
}
|
||
|
||
function Degraded({ reason }) {
|
||
return (
|
||
<div style={degradedStyle}>
|
||
<p style={{ margin: 0, fontSize: '0.85rem' }}>{reason}</p>
|
||
</div>
|
||
)
|
||
}
|
||
|
||
function clamp(v, min, max) {
|
||
return Math.min(Math.max(v, min), max)
|
||
}
|
||
|
||
function formatPercent(pct) {
|
||
return `${pct.toLocaleString(undefined, { minimumFractionDigits: 1, maximumFractionDigits: 1 })}%`
|
||
}
|
||
|
||
function formatDelta(v) {
|
||
const sign = v > 0 ? '+' : ''
|
||
return `${sign}${v.toLocaleString(undefined, { minimumFractionDigits: 2, maximumFractionDigits: 2 })}`
|
||
}
|
||
|
||
function formatTick(v) {
|
||
if (v === 0) return '0'
|
||
const abs = Math.abs(v)
|
||
return v.toLocaleString(undefined, {
|
||
minimumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
|
||
maximumFractionDigits: abs < 1 ? 2 : abs < 10 ? 1 : 0,
|
||
})
|
||
}
|
||
|
||
const cardTitle = { fontSize: '0.85rem', marginBottom: 'var(--space-1)', color: 'var(--cv-ink-2)' }
|
||
const headerStyle = {
|
||
display: 'flex',
|
||
justifyContent: 'space-between',
|
||
alignItems: 'flex-start',
|
||
flexWrap: 'wrap',
|
||
gap: 'var(--space-2)',
|
||
}
|
||
const subtitleStyle = { fontSize: '0.8rem', color: 'var(--cv-ink-3)', margin: 0, maxWidth: '32rem' }
|
||
const emptyStyle = { color: 'var(--cv-ink-3)', fontSize: '0.85rem', padding: 'var(--space-2) 0' }
|
||
const captionStyle = { fontSize: '0.74rem', color: 'var(--cv-ink-3)', margin: '0.4rem 0 0', lineHeight: 1.5 }
|
||
const readoutStyle = {
|
||
display: 'flex',
|
||
alignItems: 'center',
|
||
gap: 'var(--space-3)',
|
||
flexWrap: 'wrap',
|
||
padding: 'var(--space-2) 0',
|
||
borderTop: '3px double var(--cv-ink)',
|
||
borderBottom: '1px solid var(--cv-rule)',
|
||
margin: 'var(--space-2) 0',
|
||
}
|
||
const readoutNumber = {
|
||
fontFamily: "'Source Serif 4', Georgia, serif",
|
||
fontWeight: 700,
|
||
fontSize: '2.75rem',
|
||
lineHeight: 1,
|
||
letterSpacing: '-0.02em',
|
||
color: 'var(--cv-ink)',
|
||
fontVariantNumeric: 'tabular-nums',
|
||
display: 'block',
|
||
}
|
||
const readoutLabel = {
|
||
fontSize: '0.72rem',
|
||
fontWeight: 600,
|
||
textTransform: 'uppercase',
|
||
letterSpacing: '0.08em',
|
||
color: 'var(--cv-ink-3)',
|
||
display: 'block',
|
||
marginTop: '0.3rem',
|
||
}
|
||
const degradedStyle = {
|
||
background: 'var(--cv-paper-2)',
|
||
border: '1px solid var(--cv-rule)',
|
||
borderLeft: '3px solid var(--cv-ink-4)',
|
||
borderRadius: 'var(--radius-md)',
|
||
padding: 'var(--space-2)',
|
||
marginTop: 'var(--space-2)',
|
||
color: 'var(--cv-ink-2)',
|
||
}
|