Files
jared 52a0c77e31
Deploy to Git Pages / build-and-deploy (push) Failing after 19s
Initial commit: CRDC Arrests API demo app
React + Vite static site demonstrating the CRDC School Arrest Rate API.
Features 6 interactive charts comparing observed arrest data against Bayesian
model estimates across U.S. school districts, with Civilytics visual identity.

- State selector and district search with 'interesting' suggestions (top arrests)
- Animated histogram loading grid showing posterior draw progress
- Charts: time series, rate by group, district vs national, model comparison quadrants, density proxy, exceedance probability
- Static JSON fixture for national rates (no API changes needed)
- Dockerfile + docker-compose.yml for self-hosted deployment
- GitHub Actions workflow and _config.yml for Git Pages
2026-08-10 10:54:51 -04:00

361 lines
18 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CRDC Arrests API Demonstration App — Design Proposal
## 1. Overview
A web application hosted at `civilytics.org` that demonstrates the [CRDC School
Arrest Rate API](https://crdc-api.civilytics.org/api/v1/) by letting any visitor
explore school-based arrest estimates for any U.S. school district, with a strong
visual narrative built around Bayesian model comparisons.
The app is a **frontend-only static application** that calls the public read-only
API directly from the browser. No server-side code, no database, no auth — just
HTML/CSS/JS (with optional Docker container for local dev or self-hosting behind
a reverse proxy).
---
## 2. User Flow
```
┌─────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ State │ → │ District Search │ → │ Loading (anim) │ → ╔══════════════╗
│ selector │ │ w/ suggestions │ │ patience msg │ ║ Chart panel ║
└─────────────┘ └──────────────────┘ └─────────────────┘ ║ 6 panels ║
║ see §4 ║
╚══════════════╝
```
### Step 1 — State prompt
- Full-screen landing card with Civilytics branding.
- Dropdown / typeahead for U.S. state (50 states + DC). Two-letter codes map to
the API `state` parameter directly.
- On select → slide to Step 2.
### Step 2 — District search
- Search-as-you-type input, hitting `/api/v1/districts?q=<partial>&state=XX`.
- **Suggested districts** appear as a horizontal carousel of "interesting" picks:
- Top N by total arrest count in the most recent wave (2021-22).
- Fetched once per state via `limit=500` from `/estimates?state=XX&year=21-22`
sorted client-side by `observed_arrests`.
- Each suggestion card shows: district name, enrollment (in small), observed
arrests. Clicking a suggestion jumps straight to loading + charts.
- Typing filters the list live; selecting an item from either source proceeds.
### Step 3 — Loading / patience animation
- A full-screen overlay with an animated visual that conveys "data is being
fetched across multiple endpoints, please be patient."
- The animation should reflect the **multi-model comparison** theme: e.g., a grid
of small histogram-like bars (representing posterior draws) that animate in and
out row by row as each model's endpoint responds.
- Estimated total API calls per district lookup: ~12–15 (one per model × race ×
sex combination, filtered to the selected district). The animation should scale
visually with progress.
### Step 4 — Chart panel (6 charts in a 3×2 grid)
All charts follow Civilytics visual identity: warm paper background (`#FAF7F2`),
civic navy text (`#0E1A2B`), ember accent (`#C25311`). Data-viz palette uses the
supporting colors from `_tokens.scss` (teal, plum, moss, brass).
#### Row 1 — Observed & descriptive
**Chart 1: Arrests over time by CRDC wave (line chart)**
- X-axis: CRDC waves (`2015-16`, `2017-18`, `2021-22`).
- Y-axis: total observed arrest count.
- Line + points; each point labeled with the rate per 1,000 students.
- Source: `/estimates?leaid=XXXXX&year=...` across all three years (or a single
call to `/districts/{leaid}` which returns all demographics for one district).
**Chart 2: Arrest rate by student group — most recent year (bar chart)**
- Bars for each race×sex combination in `AM/BL/HI/WH × F/M`.
- Y-axis: arrests per 1,000 students.
- Color-coded by the supporting palette; legend shows full labels.
**Chart 3: District vs. national — top student group (comparison chart)**
- Identify the district's highest-arrest-rate group.
- Side-by-side bars or a small multiples comparison against the corresponding
**national** rate for that same group.
- National rates fetched via `/states?state=XX&race=&sex=&year=21-22` aggregated,
or more precisely from the national summary (which may need to be computed as
an aggregate across all states).
#### Row 2 — Bayesian model distribution comparisons
All four quadrants show results for the **selected district**, comparing:
- **Column A**: One-year models (`unified_m1_mod`, `unified_m2_mod`) vs.
- **Column B**: Three-year models (`unified_m3_mod` through `unified_m5_mod`).
- Within each column, rows differentiate baseline (no covariate) from covariate
models.
**Chart 4: Predicted draws by year vs. observed (scatter / point-range)**
- For three-year models only: for each of the 3 waves, show the model's predicted
median and 95% interval alongside the observed value as a separate marker.
- Layout: x-axis = wave; y-axis = arrest count; points dodge left (model) vs.
right (observed).
**Chart 5: Observed rate per group against model distribution (density / ridge)**
- For each student group, show the observed rate and overlay the posterior draw
distributions from all four model types as ridgeline or violin plots.
- Mirrors `wp_fig_group_density()` in the white paper.
**Chart 6: Probability district exceeds national rate per group (bar chart)**
- For each race×sex group, compute P(district rate > national rate) using the
posterior draws from a chosen model (e.g., the default `unified_m2_mod`).
- Bars colored by threshold crossing; annotated with exact probability.
---
## 3. Data Sources & API Endpoints Used
| Purpose | Endpoint | Frequency |
|---|---|---|
| State list / validation | Hardcoded enum (`ALLOWED_STATES`) | Once |
| District search | `/api/v1/districts?q=&state=` | On keystroke (debounced) |
| "Interesting" suggestions | `/api/v1/estimates?state=XX&year=21-22` sorted by `observed_arrests desc` | Once per state |
| Single district, all demographics | `/api/v1/distimates/{leaid}` or `/estimates?leaid=` with year/model filters | ~6 calls × 3 years = 18+ |
| National comparison rates | `/api/v1/states?state=&race=&sex=&year=21-22` aggregated across states, OR a dedicated national endpoint if available | Once per group |
| Model metadata | `/api/v1/models` | Once (cache) |
**Key data structures returned by the API:**
The `/estimates/{leaid}` endpoint returns one row per `race × sex × year × model` with:
```json
{
"leaid": "...",
"lea_name": "...",
"state": "TX",
"race": "BL",
"sex": "M",
"year": "21-22",
"model": "unified_m2_mod",
"stu_enroll": 35963,
"observed_arrests": 1,
"rate_median": 0.42,
"rate_lower": 0.18,
"rate_upper": 0.91,
"count_median": 15.3,
"count_lower": 6.7,
"count_upper": 28.9
}
```
The `/draws` endpoint returns a Hugging Face Parquet shard URL + DuckDB SQL for
bulk posterior draw access — useful if we need to compute custom quantities like
P(district > national) without round-tripping through multiple API calls. However,
for a demo app that runs in the browser, relying on 10 HF Parquet shards per model
is impractical (80 GB total). The summary endpoints (`/estimates`) provide enough
aggregated information for all six charts using only `rate_median`, `count_median`,
and interval bounds — **no raw draw access needed**.
---
## 4. Visual Design Language (from reference materials)
### Color palette
```css
:root {
--cv-paper: #FAF7F2; /* warm paper background */
--cv-ink: #0E1A2B; /* primary text — civic navy/black */
--cv-navy-600:#22406A; /* links, accents */
--cv-accent: #C25311; /* ember — alerts, highlights, eyebrows */
/* Data-viz supporting colors (from _tokens.scss) */
--teal-600: #1F6F70;
--plum-600: #6B3A5E;
--moss-600: #4A6B2F;
--brass-600: #B8751C;
}
```
### Typography
- **Display**: Libre Franklin / Inter — headings, stat callouts.
- **Body**: Inter — UI text, labels, captions.
- **Mono**: JetBrains Mono — code snippets, API URLs in footers.
### Chart patterns (from social media posts & white paper)
1. **Ridgeline density plots** for posterior draws (`geom_density_ridges`).
2. **Pointrange / error bar charts** comparing model intervals to observed values.
3. **Faceted small multiples** — always split by `covariate × time` (baseline vs.
covariate; one-year vs. three-year).
4. **Transparent PNG export with watermark logo** in bottom-right corner.
5. **"Stat callout" hero numbers** for key metrics (e.g., total arrests, rate per 1k).
---
## 5. Technology Stack Recommendation
### Recommended: React + Vite + vanilla CSS (static site)
| Layer | Choice | Rationale |
|---|---|---|
| Framework | **React 19** (no framework overhead; Vite dev server) | Component model for charts, built-in state management via hooks |
| Build tool | **Vite** | Fast HMR, native ES modules, trivial static export (`npm run build`) |
| Charting | **`@visx/visx`** or plain SVG/CSS animations | Lightweight, no heavy deps; we control every pixel to match Civilytics style |
| HTTP client | `fetch` with AbortController + exponential backoff | No extra dependency; browser-native |
| Styling | Plain CSS custom properties (no Tailwind) | Zero-runtime, matches the existing SCSS token system exactly |
| Deployment | Static site on any host (GitHub Pages, Vercel, Netlify) or Docker nginx container | Self-hostable; no backend required |
#### Why not Shiny / Quarto?
The API returns JSON and we need a dynamic SPA with loading states, debounced search, and animated transitions. A static React app is the most natural fit. R/Shiny would add unnecessary server-side complexity for what is fundamentally a browser-based data visualization demo. The existing `crdc-arrests` project uses Quarto + ggplot2 for reports; this demo app is a different artifact with different requirements.
#### Why not Svelte or Vue?
React has the largest ecosystem, best tooling (Vite), and most team familiarity. For a 6-chart SPA it's more than sufficient without being overkill.
### Docker option
A minimal `nginx:alpine` container serves the built static files. ~20 MB image. Can be run on your self-hosted fleet (`efron`, `maxwell`, etc.) behind Caddy/Traefik with TLS via Let's Encrypt — consistent with how you host other civilytics.org services.
---
## 6. Build & Deployment Strategy
### Local development
```bash
npm install # one-time: installs React, Vite, dev deps
npm run dev # starts Vite on localhost:5173
# Edit src/App.jsx / src/components/*.jsx — hot reload
```
### Production build (static)
```bash
npm run build # outputs dist/ with index.html + assets
# Upload dist/ to any static host, or:
docker build -t crdc-demo . && docker run -p 8080:80 crdc-demo
```
### Docker deployment (self-hosted on your fleet)
1. Build image locally or via CI: `docker buildx build --platform linux/amd64 -t registry.civilytics.org/crdc-demo:latest .`
2. Push to local registry on `maxwell`.
3. Deploy via docker-compose or a simple systemd service with nginx container + Caddy reverse proxy for TLS termination.
```yaml
# docker-compose.yml (minimal)
services:
app:
image: crdc-demo:latest
ports: ["8080:80"]
restart: unless-stopped
```
### CI/CD (optional, GitHub Actions)
- On push to `main`: run lint + build, publish Docker image to local registry.
- No automated deployment — you control when new versions go live on the fleet.
---
## 7. File Structure
```
crdc-demo/
├── public/ # Static assets (logo, favicon)
│ └── civilytics-logo.svg
├── src/
│ ├── components/ # Reusable UI pieces
│ │ ├── StateSelector.jsx
│ │ ├── DistrictSearch.jsx
│ │ ├── LoadingAnimation.jsx
│ │ └── StatCallout.jsx
│ ├── charts/ # The 6 chart components
│ │ ├── ArrestsOverTime.jsx
│ │ ├── RateByGroupBar.jsx
│ │ ├── DistrictVsNational.jsx
│ │ ├── ModelDrawsComparison.jsx
│ │ ├── ObservedRateDensity.jsx
│ │ └── ExceedanceProbability.jsx
│ ├── hooks/
│ │ ├── useApi.js # fetch wrapper with retry/backoff
│ │ └── useDistrictData.js # orchestrates all 6 charts' data needs
│ ├── styles/
│ │ ├── tokens.css # Civilytics design tokens
│ │ └── main.css
│ ├── App.jsx # Main router: state → search → loading → charts
│ └── main.jsx # React entry point
├── Dockerfile
├── vite.config.js
└── package.json
```
---
## 8. Key Implementation Notes & Risks
### Risk 1 — API response time for multi-model data
Fetching all models × race×sex combinations for a single district requires ~40
individual `/estimates` calls (5 unified + 5 stratified models × 8 groups). Each
API call may take 200–500ms. **Mitigation**: use `Promise.allSettled()` to fire
all requests in parallel; show progress as batches resolve. The loading animation
should reflect this batching visually (e.g., rows of bars filling left-to-right).
### Risk 2 — National rate computation
The `/states` endpoint returns per-state aggregates, not a national total. To get
national arrest rates by student group for Chart 3 and Chart 6:
- **Option A**: Fetch all states (`limit=100`, iterate through ~50 pages) and sum.
Too slow for a browser app.
- **Option B**: Add a `/api/v1/national` endpoint to the API (server-side aggregate).
Requires modifying `crdc-arrests/api/` — quick R/SQL change but needs your approval.
- **Recommended short-term**: Cache national rates as a static JSON file generated
during data release and committed alongside this app. The white paper already
computes these values; we can extract them into a small fixture.
### Risk 3 — Chart complexity (Chart 5 & 6)
Charts 5 (density/ridge) and 6 (exceedance probability) require either raw posterior
draws or sufficient summary statistics to reconstruct distributions. The `/estimates`
endpoint provides `count_median`, `count_lower`, `count_upper` but not the full draw
distribution. **Mitigation**: Use interval bounds + median as a proxy for ridge plots
(showing just the 95% HPD region), and compute exceedance probability using a normal
approximation to the posterior (mean=median, sd derived from interval width). This is
a reasonable approximation for demonstration purposes but should be clearly labeled.
### Risk 4 — Mobile responsiveness
The social media figures are desktop-first PNGs. The web app must work on mobile:
- Use CSS Grid with `auto-fit` for the chart panel (1 column on mobile, 3 on desktop).
- Make the loading animation responsive.
- Ensure touch targets in district search are ≥44px.
---
## 9. Open Questions for You
Before implementation begins, I need your input on these decisions:
### Q1 — National rates source
Chart 3 and Chart 6 compare a district's arrest rate to the **national average**
for each student group. The API doesn't have a `/national` endpoint. How should we handle this?
- **(A)** Cache national rates as a static JSON fixture (generated from the white paper / release data). [Recommended — simplest, no API changes]
- **(B)** Add a `/api/v1/national` aggregate endpoint to `crdc-arrests/api/` and have the demo call it live.
- **(C)** Approximate by fetching the top 5 largest states' rates weighted by enrollment (fast but imprecise).
### Q2 — Model selection for distribution charts
Charts 4–6 show results from multiple model types (one-year, three-year, baseline, covariate). Should users be able to toggle which models are displayed, or should we always show all four quadrants as specified in your requirements?
- **(A)** Always show all four quadrants (1yr no-cov, 1yr cov, 3yr no-cov, 3yr cov) — matches the white paper's `wp_fig_district_intervals()` layout. [Recommended]
- **(B)** Let users toggle between unified vs. stratified models via a dropdown.
- **(C)** Default to showing only the recommended model (`unified_m2_mod`) with an "advanced" expando for all four.
### Q3 — Loading animation style
The patience animation should convey that data is being fetched across multiple endpoints/models. What visual approach do you prefer?
- **(A)** Animated grid of histogram bars (one per model×group) that fill sequentially as API calls resolve, with a counter showing "X of ~40 datasets loaded". [Recommended — matches the Bayesian posterior theme]
- **(B)** A simple spinner + text message ("Fetching 12 model results from the CRDC API…").
- **(C)** An abstract animation (e.g., particles converging) that loops while loading, with no per-request feedback.
### Q4 — Deployment target
Where should this be hosted? Your infrastructure is self-hosted (`efron`, `maxwell`, etc.), but you also mentioned civilytics.org. Which deployment approach do you want me to implement first?
- **(A)** Docker container ready for your fleet (nginx + static files) — I'll write the Dockerfile and docker-compose.yml, deploy on a dev host. [Recommended]
- **(B)** Static site optimized for GitHub Pages / Vercel with CI/CD via GitHub Actions.
- **(C)** Both A and B (Docker for self-hosting, plus GH Pages config).
---
## 10. Next Steps
Once you answer the four questions above, I'll begin implementation:
1. **Week 1**: Scaffold project (Vite + React), implement design tokens & typography, build state selector + district search with "interesting" suggestions.
2. **Week 2**: Implement loading animation; fetch all model data in parallel; build Charts 1–3 (observed/descriptive).
3. **Week 3**: Build Charts 4–6 (model comparison distributions); wire up national rate fixture if Q1=A.
4. **Week 4**: Polish transitions, mobile responsiveness, Dockerfile + deployment configs; write README with usage/deployment docs.
The app will be fully self-contained — no external dependencies beyond the public CRDC API and optionally a cached JSON fixture for national rates. All code follows Civilytics visual identity as documented in `theme/_tokens.scss` and demonstrated in `social_media_posts.md`.