# CRDC Arrests API Demonstration App — Design Proposal ## 1. Overview A web application hosted at `civilytics.org` that demonstrates the [CRDC School Arrest Rate API](https://crdc-api.civilytics.org/api/v1/) by letting any visitor explore school-based arrest estimates for any U.S. school district, with a strong visual narrative built around Bayesian model comparisons. The app is a **frontend-only static application** that calls the public read-only API directly from the browser. No server-side code, no database, no auth — just HTML/CSS/JS (with optional Docker container for local dev or self-hosting behind a reverse proxy). --- ## 2. User Flow ``` ┌─────────────┐ ┌──────────────────┐ ┌─────────────────┐ │ State │ → │ District Search │ → │ Loading (anim) │ → ╔══════════════╗ │ selector │ │ w/ suggestions │ │ patience msg │ ║ Chart panel ║ └─────────────┘ └──────────────────┘ └─────────────────┘ ║ 6 panels ║ ║ see §4 ║ ╚══════════════╝ ``` ### Step 1 — State prompt - Full-screen landing card with Civilytics branding. - Dropdown / typeahead for U.S. state (50 states + DC). Two-letter codes map to the API `state` parameter directly. - On select → slide to Step 2. ### Step 2 — District search - Search-as-you-type input, hitting `/api/v1/districts?q=&state=XX`. - **Suggested districts** appear as a horizontal carousel of "interesting" picks: - Top N by total arrest count in the most recent wave (2021-22). - Fetched once per state via `limit=500` from `/estimates?state=XX&year=21-22` sorted client-side by `observed_arrests`. - Each suggestion card shows: district name, enrollment (in small), observed arrests. Clicking a suggestion jumps straight to loading + charts. - Typing filters the list live; selecting an item from either source proceeds. ### Step 3 — Loading / patience animation - A full-screen overlay with an animated visual that conveys "data is being fetched across multiple endpoints, please be patient." - The animation should reflect the **multi-model comparison** theme: e.g., a grid of small histogram-like bars (representing posterior draws) that animate in and out row by row as each model's endpoint responds. - Estimated total API calls per district lookup: ~12–15 (one per model × race × sex combination, filtered to the selected district). The animation should scale visually with progress. ### Step 4 — Chart panel (6 charts in a 3×2 grid) All charts follow Civilytics visual identity: warm paper background (`#FAF7F2`), civic navy text (`#0E1A2B`), ember accent (`#C25311`). Data-viz palette uses the supporting colors from `_tokens.scss` (teal, plum, moss, brass). #### Row 1 — Observed & descriptive **Chart 1: Arrests over time by CRDC wave (line chart)** - X-axis: CRDC waves (`2015-16`, `2017-18`, `2021-22`). - Y-axis: total observed arrest count. - Line + points; each point labeled with the rate per 1,000 students. - Source: `/estimates?leaid=XXXXX&year=...` across all three years (or a single call to `/districts/{leaid}` which returns all demographics for one district). **Chart 2: Arrest rate by student group — most recent year (bar chart)** - Bars for each race×sex combination in `AM/BL/HI/WH × F/M`. - Y-axis: arrests per 1,000 students. - Color-coded by the supporting palette; legend shows full labels. **Chart 3: District vs. national — top student group (comparison chart)** - Identify the district's highest-arrest-rate group. - Side-by-side bars or a small multiples comparison against the corresponding **national** rate for that same group. - National rates fetched via `/states?state=XX&race=&sex=&year=21-22` aggregated, or more precisely from the national summary (which may need to be computed as an aggregate across all states). #### Row 2 — Bayesian model distribution comparisons All four quadrants show results for the **selected district**, comparing: - **Column A**: One-year models (`unified_m1_mod`, `unified_m2_mod`) vs. - **Column B**: Three-year models (`unified_m3_mod` through `unified_m5_mod`). - Within each column, rows differentiate baseline (no covariate) from covariate models. **Chart 4: Predicted draws by year vs. observed (scatter / point-range)** - For three-year models only: for each of the 3 waves, show the model's predicted median and 95% interval alongside the observed value as a separate marker. - Layout: x-axis = wave; y-axis = arrest count; points dodge left (model) vs. right (observed). **Chart 5: Observed rate per group against model distribution (density / ridge)** - For each student group, show the observed rate and overlay the posterior draw distributions from all four model types as ridgeline or violin plots. - Mirrors `wp_fig_group_density()` in the white paper. **Chart 6: Probability district exceeds national rate per group (bar chart)** - For each race×sex group, compute P(district rate > national rate) using the posterior draws from a chosen model (e.g., the default `unified_m2_mod`). - Bars colored by threshold crossing; annotated with exact probability. --- ## 3. Data Sources & API Endpoints Used | Purpose | Endpoint | Frequency | |---|---|---| | State list / validation | Hardcoded enum (`ALLOWED_STATES`) | Once | | District search | `/api/v1/districts?q=&state=` | On keystroke (debounced) | | "Interesting" suggestions | `/api/v1/estimates?state=XX&year=21-22` sorted by `observed_arrests desc` | Once per state | | Single district, all demographics | `/api/v1/distimates/{leaid}` or `/estimates?leaid=` with year/model filters | ~6 calls × 3 years = 18+ | | National comparison rates | `/api/v1/states?state=&race=&sex=&year=21-22` aggregated across states, OR a dedicated national endpoint if available | Once per group | | Model metadata | `/api/v1/models` | Once (cache) | **Key data structures returned by the API:** The `/estimates/{leaid}` endpoint returns one row per `race × sex × year × model` with: ```json { "leaid": "...", "lea_name": "...", "state": "TX", "race": "BL", "sex": "M", "year": "21-22", "model": "unified_m2_mod", "stu_enroll": 35963, "observed_arrests": 1, "rate_median": 0.42, "rate_lower": 0.18, "rate_upper": 0.91, "count_median": 15.3, "count_lower": 6.7, "count_upper": 28.9 } ``` The `/draws` endpoint returns a Hugging Face Parquet shard URL + DuckDB SQL for bulk posterior draw access — useful if we need to compute custom quantities like P(district > national) without round-tripping through multiple API calls. However, for a demo app that runs in the browser, relying on 10 HF Parquet shards per model is impractical (80 GB total). The summary endpoints (`/estimates`) provide enough aggregated information for all six charts using only `rate_median`, `count_median`, and interval bounds — **no raw draw access needed**. --- ## 4. Visual Design Language (from reference materials) ### Color palette ```css :root { --cv-paper: #FAF7F2; /* warm paper background */ --cv-ink: #0E1A2B; /* primary text — civic navy/black */ --cv-navy-600:#22406A; /* links, accents */ --cv-accent: #C25311; /* ember — alerts, highlights, eyebrows */ /* Data-viz supporting colors (from _tokens.scss) */ --teal-600: #1F6F70; --plum-600: #6B3A5E; --moss-600: #4A6B2F; --brass-600: #B8751C; } ``` ### Typography - **Display**: Libre Franklin / Inter — headings, stat callouts. - **Body**: Inter — UI text, labels, captions. - **Mono**: JetBrains Mono — code snippets, API URLs in footers. ### Chart patterns (from social media posts & white paper) 1. **Ridgeline density plots** for posterior draws (`geom_density_ridges`). 2. **Pointrange / error bar charts** comparing model intervals to observed values. 3. **Faceted small multiples** — always split by `covariate × time` (baseline vs. covariate; one-year vs. three-year). 4. **Transparent PNG export with watermark logo** in bottom-right corner. 5. **"Stat callout" hero numbers** for key metrics (e.g., total arrests, rate per 1k). --- ## 5. Technology Stack Recommendation ### Recommended: React + Vite + vanilla CSS (static site) | Layer | Choice | Rationale | |---|---|---| | Framework | **React 19** (no framework overhead; Vite dev server) | Component model for charts, built-in state management via hooks | | Build tool | **Vite** | Fast HMR, native ES modules, trivial static export (`npm run build`) | | Charting | **`@visx/visx`** or plain SVG/CSS animations | Lightweight, no heavy deps; we control every pixel to match Civilytics style | | HTTP client | `fetch` with AbortController + exponential backoff | No extra dependency; browser-native | | Styling | Plain CSS custom properties (no Tailwind) | Zero-runtime, matches the existing SCSS token system exactly | | Deployment | Static site on any host (GitHub Pages, Vercel, Netlify) or Docker nginx container | Self-hostable; no backend required | #### Why not Shiny / Quarto? The API returns JSON and we need a dynamic SPA with loading states, debounced search, and animated transitions. A static React app is the most natural fit. R/Shiny would add unnecessary server-side complexity for what is fundamentally a browser-based data visualization demo. The existing `crdc-arrests` project uses Quarto + ggplot2 for reports; this demo app is a different artifact with different requirements. #### Why not Svelte or Vue? React has the largest ecosystem, best tooling (Vite), and most team familiarity. For a 6-chart SPA it's more than sufficient without being overkill. ### Docker option A minimal `nginx:alpine` container serves the built static files. ~20 MB image. Can be run on your self-hosted fleet (`efron`, `maxwell`, etc.) behind Caddy/Traefik with TLS via Let's Encrypt — consistent with how you host other civilytics.org services. --- ## 6. Build & Deployment Strategy ### Local development ```bash npm install # one-time: installs React, Vite, dev deps npm run dev # starts Vite on localhost:5173 # Edit src/App.jsx / src/components/*.jsx — hot reload ``` ### Production build (static) ```bash npm run build # outputs dist/ with index.html + assets # Upload dist/ to any static host, or: docker build -t crdc-demo . && docker run -p 8080:80 crdc-demo ``` ### Docker deployment (self-hosted on your fleet) 1. Build image locally or via CI: `docker buildx build --platform linux/amd64 -t registry.civilytics.org/crdc-demo:latest .` 2. Push to local registry on `maxwell`. 3. Deploy via docker-compose or a simple systemd service with nginx container + Caddy reverse proxy for TLS termination. ```yaml # docker-compose.yml (minimal) services: app: image: crdc-demo:latest ports: ["8080:80"] restart: unless-stopped ``` ### CI/CD (optional, GitHub Actions) - On push to `main`: run lint + build, publish Docker image to local registry. - No automated deployment — you control when new versions go live on the fleet. --- ## 7. File Structure ``` crdc-demo/ ├── public/ # Static assets (logo, favicon) │ └── civilytics-logo.svg ├── src/ │ ├── components/ # Reusable UI pieces │ │ ├── StateSelector.jsx │ │ ├── DistrictSearch.jsx │ │ ├── LoadingAnimation.jsx │ │ └── StatCallout.jsx │ ├── charts/ # The 6 chart components │ │ ├── ArrestsOverTime.jsx │ │ ├── RateByGroupBar.jsx │ │ ├── DistrictVsNational.jsx │ │ ├── ModelDrawsComparison.jsx │ │ ├── ObservedRateDensity.jsx │ │ └── ExceedanceProbability.jsx │ ├── hooks/ │ │ ├── useApi.js # fetch wrapper with retry/backoff │ │ └── useDistrictData.js # orchestrates all 6 charts' data needs │ ├── styles/ │ │ ├── tokens.css # Civilytics design tokens │ │ └── main.css │ ├── App.jsx # Main router: state → search → loading → charts │ └── main.jsx # React entry point ├── Dockerfile ├── vite.config.js └── package.json ``` --- ## 8. Key Implementation Notes & Risks ### Risk 1 — API response time for multi-model data Fetching all models × race×sex combinations for a single district requires ~40 individual `/estimates` calls (5 unified + 5 stratified models × 8 groups). Each API call may take 200–500ms. **Mitigation**: use `Promise.allSettled()` to fire all requests in parallel; show progress as batches resolve. The loading animation should reflect this batching visually (e.g., rows of bars filling left-to-right). ### Risk 2 — National rate computation The `/states` endpoint returns per-state aggregates, not a national total. To get national arrest rates by student group for Chart 3 and Chart 6: - **Option A**: Fetch all states (`limit=100`, iterate through ~50 pages) and sum. Too slow for a browser app. - **Option B**: Add a `/api/v1/national` endpoint to the API (server-side aggregate). Requires modifying `crdc-arrests/api/` — quick R/SQL change but needs your approval. - **Recommended short-term**: Cache national rates as a static JSON file generated during data release and committed alongside this app. The white paper already computes these values; we can extract them into a small fixture. ### Risk 3 — Chart complexity (Chart 5 & 6) Charts 5 (density/ridge) and 6 (exceedance probability) require either raw posterior draws or sufficient summary statistics to reconstruct distributions. The `/estimates` endpoint provides `count_median`, `count_lower`, `count_upper` but not the full draw distribution. **Mitigation**: Use interval bounds + median as a proxy for ridge plots (showing just the 95% HPD region), and compute exceedance probability using a normal approximation to the posterior (mean=median, sd derived from interval width). This is a reasonable approximation for demonstration purposes but should be clearly labeled. ### Risk 4 — Mobile responsiveness The social media figures are desktop-first PNGs. The web app must work on mobile: - Use CSS Grid with `auto-fit` for the chart panel (1 column on mobile, 3 on desktop). - Make the loading animation responsive. - Ensure touch targets in district search are ≥44px. --- ## 9. Open Questions for You Before implementation begins, I need your input on these decisions: ### Q1 — National rates source Chart 3 and Chart 6 compare a district's arrest rate to the **national average** for each student group. The API doesn't have a `/national` endpoint. How should we handle this? - **(A)** Cache national rates as a static JSON fixture (generated from the white paper / release data). [Recommended — simplest, no API changes] - **(B)** Add a `/api/v1/national` aggregate endpoint to `crdc-arrests/api/` and have the demo call it live. - **(C)** Approximate by fetching the top 5 largest states' rates weighted by enrollment (fast but imprecise). ### Q2 — Model selection for distribution charts Charts 4–6 show results from multiple model types (one-year, three-year, baseline, covariate). Should users be able to toggle which models are displayed, or should we always show all four quadrants as specified in your requirements? - **(A)** Always show all four quadrants (1yr no-cov, 1yr cov, 3yr no-cov, 3yr cov) — matches the white paper's `wp_fig_district_intervals()` layout. [Recommended] - **(B)** Let users toggle between unified vs. stratified models via a dropdown. - **(C)** Default to showing only the recommended model (`unified_m2_mod`) with an "advanced" expando for all four. ### Q3 — Loading animation style The patience animation should convey that data is being fetched across multiple endpoints/models. What visual approach do you prefer? - **(A)** Animated grid of histogram bars (one per model×group) that fill sequentially as API calls resolve, with a counter showing "X of ~40 datasets loaded". [Recommended — matches the Bayesian posterior theme] - **(B)** A simple spinner + text message ("Fetching 12 model results from the CRDC API…"). - **(C)** An abstract animation (e.g., particles converging) that loops while loading, with no per-request feedback. ### Q4 — Deployment target Where should this be hosted? Your infrastructure is self-hosted (`efron`, `maxwell`, etc.), but you also mentioned civilytics.org. Which deployment approach do you want me to implement first? - **(A)** Docker container ready for your fleet (nginx + static files) — I'll write the Dockerfile and docker-compose.yml, deploy on a dev host. [Recommended] - **(B)** Static site optimized for GitHub Pages / Vercel with CI/CD via GitHub Actions. - **(C)** Both A and B (Docker for self-hosting, plus GH Pages config). --- ## 10. Next Steps Once you answer the four questions above, I'll begin implementation: 1. **Week 1**: Scaffold project (Vite + React), implement design tokens & typography, build state selector + district search with "interesting" suggestions. 2. **Week 2**: Implement loading animation; fetch all model data in parallel; build Charts 1–3 (observed/descriptive). 3. **Week 3**: Build Charts 4–6 (model comparison distributions); wire up national rate fixture if Q1=A. 4. **Week 4**: Polish transitions, mobile responsiveness, Dockerfile + deployment configs; write README with usage/deployment docs. The app will be fully self-contained — no external dependencies beyond the public CRDC API and optionally a cached JSON fixture for national rates. All code follows Civilytics visual identity as documented in `theme/_tokens.scss` and demonstrated in `social_media_posts.md`.