fix: rank district suggestions over the full state, not the first 500 rows
DistrictSearch built its "most arrests" list from /estimates?state=XX&year=21-22&limit=500. That endpoint returns rows ORDER BY LEAID, RACE, SEX at eight rows per district, so a 500-row cap is not a sample of the state — it is the ~62 lowest-LEAID districts in it. Measured against California (11,488 rows, 1,715 districts): the old read covered 68 districts, and 6 of the true top 8 were invisible to it. It suggested districts with 1 and 2 arrests as the state's most notable, while San Diego Unified (178), Fresno Unified (77) and Kern High (69) never appeared. Replaced with a committed fixture, public/data/top_districts.json, generated by scripts/build-top-districts.mjs. The script pages each state to completion using meta.total from the response envelope and fails loudly on a short read, since a silent truncation there would reintroduce exactly this bug. 51 states, 135 requests, ~104KB, following the national_rates.json precedent. Re-run it only when a new CRDC wave lands. The search screen also loses a multi-second fetch on every visit, and the hardcoded "Try Derby (KS), Paterson (NJ)" hint goes with it — the real list supersedes it. Degrades to search-only if the fixture is missing. fetchStateDistricts() is kept for scripts and ad-hoc use, with its JSDoc now warning that any short read ranks by LEAID. pages.yml gains public/** in its paths filter: the fixture ships with the build, so regenerating it has to be able to trigger a deploy on its own.
This commit is contained in: