diff --git a/README.md b/README.md index d209f11..d017560 100644 --- a/README.md +++ b/README.md @@ -25,6 +25,22 @@ package implements. - `USCOGDATA_CACHE_DIR` — optional override for the manifest cache directory - `USCOGDATA_MANIFEST_TTL_SECS` — optional manifest re-fetch TTL (default 3600) +## Raw-parquet caveat: `survey_weight` is not an aggregation weight + +Users reading the corpus parquet directly (DuckDB, arrow) will see a +`survey_weight` column (schema v5, col 26). It is legacy Census IndFin +sample-design **metadata passed through verbatim** — the Census Bureau's own +source documentation says it "is for informational purposes only and should +not be used to derive any other statistics" (`_ReadMe_First_IndFin.txt`; +likewise `UserGuide.xls` Data User Note 8: "Do not use the weight field to +derive state or national totals"). The raw encoding is also inconsistent +across vintages (reciprocal scale most years, direct scale in 2003, a `1` +placeholder in 1967/70/71/73/2001, all-`0` in 2007–2012, `NA` for all +modern-source rows), so `sum(amt * survey_weight/10000)`-style expressions +produce silently wrong totals — including exact zeros for 2007–2012. Sum +`amt` unweighted; no uscogdata function reads this column. Full evidence: +`cog_pipeline/.superpowers/sdd/weight-semantics-findings.md`. + ## Developer notes ### Testing