.build_series_break_refs() matches `fin_code IN (<codes in the result>)`. No row's item_code is ever the literal "ALL", so the four corpus-wide entries could never match and reached no user: SB085 1977 dollar precision across the 1976/1977 boundary SB087 2002 imputation exclusion FY2002-2006 SB194 2012 dense -> sparse representation change SB086 2017 government id scheme change SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv carries an ALL @ 2012 entry describing the representation change, SO cog_explain() surfaces it". The entry shipped; the reader dropped it. A query spanning FY2011 -> FY2012 crosses the boundary where an absent cell stops meaning "Census published $0" and starts meaning "not reported", and nothing said so. Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs() on the break_year window alone -- which codes a result happens to contain is irrelevant to a caveat about the corpus. A separate field rather than more entries in series_break_refs, because an ALL caveat qualifies the whole result and folding the two together invites reading it as a caveat about one series; .build_series_break_refs() now excludes 'ALL' explicitly so the two stay disjoint by construction. cog_explain() prints them under their own "Corpus-wide caveats" heading, and cog-api passes provenance through verbatim, so the field reaches the API with no change there. On the year rule: all four entries are BOUNDARY caveats -- their own join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not being comparable across FY2012, of pre- vs post-2017 ids -- so the same `break_year BETWEEN min(years) AND max(years)` rule the code-specific path uses is the right one, and matches the issue's DoD 1. The issue's DoD 3 also asks that a FY2011 query surface SB085; that cannot hold under DoD 1 and does not hold under any reading of SB085's text, whose boundary is 1976/1977. Tested with a range that actually spans it, and flagged on the issue. Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's bundled fixture, which predates the break being catalogued. Suite: 606 pass / 0 fail / 6 skip (was 594/0/6). cog-api 357 / 0 / 8, unchanged.
45 lines
2.7 KiB
JSON
45 lines
2.7 KiB
JSON
{
|
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
"$id": "https://civilytics.org/schemas/uscogdata/provenance-v1.json",
|
|
"title": "uscogdata provenance v1",
|
|
"type": "object",
|
|
"required": ["verb", "target", "years", "scope", "manifest", "sql_query"],
|
|
"properties": {
|
|
"verb": { "type": "string" },
|
|
"call": { "type": "string" },
|
|
"target": { "type": "object" },
|
|
"years": { "type": "array", "items": { "type": "integer" } },
|
|
"category": { "type": ["string", "array", "null"] },
|
|
"basis": { "type": ["string", "null"] },
|
|
"basis_note": { "type": ["string", "null"] },
|
|
"expenditure_concept": {
|
|
"type": "string",
|
|
"enum": ["direct", "total"],
|
|
"description": "Which spending concept produced this result. 'direct' is the government's own E/F/G spending; 'total' adds its intergovernmental payments (M to local governments, L to state governments). Only 'direct' is valid for results combined across governments."
|
|
},
|
|
"expenditure_concept_note": {
|
|
"type": ["string", "null"],
|
|
"description": "How the intergovernmental leg was assembled; null for 'direct'."
|
|
},
|
|
"expenditure_concept_direct_suppressed": {
|
|
"type": "boolean",
|
|
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
|
|
},
|
|
"harmonization": { "type": "object" },
|
|
"recipe": { "type": ["object", "null"] },
|
|
"suggestions": { "type": "array" },
|
|
"scope": { "type": "object" },
|
|
"codes_summed": { "type": "object" },
|
|
"aggregate_fallback": { "type": ["object", "null"] },
|
|
"transformations":{ "type": "object" },
|
|
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
|
"corpus_break_refs": {
|
|
"type": "array",
|
|
"items": { "type": "string" },
|
|
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
|
},
|
|
"manifest": { "type": "object" },
|
|
"sql_query": { "type": "string" }
|
|
}
|
|
}
|