F-7: inst/schemas/provenance-v1.json described coverage_window as mapping
each *observed* balance_subtype, but the query at R/balance_caveats.R has
no predicate tied to the query's codes and always returns every subtype in
the mounted corpus. Took option (b) of the two the review offered -- change
the doc, not the code. Reporting all windows is the better product
behaviour (it answers 'is there a family I missed?'), it is what cog-api#26
already forwards verbatim, and option (a) would make the block empty for a
0-row result. Reworded to say the windows are corpus-wide and that
'truncated' is the query-scoped field. Pinned by a new test either way.
F-10: the 'Current State' block was self-contradictory after a partial
update -- headed 2026-04-27, claiming branch main @ d65e9fe, with a
2026-08-03 test count measured on feat/cog-balances-25 underneath it, and
listing README.md / _pkgdown.yml as outstanding when both exist and
_pkgdown.yml was edited by this branch. All numbers below re-measured on
the final tree after every other fix in this wave, not before:
788 tests (testthat::test_local()), 14 exports (NAMESPACE), 14 man/*.Rd,
2 vignettes, no docs/ (pkgdown::build_site() genuinely still outstanding,
as is the .Rbuildignore fixture entry -- both kept in the list).
The related deferred README.md item is closed with no change, per the
review's ruling: README.md enumerates no verbs at all, so naming
cog_balances would make it the only non-cog_spending verb mentioned.
70 lines
5.5 KiB
JSON
70 lines
5.5 KiB
JSON
{
|
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
"$id": "https://civilytics.org/schemas/uscogdata/provenance-v1.json",
|
|
"title": "uscogdata provenance v1",
|
|
"type": "object",
|
|
"required": ["verb", "target", "years", "scope", "manifest", "sql_query"],
|
|
"properties": {
|
|
"verb": { "type": "string" },
|
|
"call": { "type": "string" },
|
|
"target": { "type": "object" },
|
|
"years": { "type": "array", "items": { "type": "integer" } },
|
|
"category": { "type": ["string", "array", "null"] },
|
|
"basis": { "type": ["string", "null"] },
|
|
"basis_note": { "type": ["string", "null"] },
|
|
"expenditure_concept": {
|
|
"type": "string",
|
|
"enum": ["primary", "direct", "total"],
|
|
"description": "Which spending concept produced this result, defined as crosswalk spend_subtype sets (never item-code prefixes). 'primary' (the default) is the government's own service provision: operations + capital + assistance. 'direct' adds interest on debt and insurance trust benefit payments (Census's published Direct Expenditure). 'total' adds intergovernmental payments (M to local governments, L to state government, Q11/Q12/Q18 to school systems). Only 'primary' and 'direct' are valid for results combined across governments."
|
|
},
|
|
"expenditure_concept_note": {
|
|
"type": ["string", "null"],
|
|
"description": "How the intergovernmental leg was assembled; null for 'primary' and 'direct'."
|
|
},
|
|
"expenditure_concept_direct_suppressed": {
|
|
"type": "boolean",
|
|
"description": "TRUE when expenditure_concept = 'total' and at least one requested (year, category) has intergovernmental rows but NO Direct rows in this corpus (typically a legacy aggregate-only family) -- those result rows report the intergovernmental leg alone, not Direct + IG. Always FALSE for expenditure_concept = 'primary' or 'direct'. See the affected rows' `notes` for the recovering recipe, if any."
|
|
},
|
|
"revenue_concept": {
|
|
"type": "string",
|
|
"enum": ["general", "total"],
|
|
"description": "Which revenue concept produced this result, defined as crosswalk revenue_subtype sets (never item-code prefixes). 'general' (the default) is Census General Revenue: own_source + federal + state + local_aid. 'total' is Census Total Revenue: general plus utility, liquor store and insurance trust revenue. Census defines the first by subtracting the other three from the second (manual section 4.3). Meaningful for cog_revenue() results; spending results carry the default.",
|
|
"$comment": "The employee-retirement (X) codes inside insurance_trust stop at FY2016, so a 'total' series steps at the FY2016/FY2017 seam for collection-scope reasons (series breaks SB197-SB202)."
|
|
},
|
|
"harmonization": { "type": "object" },
|
|
"recipe": { "type": ["object", "null"] },
|
|
"suggestions": { "type": "array" },
|
|
"scope": { "type": "object" },
|
|
"codes_summed": { "type": "object" },
|
|
"aggregate_fallback": { "type": ["object", "null"] },
|
|
"transformations":{ "type": "object" },
|
|
"series_break_refs": { "type": "array", "items": { "type": "string" } },
|
|
"completion": {
|
|
"type": "object",
|
|
"description": "What `complete = TRUE` filled. `applied` is FALSE on an ordinary query. `rows_filled` counts cells added to the requested grid, and `absence_means` maps each requested year to the meaning of an absent cell there ('census_zero' in a dense_source year, 'not_reported' in a sparse_source one). Filled rows carry `value_source` in the result: 'reported', 'census_zero' (amount 0 -- Census published $0), or 'not_reported' (amount NA -- unknown).",
|
|
"properties": {
|
|
"applied": { "type": "boolean" },
|
|
"rows_filled": { "type": "integer" },
|
|
"absence_means": { "type": "object" }
|
|
}
|
|
},
|
|
"corpus_break_refs": {
|
|
"type": "array",
|
|
"items": { "type": "string" },
|
|
"description": "Ids of catalogued series breaks whose fin_code is the literal 'ALL' -- caveats about the corpus as a whole (dollar precision across 1976/1977, imputation exclusion from 2002, the dense -> sparse representation change at 2012, the government id scheme change at 2017) rather than about one item code. Selected on the break_year window alone, so they do not depend on which codes a result contains. Disjoint from series_break_refs by construction: an entry qualifies the whole result, not one series."
|
|
},
|
|
"balance_caveats": {
|
|
"type": ["object", "null"],
|
|
"description": "Present only on cog_balances() results (null/absent for cog_spending()/cog_revenue()). `not_gaap` is always TRUE and `not_gaap_note` explains that Census holdings are gross -- no liabilities are netted -- so they are NOT comparable to a GAAP fund balance. `coverage_window` maps EVERY balance_subtype present in the mounted corpus -- not only the ones this query observed -- to its measured [min year, max year] there (never hardcoded), so a caller can see which families exist and over what span before deciding they missed one. `truncated` is the query-scoped field: it lists only the subtypes this result actually observed whose coverage_window does not fully span the requested years.",
|
|
"properties": {
|
|
"not_gaap": { "type": "boolean" },
|
|
"not_gaap_note": { "type": "string" },
|
|
"coverage_window": { "type": "object" },
|
|
"truncated": { "type": "array", "items": { "type": "string" } }
|
|
}
|
|
},
|
|
"manifest": { "type": "object" },
|
|
"sql_query": { "type": "string" }
|
|
}
|
|
}
|