2c532bde19cab30092f3620a5f9af69af0447c11
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4b23dbd9f4
|
feat: revenue_concept = c("general", "total") off the crosswalk (#12)
Closes the last blocked test in the suite. Owner ruled both halves of the open question yes on 2026-07-30. `cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`, with Census's two published concepts defined as crosswalk `revenue_subtype` sets rather than item-code prefixes: general = own_source + federal + state + local_aid (the default) total = general + utility + liquor_store + insurance_trust The manual defines the first by subtracting the other three from the second (4.3), so both are computable only once all four families are named -- which cog_pipeline#79 does. Insurance trust now includes the employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes. - inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue subtype; the concept narrows in R via the existing subtype_scope machinery, exactly as expenditure_concept narrows spending_long. - cog_explain() now prints each verb's OWN concept. It previously printed `expenditure_concept` unconditionally, so a cog_revenue() caller was told "Concept: primary" -- a spending concept their result has nothing to do with. - Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows). Corrected two stale expectations in the blocked test while un-skipping it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state governments and is nonzero for Wisconsin; X04 is an exhibit code for an INTRAgovernmental transfer that Census's own "Total Emp Ret Rev" excludes. Verified against that Census field: the right set is X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of $33,377,093k predated the Y codes being classified -- complete Total Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785 + X 2,283,883). Behaviour change worth knowing: `general` is now STRICT Census General Revenue, so utility and liquor store revenue leave the default. Measured on the fixture that is 15.9% of what cog_revenue() returned for cities, vs 1.2% for states and 1.7% for counties. Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had no skipped tests. Closes #12 |
||
|
|
93300ae0c1
|
feat: three-concept expenditure model classified by crosswalk membership (#11)
R-CMD-check / check (push) Successful in 3m5s
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:
primary = operations + capital + assistance
direct = primary + interest + insurance_benefits (Census Direct)
total = direct + intergovernmental (M/L/Q via ig views)
- inst/sql: flow views (20-25) select by crosswalk membership;
summary_categories moves to 11- so it registers before them (DuckDB
binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
harmonization exclusion count, and the complete = TRUE grid;
flow_prefixes survives only to scope recipe suggestions.
cog_geographic_rollup/cog_peer_compare accept primary|direct, still
refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
(uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
proof is asserted against the crosswalk, not the default
cog_revenue() call (which stays General Revenue pending #12).
Suite: 696 pass / 0 fail / 1 skip (#12, expected).
Closes #11
|
||
|
|
af85a23ea7
|
feat: complete = TRUE fills absent cells with their meaning (#18)
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide era's explicit zeros, which made absence ambiguous: <= FY2011 dense_source absent => Census published $0 >= FY2012 sparse_source absent => not reported, unknown A wide-era query whose cells were all $0 had begun returning nothing at all, with no way to get them back -- strictly less than the reader exposed before, which is why #64 filed this follow-on. complete = TRUE fills the requested grid from `code_set` and stamps every row with value_source: "reported", "census_zero" (amt 0), or "not_reported" (amt NA). The NA is the point. Filling a modern absence with 0 would invent data, which is exactly the error the representation contract exists to prevent -- and it makes this strictly MORE informative than the pre-sparsification corpus, which could not tell a published zero from an unreported cell either. Measured on the fixture, Broward County: FY2011 returns 28 reported + 16 census_zero; FY2019 returns 30 reported + 14 not_reported. The five categories that walkthrough finding F-006 read as "retired at FY2012" now report themselves correctly as census_zero before and not_reported after. Scoping decisions, each of which would invent rows if taken loosely: - The grid is per government TYPE (code_set.type). Filling against the union of all types would give a county cells like "state IG transfer to school districts", indistinguishable from real census zeros. - NOT is_aggregate, mirroring spending_long/revenue_long. Without it the grid offers cells those views never return, so each would fill as a phantom $0. - Filling happens BEFORE per_capita and inflation, so a census_zero stays 0 through both and a not_reported stays NA rather than becoming 0. Two new views (36-representation, 37-code_set) are gated on the manifest LISTING those tables, not on schema_version. Sparsification did not bump the version -- the fixture this package shipped against until 2026-07-30 was already v6 and carried neither table -- so a version gate would register a view over a missing file and fail at CREATE VIEW time on exactly the corpora the check exists to tolerate. with_corpus_missing_representation() models that corpus and asserts the abort. Refused where the fill would be guesswork, both classed uscogdata_complete_unsupported: a recipe defines its own component codes and never touches summary_categories; the intergovernmental leg deliberately keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones code_set describes. Expected cell sets in the tests are computed from the corpus parquet directly, never through the verb -- verifying what a filter does through that same filter proves nothing. Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed separately. Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean. |
||
|
|
4de915b557
|
feat: cog_recipes + recipe= + signposting suggestions
Adds cog_recipes() to list the curated harmonization_recipes catalog (24 recipes / schema_version >= 5), and a recipe= argument on cog_spending()/ cog_revenue() that runs a recipe's generic multi-code join instead of the category view: SUM(amt * weight) across whichever component codes are present for a (year, canonical_govid), scoped by gov_type_scope. The join deliberately does not filter is_aggregate -- the wide era (<= 2011) exposes these split families (corrections 04+05, IG *89/*47, U4- rents, etc.) ONLY as aggregate rows, with leaf codes first appearing in 2012, so excluding aggregates would zero out the wide-era half of every recipe. This is safe by corpus construction: wide-era rows are aggregate-only, modern rows are leaf-only, and every component is year-scoped, so there is no double-counting. recipe= is mutually exclusive with category=; the result's subtype column reads "recipe" and category reads the recipe's label. Adds recipe-component-driven signposting: when a basis="harmonized" + category query comes back with zero rows in a requested year, and a harmonization recipe covering that category would actually produce rows for this government in that year (via the same join .run_recipe() uses), the recipe is surfaced in provenance$suggestions plus one cli::cli_inform() message. This is deliberately keyed off recipe components rather than harmonization_map's suggested_recipe_id column (which is empty on every live row -- the wide era's split families are NA-by-construction via aggregate exclusion, not an NA ruling to hang a suggestion off of). Also populates the previously-always-empty provenance$series_break_refs (schema v5 only: series_breaks_pq rows whose fin_code is among the observed codes and whose break_year falls in the requested span), and extends cog_explain() with Basis/Harmonization/Recipe/Suggestions/Series breaks sections. |
||
|
|
7818cd2b1a
|
feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
|
||
|
|
cbc867bed1 | docs(per-capita): refresh roxygen for per-year denominator + pop_source | ||
|
|
c682e6547d
|
feat: cog_spending + cog_revenue + cog_explain
Three core query verbs over the spending_annotated / revenue_annotated DuckDB views. Each verb accepts vector govid, vector years, optional category filter, per_capita flag, and adjust_to_year for CPI-U real-dollar conversion (bundled index). Amounts are returned in full USD (SUM(amt) * 1000) so callers can freely rescale to millions/billions. The $1,000s -> $USD conversion is recorded in provenance$transformations$units_conversion. Every result carries an attr(., "provenance") list matching inst/schemas/provenance-v1.json. cog_explain() prints the structured form via cli or returns the raw list for MCP/JSON consumers. Also: .fetch_or_cache_manifest() now handles local fixture paths so tests can point USCOGDATA_FIXTURE_URL at the pipeline publish_cache/ without a working HTTP server. Tests: 80 pass / 0 fail. devtools::check() 0E/0W/2N (both notes pre-existing / environmental). |