Commit Graph
16 Commits
Author SHA1 Message Date
jared c7260cb20c fix: scope total's coverage-gap detection to the Direct leg; require IG category rows (C1, C2)
C1: spending_long/spending_long_harmonized filter NOT is_aggregate but
ig_long deliberately doesn't (legacy IG lives on aggregate rows), so a
legacy aggregate-only family (e.g. Corrections pre-2012) can survive on
the IG leg while Direct is suppressed. expenditure_concept = "total"
then UNIONs an IG-only figure that reads as a plausible Total, and the
coverage-gap suggestion machinery -- fed the UNION'd result -- saw the
surviving IG row as coverage and stayed silent.

  (a) .build_suggestions() is now fed a Direct-leg-only view of the
      result (IG rows filtered out before the gap-years computation),
      so the recipe hints fire for "total" exactly as they do for
      "direct".
  (b) Any row where IG has dollars but Direct has none for the same
      (year, canonical_govid, category) is now flagged: the row's
      `notes` name the recovering recipe (drawn from the Direct-leg
      suggestions), and provenance gains an explicit
      `expenditure_concept_direct_suppressed` boolean plus an appended
      warning on `expenditure_concept_note` -- both cheap for a
      downstream consumer (cog-api passes provenance through verbatim)
      to test, rather than silently asserting Direct + IG when that
      arithmetic didn't happen.

Measured before/after on AL state government, Corrections, 2011:
"total" already correctly returns the corpus's actual IG-only figure
($31,358,000, vs. true Direct of $521,651,000 via recipe =
"corrections_combined"), but before this fix it did so with 0
suggestions and an unqualified "Total = Direct + IG" note; after, it
fires 3 recipe hints and both the row notes and provenance say plainly
that Direct is unavailable through this basis.

C2: the 66 M/L summary_categories rows arrived via cog_pipeline PR #59
with no schema_version bump, so schema_version can't gate "total" --
a pre-#59 corpus can report any supported schema_version and still
have zero M/L category rows, in which case ig_annotated's LEFT JOIN
silently produces NA category/spend_subtype (0 rows for a specific
category, or one invisible NA-subtype group for category = NULL). New
.require_ig_categories() checks summary_categories directly and aborts
with class uscogdata_ig_categories_unsupported, naming PR #59 and
directing the user to a newer corpus.

Reconciles tests/testthat/test-views.R's v4-shaped-corpus test (whose
synthetic summary_categories carries only one E36 row) by asserting
the new guard fires against that same connection, rather than leaving
the two silently contradictory.
2026-07-27 12:06:03 -04:00
jared 7913b0f664 feat: record expenditure_concept in provenance and its JSON schema
Always populated, never implicit, so a downstream artifact says which concept
produced it. cog-api passes provenance through verbatim.
2026-07-27 10:27:12 -04:00
jared e2088458e1 fix: address Task 3 code review (bool_or, invariant tests, guards, docs)
Nine review items on the expenditure_concept = direct|total feature:

- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
  bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
  dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
  wide-era data puts a $0 leaf row in the same group as the real aggregate
  row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
  across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
  aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
  scoped to the aggregate flag rather than bare code presence (M89/L89
  continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
  pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
  retained, NULL harmonized_code coalesced, L-- excluded), rather than
  relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
  necessary (not just incidental) against a corpus whose `long` genuinely
  lacks a harmonized_code column, and rewrote the misleading "v5-only
  parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
  test, guarded .verb_spendrev() against total on a non-spending
  view_base, added a roxygen caveat against summing total across levels
  of government, and replaced an uncheckable corpus-wide SQL comment
  figure with a fixture-verifiable one.

Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
2026-07-27 10:00:07 -04:00
jared fefd4fe969 feat: expenditure_concept = direct|total in cog_spending()
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
2026-07-27 09:28:17 -04:00
jared 7ed1da9b79 fix: drop the inert K prefix from the spending flow prefixes
K matches zero rows corpus-wide (audited pipeline-side). Numerically inert;
removed so the code stops implying a prefix the data never had.
2026-07-27 09:16:41 -04:00
jared 9240a18ea3 chore: regenerate fixture corpus with the intergovernmental category rows
Picks up pipeline PR #59: summary_categories now carries 66 M/L rows under
spend_subtype = intergovernmental (194 -> 260 rows).

cog_categories() (R/categories.R) has no item-code prefix filter, so the new
IG rows surface immediately as a third spend subtype; this broke
test-categories.R:18's closed enumeration. Adjudicated (2026-07-27): this is
correct behavior, not a regression -- cog_categories() is a discovery verb
documented to surface valid category values, and after Task 3 lands users
will see spend_subtype = intergovernmental in cog_spending(expenditure_concept
= total) results. Widened the subtype assertion and added positive coverage
asserting the intergovernmental subtype and that it reuses existing functional
categories (plus Other Education, pipeline #58). R/categories.R itself is
unchanged -- its behavior was already right.

Suite: PASS 469, FAIL 0 (baseline 467 + widened assertion + 2 new
expectations).
2026-07-27 09:11:10 -04:00
jared 3583c05852 chore: regenerate fixture corpus from the Option B (single-flavor aggregate) publish tree
R-CMD-check / check (push) Successful in 2m49s
R-CMD-check / check (pull_request) Successful in 2m38s
Source: cog_pipeline publish_cache built 2026-07-23T16:06:45Z at 4f992a0
(pipeline PR #43, issue #28 Option B ruling: legacy aggregate families now
publish only the H2-designated Direct flavor; Census Total = code + M-code).

Fixture delta, verified against the prior partition: year=2011 loses
173,394 non-designated aggregate rows (3,037,606 -> 2,864,212; -05 max
rows/gov 2 -> 1), leaf rows byte-identical; 2012/2019/2020 partitions,
metadata parquets, and docs unchanged. manifest.json resyncs sha256 /
row_count / size_bytes for the changed partition.

Full suite vs the regenerated fixture: 467 PASS / 0 FAIL / 0 WARN / 0 SKIP
— zero pin adjudications needed (reader verbs filter is_aggregate rows and
no main test pins legacy aggregate counts).
2026-07-23 12:16:05 -04:00
jaredandClaude Opus 4.8 e813ffd3aa feat: accept corpus schema v6 (FIPS geography harmonization); v6 fixture
Schema v6 (cog_pipeline 2026-07-22) renamed the long table's
fips_state_code/fips_county_code to fips_state_asof/fips_county_asof and added
cog_legacy_state/cog_legacy_county (26 -> 28 cols). This package references
none of those columns and its geography always came from
canonical_fips_xwalk (already present-based), so acceptance is a version-set
bump: supported = c(4L, 5L) -> c(4L, 5L, 6L) in .validate_schema() and
cog_open(). A prominent note in .validate_schema() documents the SILENT
semantic change for raw-long readers: long fips_state/fips_county are now
PRESENT/harmonized geography (carried back per government), not as-of-year.

Fixture regenerated from the published v6 tree (schema_version 6, 28 cols).
Test updates:
  * test-manifest.R: v6 accepted; boundary rejection moves to v7.
  * test-spending.R: the na_rows_excluded pin (0) predated the Task 18 map
    extension, which added E/F/G-prefix discontinued_na rulings (E21/F21/G21,
    Education NEC local, SB184-186). Broward's 2011 partition zero-pads
    exactly those codes: 3 NA-harmonized rows excluded, all amt=0, so the
    excluded AMOUNT pin stays 0. Data-verified against the v6 fixture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 13:16:11 -04:00
jared 7818cd2b1a feat: basis= harmonized/raw with v4/v5 dual-accept
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.

Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.

Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
2026-07-18 23:19:17 -04:00
jared 70cf553828 chore: regenerate fixture corpus from cog_pipeline Phase Q4 publish (a082b26)
R-CMD-check / check (push) Successful in 2m18s
Metadata tables refreshed from the Phase Q4 corpus (schema 4, pipeline_commit
a082b26): canonical_fips_xwalk.parquet now carries the extended population
bridge (pop_confidence exact 15.7% -> 97.3%); canonical_alias.parquet reflects
the Q3 rename continuations. 2019/2020 long partitions unchanged (continuations
remap only pre-2017 predecessor rows). Suite 336 PASS / 0 FAIL / 0 WARN.
2026-07-13 19:36:10 -04:00
jared 570a9408a2 feat: regenerate fixture corpus from Phase P publish tree + committed regen script
Adds data-raw/regenerate_fixture_corpus.R, parameterized by publish-cache
path, so the fixture is never a manual rebuild again. Regenerates the
2019/2020 long partitions (byte-for-byte copy), the full 39,377-row
canonical_fips_xwalk master, the new 117,503-row canonical_alias lookup
table, and summary_categories from the Phase P publish tree; resyncs the
four fixture docs; and hand-builds manifest.json with schema_version 4
and freshly computed sha256/row_count/size_bytes for every shipped file.
Fixture grows from 3.7MB to 5.5MB, well under the 25MB budget.
2026-07-11 09:31:49 -04:00
jared ed9658d267 feat(sql): add gov_population_yearly view
Exposes one row per (year, canonical_govid) drawn from long.population.
Used by per-capita denominators and peer matching.
2026-04-29 15:12:53 -04:00
jared a28fb2e19b chore(fixture): regenerate against cog_pipeline Layer 1 + Layer 2 fixes
Refreshes the bundled fixture corpus against the upstream resolver fix
(gate place-less fallback to states/counties only) and the Phase O
extended FIPS xwalk (post-2012 incorporations, Utah metro townships,
Connecticut planning regions). After regeneration:

- Salt Lake County now resolves to 20 distinct cities/townships instead
  of collapsing six into Midvale's canonical_govid.
- Zero govs with conflicting populations within (year, canonical_govid).
- 593 new fips_extended canonical_govids in the xwalk (377 cities, 214
  townships, 2 counties).
- Sentinel count drops from 20-39/year to 2-3/year (the residual
  reflects type-2/3 entities with GOVS legacy_id but no FIPS triplet —
  a separate gap, documented in cog_pipeline).

manifest.json updated with fresh SHAs, sizes, row counts, and pipeline
commit reference.

All 283 uscogdata tests pass against the new fixture.
2026-04-29 15:05:38 -04:00
jared a640c9cc21 test: bundle fixture corpus + wire local-path test helpers
Adds inst/extdata/fixture_corpus/ — a 3.6 MB two-year (2019/2020) slice
of the published corpus (OH+VT+WY fixture from cog_pipeline test profile
plus all 50 states). Includes canonical_fips_xwalk.parquet,
summary_categories.parquet, docs/, and a trimmed manifest.json.

setup.R now points USCOGDATA_URL at the bundled fixture automatically,
bypassing HTTP / Nextcloud entirely. DuckDB reads local parquet via the
existing .is_local_path() fast-path in manifest.R; no httpfs required.
session is reset between test files via withr::defer(cog_close()).

helper-fixture.R gains fixture_corpus_path(), a richer skip_if_no_corpus()
that checks the bundled fixture first, and with_fixture_corpus() for
tests that need explicit session isolation.

test-spending.R: adjust the inflate-column test to use 2019 (fixture year)
instead of 2015 (absent from fixture).

Result: 181 PASS / 0 FAIL / 0 SKIP — all tests run against real parquet
data with real DuckDB queries and no network dependency.
2026-04-27 12:48:58 -04:00
jared 8d2a8d341f feat: view registration from inst/sql/
Seven DuckDB views register on session open: long (raw), spending_long
and revenue_long (prefix-filtered, NOT is_aggregate per reader-spec §4),
canonical_fips_xwalk and summary_categories (identity), spending_annotated
and revenue_annotated (LEFT JOIN xwalk + categories for verb composition).

File prefix `NN-` enforces creation order so *_annotated views resolve
their *_long dependencies.

Deviation from plan: series_breaks/cpi_annual/legacy_aggregate_map and
*_with_transforms views are deferred — their backing parquets are not
in the v0.1 corpus (manifest.files.metadata only lists canonical_fips_xwalk
and summary_categories). The verbs will compute CPI adjustment verb-side
against a bundled cpi table in a later task.
2026-04-23 10:21:36 -04:00
jared f7035049e7 feat: package skeleton — DESCRIPTION, NAMESPACE, session/manifest/cache/views
Minimal skeleton for uscogdata v0.1. Internal session layer with lazy
cog_open(), manifest fetch+validate+cache, view registration placeholder.
Depends on DuckDB >=1.0, httr2, jsonlite. inst/schemas/provenance-v1.json
ships the structured provenance JSON Schema.
2026-04-23 09:05:54 -04:00