DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.
Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.
Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
Closes the last blocked test in the suite. Owner ruled both halves of the
open question yes on 2026-07-30.
`cog_revenue()` gains `revenue_concept`, mirroring `expenditure_concept`,
with Census's two published concepts defined as crosswalk
`revenue_subtype` sets rather than item-code prefixes:
general = own_source + federal + state + local_aid (the default)
total = general + utility + liquor_store + insurance_trust
The manual defines the first by subtracting the other three from the
second (4.3), so both are computable only once all four families are
named -- which cog_pipeline#79 does. Insurance trust now includes the
employee-retirement X codes (X01/X02/X05/X08) alongside the Y codes.
- inst/sql: revenue_long / revenue_long_harmonized carry EVERY revenue
subtype; the concept narrows in R via the existing subtype_scope
machinery, exactly as expenditure_concept narrows spending_long.
- cog_explain() now prints each verb's OWN concept. It previously
printed `expenditure_concept` unconditionally, so a cog_revenue()
caller was told "Concept: primary" -- a spending concept their result
has nothing to do with.
- Fixture regenerated at pipeline_commit aadb46b (330 crosswalk rows).
Corrected two stale expectations in the blocked test while un-skipping
it. It asserted X01+X04+X05+X08 and omitted X02, which applies to state
governments and is nonzero for Wisconsin; X04 is an exhibit code for an
INTRAgovernmental transfer that Census's own "Total Emp Ret Rev"
excludes. Verified against that Census field: the right set is
X01+X02+X05+X08 = $2,283,883k, exactly. And its expected `total` of
$33,377,093k predated the Y codes being classified -- complete Total
Revenue for WI FY2012 is $34,881,961k (general 31,338,293 + Y 1,259,785
+ X 2,283,883).
Behaviour change worth knowing: `general` is now STRICT Census General
Revenue, so utility and liquor store revenue leave the default. Measured
on the fixture that is 15.9% of what cog_revenue() returned for cities,
vs 1.2% for states and 1.7% for counties.
Suite: 716 pass / 0 fail / 0 skip -- the first time this package has had
no skipped tests.
Closes#12
Rewrites expenditure/revenue classification off item-code first-letter
prefixes and onto summary_categories membership (F-018: prefix Y spans
revenue, expenditure, and balance codes), and exposes
expenditure_concept = c("primary", "direct", "total") with primary as
the new default:
primary = operations + capital + assistance
direct = primary + interest + insurance_benefits (Census Direct)
total = direct + intergovernmental (M/L/Q via ig views)
- inst/sql: flow views (20-25) select by crosswalk membership;
summary_categories moves to 11- so it registers before them (DuckDB
binds view sources eagerly). The IG leg gains Q11/Q12/Q18 state
school-system payments (F-017).
- R: one subtype scope per verb call drives the verb SQL, the
harmonization exclusion count, and the complete = TRUE grid;
flow_prefixes survives only to scope recipe suggestions.
cog_geographic_rollup/cog_peer_compare accept primary|direct, still
refuse total, and now actually pass the concept through.
- Balance codes can never reach a spending or revenue result
(uscogdata#25), asserted at both view and verb level.
- Deletes the #11 skip; per the 2026-07-30 owner ruling the F-018 Y01
proof is asserted against the crosswalk, not the default
cog_revenue() call (which stays General Revenue pending #12).
Suite: 696 pass / 0 fail / 1 skip (#12, expected).
Closes#11
Sparsification (cog_pipeline#64, SB194) stopped the corpus storing the wide
era's explicit zeros, which made absence ambiguous:
<= FY2011 dense_source absent => Census published $0
>= FY2012 sparse_source absent => not reported, unknown
A wide-era query whose cells were all $0 had begun returning nothing at all,
with no way to get them back -- strictly less than the reader exposed before,
which is why #64 filed this follow-on.
complete = TRUE fills the requested grid from `code_set` and stamps every row
with value_source: "reported", "census_zero" (amt 0), or "not_reported"
(amt NA). The NA is the point. Filling a modern absence with 0 would invent
data, which is exactly the error the representation contract exists to
prevent -- and it makes this strictly MORE informative than the
pre-sparsification corpus, which could not tell a published zero from an
unreported cell either.
Measured on the fixture, Broward County: FY2011 returns 28 reported + 16
census_zero; FY2019 returns 30 reported + 14 not_reported. The five
categories that walkthrough finding F-006 read as "retired at FY2012" now
report themselves correctly as census_zero before and not_reported after.
Scoping decisions, each of which would invent rows if taken loosely:
- The grid is per government TYPE (code_set.type). Filling against the
union of all types would give a county cells like "state IG transfer to
school districts", indistinguishable from real census zeros.
- NOT is_aggregate, mirroring spending_long/revenue_long. Without it the
grid offers cells those views never return, so each would fill as a
phantom $0.
- Filling happens BEFORE per_capita and inflation, so a census_zero stays
0 through both and a not_reported stays NA rather than becoming 0.
Two new views (36-representation, 37-code_set) are gated on the manifest
LISTING those tables, not on schema_version. Sparsification did not bump the
version -- the fixture this package shipped against until 2026-07-30 was
already v6 and carried neither table -- so a version gate would register a
view over a missing file and fail at CREATE VIEW time on exactly the corpora
the check exists to tolerate. with_corpus_missing_representation() models
that corpus and asserts the abort.
Refused where the fill would be guesswork, both classed
uscogdata_complete_unsupported: a recipe defines its own component codes and
never touches summary_categories; the intergovernmental leg deliberately
keeps aggregate rows (inst/sql/24-ig_long.sql) so its cells are not the ones
code_set describes.
Expected cell sets in the tests are computed from the corpus parquet
directly, never through the verb -- verifying what a filter does through
that same filter proves nothing.
Closes DoD 2, 3 and 4 of #18. DoD 5 (the cog-api follow-on) is filed
separately.
Suite: 658 pass / 0 fail / 3 skip (was 629/0/3). rcmdcheck clean.
Nine review items on the expenditure_concept = direct|total feature:
- bool_and(is_aggregate) -> bool_or(is_aggregate) for aggregate_fallback:
bool_and silently misreported $5,740,775,000 of aggregate-sourced IG
dollars (AL state 2011) as aggregate_fallback = FALSE, because the dense
wide-era data puts a $0 leaf row in the same group as the real aggregate
row. bool_or is a no-op for Direct/Revenue (verified: 0 mismatched groups
across both tables) and correct for the IG leg.
- Added a year-disjointness invariant test for the four legacy
aggregate/leaf IG pairs (M47/M94, M89/M91-93, L47/L94, L89/L91-93),
scoped to the aggregate flag rather than bare code presence (M89/L89
continue past 2011 as independent, non-aggregate leaves).
- Extended the real-SQL-text/synthetic-parquet harness in test-views.R to
pin ig_long/ig_long_harmonized's predicates directly (aggregate rows
retained, NULL harmonized_code coalesced, L-- excluded), rather than
relying on one fixture row's incidental shape.
- Added a test proving the .harmonization_view_files schema-v5 guard is
necessary (not just incidental) against a corpus whose `long` genuinely
lacks a harmonized_code column, and rewrote the misleading "v5-only
parquet files" comment to name both real reasons a file is gated.
- Fixed an NA-fragile subtype filter, extended the expected-view-list
test, guarded .verb_spendrev() against total on a non-spending
view_base, added a roxygen caveat against summing total across levels
of government, and replaced an uncheckable corpus-wide SQL comment
figure with a fixture-verifiable one.
Full suite: 485/0/0 -> 503/0/0 (18 new expectations, zero pre-existing
value changed).
total adds an intergovernmental leg (M = to local, L = to state) as a UNION ALL
over new ig_annotated views. The IG leg deliberately skips NOT is_aggregate --
legacy IG lives almost entirely on aggregate rows, and the aggregate codes are
year-disjoint from their modern leaf components, so nothing double-counts.
L-- (the IG-to-state family total) is excluded. direct is the default and is
numerically unchanged.
Adds schema_version 5 support alongside the existing v4 corpus:
.validate_schema() now accepts a supported set (4, 5) instead of a single
expected version, and cog_spending()/cog_revenue() gain basis =
c("harmonized", "raw"). Harmonized basis routes to new
spending_annotated_harmonized / revenue_annotated_harmonized views built on
spending_long_harmonized / revenue_long_harmonized (REPLACE(harmonized_code
AS item_code), excluding aggregate and NA-harmonized rows); raw basis is
byte-identical to the pre-Phase-R2 behavior. On a v4 corpus, an unspecified
basis silently resolves to "raw" with a provenance note; an explicit
basis = "harmonized" aborts with an actionable message.
Provenance gains basis, basis_note, and a harmonization block
(applied/na_rows_excluded/na_amount_excluded). The five new schema-v5-only
SQL views (harmonized long/annotated views, harmonization_map,
harmonization_recipes, series_breaks_pq) are registered conditionally on
manifest$schema_version >= 5, since DuckDB's read_parquet() errors eagerly
at CREATE VIEW time when the backing file doesn't exist on a v4 corpus.
Fixture corpus regenerated to schema_version 5 / years 2011, 2012, 2019,
2020 (2011->2012 spans the wide-aggregate -> modern-leaf format boundary
needed for the harmonization/recipe work), with the harmonization_map /
harmonization_recipes / series_breaks parquet tables bundled alongside the
existing metadata registries.
Seven DuckDB views register on session open: long (raw), spending_long
and revenue_long (prefix-filtered, NOT is_aggregate per reader-spec §4),
canonical_fips_xwalk and summary_categories (identity), spending_annotated
and revenue_annotated (LEFT JOIN xwalk + categories for verb composition).
File prefix `NN-` enforces creation order so *_annotated views resolve
their *_long dependencies.
Deviation from plan: series_breaks/cpi_annual/legacy_aggregate_map and
*_with_transforms views are deferred — their backing parquets are not
in the v0.1 corpus (manifest.files.metadata only lists canonical_fips_xwalk
and summary_categories). The verbs will compute CPI adjustment verb-side
against a bundled cpi table in a later task.