.build_series_break_refs() matches `fin_code IN (<codes in the result>)`. No row's item_code is ever the literal "ALL", so the four corpus-wide entries could never match and reached no user: SB085 1977 dollar precision across the 1976/1977 boundary SB087 2002 imputation exclusion FY2002-2006 SB194 2012 dense -> sparse representation change SB086 2017 government id scheme change SB194 is why this matters now. cog_pipeline#64 DoD 4 was "series_breaks.csv carries an ALL @ 2012 entry describing the representation change, SO cog_explain() surfaces it". The entry shipped; the reader dropped it. A query spanning FY2011 -> FY2012 crosses the boundary where an absent cell stops meaning "Census published $0" and starts meaning "not reported", and nothing said so. Provenance gains `corpus_break_refs`, built by .build_corpus_break_refs() on the break_year window alone -- which codes a result happens to contain is irrelevant to a caveat about the corpus. A separate field rather than more entries in series_break_refs, because an ALL caveat qualifies the whole result and folding the two together invites reading it as a caveat about one series; .build_series_break_refs() now excludes 'ALL' explicitly so the two stay disjoint by construction. cog_explain() prints them under their own "Corpus-wide caveats" heading, and cog-api passes provenance through verbatim, so the field reaches the API with no change there. On the year rule: all four entries are BOUNDARY caveats -- their own join_advice speaks of crossing 1976/1977, of FY2002-2006, of absence not being comparable across FY2012, of pre- vs post-2017 ids -- so the same `break_year BETWEEN min(years) AND max(years)` rule the code-specific path uses is the right one, and matches the issue's DoD 1. The issue's DoD 3 also asks that a FY2011 query surface SB085; that cannot hold under DoD 1 and does not hold under any reading of SB085's text, whose boundary is 1976/1977. Tested with a range that actually spans it, and flagged on the issue. Stacked on fix/regen-fixture-corpus-18: SB194 does not exist in main's bundled fixture, which predates the break being catalogued. Suite: 606 pass / 0 fail / 6 skip (was 594/0/6). cog-api 357 / 0 / 8, unchanged.
55 lines
2.4 KiB
R
55 lines
2.4 KiB
R
# R/series_breaks.R
|
|
# Populates prov$series_break_refs (schema in inst/schemas/provenance-v1.json
|
|
# defines the field; it was always present but always empty pre-Phase-R2)
|
|
# with the ids of any catalogued series break whose fin_code appears among
|
|
# the result's observed item codes and whose break_year falls inside the
|
|
# requested year span -- the "break warnings in the provenance envelope"
|
|
# spec § 5 promises downstream consumers (cog-api passes provenance through
|
|
# verbatim). schema_version >= 5 only: series_breaks_pq isn't registered on
|
|
# an older corpus.
|
|
|
|
#' @noRd
|
|
.build_series_break_refs <- function(con, codes_observed, years, schema_version) {
|
|
if (schema_version < 5L || length(codes_observed) == 0L) return(character(0))
|
|
sql <- sprintf(
|
|
"SELECT DISTINCT break_id
|
|
FROM series_breaks_pq
|
|
WHERE fin_code IN (%s) AND fin_code <> 'ALL'
|
|
AND break_year BETWEEN %d AND %d
|
|
ORDER BY break_id",
|
|
.sql_lit_chr(codes_observed), min(as.integer(years)), max(as.integer(years))
|
|
)
|
|
DBI::dbGetQuery(con, sql)$break_id
|
|
}
|
|
|
|
#' Corpus-wide caveats: catalogued breaks whose `fin_code` is the literal
|
|
#' `"ALL"` rather than an item code. They qualify the whole result, so they
|
|
#' cannot be matched the way `.build_series_break_refs()` matches -- no row's
|
|
#' `item_code` is ever `"ALL"`, which is exactly why they reached no user
|
|
#' before uscogdata#19. Selection is on the break_year window alone: which
|
|
#' codes a result happens to contain is irrelevant to a caveat about the
|
|
#' corpus.
|
|
#'
|
|
#' All four catalogued entries are *boundary* caveats (dollar precision
|
|
#' across 1976/1977, imputation exclusion from 2002, the dense -> sparse
|
|
#' representation change at 2012, the id scheme change at 2017), so the same
|
|
#' `break_year BETWEEN min(years) AND max(years)` rule the code-specific
|
|
#' path uses is the right one -- a request that never crosses the boundary
|
|
#' is not affected by it.
|
|
#'
|
|
#' Returned separately from `series_break_refs` so a consumer can tell a
|
|
#' whole-result caveat from a break in one series; the two are disjoint by
|
|
#' construction.
|
|
#' @noRd
|
|
.build_corpus_break_refs <- function(con, years, schema_version) {
|
|
if (schema_version < 5L || length(years) == 0L) return(character(0))
|
|
sql <- sprintf(
|
|
"SELECT DISTINCT break_id
|
|
FROM series_breaks_pq
|
|
WHERE fin_code = 'ALL' AND break_year BETWEEN %d AND %d
|
|
ORDER BY break_id",
|
|
min(as.integer(years)), max(as.integer(years))
|
|
)
|
|
DBI::dbGetQuery(con, sql)$break_id
|
|
}
|