fix(suggestions): scope candidate recipes by category_type (#34)
R-CMD-check / check (push) Successful in 4m15s
R-CMD-check / check (pull_request) Successful in 4m25s

.query_candidate_recipes() (extracted in #33) now filters candidates by
category_type ('expenditure' vs 'revenue'), derived from the calling
verb's own flow_prefixes (E/F/G -> 'expenditure', else 'revenue').

Without this, a category shared across both flow families in
summary_categories leaked cross-family recipes: cog_revenue(category =
"Corrections") surfaced the expenditure-only corrections_combined recipe
(E04/E05) merely because "Corrections" is also a spending category name,
and cog_spending(category = "IG Federal") surfaced the revenue-only
ig_federal_b47_wide recipe. Both are wrong: following either hint would
attribute dollars to the wrong flow, or (IG Federal) fire the
coverage-gap machinery for a category the calling verb structurally
cannot report on at all.

Updates the two tests this changes the expected behavior of:
- "a mis-scoped cog_spending() call never attaches an M/L counterpart to
  a revenue-flavored recipe" (test-expenditure-concept.R): IG Federal is
  revenue-only, so a spending call now finds zero candidates outright
  rather than firing the suggestion and then blocking its M/L
  counterpart as a second-order check.
- "cog_revenue never suggests expenditure-only recipes"
  (test-recipes.R, was "I1: ... never fabricates suppressed dollars"):
  corrections_combined is expenditure-only, so a revenue call now never
  considers it as a candidate, rather than considering it and reporting
  zero suppressed dollars.

All 1078 tests pass (2 skipped live-corpus), measured devtools::test()
against this commit in a clean worktree stacked on the #33 refactor.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-09 11:11:16 -04:00
co-authored by Claude Sonnet 5
parent 0c7c7eb299
commit 392643bd74
3 changed files with 58 additions and 35 deletions
+38 -8
View File
@@ -89,7 +89,7 @@
#' project's "functions under 50 lines" convention:
#' \itemize{
#' \item `.query_candidate_recipes()` -- candidate recipe lookup by
#' category/subtype scope + M/L exclusion.
#' category/subtype scope + `category_type` filter (#34) + M/L exclusion.
#' \item `.query_recipe_meta()` -- metadata (label, year spans).
#' \item `.query_covered_years()` -- Path 1 gap-year coverage via the
#' recipe's own generic join.
@@ -126,8 +126,16 @@
# by `category` (`.ALL_CATEGORIES` is never a row in
# `summary_categories.category`, so a category-keyed sub-select always
# came back empty here). The M/L exclusion below is unchanged either way.
candidates <- .query_candidate_recipes(con, category, all_categories,
subtype_col, subtype_scope)
#
# Issue #34: scope the candidate query by `category_type` ('expenditure'
# vs 'revenue') to prevent cross-flow-family leakage -- e.g.
# `cog_revenue(category = "Corrections")` must not surface
# expenditure-only recipes (E04/E05) merely because they share the same
# category name in summary_categories. The type is derived from
# flow_prefixes: E/F/G -> 'expenditure', anything else -> 'revenue'.
candidates <- .query_candidate_recipes(con, category, flow_prefixes,
all_categories, subtype_col,
subtype_scope)
if (length(candidates) == 0L) return(list())
result_years <- if (is.null(result) || nrow(result) == 0L) {
@@ -217,8 +225,17 @@
#' `summary_categories.category`, so a category-keyed sub-select always
#' returns zero candidates and silently disables signposting.
#'
#' Scope is also by `category_type` ('expenditure' vs 'revenue', Issue #34)
#' to prevent cross-flow-family leakage: `cog_revenue(category =
#' "Corrections")` must not surface expenditure-only recipes (E04/E05)
#' merely because they share the same category name in summary_categories.
#' The type is derived from flow_prefixes: E/F/G -> 'expenditure', anything
#' else -> 'revenue'.
#'
#' @param con Active DuckDB connection.
#' @param category Category name, or `NULL`.
#' @param flow_prefixes The calling verb's own flow-type prefixes (see
#' `.build_suggestions()`). Used to derive `category_type` (#34).
#' @param all_categories `TRUE` when the caller used `.ALL_CATEGORIES`.
#' @param subtype_col Name of the summary_categories subtype column to
#' scope by when `all_categories = TRUE`; ignored otherwise.
@@ -226,18 +243,31 @@
#' when `all_categories = TRUE`; ignored otherwise.
#' @return Character vector of recipe IDs (possibly empty).
#' @noRd
.query_candidate_recipes <- function(con, category, all_categories = FALSE,
.query_candidate_recipes <- function(con, category, flow_prefixes,
all_categories = FALSE,
subtype_col = NULL,
subtype_scope = NULL) {
# Issue #34: derive category_type from flow_prefixes to prevent
# cross-flow-family leakage -- e.g. cog_revenue(category = "Corrections")
# must not surface expenditure-only recipes merely because they share the
# same category name in summary_categories.
category_type <- if (all(flow_prefixes %in% c("E", "F", "G"))) {
"expenditure"
} else {
"revenue"
}
candidate_scope_sql <- if (isTRUE(all_categories)) {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories WHERE %s IN (%s)",
subtype_col, .sql_lit_chr(subtype_scope)
"SELECT DISTINCT item_code FROM summary_categories
WHERE %s IN (%s) AND category_type = '%s'",
subtype_col, .sql_lit_chr(subtype_scope), category_type
)
} else {
sprintf(
"SELECT DISTINCT item_code FROM summary_categories WHERE category IN (%s)",
.sql_lit_chr(category)
"SELECT DISTINCT item_code FROM summary_categories
WHERE category IN (%s) AND category_type = '%s'",
.sql_lit_chr(category), category_type
)
}