Returned amounts are full US dollars; the x1000 conversion is undocumented in reader-facing docs #15

Closed
opened 2026-07-29 00:05:29 -04:00 by jared · 2 comments
Owner

Filed from the Madison walkthrough audit (cog_explorer/docs/walkthroughs/FINDINGS.md, 2026-07-28). Verdict: definitional — the package's behaviour is the friendlier choice and is not wrong; the conversion is simply not stated where most readers meet the package.

Root cause

The raw Census source files report thousands of dollars. uscogdata multiplies by 1000 and returns amt_nominal, amt_real, amt_per_capita_nominal, and amt_per_capita_real in full US dollars. The conversion is real, correct, and recorded in the return value's provenance:

attr(cog_spending("552025209777", 2020L), "provenance")$transformations$units_conversion
#> $applied      TRUE
#> $source_unit  "$1,000s (raw Census)"
#> $target_unit  "$USD"
#> $multiplier   1000

Meanwhile cog_explorer's CLAUDE.md states, correctly of the raw amt column and misleadingly of everything else: "All raw amt values are in $1,000s." A reader who applies that rule to amt_nominal overstates every figure by 1000x.

Where it is and isn't already documented (checked 2026-07-29, after the audit was written)

Surface States the conversion?
?cog_spending / man/cog_spending.Rd @return Yes — "Amounts are returned in full U.S. dollars (the raw corpus stores them in $1,000s; this verb multiplies by 1000…)"
?cog_revenue / man/cog_revenue.Rd @return Yes — same wording
provenance$transformations$units_conversion Yes
README.md No
vignettes/total-spending.Rmd, vignettes/population-denominators.Rmd No
cog-api's data-dictionary.md Yes (since 2b71b41, 2026-07-18)
cog-api's llms.txt — the agent-facing surface No
cog_explorer/CLAUDE.md No, and states the opposite rule

So the finding's proposed data-dictionary sentence has partly landed already. What remains is that a reader who meets this package through its README, a vignette, the agent-facing llms.txt, or cog_explorer's own conventions doc gets no signal at all — and one of those actively points the wrong way.

Findings resolved

Finding Severity Summary
F-004 high amt_nominal is full dollars, not the raw $1,000s

Reproduction (verbatim from FINDINGS.md, verified against the live corpus)

cog_spending(govid = "552025209777", years = 2022L) |>
  dplyr::filter(category == "Fire", spend_subtype == "capital")
#> amt_nominal = 1317000   (i.e. $1.317M, not $1.317B)

Why it matters

This is the highest-consequence definitional gap the audit found. A 1000x overstatement produces a plausible-looking dollar amount, not an obvious error — nothing downstream will flag it. The correction is entirely available to the reader; it just isn't where the reader is.

Definition of done

  1. README.md states the units once, in the first place a returned data frame is shown. The audit's proposed wording:

    amt_nominal, amt_real, amt_per_capita_nominal and amt_per_capita_real are expressed in full US dollars. The underlying Census source files report thousands of dollars; uscogdata performs this conversion for you.

  2. Both vignettes carry the same statement where they first print amounts.

  3. Secondary task, cog_explorer/CLAUDE.md: its Data Conventions section needs a clause distinguishing raw amt (thousands) from amt_nominal and friends (full dollars). cog_explorer has no git remote, so it cannot carry its own issue; recorded here because this package performs the conversion. Whoever fixes (1) should make that edit in the same pass.

  4. Secondary task, cog-api: llms.txt should carry the same sentence — data-dictionary.md already does, but llms.txt is the surface this API names for agent consumption and it is silent on units.

  5. Test goes green: tests/testthat/test-amount-units-documented.R → test_that("returned amounts are documented as full US dollars where readers meet the package", ...). It asserts README.md and both vignettes state the conversion, and pins it numerically against the fixture — Madison FY2020 amt_nominal sums to exactly 1000 × the raw corpus's amt for the same government-year, read straight from the parquet partitions rather than through the verb — so the documented claim and the behaviour cannot drift apart. It deliberately does not assert on the .Rd files, which already carry the statement. Remove the skip() on line 1 of the test body to activate.

Severity: high. Verdict: definitional (documentation, plus a disclosure already present in provenance that the prose should point at).

Filed from the **Madison walkthrough audit** (`cog_explorer/docs/walkthroughs/FINDINGS.md`, 2026-07-28). Verdict: **definitional** — the package's behaviour is the friendlier choice and is not wrong; the conversion is simply not stated where most readers meet the package. ## Root cause The raw Census source files report **thousands of dollars**. `uscogdata` multiplies by 1000 and returns `amt_nominal`, `amt_real`, `amt_per_capita_nominal`, and `amt_per_capita_real` in **full US dollars**. The conversion is real, correct, and recorded in the return value's provenance: ```r attr(cog_spending("552025209777", 2020L), "provenance")$transformations$units_conversion #> $applied TRUE #> $source_unit "$1,000s (raw Census)" #> $target_unit "$USD" #> $multiplier 1000 ``` Meanwhile `cog_explorer`'s `CLAUDE.md` states, correctly of the raw `amt` column and misleadingly of everything else: *"All raw `amt` values are in $1,000s."* A reader who applies that rule to `amt_nominal` overstates every figure by 1000x. ## Where it is and isn't already documented (checked 2026-07-29, after the audit was written) | Surface | States the conversion? | |---|---| | `?cog_spending` / `man/cog_spending.Rd` `@return` | **Yes** — "Amounts are returned in **full U.S. dollars** (the raw corpus stores them in $1,000s; this verb multiplies by 1000…)" | | `?cog_revenue` / `man/cog_revenue.Rd` `@return` | **Yes** — same wording | | `provenance$transformations$units_conversion` | **Yes** | | `README.md` | **No** | | `vignettes/total-spending.Rmd`, `vignettes/population-denominators.Rmd` | **No** | | `cog-api`'s `data-dictionary.md` | **Yes** (since 2b71b41, 2026-07-18) | | `cog-api`'s `llms.txt` — the agent-facing surface | **No** | | `cog_explorer/CLAUDE.md` | **No, and states the opposite rule** | So the finding's proposed data-dictionary sentence has partly landed already. What remains is that a reader who meets this package through its README, a vignette, the agent-facing `llms.txt`, or `cog_explorer`'s own conventions doc gets no signal at all — and one of those actively points the wrong way. ## Findings resolved | Finding | Severity | Summary | |---|---|---| | **F-004** | high | `amt_nominal` is full dollars, not the raw $1,000s | ## Reproduction (verbatim from FINDINGS.md, verified against the live corpus) ```r cog_spending(govid = "552025209777", years = 2022L) |> dplyr::filter(category == "Fire", spend_subtype == "capital") #> amt_nominal = 1317000 (i.e. $1.317M, not $1.317B) ``` ## Why it matters **This is the highest-consequence definitional gap the audit found.** A 1000x overstatement produces a plausible-looking dollar amount, not an obvious error — nothing downstream will flag it. The correction is entirely available to the reader; it just isn't where the reader is. ## Definition of done 1. `README.md` states the units once, in the first place a returned data frame is shown. The audit's proposed wording: > `amt_nominal`, `amt_real`, `amt_per_capita_nominal` and `amt_per_capita_real` are expressed in **full US dollars**. The underlying Census source files report thousands of dollars; uscogdata performs this conversion for you. 2. Both vignettes carry the same statement where they first print amounts. 3. **Secondary task, `cog_explorer/CLAUDE.md`:** its Data Conventions section needs a clause distinguishing raw `amt` (thousands) from `amt_nominal` and friends (full dollars). `cog_explorer` has **no git remote**, so it cannot carry its own issue; recorded here because this package performs the conversion. Whoever fixes (1) should make that edit in the same pass. 4. **Secondary task, `cog-api`:** `llms.txt` should carry the same sentence — `data-dictionary.md` already does, but `llms.txt` is the surface this API names for agent consumption and it is silent on units. 5. Test goes green: **`tests/testthat/test-amount-units-documented.R`** → `test_that("returned amounts are documented as full US dollars where readers meet the package", ...)`. It asserts `README.md` and both vignettes state the conversion, and pins it numerically against the fixture — Madison FY2020 `amt_nominal` sums to exactly `1000 ×` the raw corpus's `amt` for the same government-year, read straight from the parquet partitions rather than through the verb — so the documented claim and the behaviour cannot drift apart. It deliberately does **not** assert on the `.Rd` files, which already carry the statement. Remove the `skip()` on line 1 of the test body to activate. **Severity: high. Verdict: definitional** (documentation, plus a disclosure already present in provenance that the prose should point at).
jared added the severity/highverdict/definitionalmadison-walkthrough labels 2026-07-29 00:05:29 -04:00
kodor was assigned by jared 2026-07-29 08:13:49 -04:00
jared added the kodor/fix label 2026-07-29 08:13:49 -04:00
Author
Owner

Taking this over. No kodor activity since assignment on 2026-07-29 — no branch, no PR, no comment — so I am picking it up rather than leaving it queued. Reassign back if kodor is simply slow and you would rather it land there; I will not push until this repo is otherwise quiet.

Being fixed together with the other two kodor/fix issues in one PR: they are all single-file, all have committed acceptance tests, and two of them touch the same documentation surfaces.

Taking this over. No kodor activity since assignment on 2026-07-29 — no branch, no PR, no comment — so I am picking it up rather than leaving it queued. Reassign back if kodor is simply slow and you would rather it land there; I will not push until this repo is otherwise quiet. Being fixed together with the other two `kodor/fix` issues in one PR: they are all single-file, all have committed acceptance tests, and two of them touch the same documentation surfaces.
Author
Owner

Fixed and merged in PR #22 (8db944e on main). Closing manually: that PR said "Closes #16, #15, #14", and Gitea only parsed the first reference, so #16 auto-closed while this one stayed open. The keyword needs repeating per issue.

Verified present on main, not just claimed.

Fixed and merged in PR #22 (`8db944e` on `main`). Closing manually: that PR said "Closes #16, #15, #14", and Gitea only parsed the first reference, so #16 auto-closed while this one stayed open. The keyword needs repeating per issue. Verified present on `main`, not just claimed.
jared closed this issue 2026-07-30 12:01:43 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#15