Files
uscogdata/CLAUDE.md
T
jared dfda39051e docs: point agents at the Civilytics values file before they start
Mirrors the same block in cog_explorer/CLAUDE.md. A project-level @ import
does not preload, so the instruction to read it is the mechanism.
2026-08-08 17:03:39 -04:00

134 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md — Agent Instructions for uscogdata
## Project Overview
R package providing a curated reader API for the Civilytics US Census of
Governments finance corpus. Reads Hive-partitioned parquet + `manifest.json`
published by `cog_pipeline` via DuckDB (local path or remote URL). This is a
standalone Gitea repo, sibling to `cog_explorer/cog_pipeline/`.
**Gitea remote:** `gitea.civilytics.org/Civilytics/uscogdata`
**Full implementation plan:** `../cog_pipeline/docs/plan_phase_n_tasks.md` (Tasks 2.1–2.8 + Phase 3)
**Reader contract spec:** `../cog_pipeline/docs/reader-specification.md`
## Architecture
### Data flow
```
USCOGDATA_URL (local path or https://)
→ cog_open() fetches manifest.json, registers DuckDB views from inst/sql/
→ verbs query views via DBI
→ results carry provenance attr
```
### Key files
- `R/config.R` — `.cfg()`, `.resolve_url()`, `.uscogdata_env` (mutable state)
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
`revenue_annotated`/`ig_annotated`), and metadata views
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
`representation`, `code_set`)
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
**stock** (a balance at a point in time) rather than a **flow** (activity
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
disclosure + measured per-subtype coverage windows).
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
- `R/mirror.R` — `cog_mirror()` (local corpus copy via httr2 WebDAV)
- `R/explain.R` — `cog_explain()` (prints/returns provenance attr)
- `R/categories.R` — `cog_categories()` (queries `summary_categories` view)
- `R/adjust.R` — `.inflate()` helper using bundled CPIAUCSL in `sysdata.rda`
- `R/provenance.R` — `.build_provenance()` attached as attr to all verb results
### Corpus URL resolution
`USCOGDATA_URL` env var → `uscogdata.url` option → default placeholder in `config.R`.
Any value without `://` is treated as a local path by `.is_local_path()` and reads
`manifest.json` directly from disk (no HTTP, no TTL cache).
## Current State (2026-08-03)
**Version:** 0.1.0 (pre-release)
**Branch:** `feat/cog-balances-25`, commit `fde62eb`
**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
### Completed (Tasks 2.1–2.7)
All **14** exports implemented and tested (measured from `NAMESPACE`):
`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
### Remaining to v0.1 release
1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
`README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
`total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
`pkgdown::build_site()` has never been run (no `docs/`).
2. **Phase 3 — cog_explorer bridge:** create
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
`cog_spending()` against live Nextcloud corpus, renders a chart). Gate N14.
3. **Release gates:** run test suite against live corpus
(`USCOGDATA_URL=<live-url> devtools::test()`), add
`^inst/extdata/fixture_corpus$` to `.Rbuildignore`, tag `v0.1.0`, cut
Gitea release.
## Testing
```r
devtools::test() # uses bundled fixture, fully offline
testthat::test_local() # same, used by CI
```
`tests/testthat/setup.R` sets `USCOGDATA_URL` to the bundled fixture automatically.
`helper-fixture.R` provides `skip_if_no_corpus()`, `fixture_corpus_path()`,
`with_fixture_corpus()`.
To run against the live corpus:
```r
Sys.setenv(USCOGDATA_URL = "<live-url-with-trailing-slash>")
devtools::test()
```
## Coding Conventions
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
- Return value is always a `tbl_df` with a `provenance` attribute
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
- SQL has two layers. **View definitions** live in `inst/sql/` and are
registered by `.register_views()`, which globs the directory in sorted order
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
a numbered `.sql` file; build a query in R.
- No arrow dependency — DuckDB reads parquet natively
- `withr` is a Suggests-only dep; only used in tests
## Domain context — read this first
**Before doing any work in this repo, read `~/.claude/memory/values/civilytics.md`.**
It carries the purpose, direction, and constraints for this domain. It is not optional
context — read it before planning or writing code, not after. (An `@` import will not
work here; project-level imports don't preload. The read is the mechanism.)