Files
uscogdata/CLAUDE.md
T
jaredandClaude Opus 5.5 63b421aea2 docs: point CLAUDE.md at the pipeline clone's paths
The repo moved out of Nextcloud today, and its sibling pipeline clone is
now census_of_governments_finance_pipeline, not cog_explorer/cog_pipeline.
The phase N task plan had already moved to docs/archive/ on 2026-07-23
(pipeline acfacf2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 11:15:00 -04:00

135 lines
6.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md — Agent Instructions for uscogdata
## Project Overview
R package providing a curated reader API for the Civilytics US Census of
Governments finance corpus. Reads Hive-partitioned parquet + `manifest.json`
published by `cog_pipeline` via DuckDB (local path or remote URL). This is a
standalone Gitea repo, sibling to the pipeline repo,
`Civilytics/census_of_governments_finance_pipeline`, cloned beside this one.
**Gitea remote:** `gitea.civilytics.org/Civilytics/uscogdata`
**Full implementation plan (archived):** `../census_of_governments_finance_pipeline/docs/archive/plan_phase_n_tasks.md` (Tasks 2.1–2.8 + Phase 3)
**Reader contract spec:** `../census_of_governments_finance_pipeline/docs/reader-specification.md`
## Architecture
### Data flow
```
USCOGDATA_URL (local path or https://)
→ cog_open() fetches manifest.json, registers DuckDB views from inst/sql/
→ verbs query views via DBI
→ results carry provenance attr
```
### Key files
- `R/config.R` — `.cfg()`, `.resolve_url()`, `.uscogdata_env` (mutable state)
- `R/session.R` — `cog_open()`, `cog_close()`, `.ensure_session()`, `.coerce_govid_input()`
- `R/manifest.R` — `.fetch_or_cache_manifest()`, `.is_local_path()` (local paths bypass HTTP/cache)
- `R/views.R` — `.register_views()` (substitutes `{url}` into SQL files at `inst/sql/`)
- `inst/sql/` — **23** SQL view definitions (measured), numbered by load order
(`10-` through `46-`): the `*_long` layer (`long`, `spending_long`,
`revenue_long`, `ig_long`, `balance_long`, plus `_harmonized` variants of
`spending_long`/`revenue_long`/`ig_long`), the `*_annotated` layer
(`spending_annotated`, `revenue_annotated`, `ig_annotated`,
`balance_annotated`, plus `_harmonized` variants of `spending_annotated`/
`revenue_annotated`/`ig_annotated`), and metadata views
(`canonical_fips_xwalk`, `summary_categories`, `gov_population_yearly`,
`harmonization_map`, `harmonization_recipes`, `series_breaks_pq`,
`representation`, `code_set`)
- `R/spending.R` / `R/revenue.R` — `cog_spending()` / `cog_revenue()` via shared `.verb_spendrev()`
- `R/balances.R` — `cog_balances()`. A third money-adjacent verb, but returns a
**stock** (a balance at a point in time) rather than a **flow** (activity
over a fiscal year), so it does NOT route through `.verb_spendrev()` and has
no `expenditure_concept`/`revenue_concept`/`complete`/`subtype` arguments.
`R/balance_caveats.R` attaches `provenance$balance_caveats` (GAAP-vs-gross
disclosure + measured per-subtype coverage windows).
- `R/rollup.R` — `cog_geographic_rollup()` (accepts named list of govids by layer)
- `R/peers.R` — `cog_find_peers()` + `cog_peer_compare()`
- `R/search.R` — `cog_gov_search()` (name pattern, state, type filters)
- `R/mirror.R` — `cog_mirror()` (local corpus copy via httr2 WebDAV)
- `R/explain.R` — `cog_explain()` (prints/returns provenance attr)
- `R/categories.R` — `cog_categories()` (queries `summary_categories` view)
- `R/adjust.R` — `.inflate()` helper using bundled CPIAUCSL in `sysdata.rda`
- `R/provenance.R` — `.build_provenance()` attached as attr to all verb results
### Corpus URL resolution
`USCOGDATA_URL` env var → `uscogdata.url` option → default placeholder in `config.R`.
Any value without `://` is treated as a local path by `.is_local_path()` and reads
`manifest.json` directly from disk (no HTTP, no TTL cache).
## Current State (2026-08-03)
**Version:** 0.1.0 (pre-release)
**Branch:** `feat/cog-balances-25`, commit `fde62eb`
**Tests:** 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured `testthat::test_local()`, 2026-08-03, after the final-review fix wave)
**CI:** Gitea Actions green (`.gitea/workflows/ci.yml`)
### Completed (Tasks 2.1–2.7)
All **14** exports implemented and tested (measured from `NAMESPACE`):
`cog_spending`, `cog_revenue`, `cog_balances`, `cog_explain`,
`cog_geographic_rollup`, `cog_find_peers`, `cog_peer_compare`,
`cog_gov_search`, `cog_mirror`, `cog_categories`, `cog_recipes`,
`cog_manifest`, `cog_basket_resolution`, `cog_basket_unresolved`.
Bundled fixture corpus at `inst/extdata/fixture_corpus/` (years
2011, 2012, 2019, 2020 — measured via DuckDB `read_parquet(hive_partitioning=1)`,
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
### Remaining to v0.1 release
1. **Task 2.8 — Docs:** mostly done — all 14 exports have a `man/*.Rd`,
`README.md` and `_pkgdown.yml` exist, and `vignettes/` carries
`total-spending.Rmd` + `population-denominators.Rmd`. Outstanding:
`pkgdown::build_site()` has never been run (no `docs/`).
2. **Phase 3 — cog_explorer bridge:** create
`cog_explorer/examples/hello_world_uscogdata.Rmd` (installs from Gitea, runs
`cog_spending()` against live Nextcloud corpus, renders a chart). Gate N14.
3. **Release gates:** run test suite against live corpus
(`USCOGDATA_URL=<live-url> devtools::test()`), add
`^inst/extdata/fixture_corpus$` to `.Rbuildignore`, tag `v0.1.0`, cut
Gitea release.
## Testing
```r
devtools::test() # uses bundled fixture, fully offline
testthat::test_local() # same, used by CI
```
`tests/testthat/setup.R` sets `USCOGDATA_URL` to the bundled fixture automatically.
`helper-fixture.R` provides `skip_if_no_corpus()`, `fixture_corpus_path()`,
`with_fixture_corpus()`.
To run against the live corpus:
```r
Sys.setenv(USCOGDATA_URL = "<live-url-with-trailing-slash>")
devtools::test()
```
## Coding Conventions
- All verbs call `.ensure_session()` first, then query via `DBI::dbGetQuery()`
- Return value is always a `tbl_df` with a `provenance` attribute
- govid inputs always go through `.coerce_govid_input()` (accepts character or data frame)
- SQL has two layers. **View definitions** live in `inst/sql/` and are
registered by `.register_views()`, which globs the directory in sorted order
and substitutes `{url}`. **Query construction** is inline `sprintf()` in R
(`.build_verb_sql()`, `.run_recipe()`, `.attach_per_capita()`). Add a view as
a numbered `.sql` file; build a query in R.
- No arrow dependency — DuckDB reads parquet natively
- `withr` is a Suggests-only dep; only used in tests
## Domain context — read this first
**Before doing any work in this repo, read `~/.claude/memory/values/civilytics.md`.**
It carries the purpose, direction, and constraints for this domain. It is not optional
context — read it before planning or writing code, not after. (An `@` import will not
work here; project-level imports don't preload. The read is the mechanism.)