The repo moved out of Nextcloud today, and its sibling pipeline clone is now census_of_governments_finance_pipeline, not cog_explorer/cog_pipeline. The phase N task plan had already moved to docs/archive/ on 2026-07-23 (pipeline acfacf2). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
6.4 KiB
CLAUDE.md — Agent Instructions for uscogdata
Project Overview
R package providing a curated reader API for the Civilytics US Census of
Governments finance corpus. Reads Hive-partitioned parquet + manifest.json
published by cog_pipeline via DuckDB (local path or remote URL). This is a
standalone Gitea repo, sibling to the pipeline repo,
Civilytics/census_of_governments_finance_pipeline, cloned beside this one.
Gitea remote: gitea.civilytics.org/Civilytics/uscogdata
Full implementation plan (archived): ../census_of_governments_finance_pipeline/docs/archive/plan_phase_n_tasks.md (Tasks 2.1–2.8 + Phase 3)
Reader contract spec: ../census_of_governments_finance_pipeline/docs/reader-specification.md
Architecture
Data flow
USCOGDATA_URL (local path or https://)
→ cog_open() fetches manifest.json, registers DuckDB views from inst/sql/
→ verbs query views via DBI
→ results carry provenance attr
Key files
R/config.R—.cfg(),.resolve_url(),.uscogdata_env(mutable state)R/session.R—cog_open(),cog_close(),.ensure_session(),.coerce_govid_input()R/manifest.R—.fetch_or_cache_manifest(),.is_local_path()(local paths bypass HTTP/cache)R/views.R—.register_views()(substitutes{url}into SQL files atinst/sql/)inst/sql/— 23 SQL view definitions (measured), numbered by load order (10-through46-): the*_longlayer (long,spending_long,revenue_long,ig_long,balance_long, plus_harmonizedvariants ofspending_long/revenue_long/ig_long), the*_annotatedlayer (spending_annotated,revenue_annotated,ig_annotated,balance_annotated, plus_harmonizedvariants ofspending_annotated/revenue_annotated/ig_annotated), and metadata views (canonical_fips_xwalk,summary_categories,gov_population_yearly,harmonization_map,harmonization_recipes,series_breaks_pq,representation,code_set)R/spending.R/R/revenue.R—cog_spending()/cog_revenue()via shared.verb_spendrev()R/balances.R—cog_balances(). A third money-adjacent verb, but returns a stock (a balance at a point in time) rather than a flow (activity over a fiscal year), so it does NOT route through.verb_spendrev()and has noexpenditure_concept/revenue_concept/complete/subtypearguments.R/balance_caveats.Rattachesprovenance$balance_caveats(GAAP-vs-gross disclosure + measured per-subtype coverage windows).R/rollup.R—cog_geographic_rollup()(accepts named list of govids by layer)R/peers.R—cog_find_peers()+cog_peer_compare()R/search.R—cog_gov_search()(name pattern, state, type filters)R/mirror.R—cog_mirror()(local corpus copy via httr2 WebDAV)R/explain.R—cog_explain()(prints/returns provenance attr)R/categories.R—cog_categories()(queriessummary_categoriesview)R/adjust.R—.inflate()helper using bundled CPIAUCSL insysdata.rdaR/provenance.R—.build_provenance()attached as attr to all verb results
Corpus URL resolution
USCOGDATA_URL env var → uscogdata.url option → default placeholder in config.R.
Any value without :// is treated as a local path by .is_local_path() and reads
manifest.json directly from disk (no HTTP, no TTL cache).
Current State (2026-08-03)
Version: 0.1.0 (pre-release)
Branch: feat/cog-balances-25, commit fde62eb
Tests: 788 PASS / 0 FAIL / 0 SKIP / 0 WARN (measured testthat::test_local(), 2026-08-03, after the final-review fix wave)
CI: Gitea Actions green (.gitea/workflows/ci.yml)
Completed (Tasks 2.1–2.7)
All 14 exports implemented and tested (measured from NAMESPACE):
cog_spending, cog_revenue, cog_balances, cog_explain,
cog_geographic_rollup, cog_find_peers, cog_peer_compare,
cog_gov_search, cog_mirror, cog_categories, cog_recipes,
cog_manifest, cog_basket_resolution, cog_basket_unresolved.
Bundled fixture corpus at inst/extdata/fixture_corpus/ (years
2011, 2012, 2019, 2020 — measured via DuckDB read_parquet(hive_partitioning=1),
2026-08-03; all 50 states). Tests run fully offline — no credentials needed.
Remaining to v0.1 release
-
Task 2.8 — Docs: mostly done — all 14 exports have a
man/*.Rd,README.mdand_pkgdown.ymlexist, andvignettes/carriestotal-spending.Rmd+population-denominators.Rmd. Outstanding:pkgdown::build_site()has never been run (nodocs/). -
Phase 3 — cog_explorer bridge: create
cog_explorer/examples/hello_world_uscogdata.Rmd(installs from Gitea, runscog_spending()against live Nextcloud corpus, renders a chart). Gate N14. -
Release gates: run test suite against live corpus (
USCOGDATA_URL=<live-url> devtools::test()), add^inst/extdata/fixture_corpus$to.Rbuildignore, tagv0.1.0, cut Gitea release.
Testing
devtools::test() # uses bundled fixture, fully offline
testthat::test_local() # same, used by CI
tests/testthat/setup.R sets USCOGDATA_URL to the bundled fixture automatically.
helper-fixture.R provides skip_if_no_corpus(), fixture_corpus_path(),
with_fixture_corpus().
To run against the live corpus:
Sys.setenv(USCOGDATA_URL = "<live-url-with-trailing-slash>")
devtools::test()
Coding Conventions
- All verbs call
.ensure_session()first, then query viaDBI::dbGetQuery() - Return value is always a
tbl_dfwith aprovenanceattribute - govid inputs always go through
.coerce_govid_input()(accepts character or data frame) - SQL has two layers. View definitions live in
inst/sql/and are registered by.register_views(), which globs the directory in sorted order and substitutes{url}. Query construction is inlinesprintf()in R (.build_verb_sql(),.run_recipe(),.attach_per_capita()). Add a view as a numbered.sqlfile; build a query in R. - No arrow dependency — DuckDB reads parquet natively
withris a Suggests-only dep; only used in tests
Domain context — read this first
Before doing any work in this repo, read ~/.claude/memory/values/civilytics.md.
It carries the purpose, direction, and constraints for this domain. It is not optional
context — read it before planning or writing code, not after. (An @ import will not
work here; project-level imports don't preload. The read is the mechanism.)