feat: uscogdata 0.3.0 — first public release #41

Merged
jared merged 15 commits from feat/public-release-0.3.0 into main 2026-08-08 18:56:15 -04:00
Owner

Summary

Makes the package installable and usable by someone who is not us, and
cuts 0.3.0 as the first public release.

The defect this fixes

uscogdata could not read a remote corpus at all. Two independent
causes, either fatal on its own:

  1. The default corpus URL was a REPLACE_WITH_SHARE_TOKEN sentinel and
    no document in the package supplied a working one — a new user got
    uscogdata_url_not_configured with nothing actionable behind it.
  2. inst/sql/10-long.sql globbed {url}data/long/**/*.parquet. DuckDB
    cannot expand a glob over generic HTTP — there is no directory listing
    — and allow_asterisks_in_http_paths does not help; it forwards the
    literal **/* as a filename and 404s.

It went unnoticed for months because every path that exercised the
package used a local corpus: the test fixture, and the production API
(a host mount). Nothing ran it the way a new user does.

Partition paths now come from the corpus manifest's long_partitions[]
instead of a glob. Measured against the published corpus: the explicit
list returns the same 46,148,034 rows the hf:// glob does, with
hive_partitioning still recovering year. Enumeration is host-agnostic
— an HTTPS mirror, a Nextcloud share and a local cog_mirror() copy all
take one code path — where hf:// would bind the reader to one vendor's
protocol and still need special-casing, since manifest fetching goes
through httr2, which cannot speak hf://.

The default is now the public Hugging Face mirror: CC-BY-4.0, no
credential, CDN-backed. USCOGDATA_URL still overrides.

Also in this release

  • Metadata: Authors@R is Jared E. Knowles aut/cre with ORCID,
    Civilytics Consulting LLC cph/fnd (it was an org with no human, so
    citation() and r-universe had nothing to render). Adds URL and
    BugReports. MaxCorpusSchema 5 -> 7 — it contradicted
    .validate_schema() and the published corpus by two versions.
  • License: adds LICENSE.md (the repo carried no license text at
    all) and corrects the copyright holder.
  • Vignettes now ship. .Rbuildignore excluded ^vignettes$, so
    vignette("total-spending") failed for every user — while the README
    told them to run it.
  • pkgdown indexed 6 of 14 exports, so the docs site did not build.
  • README rewritten for a stranger, with a quickstart that runs
    unconfigured and a "Where the data comes from" section linking the API
    docs, the live API, the HF corpus and the Census source.
  • NEWS recast as a release announcement; the 0.2.0 changelog is kept
    verbatim, the 0.1.0 development churn dropped.
  • CONTRIBUTING added, including the cog-api dependency: its CI clones
    this package at USCOGDATA_REF, defaulting to main with no pin.

Test plan

  • devtools::test() offline against the bundled fixture — 968 passing, 0 failures
  • USCOGDATA_LIVE_TEST=true devtools::test() against the published corpus — 974 passing, 0 skips
  • R CMD check --as-cran — 0 errors, 0 warnings, 0 notes
  • pkgdown::build_site() — clean, all metadata ok
  • Both vignettes resolve from an installed copy; citation() renders correctly
  • Cold start: a rocker/r-ver:4.4 container that had never seen the
    package installed it from a source tarball and ran the README
    quickstart verbatim with no environment variables set
  • cog-api suite against this branch — 739 passing, 0 failures

The cold-start check is the only one that would have caught the original
defect, and its absence is why the defect survived. It is now in the
release checklist in CONTRIBUTING.md.

Notes

  • Independent of #40 (fixture regeneration); either can merge first.
  • Distribution work — public Gitea, GitHub mirror, r-universe — is
    deliberately not in this PR. r-universe publishes check results on
    registration, so it comes after this is green and merged.
  • Design spec: specs/2026-08-08-public-release-design.md
  • Plan: plans/2026-08-08-public-release.md
## Summary Makes the package installable and usable by someone who is not us, and cuts **0.3.0** as the first public release. ### The defect this fixes **`uscogdata` could not read a remote corpus at all.** Two independent causes, either fatal on its own: 1. The default corpus URL was a `REPLACE_WITH_SHARE_TOKEN` sentinel and no document in the package supplied a working one — a new user got `uscogdata_url_not_configured` with nothing actionable behind it. 2. `inst/sql/10-long.sql` globbed `{url}data/long/**/*.parquet`. DuckDB cannot expand a glob over generic HTTP — there is no directory listing — and `allow_asterisks_in_http_paths` does not help; it forwards the literal `**/*` as a filename and 404s. It went unnoticed for months because **every** path that exercised the package used a *local* corpus: the test fixture, and the production API (a host mount). Nothing ran it the way a new user does. Partition paths now come from the corpus manifest's `long_partitions[]` instead of a glob. Measured against the published corpus: the explicit list returns the same **46,148,034 rows** the `hf://` glob does, with `hive_partitioning` still recovering `year`. Enumeration is host-agnostic — an HTTPS mirror, a Nextcloud share and a local `cog_mirror()` copy all take one code path — where `hf://` would bind the reader to one vendor's protocol and still need special-casing, since manifest fetching goes through httr2, which cannot speak `hf://`. The default is now the public Hugging Face mirror: CC-BY-4.0, no credential, CDN-backed. `USCOGDATA_URL` still overrides. ### Also in this release - **Metadata**: `Authors@R` is Jared E. Knowles `aut`/`cre` with ORCID, Civilytics Consulting LLC `cph`/`fnd` (it was an org with no human, so `citation()` and r-universe had nothing to render). Adds `URL` and `BugReports`. `MaxCorpusSchema` 5 -> 7 — it contradicted `.validate_schema()` and the published corpus by two versions. - **License**: adds `LICENSE.md` (the repo carried no license text at all) and corrects the copyright holder. - **Vignettes now ship.** `.Rbuildignore` excluded `^vignettes$`, so `vignette("total-spending")` failed for every user — while the README told them to run it. - **pkgdown** indexed 6 of 14 exports, so the docs site did not build. - **README** rewritten for a stranger, with a quickstart that runs unconfigured and a "Where the data comes from" section linking the API docs, the live API, the HF corpus and the Census source. - **NEWS** recast as a release announcement; the 0.2.0 changelog is kept verbatim, the 0.1.0 development churn dropped. - **CONTRIBUTING** added, including the cog-api dependency: its CI clones this package at `USCOGDATA_REF`, defaulting to `main` with no pin. ## Test plan - [x] `devtools::test()` offline against the bundled fixture — **968 passing**, 0 failures - [x] `USCOGDATA_LIVE_TEST=true devtools::test()` against the published corpus — **974 passing**, 0 skips - [x] `R CMD check --as-cran` — **0 errors, 0 warnings, 0 notes** - [x] `pkgdown::build_site()` — clean, all metadata ok - [x] Both vignettes resolve from an installed copy; `citation()` renders correctly - [x] **Cold start**: a `rocker/r-ver:4.4` container that had never seen the package installed it from a source tarball and ran the README quickstart verbatim with **no environment variables set** - [x] **cog-api** suite against this branch — **739 passing**, 0 failures The cold-start check is the only one that would have caught the original defect, and its absence is why the defect survived. It is now in the release checklist in `CONTRIBUTING.md`. ## Notes - Independent of #40 (fixture regeneration); either can merge first. - Distribution work — public Gitea, GitHub mirror, r-universe — is deliberately **not** in this PR. r-universe publishes check results on registration, so it comes after this is green and merged. - Design spec: `specs/2026-08-08-public-release-design.md` - Plan: `plans/2026-08-08-public-release.md`
jared added 15 commits 2026-08-08 18:42:04 -04:00
Mirrors the same block in cog_explorer/CLAUDE.md. A project-level @ import
does not preload, so the instruction to read it is the mechanism.
Covers the P0 finding that the package cannot read the corpus remotely at
all -- no working default URL, and Hive globs are unsupported over generic
HTTP by DuckDB 1.5.5. Fix is manifest-driven file enumeration (measured:
46,148,034 rows over plain https, identical to the hf:// glob) plus a
working public default.

Also: seven release-readiness fixes, a README restructured for a stranger,
NEWS rewritten as an initial release rather than a pre-release churn log,
and the Gitea-canonical/GitHub-mirror/r-universe distribution mechanics.
Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest
enumeration, working default URL, and the live-corpus test whose absence
let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite
README, NEWS and CONTRIBUTING.

Distribution mechanics stay out of scope -- r-universe publishes check
results on registration, so it comes after final verification is green.
Spec and plan were written against a branch 25 commits behind main, where
the package still read 0.1.0. It is 0.2.0, with a real 0.2.0 changelog in
NEWS that the plan would have deleted.

0.3.0 rather than 0.2.0 because remote corpus reads go from broken to
working and the default URL from placeholder to live -- user-visible
behaviour, so a minor bump. Not 1.0.0: types 4 and 5 remain out of scope.

Task 9 now prepends a 0.3.0 section, keeps 0.2.0 verbatim with a diff
check to prove it, and drops only the 0.1.0 development churn. Task 4
gains the Version bump.
DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.

Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.

Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in
the package supplied a working URL, so a new user installing uscogdata
had no path to a session at all -- just an actionable-looking error with
nothing actionable behind it.

The default is now the public HuggingFace mirror: CC-BY-4.0, no
credential, CDN-backed, and it keeps the origin's uplink out of the read
path. USCOGDATA_URL and options(uscogdata.url=) still override, so
Nextcloud and cog_mirror() copies are unaffected.

The sentinel guard stays for half-edited configs; the two tests covering
it set the URL explicitly, so they only needed renaming to stop calling
it 'the default'.
The remote-read defect survived because every test path used a local
corpus, and so did the API in production. This is the only test that runs
the package the way a new user does: no USCOGDATA_URL, no option, no
fixture -- just install and call a verb.

Gated on USCOGDATA_LIVE_TEST so offline CI skips rather than fails.

Measured against the live corpus while writing this: cog_spending for one
government is 3.9s for a single year and 5.9s across 2000-2022. Well above
the 1.5-2.8s raw parquet scan, because the verbs also join crosswalks,
resolve categories and assemble provenance.
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.

MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.

0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
LICENSE held only the two-line stub and no LICENSE.md existed, so the
repo carried no license text for a human browsing it or for GitHub's
license detector.

usethis::use_mit_license() writes LICENSE.md but leaves an existing
LICENSE alone, so the stub kept saying 'Civilytics' while the full text
said 'Civilytics Consulting LLC'. Corrected by hand, with a test pinning
the two together.
.Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed
uscogdata had no vignettes at all -- while the README instructed users to
run vignette("total-spending"), which failed for every one of them.

Both build offline: total-spending points USCOGDATA_URL at the bundled
fixture, population-denominators is eval = FALSE. Confirmed present in the
built tarball as both source and rendered inst/doc/.

The test also pins the fixture as never-excluded -- it is what lets
R CMD check run with no credentials on r-universe and GitHub Actions.
The reference index covered 6 of 14 exports, so pkgdown errored on the
eight missing topics and the docs site did not build at all. Adds a
Comparison & aggregation section and a Corpus metadata section, lists both
vignettes as articles, and sets url so canonical links and search resolve.

build_site() now completes clean: reference metadata ok, no problems.
Reordered around a new user: what the data is, where it comes from,
install, a quickstart that runs with no configuration, then the
full-dollars warning and the concepts that decide whether a published
number is right.

Adds a 'Where the data comes from' section linking the API documentation
site, the live API, the Hugging Face corpus and the Census source, so
attribution and provenance are reachable from the top rather than implied.

Drops the sibling-repo path, the commented-out install line, the Status
block, and the release advice telling you to strip the fixture -- which
would break the vignette and leave public CI unable to check without
credentials.

The quickstart passes years=; cog_spending() has no full-history default,
so the obvious one-liner errors on a reader's first call.
NEWS described changes relative to states no user had ever seen --
'Breaking: corpus schema_version 4', 'the package now requires...' --
across the whole pre-release development. To someone deciding whether to
depend on this, that reads as instability.

0.3.0 is written as an announcement: what it covers, the verbs, that
reading the corpus now works out of the box, four things to know before a
first query, and the known limits. The 0.2.0 changelog is kept verbatim.
The 0.1.0 development log is dropped; that history is in git.

cog_explain() now documents what provenance actually holds, since the
README points readers at it -- in particular why series_break_refs and
corpus_break_refs are separate fields rather than one list.
Moves developer, testing and release instructions out of the README,
minus the fixture-stripping advice, which was wrong.

Explains that a GitHub PR closes itself as merged once the mirror syncs,
because the merge preserves the contributor's SHAs -- so a PR closing
without a visible Merge click reads as success rather than rejection.

Documents the cog-api dependency: its CI clones this package at
USCOGDATA_REF, defaulting to main with no pin, so anything merged here
reaches the API's next build. Includes the commands to run its suite
against a branch first.
fix: keep doc/ and Meta/ out of the build
R-CMD-check / check (push) Successful in 3m58s
R-CMD-check / check (pull_request) Successful in 3m52s
44e4953f94
Removing ^vignettes$ was right; removing ^doc$ and ^Meta$ with it was
not. Those are devtools::build_vignettes() artefacts, not sources -- R CMD
build regenerates inst/doc/ from vignettes/ by itself, and shipping the
local copies earned a 'non-standard file/directory found at top level'
NOTE.

R CMD check --as-cran is now 0 errors, 0 warnings, 0 notes.
jared merged commit de2ba0cfb9 into main 2026-08-08 18:56:15 -04:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Civilytics/uscogdata#41