Makes the package installable and usable by someone who is not us, and
cuts 0.3.0 as the first public release.
The defect this fixes
uscogdata could not read a remote corpus at all. Two independent
causes, either fatal on its own:
The default corpus URL was a REPLACE_WITH_SHARE_TOKEN sentinel and
no document in the package supplied a working one — a new user got uscogdata_url_not_configured with nothing actionable behind it.
inst/sql/10-long.sql globbed {url}data/long/**/*.parquet. DuckDB
cannot expand a glob over generic HTTP — there is no directory listing
— and allow_asterisks_in_http_paths does not help; it forwards the
literal **/* as a filename and 404s.
It went unnoticed for months because every path that exercised the
package used a local corpus: the test fixture, and the production API
(a host mount). Nothing ran it the way a new user does.
Partition paths now come from the corpus manifest's long_partitions[]
instead of a glob. Measured against the published corpus: the explicit
list returns the same 46,148,034 rows the hf:// glob does, with hive_partitioning still recovering year. Enumeration is host-agnostic
— an HTTPS mirror, a Nextcloud share and a local cog_mirror() copy all
take one code path — where hf:// would bind the reader to one vendor's
protocol and still need special-casing, since manifest fetching goes
through httr2, which cannot speak hf://.
The default is now the public Hugging Face mirror: CC-BY-4.0, no
credential, CDN-backed. USCOGDATA_URL still overrides.
Also in this release
Metadata: Authors@R is Jared E. Knowles aut/cre with ORCID,
Civilytics Consulting LLC cph/fnd (it was an org with no human, so citation() and r-universe had nothing to render). Adds URL and BugReports. MaxCorpusSchema 5 -> 7 — it contradicted .validate_schema() and the published corpus by two versions.
License: adds LICENSE.md (the repo carried no license text at
all) and corrects the copyright holder.
Vignettes now ship..Rbuildignore excluded ^vignettes$, so vignette("total-spending") failed for every user — while the README
told them to run it.
pkgdown indexed 6 of 14 exports, so the docs site did not build.
README rewritten for a stranger, with a quickstart that runs
unconfigured and a "Where the data comes from" section linking the API
docs, the live API, the HF corpus and the Census source.
NEWS recast as a release announcement; the 0.2.0 changelog is kept
verbatim, the 0.1.0 development churn dropped.
CONTRIBUTING added, including the cog-api dependency: its CI clones
this package at USCOGDATA_REF, defaulting to main with no pin.
Test plan
devtools::test() offline against the bundled fixture — 968 passing, 0 failures
USCOGDATA_LIVE_TEST=true devtools::test() against the published corpus — 974 passing, 0 skips
Both vignettes resolve from an installed copy; citation() renders correctly
Cold start: a rocker/r-ver:4.4 container that had never seen the
package installed it from a source tarball and ran the README
quickstart verbatim with no environment variables set
cog-api suite against this branch — 739 passing, 0 failures
The cold-start check is the only one that would have caught the original
defect, and its absence is why the defect survived. It is now in the
release checklist in CONTRIBUTING.md.
Notes
Independent of #40 (fixture regeneration); either can merge first.
Distribution work — public Gitea, GitHub mirror, r-universe — is
deliberately not in this PR. r-universe publishes check results on
registration, so it comes after this is green and merged.
## Summary
Makes the package installable and usable by someone who is not us, and
cuts **0.3.0** as the first public release.
### The defect this fixes
**`uscogdata` could not read a remote corpus at all.** Two independent
causes, either fatal on its own:
1. The default corpus URL was a `REPLACE_WITH_SHARE_TOKEN` sentinel and
no document in the package supplied a working one — a new user got
`uscogdata_url_not_configured` with nothing actionable behind it.
2. `inst/sql/10-long.sql` globbed `{url}data/long/**/*.parquet`. DuckDB
cannot expand a glob over generic HTTP — there is no directory listing
— and `allow_asterisks_in_http_paths` does not help; it forwards the
literal `**/*` as a filename and 404s.
It went unnoticed for months because **every** path that exercised the
package used a *local* corpus: the test fixture, and the production API
(a host mount). Nothing ran it the way a new user does.
Partition paths now come from the corpus manifest's `long_partitions[]`
instead of a glob. Measured against the published corpus: the explicit
list returns the same **46,148,034 rows** the `hf://` glob does, with
`hive_partitioning` still recovering `year`. Enumeration is host-agnostic
— an HTTPS mirror, a Nextcloud share and a local `cog_mirror()` copy all
take one code path — where `hf://` would bind the reader to one vendor's
protocol and still need special-casing, since manifest fetching goes
through httr2, which cannot speak `hf://`.
The default is now the public Hugging Face mirror: CC-BY-4.0, no
credential, CDN-backed. `USCOGDATA_URL` still overrides.
### Also in this release
- **Metadata**: `Authors@R` is Jared E. Knowles `aut`/`cre` with ORCID,
Civilytics Consulting LLC `cph`/`fnd` (it was an org with no human, so
`citation()` and r-universe had nothing to render). Adds `URL` and
`BugReports`. `MaxCorpusSchema` 5 -> 7 — it contradicted
`.validate_schema()` and the published corpus by two versions.
- **License**: adds `LICENSE.md` (the repo carried no license text at
all) and corrects the copyright holder.
- **Vignettes now ship.** `.Rbuildignore` excluded `^vignettes$`, so
`vignette("total-spending")` failed for every user — while the README
told them to run it.
- **pkgdown** indexed 6 of 14 exports, so the docs site did not build.
- **README** rewritten for a stranger, with a quickstart that runs
unconfigured and a "Where the data comes from" section linking the API
docs, the live API, the HF corpus and the Census source.
- **NEWS** recast as a release announcement; the 0.2.0 changelog is kept
verbatim, the 0.1.0 development churn dropped.
- **CONTRIBUTING** added, including the cog-api dependency: its CI clones
this package at `USCOGDATA_REF`, defaulting to `main` with no pin.
## Test plan
- [x] `devtools::test()` offline against the bundled fixture — **968 passing**, 0 failures
- [x] `USCOGDATA_LIVE_TEST=true devtools::test()` against the published corpus — **974 passing**, 0 skips
- [x] `R CMD check --as-cran` — **0 errors, 0 warnings, 0 notes**
- [x] `pkgdown::build_site()` — clean, all metadata ok
- [x] Both vignettes resolve from an installed copy; `citation()` renders correctly
- [x] **Cold start**: a `rocker/r-ver:4.4` container that had never seen the
package installed it from a source tarball and ran the README
quickstart verbatim with **no environment variables set**
- [x] **cog-api** suite against this branch — **739 passing**, 0 failures
The cold-start check is the only one that would have caught the original
defect, and its absence is why the defect survived. It is now in the
release checklist in `CONTRIBUTING.md`.
## Notes
- Independent of #40 (fixture regeneration); either can merge first.
- Distribution work — public Gitea, GitHub mirror, r-universe — is
deliberately **not** in this PR. r-universe publishes check results on
registration, so it comes after this is green and merged.
- Design spec: `specs/2026-08-08-public-release-design.md`
- Plan: `plans/2026-08-08-public-release.md`
Covers the P0 finding that the package cannot read the corpus remotely at
all -- no working default URL, and Hive globs are unsupported over generic
HTTP by DuckDB 1.5.5. Fix is manifest-driven file enumeration (measured:
46,148,034 rows over plain https, identical to the hf:// glob) plus a
working public default.
Also: seven release-readiness fixes, a README restructured for a stranger,
NEWS rewritten as an initial release rather than a pre-release churn log,
and the Gitea-canonical/GitHub-mirror/r-universe distribution mechanics.
Ten tasks, 58 steps, TDD throughout. Tasks 1-3 fix the P0 (manifest
enumeration, working default URL, and the live-corpus test whose absence
let the defect survive); 4-7 are metadata and packaging; 8-10 rewrite
README, NEWS and CONTRIBUTING.
Distribution mechanics stay out of scope -- r-universe publishes check
results on registration, so it comes after final verification is green.
Spec and plan were written against a branch 25 commits behind main, where
the package still read 0.1.0. It is 0.2.0, with a real 0.2.0 changelog in
NEWS that the plan would have deleted.
0.3.0 rather than 0.2.0 because remote corpus reads go from broken to
working and the default URL from placeholder to live -- user-visible
behaviour, so a minor bump. Not 1.0.0: types 4 and 5 remain out of scope.
Task 9 now prepends a 0.3.0 section, keeps 0.2.0 verbatim with a diff
check to prove it, and drops only the 0.1.0 development churn. Task 4
gains the Version bump.
DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.
Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.
Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
The default was a REPLACE_WITH_SHARE_TOKEN sentinel and no document in
the package supplied a working URL, so a new user installing uscogdata
had no path to a session at all -- just an actionable-looking error with
nothing actionable behind it.
The default is now the public HuggingFace mirror: CC-BY-4.0, no
credential, CDN-backed, and it keeps the origin's uplink out of the read
path. USCOGDATA_URL and options(uscogdata.url=) still override, so
Nextcloud and cog_mirror() copies are unaffected.
The sentinel guard stays for half-edited configs; the two tests covering
it set the URL explicitly, so they only needed renaming to stop calling
it 'the default'.
The remote-read defect survived because every test path used a local
corpus, and so did the API in production. This is the only test that runs
the package the way a new user does: no USCOGDATA_URL, no option, no
fixture -- just install and call a verb.
Gated on USCOGDATA_LIVE_TEST so offline CI skips rather than fails.
Measured against the live corpus while writing this: cog_spending for one
government is 3.9s for a single year and 5.9s across 2000-2022. Well above
the 1.5-2.8s raw parquet scan, because the verbs also join crosswalks,
resolve categories and assemble provenance.
Authors@R was an org with no human, so citation() and the r-universe
maintainer page had nothing to render and ORCID could not collate this
with merTools. The given-name vector c("Jared", "E.") matches merTools
exactly; person("Jared", "E. Knowles") would render the same but put the
middle initial in the family-name slot.
MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the
published corpus is 7 -- metadata contradicting code by two versions.
0.3.0 rather than 0.2.0: remote reads go from broken to working and the
default URL from placeholder to live, which is user-visible behaviour.
LICENSE held only the two-line stub and no LICENSE.md existed, so the
repo carried no license text for a human browsing it or for GitHub's
license detector.
usethis::use_mit_license() writes LICENSE.md but leaves an existing
LICENSE alone, so the stub kept saying 'Civilytics' while the full text
said 'Civilytics Consulting LLC'. Corrected by hand, with a test pinning
the two together.
.Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed
uscogdata had no vignettes at all -- while the README instructed users to
run vignette("total-spending"), which failed for every one of them.
Both build offline: total-spending points USCOGDATA_URL at the bundled
fixture, population-denominators is eval = FALSE. Confirmed present in the
built tarball as both source and rendered inst/doc/.
The test also pins the fixture as never-excluded -- it is what lets
R CMD check run with no credentials on r-universe and GitHub Actions.
The reference index covered 6 of 14 exports, so pkgdown errored on the
eight missing topics and the docs site did not build at all. Adds a
Comparison & aggregation section and a Corpus metadata section, lists both
vignettes as articles, and sets url so canonical links and search resolve.
build_site() now completes clean: reference metadata ok, no problems.
Reordered around a new user: what the data is, where it comes from,
install, a quickstart that runs with no configuration, then the
full-dollars warning and the concepts that decide whether a published
number is right.
Adds a 'Where the data comes from' section linking the API documentation
site, the live API, the Hugging Face corpus and the Census source, so
attribution and provenance are reachable from the top rather than implied.
Drops the sibling-repo path, the commented-out install line, the Status
block, and the release advice telling you to strip the fixture -- which
would break the vignette and leave public CI unable to check without
credentials.
The quickstart passes years=; cog_spending() has no full-history default,
so the obvious one-liner errors on a reader's first call.
NEWS described changes relative to states no user had ever seen --
'Breaking: corpus schema_version 4', 'the package now requires...' --
across the whole pre-release development. To someone deciding whether to
depend on this, that reads as instability.
0.3.0 is written as an announcement: what it covers, the verbs, that
reading the corpus now works out of the box, four things to know before a
first query, and the known limits. The 0.2.0 changelog is kept verbatim.
The 0.1.0 development log is dropped; that history is in git.
cog_explain() now documents what provenance actually holds, since the
README points readers at it -- in particular why series_break_refs and
corpus_break_refs are separate fields rather than one list.
Moves developer, testing and release instructions out of the README,
minus the fixture-stripping advice, which was wrong.
Explains that a GitHub PR closes itself as merged once the mirror syncs,
because the merge preserves the contributor's SHAs -- so a PR closing
without a visible Merge click reads as success rather than rejection.
Documents the cog-api dependency: its CI clones this package at
USCOGDATA_REF, defaulting to main with no pin, so anything merged here
reaches the API's next build. Includes the commands to run its suite
against a branch first.
Removing ^vignettes$ was right; removing ^doc$ and ^Meta$ with it was
not. Those are devtools::build_vignettes() artefacts, not sources -- R CMD
build regenerates inst/doc/ from vignettes/ by itself, and shipping the
local copies earned a 'non-standard file/directory found at top level'
NOTE.
R CMD check --as-cran is now 0 errors, 0 warnings, 0 notes.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Summary
Makes the package installable and usable by someone who is not us, and
cuts 0.3.0 as the first public release.
The defect this fixes
uscogdatacould not read a remote corpus at all. Two independentcauses, either fatal on its own:
REPLACE_WITH_SHARE_TOKENsentinel andno document in the package supplied a working one — a new user got
uscogdata_url_not_configuredwith nothing actionable behind it.inst/sql/10-long.sqlglobbed{url}data/long/**/*.parquet. DuckDBcannot expand a glob over generic HTTP — there is no directory listing
— and
allow_asterisks_in_http_pathsdoes not help; it forwards theliteral
**/*as a filename and 404s.It went unnoticed for months because every path that exercised the
package used a local corpus: the test fixture, and the production API
(a host mount). Nothing ran it the way a new user does.
Partition paths now come from the corpus manifest's
long_partitions[]instead of a glob. Measured against the published corpus: the explicit
list returns the same 46,148,034 rows the
hf://glob does, withhive_partitioningstill recoveringyear. Enumeration is host-agnostic— an HTTPS mirror, a Nextcloud share and a local
cog_mirror()copy alltake one code path — where
hf://would bind the reader to one vendor'sprotocol and still need special-casing, since manifest fetching goes
through httr2, which cannot speak
hf://.The default is now the public Hugging Face mirror: CC-BY-4.0, no
credential, CDN-backed.
USCOGDATA_URLstill overrides.Also in this release
Authors@Ris Jared E. Knowlesaut/crewith ORCID,Civilytics Consulting LLC
cph/fnd(it was an org with no human, socitation()and r-universe had nothing to render). AddsURLandBugReports.MaxCorpusSchema5 -> 7 — it contradicted.validate_schema()and the published corpus by two versions.LICENSE.md(the repo carried no license text atall) and corrects the copyright holder.
.Rbuildignoreexcluded^vignettes$, sovignette("total-spending")failed for every user — while the READMEtold them to run it.
unconfigured and a "Where the data comes from" section linking the API
docs, the live API, the HF corpus and the Census source.
verbatim, the 0.1.0 development churn dropped.
this package at
USCOGDATA_REF, defaulting tomainwith no pin.Test plan
devtools::test()offline against the bundled fixture — 968 passing, 0 failuresUSCOGDATA_LIVE_TEST=true devtools::test()against the published corpus — 974 passing, 0 skipsR CMD check --as-cran— 0 errors, 0 warnings, 0 notespkgdown::build_site()— clean, all metadata okcitation()renders correctlyrocker/r-ver:4.4container that had never seen thepackage installed it from a source tarball and ran the README
quickstart verbatim with no environment variables set
The cold-start check is the only one that would have caught the original
defect, and its absence is why the defect survived. It is now in the
release checklist in
CONTRIBUTING.md.Notes
deliberately not in this PR. r-universe publishes check results on
registration, so it comes after this is green and merged.
specs/2026-08-08-public-release-design.mdplans/2026-08-08-public-release.mdDuckDB cannot expand a glob over generic HTTP -- there is no directory listing, and allow_asterisks_in_http_paths only forwards the literal '**/*' as a filename, which 404s. So every remote corpus read failed. Only local paths worked, which is how the API (a host mount) and the test fixture run, so nothing ever caught it. Measured against the published corpus: the explicit list returns the same 46,148,034 rows the hf:// glob does, with hive_partitioning still recovering year from the paths. Building it from the manifest keeps the reader host-agnostic rather than binding it to one vendor's protocol. Also extracts .render_view_sql(). Four test sites had hand-rolled the {url} substitution -- one commented as doing it 'exactly as .register_views() does' -- and all four broke on the second token. They now share the one function that knows the vocabulary, and a new test renders every SQL file to prove no token survives.Authors@R was an org with no human, so citation() and the r-universe maintainer page had nothing to render and ORCID could not collate this with merTools. The given-name vector c("Jared", "E.") matches merTools exactly; person("Jared", "E. Knowles") would render the same but put the middle initial in the family-name slot. MaxCorpusSchema claimed 5 while .validate_schema() accepts 4-7 and the published corpus is 7 -- metadata contradicting code by two versions. 0.3.0 rather than 0.2.0: remote reads go from broken to working and the default URL from placeholder to live, which is user-visible behaviour..Rbuildignore excluded ^vignettes$, ^doc$ and ^Meta$, so an installed uscogdata had no vignettes at all -- while the README instructed users to run vignette("total-spending"), which failed for every one of them. Both build offline: total-spending points USCOGDATA_URL at the bundled fixture, population-denominators is eval = FALSE. Confirmed present in the built tarball as both source and rendered inst/doc/. The test also pins the fixture as never-excluded -- it is what lets R CMD check run with no credentials on r-universe and GitHub Actions.jared referenced this pull request2026-08-08 18:44:22 -04:00