cog_spending(), cog_revenue() and cog_balances() gain optional state/type
arguments. Both default to NULL, so every existing govid-based call is
unchanged.
The verbs took a cohort only as a govid vector, which .sql_lit_chr()
rendered into a quoted IN list and .verb_spendrev() embedded into 5-8
separate statements per call: the scope check, the main aggregate, the
per-capita join, the harmonization block, and the suggestion and
suppression queries. For type = "city" that list is 301,589 characters,
parsed and planned from scratch every time it appears.
Passing state/type instead expresses the cohort as a subquery against
canonical_fips_xwalk, so its size never enters the SQL string at all.
Measured on the production corpus, same FY2022 aggregate over the
20,106-government city cohort, DUCKDB_THREADS=2, median of 5:
IN (20,106 literals) -- 0.3.0 432 ms
join against a temp cohort table 132 ms
predicate on canonical_fips_xwalk 102 ms
no cohort filter at all (the floor) 105 ms
The predicate reaches the no-filter floor: the cohort restriction is
now free. End to end through cog_spending(category = "Police"),
1080 ms -> 271 ms, 3.99x -- larger than the single-query saving,
because the repetition across statements is what actually cost.
Design decisions, both made explicitly rather than left implicit:
- govid AND state/type INTERSECT. "These ids, narrowed to that
state/type" is a real query, and an error here could never be
relaxed later without breaking callers.
- A predicate cohort has no id list to report, so
provenance$scope$govids_found/govids_missing stay empty and a new
scope$cohort block carries state, type and n_governments. Resolving
the ids just to report them would put 20,000 govids in every
fleet-scale response body -- the cost this change removes. A
govid-named cohort's provenance is untouched.
state/type are coerced with .coerce_state_to_fips()/.coerce_type(), the
same helpers cog_gov_search() uses. That is load-bearing: the argument
is a postal abbreviation ("WI") while fips_state holds a FIPS code
("55"), and a predicate on the raw parameter matches nothing and returns
an empty result indistinguishable from "reported nothing". cog-api hit
exactly this trap optimizing the same path.
.attach_per_capita() now keys its population lookup on the govids present
in the result rather than the requested cohort. Those are the only ones
its LEFT JOIN can match, so the output is identical -- but it needs no id
list, and on a paginated call it looks up one page instead of the fleet.
Fixes uscogdata#58.
158 lines
7.2 KiB
Markdown
158 lines
7.2 KiB
Markdown
# uscogdata 0.4.0
|
|
|
|
## Cohorts can be named by predicate, not just by id
|
|
|
|
`cog_spending()`, `cog_revenue()` and `cog_balances()` gain optional `state`
|
|
and `type` arguments. Both default to `NULL`, so every existing call behaves
|
|
exactly as before.
|
|
|
|
Passing them expresses the cohort as a subquery against `canonical_fips_xwalk`
|
|
inside each statement, instead of round-tripping the ids through R and
|
|
rendering them back into a literal `IN` list:
|
|
|
|
```r
|
|
# before: resolve 20,106 ids in R, then embed them in every statement
|
|
ids <- cog_gov_search(NULL, state = "CA", type = "city")$canonical_govid
|
|
cog_spending(ids, years = 2022)
|
|
|
|
# now: the cohort never leaves the database
|
|
cog_spending(years = 2022, state = "CA", type = "city")
|
|
```
|
|
|
|
Measured against the production corpus, same FY2022 aggregate over the
|
|
20,106-government `type = "city"` cohort:
|
|
|
|
| cohort expressed as | time |
|
|
|---|---:|
|
|
| `IN (20,106 literals)` | 449 ms |
|
|
| join against a temp cohort table | 99 ms |
|
|
| predicate on `canonical_fips_xwalk` | **94 ms** |
|
|
| no cohort filter at all (the floor) | 88 ms |
|
|
|
|
**4.8x, within 7% of the floor.** The rendered `IN` list was 301,591
|
|
characters and was re-parsed in 5-8 separate statements per call, so the cost
|
|
was paid repeatedly; the predicate's size is constant in the cohort.
|
|
|
|
`state` and `type` use the same vocabulary and the same internal coercion as
|
|
`cog_gov_search()` -- `state` is a postal abbreviation (`"WI"`) even though the
|
|
crosswalk column holds a FIPS code (`"55"`).
|
|
|
|
Supplying `govid` **and** `state`/`type` intersects them: the governments in
|
|
`govid` that also match the predicate. Naming no cohort at all now aborts with
|
|
class `uscogdata_no_cohort` rather than R's "argument is missing" error.
|
|
|
|
When the cohort is named by predicate there is no id list to report, so
|
|
`provenance$scope$govids_found`/`govids_missing` are empty and
|
|
`provenance$scope$cohort` carries `state`, `type` and `n_governments` instead.
|
|
A `govid`-named cohort's provenance is unchanged.
|
|
|
|
## Fixes
|
|
|
|
* An unknown `state` abbreviation now aborts with "Unknown state abbreviation"
|
|
(class `uscogdata_unknown_state`) instead of base R's "subscript out of
|
|
bounds". `.state_abbrev_to_fips` is a named character vector, so `[[` on an
|
|
absent name threw before the curated message could be reached -- making that
|
|
message unreachable dead code in every verb that takes a `state`.
|
|
|
|
# uscogdata 0.3.0
|
|
|
|
First public release.
|
|
|
|
`uscogdata` provides curated R verbs over the Civilytics US Census of
|
|
Governments finance corpus: unit-level financial profiles, geographic rollups
|
|
and peer comparisons, with auditable provenance on every result.
|
|
|
|
## What it covers
|
|
|
|
Government types 0-3 (state, county, municipality, township), FY1967-FY2024 --
|
|
56 fiscal years, 46,148,034 rows, 190.6 MB. There is no source data for FY1968
|
|
or FY1969. Special districts (type 4) and school districts (type 5) are out of
|
|
scope pending validation.
|
|
|
|
## The verbs
|
|
|
|
`cog_spending()`, `cog_revenue()` and `cog_balances()` for flows and holdings;
|
|
`cog_gov_search()` to resolve place names (including basket mode for many at
|
|
once); `cog_find_peers()` and `cog_peer_compare()` for cohorts;
|
|
`cog_geographic_rollup()` for aggregates; `cog_categories()`, `cog_recipes()`,
|
|
`cog_manifest()` and `cog_explain()` for metadata and provenance; and
|
|
`cog_mirror()` for a local copy of the corpus.
|
|
|
|
## Reading the corpus now works out of the box
|
|
|
|
* The package reads the published corpus over HTTPS **with no configuration**.
|
|
Previously the default was a placeholder sentinel and no document in the
|
|
package supplied a working URL, so a new user had no path to a session.
|
|
* Remote reads work at all. The partitioned view used a glob, and DuckDB
|
|
cannot expand a glob over generic HTTP -- there is no directory listing to
|
|
expand against. Partition paths are now enumerated from the corpus manifest,
|
|
which is host-agnostic: an HTTPS mirror, a Nextcloud share and a local
|
|
`cog_mirror()` copy all take the same path.
|
|
* Nothing is written to disk in remote mode; DuckDB fetches only the row
|
|
groups a query needs.
|
|
|
|
## Four things to know before your first query
|
|
|
|
* **Amounts are in full US dollars.** The raw Census files report thousands;
|
|
the verbs multiply by 1000 on the way out. Do not multiply again.
|
|
* **Multi-government aggregates disclose their coverage.** The Census is a
|
|
complete enumeration only in years ending in 2 and 7; every other year is a
|
|
sample. Every such result carries `provenance$coverage` with per-year
|
|
`n_units_reporting`.
|
|
* **Absence means two different things.** Before FY2012 an absent cell means
|
|
Census published $0; from FY2012 it means not reported. `complete = TRUE`
|
|
labels which.
|
|
* **Series breaks reach you unasked.** Catalogued breaks intersecting your
|
|
query appear in provenance and in `cog_explain()`.
|
|
|
|
## Known limits
|
|
|
|
* Special districts (type 4) and school districts (type 5) are out of scope.
|
|
* Per-capita rollups exclude governments with no F-33 population, which is by
|
|
design but does silently narrow a rollup.
|
|
* `n_units_reporting` is category-conditional and is not a response rate.
|
|
* Employee-retirement (`X`) codes stop at FY2016, when those systems moved to
|
|
the Annual Survey of Public Pensions.
|
|
|
|
# uscogdata 0.2.0
|
|
|
|
## New features
|
|
|
|
* `cog_spending()` and `cog_revenue()` accept the reserved category
|
|
`"All Categories"`, returning one summed row per
|
|
`(year, canonical_govid, subtype)` across every category inside the
|
|
requested concept's subtype scope. Filtering the result to
|
|
`spend_subtype == "operations"` gives an operating-expenditure total.
|
|
`cog_geographic_rollup()` inherits it,
|
|
which is the efficient way to build a geographic total — previously a
|
|
caller had to issue one rollup per category and sum the results
|
|
(cog-api#37).
|
|
|
|
`"All Categories"` is not the same thing as `expenditure_concept = "total"`.
|
|
The concept chooses which subtypes are in scope; `"All Categories"` chooses
|
|
whether the rows inside that scope are broken out or summed.
|
|
|
|
* `cog_categories()` advertises `"All Categories"` for the expenditure and
|
|
revenue vocabularies, so the reserved value is discoverable.
|
|
|
|
* Coverage signposting (see "Signposting now catches partially-suppressed
|
|
categories" below) now also works in `category = "All Categories"` mode.
|
|
The recipe-suggestion candidate query used to be scoped by `category`,
|
|
which is never a match for the reserved `"All Categories"` value, so
|
|
`provenance$suggestions` always came back empty there — the one mode whose
|
|
whole point is "you cannot sum the wrong scope" was silently unable to
|
|
signal a wrong scope. The candidate query is now scoped by the concept's
|
|
subtype allowlist instead, symmetric with how `.build_verb_sql()` itself
|
|
scopes the summed total: Los Angeles County FY2011, `category = "All
|
|
Categories"` still excludes $271,589,000 of aggregate-published Public
|
|
Welfare (`E68`), but now names `recipe = "welfare_cash_e68_wide"` to
|
|
recover it instead of reporting zero suggestions.
|
|
|
|
## Documentation
|
|
|
|
* `cog_geographic_rollup()` and `cog_peer_compare()` now document that
|
|
`provenance$coverage`'s `n_units_reporting` is **category-conditional** and
|
|
is not a response rate: a government that was surveyed and genuinely spends
|
|
nothing in the requested category is indistinguishable from one never
|
|
surveyed (uscogdata#36).
|