fix: enumerate long partitions from the manifest instead of globbing

DuckDB cannot expand a glob over generic HTTP -- there is no directory
listing, and allow_asterisks_in_http_paths only forwards the literal
'**/*' as a filename, which 404s. So every remote corpus read failed.
Only local paths worked, which is how the API (a host mount) and the test
fixture run, so nothing ever caught it.

Measured against the published corpus: the explicit list returns the same
46,148,034 rows the hf:// glob does, with hive_partitioning still
recovering year from the paths. Building it from the manifest keeps the
reader host-agnostic rather than binding it to one vendor's protocol.

Also extracts .render_view_sql(). Four test sites had hand-rolled the
{url} substitution -- one commented as doing it 'exactly as
.register_views() does' -- and all four broke on the second token. They
now share the one function that knows the vocabulary, and a new test
renders every SQL file to prove no token survives.
This commit is contained in:
2026-08-08 17:11:44 -04:00
parent c042ee0b90
commit 99e1e86e37
5 changed files with 123 additions and 6 deletions
+5 -1
View File
@@ -1,3 +1,7 @@
CREATE OR REPLACE VIEW long AS
SELECT *
FROM read_parquet('{url}data/long/**/*.parquet', hive_partitioning = true);
-- {long_files} carries its own quoting: a bracketed list of every partition
-- the manifest enumerates, or a single quoted glob on fallback. Do NOT wrap
-- it in quotes. See .long_files_sql() in R/views.R for why a glob alone
-- cannot work over HTTP.
FROM read_parquet({long_files}, hive_partitioning = true);