How it works
- One writer, any number of readers. The first write (a file create, delete or rename) takes a session advisory lock on a dedicated connection. A second writer fails at once, and a crashed writer’s lock goes with its connection. Readers take no locks.
- Readers need
SELECTonly. They run no DDL, work on hot standbys, and see an empty store until the writer has installed the layout. - Files are immutable. A write streams rows under a fresh
file_idand publishes(volume, path) → file_idin the same transaction when the file closes.file_idis DuckDB’s cache version tag, so the external file cache never serves stale bytes. - Reads are one primary-key range query per 8 MiB piece. A large read fetches its pieces in parallel on pooled connections, streaming rows straight into DuckDB’s buffer.
- Deletes queue the old
file_id. Its rows outlive the delete by 10 minutes, so queries that already opened the file can finish it. The writer reaps in the background, at most once a minute.
Storage (schema.sql) is one table of 8120-byte inline rows (one tuple per
8 KB page, no TOAST). It needs PostgreSQL 11+ and no
extensions: no pg_cron, no superuser. The contract tests run on 11, 13, 15,
17 and 18.