Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

How it works

  • One writer, any number of readers. The first write (a file create, delete or rename) takes a session advisory lock on a dedicated connection. A second writer fails at once, and a crashed writer’s lock goes with its connection. Readers take no locks.
  • Readers need SELECT only. They run no DDL, work on hot standbys, and see an empty store until the writer has installed the layout.
  • Files are immutable. A write streams rows under a fresh file_id and publishes (volume, path) → file_id in the same transaction when the file closes. file_id is DuckDB’s cache version tag, so the external file cache never serves stale bytes.
  • Reads are one primary-key range query per 8 MiB piece. A large read fetches its pieces in parallel on pooled connections, streaming rows straight into DuckDB’s buffer.
  • Deletes queue the old file_id. Its rows outlive the delete by 10 minutes, so queries that already opened the file can finish it. The writer reaps in the background, at most once a minute.

Storage (schema.sql) is one table of 8120-byte inline rows (one tuple per 8 KB page, no TOAST). It needs PostgreSQL 11+ and no extensions: no pg_cron, no superuser. The contract tests run on 11, 13, 15, 17 and 18.