Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Development

Layout

  • schema.sql: storage layout.
  • src/store.rs: reads, the COPY writer, the writer lease and reaping.
  • src/pg.rs: the connection pool (deadpool-postgres, most recently used first) and TLS.
  • src/lib.rs: the C ABI (extension/src/include/pgvfs.h).
  • extension/: the thin C++ DuckDB FileSystem adapter. The adapter has to be C++ because DuckDB’s stable C API can use filesystems but cannot register one.
  • bench/: benchmarks. lake.py is the harness (lockstep readers, cold, warm and new-parameter passes, per-query profiling), with two datasets: city.py (Overture Houston, 1M rows) and hits.py (ClickBench, 100M rows, 105 columns). readers.py (the weekly ClickBench run) and lookups.py (1,000-geometry lookups, used by scripts/sweep.sh to split cores between PostgreSQL and readers) are standalone.
  • scripts/: DuckDB input fetching, runners using disposable PostgreSQL containers, and the Pages site builder.
  • docs/: the mdbook site, built from this README plus generated build and benchmark tables.

Build and test

just check                   # fmt, clippy, unit tests
just ext                     # stable DuckDB (1.5.6) -> target/ext/release/pgvfs.duckdb_extension
just ext nightly [WHEEL]     # DuckDB 2.0 dev wheel (default: newest on PyPI) -> target/ext/nightly/
just contract                # storage contract on a disposable PostgreSQL (PG_IMAGE=postgres:11..18)
just compat                  # the contract on every supported major
just e2e [release|nightly]   # contract + DuckDB end-to-end through the built extension
just bench --parts 20 --readers 1,4,16
scripts/lake.sh city                # Houston: 13 readers, 10 passes (~1 min incl. download)
scripts/lake.sh hits --passes 5     # 100M rows (14 GB download, cached)
MODE=profile scripts/lake.sh city   # where one warm query of each kind goes

A C++ DuckDB extension must be statically linked against duckdb_static of the exact DuckDB build that loads it. Python loads DuckDB with RTLD_LOCAL, so the host’s symbols are out of reach, and DuckDB’s own extensions are built the same way. DuckDB is never compiled here. scripts/duckdb.sh fetches its headers and prebuilt static libraries, and the container only links (extension/build.sh, a few seconds):

  • release: the release’s static-libs-linux-amd64.zip and source tarball, pinned by SHA-256 in scripts/duckdb.sh. To change versions, update the version and both digests together.
  • nightly: each 2.0 dev wheel is cut by a manual run of DuckDB’s Main CI on its exact commit (PRAGMA version). That run keeps duckdb-static-libs-linux-amd64.tar.gz for 90 days, and the download is checked against GitHub’s recorded digest. Fetching it needs a GitHub token (gh auth login or GH_TOKEN). The extension footer carries DuckDB’s version tag, or the commit id for -dev builds.

target/ext/<target>/DUCKDB_PY and DUCKDB_VERSION record the wheel each build loads into and its extension-repository directory.

CI (.github/workflows/):

  • Each push: fmt, clippy, unit tests, then the contract on PostgreSQL 18 and e2e for the stable build.
  • Weekly (Monday), or by hand:
    • the contract on PostgreSQL 11;
    • e2e for the stable and newest 2.0 dev builds;
    • publishing each build as the pre-release duckdb-<version> (scripts/release.sh, keeping every stable build and the last 4 dev builds);
    • a 10M-row, 1–4-reader benchmark appended to the bench pre-release’s history.jsonl.
  • Pages (pages.yml: weekly and on docs changes): https://pgvfs.adonm.dev is assembled from the releases by scripts/site.py. It is an mdbook (docs/) whose pages include this README’s sections, plus the INSTALL ... FROM tree. Preview it with just site.