In the review before pgoverlay 1.0, one SELECT count(*) copied a 488.5 MiB table into a copy-on-write branch. The branch had never written to that table. It read it once, and now it owned a copy of it.
Preventing exactly that is what pgoverlay is for. It gives you copy-on-write branches of a Postgres database: seed once from a running server, then create disposable, writable copies that share the seeded data and store only what they change. On the default backend, each branch is a stock postgres container whose PGDATA is an OverlayFS mount: the seed read-only in the lower layer, an empty writable layer on top of it.
In July I wrote about the first time that copy-on-write turned out to be a copy, and fixed it. What I believed was left was the cost of the first write to a file, so branches that mostly read would stay in the tens of megabytes. That was wrong, for the same reason as the first bug.
v1.0.0, released on 2026-09-29, fixes it on the default setup: a read copies no table data, and on XFS or btrfs a write copies blocks rather than files. This post replaces the July one. The two bugs are one mechanism, and the second is the one that mattered.
The first copy: before recovery started
The first real benchmark said branching a 5 GiB database took 61.9 seconds and left a 5.05 GiB writable layer behind; at 1 GiB it was 7.66 s and 1.05 GiB. That was hack/benchmark.sh in June, p50 of five runs on a Colima VM (4 CPUs, 8 GiB RAM, kernel 6.8, Docker 28.4.0 with its volumes on a VM-local ext4 disk) on an M1 Pro, with databases generated by pgbench -i. The writable layer was the size of the database at both scales. That is not overhead. That is a full copy.
The branch's own log said redo done ... elapsed: 0.00 s, so the minute went somewhere before WAL replay. A branch's seed came from pg_basebackup, so its first boot was crash recovery. Before replaying WAL, Postgres runs SyncDataDirectory(), which walks the data directory and fsyncs every file in it, so that nothing dirty from before the crash is left unflushed. To do that, fsync_fname_ext() in fd.c opens every regular file O_RDWR, because not every platform allows fsync() on a read-only descriptor.
It writes nothing. On OverlayFS that does not matter. The kernel's overlayfs documentation says that when a lower-layer file is accessed in a way that requires write access, "such as opening for write access", it is first copied up: copied whole from the lower layer into the upper one. The copy is triggered by the open, not by the write. With Postgres out of the frame, as root on a Linux box whose /var/tmp is a regular filesystem (ext4, xfs or tmpfs, not itself an overlay, as it is inside most containers):
mkdir -p /var/tmp/ovl/{lower,upper,work,merged}
dd if=/dev/zero of=/var/tmp/ovl/lower/big bs=1M count=100 status=none
mount -t overlay overlay \
-o lowerdir=/var/tmp/ovl/lower,upperdir=/var/tmp/ovl/upper,workdir=/var/tmp/ovl/work \
/var/tmp/ovl/merged
# open read-only: the upper layer stays empty
python3 -c 'import os; os.close(os.open("/var/tmp/ovl/merged/big", os.O_RDONLY))'
find /var/tmp/ovl/upper -type f # nothing
# open read-write, write nothing, close
python3 -c 'import os; os.close(os.open("/var/tmp/ovl/merged/big", os.O_RDWR))'
find /var/tmp/ovl/upper -type f # /var/tmp/ovl/upper/big
du -sh /var/tmp/ovl/upper/big # 100M
So the durability pass copied the whole dataset into the branch before it served a query. recovery_init_sync_method=syncfs (Postgres 14 and later) replaces the per-file pass with one syncfs() call per filesystem. With only that flag changed, the writable layer was 16 KiB when recovery finished, and the flag became the fix:
| 5.00 GiB database | Branch create (p50 of 5) | Writable layer after create |
|---|---|---|
| before | 61.85 s | 5.05 GiB |
| after | 1.89 s | 33.1 MiB |
The 33.1 MiB, measured once the branch was up, is almost all WAL written at the end of recovery. It comes back later.
syncfs flushes the whole filesystem under the branch's writable layer, which covers everything the branch could have left dirty. In July I also wrote that a branch's filesystem was "its own disposable overlay mount holding nothing else". It is not: on a Docker host that filesystem is shared with every other branch and container, so the cost is extra flush time when they write a lot, and the Postgres docs warn that syncfs errors may not reach Postgres before Linux 5.8. For a disposable branch that is fine. The hard costs are Postgres 13 and older, which pgoverlay refuses, and anything that is not Linux.
A SELECT is a read-write open
Postgres keeps each table and index in segment files of up to 1 GiB. In July I had measured creation and a bulk UPDATE, and concluded that "the cost moved from pay-at-create to pay-per-file-written": the first write to a segment copies the segment, and for the dev, CI and PR-review branches pgoverlay is built for, which read a lot and write a little, "branches stay in the tens of megabytes". The README went further: a branch "stores only the blocks it actually changes". I had not measured a branch that only reads. The claim rested on an assumption I never checked: that Postgres opens a file read-write only when it is going to write to it.
The review before 1.0 measured it. Commit c55cab2, Docker 29.6.2 on Linux amd64 with ext4, Postgres 17, and a table of 62,528 pages (488.5 MiB) that was VACUUM (FREEZE)d on the source before seeding:
| Branch state | Writable layer |
|---|---|
| fresh branch | 33.1 MiB |
after a 1-row UPDATE on a small table |
33.5 MiB |
after SELECT count(*) on a tiny table, then CHECKPOINT |
34.6 MiB |
right after SELECT count(*) on the big table |
523.1 MiB |
The file that appeared in the writable layer, base/5/16384, was 488.5 MiB: the whole table. 62,500 of its 62,528 pages were all-frozen, so hint bits (the per-row commit flags a first read can set; more on those below) cannot explain the copy.
It is the July mechanism on the ordinary read path. Postgres's storage manager opens every table and index segment read-write, whatever the query is going to do with it:
/* src/backend/storage/smgr/md.c, REL_17_STABLE */
static inline int
_mdfd_open_flags(void)
{
int flags = O_RDWR | PG_BINARY;
if (io_direct_flags & IO_DIRECT_DATA)
flags |= PG_O_DIRECT;
return flags;
}
/* mdopenfork(): the first segment of every relation */
fd = PathNameOpenFile(path, _mdfd_open_flags());
Postgres 14 has the same O_RDWR | PG_BINARY inline. On a branch, the first query of any kind that touches a table copies every segment it opens, up to 1 GiB each, and waits for the copy. On the evaluation host described below, with a warm page cache and a similar table plus its primary key, the first SELECT count(*) took 19.4 s and the second 116 ms: the difference is the copy. syncfs cannot help, because these are not recovery's flags. They are how stock Postgres reads. A branch that runs a test suite over a few large tables copies those tables, once per branch.
It nearly shipped as a documented limitation
The first plan was to document it: correct the claim, point read-heavy users at the ZFS and Kubernetes CSI backends, and prototype a fix for 1.1. The first v1.0.0 tag was cut that way. The release pipeline stopped it at an unrelated preflight check (the Helm chart still pointed at the release candidate's image) before anything was built or published. Before I cut it again, I decided it could not ship like that.
A copy-on-write tool whose branches copy every table they read does not have a limitation. It is the limitation, with a tool around it. So v1.0.0 waited for issue #49, which set the bar on the default setup (plain Docker on ext4, stock postgres:14 to 18 images): reads add approximately nothing, a write copies as little as the mechanism allows, create time stays independent of database size, no correctness regressions, no new privilege requirements. It also asked for Docker Desktop, Colima and OrbStack, which I did not get to measure.
The options, measured before building one
The candidates ran on one host first: Linux 7.0, Docker 29.6.2 with volumes on ext4, Postgres 17, and a table of about 520 MiB (1.5M rows: 488.5 MiB of heap plus a 32.2 MiB primary key) frozen on the source. On XFS (with reflink=1) and btrfs, OverlayFS can clone a file instead of copying it; more on that below.
| Mechanism | A SELECT count(*) adds |
A write copies | Verdict |
|---|---|---|---|
| OverlayFS, ext4 (before v1.0.0) | 522 MiB, table and index | the whole file, on open | replaced |
OverlayFS, metacopy=on |
the same bytes | the whole file, on open | rejected |
| fuse-overlayfs | the same bytes | the whole file, on open | rejected |
| LD_PRELOAD shim, ext4 | 0 bytes | the touched segment, once | default |
| OverlayFS on XFS or btrfs | about 0 | blocks | where the disk allows |
| loopback XFS pool | about 0 | blocks | deferred |
| btrfs, dm-thin, qcow2 | about 0 | blocks | rejected |
metacopy=onand fuse-overlayfs follow the same rule.metacopyonly defers the data copy for metadata changes such aschown; both copy the data on any open with write access, which is exactly Postgres's open. Byte for byte the same as before, and fuse-overlayfs adds a daemon.- A loopback XFS pool is the way I found to keep OverlayFS and stock images and still get block-level copy-on-write on a plain ext4 host, and it worked (a loopback btrfs pool would too, with a slower
fsync). But synchronous 8 kB writes on XFS on a loop device (pg_test_fsync, write plusfdatasync) ran at a third to two thirds of ext4: 429, 275 and 251 ops/s against 670, 757 and 621. It also needs privileges to manage a loop device across reboots, and its fixed capacity fails writes withENOSPCwhen the pool fills. Deferred, not rejected. - btrfs snapshots, dm-thin and qcow2 are block-level too, but need privileges, kernel modules that desktop VMs may lack, and fragile set-up; btrfs on a loop device had the slowest
fsyncof all. - A patched Postgres, a custom FUSE filesystem, a seccomp supervisor and a hardlink farm were rejected on paper, not measured: a fork of every supported major, or more moving parts in the I/O path than the problem needs.
- A per-file
FICLONEbackend with no overlay at all is left for later: it needs a reflink filesystem, and its create time grows with the number of files.
Open read-only, reopen on the first write
OverlayFS decides from the open flags. So the fix changes the flags Postgres's opens reach the kernel with, without changing Postgres.
A branch's Postgres now runs with a small LD_PRELOAD library, the lazyrw shim. When Postgres opens a table, index or transaction-status file under PGDATA read-write (without O_CREAT or O_TRUNC), the shim opens it O_RDONLY instead and remembers the descriptor, the path, the flags Postgres asked for and the file's identity. OverlayFS sees a read-only open and copies nothing.
The first write-class call on that descriptor (write, pwrite, pwritev, ftruncate, fallocate, copy_file_range, a shared writable mmap) reopens the path with the original flags, which is when OverlayFS copies the file up, and moves the new open file onto the same descriptor number, keeping its offset and flags, before the call goes ahead. Postgres never sees a different descriptor. A backend that only reads never upgrades.
/* internal/cow/lazyrw/lazyrw.c, simplified: locking, the identity check,
restoring the status flags and every error path left out (the real code
fails the write if the reopen, lseek or dup3 fails) */
ssize_t write(int fd, const void *buf, size_t n) {
if (slot_state(fd) == ST_DOWNGRADED && upgrade(fd) < 0)
return -1; /* never write through a read-only fd */
return real_write(fd, buf, n);
}
static int upgrade(int fd) {
struct slot *s = slot_for(fd);
off_t pos = lseek(fd, 0, SEEK_CUR);
int cloexec = fcntl(fd, F_GETFD) & FD_CLOEXEC;
/* reopen with the flags Postgres asked for: OverlayFS copies up here */
int nfd = raw_openat(AT_FDCWD, s->path, s->orig_flags | O_CLOEXEC, 0);
lseek(nfd, pos, SEEK_SET);
raw_dup3(nfd, fd, cloexec ? O_CLOEXEC : 0); /* same fd number */
raw_close(nfd);
s->state = ST_UPGRADED;
return 0;
}
Three details made it safe to turn on by default:
- Other backends. Postgres runs a process per connection, so after one backend's write copies a file up, others may still hold read-only descriptors for the lower file. Since Linux 4.19, OverlayFS points those at the copied-up file (stacked file operations); before 4.19 they would keep reading stale data. Every branch self-tests this on a scratch overlay when it starts, falls back to copying on open if the test fails, and reports which mode it is in (
pgoverlay_branch_cow_mode). - Truncation.
TRUNCATE,DROPandVACUUM FULLtruncate files to zero, and truncating a lower file copies all of it first. The shim turnstruncate(path, 0), andftruncate(fd, 0)on a descriptor it downgraded, into anO_TRUNCopen, for which OverlayFS copies nothing. - Scope. It is active only inside the
postgresserver binary withPGDATAset; the entrypoint shell,archive_commandandCOPY ... PROGRAMget pass-through wrappers. A descriptor is upgraded only if it is still the same file, and a failed upgrade fails the write rather than writing through a read-only descriptor. CI checks that every write-class libc function thepostgresbinary of each supported image imports is either interposed or reviewed. On Postgres 18, a branch running the shim pinsio_method=worker(18's default), so a source configured forio_uringcannot go around the shim, which sees libc calls, notio_uringsubmissions.
The four builds (glibc and musl, x86_64 and aarch64) are rebuilt byte for byte in CI and embedded in the binary.
A SELECT can still write
The 0 bytes in the options table was for a table frozen on the source. With a seed whose table had not been frozen, the shim's first SELECT in the evaluation still copied 68 MiB on Postgres 14 and 17, and 119 MiB on Postgres 18. Reads in Postgres are not always read-only, and on a fresh pg_basebackup seed there are three ways a read turns into a write:
- the copy ends with a
backup_label, so every branch would start with crash recovery, which writes the pages it replays; - the first read of a row whose inserting transaction has committed records that on the page, as a hint bit, and dirties it;
- reads prune dead row versions, and tables with old unfrozen transaction IDs get an anti-wraparound autovacuum in every branch.
Each is a real write, so each would copy the segment it touches into every branch. Seed settle does that work once, in the seed, right after pg_basebackup: start Postgres on it privately, let recovery finish, freeze and analyze every database (and pg_statistic once more, since ANALYZE wrote its rows after it was vacuumed), checkpoint, switch to a fresh WAL segment, and shut down cleanly. Branches then start without WAL replay, with nothing left to set. A fresh branch now holds 0.4 MiB on the amd64 GitHub runner below, where the June runs showed 33.1 MiB, almost all of it WAL that recovery wrote. And since the seed was shut down cleanly, Postgres skips SyncDataDirectory() altogether: syncfs is still set, but only a branch restarting after a crash needs it.
The last constant was the WAL segment. A branch appends its WAL to the segment holding the seed's shutdown checkpoint, so its first write copied that 16 MiB file. Because of the switch, the checkpoint now sits at the start of a zero-filled segment, and the settle turns everything after the segment's first two pages into a hole, after checking the tail is zeros and comparing the copy with cmp. Across five fresh branches each, the first write now takes 12 to 27 ms instead of 46 to 119 ms, and copies about 1 MiB instead of 16 MiB.
Blocks, where the filesystem allows it
The shim moves the copy from the first read to the first write. What that copy costs is up to the filesystem under the volumes.
On ext4 the kernel copies the data: a one-row UPDATE in a 1 GiB segment copies 1 GiB. The copy also lands somewhere other than where you would look for it. The UPDATE only changes a page in shared buffers; the file is first written, and copied, when that page is written out, usually at the next checkpoint. On the amd64 GitHub runner (medians of five rounds), the first one-row UPDATE of a 949 MiB table took 0.05 s, and the CHECKPOINT after it 5.51 s.
Before OverlayFS copies a file's data up, it tries to clone it:
/* fs/overlayfs/copy_up.c, Linux 6.12, abridged */
/* Try to use clone_file_range to clone up within the same fs */
cloned = vfs_clone_file_range(old_file, 0, new_file, 0, len, 0);
if (cloned == len)
goto out_fput;
/* Couldn't clone, so now we try to copy the data */
On XFS with reflink=1 and on btrfs, that clone (a reflink) shares every extent with the seed, takes milliseconds and uses no space, and a block rewritten later is copied by the filesystem, not the whole file. The same CHECKPOINT on XFS, on a loop device on the runner, took 0.13 s. On my test machine, a one-row UPDATE of a 113 MB table on XFS added 40,960 bytes that the branch does not share with the seed.
pgoverlay probes which one it has at startup, by copying up a 64 MiB file through a scratch overlay and checking with FIEMAP whether the copy shares its extents, and exports the answer as pgoverlay_cow_copyup_mode. du counts a clone as a copy (that same UPDATE showed +113 MB), so in clone mode branch usage counts only the extents a branch does not share. On a Docker host whose root filesystem is ext4, --volume-root puts the volumes on an XFS or btrfs disk instead.
What it costs
The release gate was throughput: with the shim, warm select-only throughput had to stay within noise of copying on open, and warm TPC-B within 5% of it. It ran on two dedicated GitHub-hosted VMs (amd64 and arm64, 4 vCPUs, kernel 6.17, Docker's volumes on ext4), five rounds with the order rotated each round, against pgbench scale 50 (about 750 MiB) plus a 949 MiB single-segment table, seeded and settled.
| amd64, medians of 5 rounds | before (copy on open) | v1.0.0 (shim) |
|---|---|---|
branch after a select-only run, du -sb |
768.0 MiB | 16.5 MiB |
| the same, allocated and not shared | 753.1 MiB | 1.6 MiB |
| select-only tps, warm | 118,064 | 116,833 |
| TPC-B tps, warm | 11,823 | 11,648 |
Paired within each round, the shim's warm throughput relative to copying on open was 1.006 for select-only and 0.984 for TPC-B on amd64, and 1.026 and 1.014 on arm64. The 16.5 MiB is the branch's WAL segment, 16 MiB long but mostly a hole. Cold, right after create, the copy moves where you would now expect it: on amd64, select-only was 7% slower without the shim, which copies during the reads, and TPC-B 6% slower with it, which copies at the first write (7% and 2% on arm64).
Beyond throughput, a torture suite: Postgres killed with SIGKILL in the middle of TPC-B and restarted (balances add up, pg_amcheck --heapallindexed is clean), TRUNCATE and DROP of seed tables (no table data copied), VACUUM FULL (the old file is not copied), branch-from-branch, reset, and a version matrix: Postgres 14, 15, 16, 17, 18 and 17-alpine (musl) all came up with the shim, and every read added 0 bytes.
What is still true
- On ext4, the first write to a segment still copies the segment, up to 1 GiB, and whatever writes the page out waits for it: once per segment, per branch. XFS or btrfs volumes remove that.
- A branch whose shim cannot run still copies on read. A failed kernel self-test, a libc or architecture with no build, or a static
postgresputs it in eager mode, andpgoverlay_branch_cow_modesays so. - XFS on a fast disk, and btrfs, were not measured under pgbench. The only reflink throughput numbers are from a spinning disk, which measures the disk, and a loop device, on which warm TPC-B was 21% below ext4.
- Not measured with the shim: Docker Desktop, Colima and OrbStack (Colima only ran the June benchmark), and throughput on Kubernetes hostPath.
- Crash testing is
SIGKILL, not power loss. A killed process leaves the page cache intact, so nothing here tests writes that never reached the disk. - Overlay mode needs
CAP_SYS_ADMINto mount. The Kubernetes CSI mode does not. - It is a dev and test data tool. Branches are disposable and point-in-time, with no backups, no replication and no merge-back.
Where this sits
Neon branches your data copy-on-write in its own cloud storage, and Supabase gives you a separate hosted instance per branch. Neither runs against the Postgres you already run. DBLab is the established self-hosted answer, with block-level copy-on-write on a ZFS or LVM pool you provision and run.
In July I called block-level copy-on-write the better primitive, and whole-file copy-up the price of running on plain Docker with stock images. With the shim and a settled seed, reads copy no table data, on Linux 4.19 or later. I measured that on ext4 and XFS; since the shim never opens a table read-write for a read, the filesystem underneath should not matter. On XFS or btrfs, writes are block-level too, with the same stock images, the same plain Docker, and no pool to run. On ext4 the price is still whole files, paid on the first write instead of the first read.
The part worth keeping
Measure the claim the product rests on. The July post measured creation and a bulk update, and inferred what reads would cost. The inference was the part that was wrong, and it was the one the whole design depends on.
Copy-on-write is triggered by intent, and Postgres declares intent on every read. In July the lesson was that O_RDWR on an overlay-backed path is a commitment, not a cheap default. What I missed is that Postgres makes that commitment for you, whatever the query does. The fix was not to change what Postgres does. It was to change what the kernel is told, until a write actually happens.
The release: v1.0.0. Every number above, with its raw runs, its host, and what was not measured: docs/benchmarks.md. The evaluation, option by option: issue #49. The project is Go and Apache-2.0: github.com/abd-ulbasit/pgoverlay.