Fwd: [RFC PATCH v1] On-demand WAL replay: accept connections before crash recovery has applied the WAL

From: Srinath Reddy Sadipiralla <srinath2133(at)gmail(dot)com>
To: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Fwd: [RFC PATCH v1] On-demand WAL replay: accept connections before crash recovery has applied the WAL
Date: 2026-10-10 11:08:58
Message-ID: CAFC+b6rZrrWN+Oa5pt22qG5=3q73qx4pKAgYBq+Tx03XMw=DpQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

---------- Forwarded message ---------
From: Srinath Reddy Sadipiralla <srinath2133(at)gmail(dot)com>
Date: Sat, Oct 10, 2026 at 4:14 PM
Subject: [RFC PATCH v1] On-demand WAL replay: accept connections before
crash recovery has applied the WAL
To: pgsql-hackers <pgsql-hackers(at)postgresql(dot)org>

Hi,

After a crash, the server cannot accept connections until all WAL written
since the last checkpoint has been replayed and, for crash recovery, until
the end-of-recovery checkpoint has written the result back to disk. Both
are proportional to the amount of WAL, so time-to-connect after a crash is
bound by WAL volume.

The attached patches let the startup process scan that WAL without
applying the records that modify relation pages. It replays everything
else eagerly (transaction status, multixacts, smgr and database
operations, ... anything without a block reference, which is cheap and has
no later page read that could trigger it), and records, per page, the LSNs
of the records that touch it, in a dshash table in dynamic shared memory.
Then the server opens. A page is recovered the first time any process
reads it from disk, and a new auxiliary process recovers the rest in the
background; once it has finished, the index is destroyed and a checkpoint
is requested.

Where replay happens turned out to be the whole design, so that part
first. On-demand replay runs inside StartReadBuffersImpl(), the common
path of ReadBuffer() and the read streams: a page with pending records is
read synchronously under its BM_IO_IN_PROGRESS flag, replayed, and only
then marked BM_VALID. Nothing new is added to the buffer manager's
protocol: other backends that want the page wait in WaitIO() as they would
behind a slow disk, and since no backend has been handed the page, none can
hold pointers into it, which is why redo may take an exclusive lock on it
where it would normally need a cleanup lock (the same reasoning as in
ZeroAndLockBuffer()). Redo is handed the buffer being filled instead of
reading the page again. A record's other blocks report BLK_DONE, which
redo routines must handle anyway, because after a crash during recovery any
block may already be past the record's LSN; blocks a record initializes
get a scratch local buffer. A page is replayed exactly once (full-page
images restore unconditionally, so replaying twice would rewind it), and
redo of an indexed page never reads another indexed page, so replays do
not nest.

I went through a series of crashes before arriving there, and they are
probably the most useful part of this mail for anyone touching the buffer
manager: a double content lock, unbounded recursion when redo re-read its
own page through the hook, "incorrect local pin count: 2" because a
cleanup-lock record saw the caller's pin and redo's, ItemIdIsNormal
failures in a parallel worker that had read the page through a read stream
(which bypassed my earlier hook) and kept line pointers under a pin while
another backend replayed a prune, and full-page images rewinding pages that
had been modified after recovery. All of them came from one thing: the
page was visible before it was current.

Other design points:

- Checkpoints. A checkpoint promises that everything before its redo point
is on disk; pending pages would break that, so CreateCheckPoint() refuses
while the LSNindex exists, unwaited requests are postponed, and a clean
shutdown waits for the worker. A request that waits for the checkpoint
(CHECKPOINT, pg_backup_start(), CREATE/DROP DATABASE) therefore fails
with "checkpoint request failed" during the window; for a base backup
that is the right answer, since it would copy pages that are not current
yet. Whether such requests should wait for the worker instead is a
question for this thread (see below). An earlier version stopped the
checkpointer altogether, which hung startup in RegisterSyncRequest(): it
is also the process that absorbs every fsync request.

- Which records are eager. The rule is "defer only if there is a trigger
and a gain": a later page read that will find the record, and page I/O
that is avoided by waiting. An previous version deferred CLOG, and the
first commit on a CLOG page that had never been created failed; the quiet
failure mode would have been committed rows as aborted.

- Relations and databases dropped or truncated later in the WAL forget
their pending pages. A truncation is the one eager record that reads
relation pages (the last FSM and VM page, to clear their tails). When
the next record for such a page arrives, the startup process writes the
page out and drops it from shared buffers before indexing the record, so
that no resident page has pending records when the server opens. An
earlier version instead kept such a page current by replaying every later
record for it at once; that read those records' heap pages too, which
then had to be kept current the same way, and one VACUUM truncation of a
table under 255 MB (one VM page) made the rest of that table's records
eager, each reading its heap page in the startup process (E below).

- Redo routines check InRecovery (visibilitymap_set() asserts it); it is set
for the duration of on-demand redo.

- The feature applies to crash recovery under the postmaster only. Archive
recovery, standby mode and single-user mode log a message and replay WAL
as usual.

Numbers. Measured with a harness that crashes a pgbench run with
pg_ctl -m immediate, recovers the same crashed files twice (stock, and with
fast_crash_recovery on), and compares per-table hashes, an index-path read,
the tpcb balance invariant and pg_amcheck --heapallindexed right after the
server opens, after the worker has drained the index, and after a clean
restart. No checkpoint runs during the workload. The machine is a 6-vCPU
VMware VM with a virtual NVMe disk backed by the host, 15 GB RAM; the build
is -O2 without assertions.

A. pgbench scale 100 , with 2.0 GB of WAL, 5.4 M records, 220,823 distinct
pages, shared_buffers 128 MB, warm page cache:

stock on-demand
time to accept connections 15.1 s 2.7 s (5.5x)
redo / scan 5.8 s 2.7 s
end-of-recovery checkpoint 9.2 s 0
first seq scan after ready 1.3 s 2.4 s
background drain - 37.2 s

B. Same scenario, cold page cache (2.7 GB of WAL, 9.0 M records,
270,127 pages):

time to accept connections 96.1 s 43.0 s (2.2x)
redo / scan 65.7 s 42.2 s
of which I/O wait 50.8 s 35.5 s
end-of-recovery checkpoint 30.4 s 0.8 s
first seq scan after ready 1.7 s 3.2 s
background drain - 73.3 s

C. pgbench scale 200, shared_buffers 4 GB, 9.9 GB of WAL, 57.6 M records,
648,284 pages, cold cache (the hardest case: a large window with a big
shared_buffers, so stock recovery is mostly CPU):

time to accept connections 199.4 s 106.4 s (1.9x)
redo / scan 173.3 s 106.4 s
of which CPU / I/O wait 52.3 / 121.1 s 47.4 / 59.0 s
end-of-recovery checkpoint 26.1 s 0
first seq scan after ready 10.4 s 72.6 s
background drain - 506.8 s

D. pgbench scale 100, full_page_writes = off, 1.8 GB of WAL, 18.4 M
records, 323,250 pages, cold cache (the best case: every record needs its
page read):

time to accept connections 197.2 s 40.0 s (4.9x)
redo / scan 165.7 s 39.9 s
of which I/O wait 150.2 s 28.2 s
end-of-recovery checkpoint 31.6 s 0
first seq scan after ready 1.5 s 3.9 s
background drain - 45.3 s

E. The truncation case from the design notes: pgbench scale 100 + a
140 MB table that VACUUM truncates three times during the run while clients
update it, 1.5 GB of WAL, shared_buffers 128 MB, files cached, so this
measures the scan's own work. Before and after the eviction fix, same
workload, different crashes:

kept current evicted
scan (time to accept connections) 4.8 s 3.6 s
startup: relation pages read / written 23,942 / 4,654 8 / 5
startup: WAL read beyond the window 905 MB 303 MB
first scan of that table after ready 95 ms 600 ms

These runs show a reliable 2x to 5x reduction in hard downtime across
scenarios A–D. The exact multiple depends on how much I/O write-back
stock recovery is forced to do and also by completely bypassing the
massive end-of-recovery checkpoint, we shift the bottleneck entirely
away from disk writes down to raw scan CPU. This is a deliberate
architectural trade-off of total background work in exchange for
immediate availability. The deferred costs behave exactly as expected:
the first full scan of a hot table post-recovery is 1.5–7x slower due to
on-first-touch replay, and the background drain ultimately spends more
total time re-decoding WAL than stock recovery did. But the primary goal
of drastically shrinking the window where the database cannot accept
connections is achieved.

Two things I learned from measuring that I had wrong before. With
full_page_writes on, stock crash recovery reads almost no data pages (the
first record for each page after the checkpoint carries the page), so its
critical path is CPU to apply + the write-back, not random reads. And
the scan can be too fast for its own good: at -O2 it outran the kernel's
readahead window for 8 kB reads and spent 40 of 46 seconds waiting for
WAL, slower than stock recovery on the same files; a posix_fadvise
(WILLNEED) on each segment as it is opened fixed that, and the cold runs
above include it.

I intend to repeat the 10 and 50 GB points on EC2 with gp3 storage, where
the write-back and the evictions cost what they cost in production rather
than what a laptop's virtual disk feels like that minute.

Known limitations and also unknown :), which I would like to discuss:

1. The index is not bounded. It is roughly 5-15% of the WAL volume and
lives in DSM, which is /dev/shm, thinking to spill when hits a cap.
2. Also the fast crash recovery worker could be parallel, which makes
replaying all the pending pages recovered , so the queries won't be slow.
3. Pending pages are read synchronously, one at a time, and look like
buffer hits to the read stream, which shrinks its look-ahead, need
to add AIO magic here, trying to figure out how.
4. A checkpoint requested during the window fails ("checkpoint request
failed") rather than waits: CHECKPOINT, pg_backup_start() and
CREATE/DROP DATABASE included. The alternative is to have such a
request wait for the worker in RequestCheckpoint(), since the caller
was going to wait for the checkpoint anyway; I have that written, but
kept the error in v1 so that we can decide here. One visible
consequence: with the fast_crash_recovery forced on, 018_wal_optimize
can fail,
because it issues CHECKPOINT right after a crash restart.

0001 is the feature
0002 adds a TAP test (src/test/recovery/t/060) and documentation.

check-world passes with the setting forced on through TEMP_CONFIG,
including the whole recovery suite, whose logs show the worker replaying
pages after every crash. The one exception is 018_wal_optimize, which
can fail for the reason given in 4 above when the machine is loaded.

--
Thanks :)
Srinath Reddy Sadipiralla
EDB: https://www.enterprisedb.com/
"Hello, friend?" That's lame. Maybe I should give you a name. But that's a
slippery slope.
You're only in my head. We have to remember that. It's actually happened.
I'm talking to an imaginary person.

Attachment Content-Type Size
v1-0001-Replay-WAL-on-demand-after-a-crash-instead-of-bef.patch application/x-patch 94.0 KB
v1-0002-Add-a-TAP-test-and-documentation-for-fast-crash-r.patch application/x-patch 9.7 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Andrew Dunstan 2026-10-10 11:29:47 Re: Allow table AMs to define their own reloptions
Previous Message Zsolt Parragi 2026-10-10 11:03:10 Re: Fix detection of truncated zstd-compressed backups