From 3b66067e1a34cfe95bae01de1bc5368c74678e15 Mon Sep 17 00:00:00 2001 From: Zsolt Parragi Date: Sat, 10 Oct 2026 01:45:04 +0100 Subject: [PATCH] Repro aid: widen the lost-FD_CLOSE window for the walreceiver NOT FOR COMMIT. This exists so the 2022 walreceiver regression -- the one that caused 29992a6a509 to revert the Windows graceful shutdown -- can be reproduced on demand, in stock in-core TAP tests, on hardware that is too fast to hit it by itself. The hang is the second libpqsrv_get_result() in libpqrcv_receive()'s end-of-stream path. It waits with timeout 0 on a freshly built WaitEventSet, and on Windows the one-shot FD_CLOSE has already been consumed by the wakeup that delivered the preceding bytes, so nothing ever signals it again. Three conditions have to hold at the same time: (A) libpq still needs bytes, so PQisBusy() is true and that second get_result actually waits. This is already true in stock PostgreSQL: the walsender's "if (got_STOPPING) proc_exit(0)" exits without sending ReadyForQuery. (B) the FD_CLOSE was consumed by an *earlier* wakeup. That is a race against the walsender's exit, not its write: proc_exit(0) does real work before the FIN goes out, and a walreceiver that wakes in between sees the FIN land on the hang site's own handle, which wakes it harmlessly. This patch widens that window. (C) the close is graceful. Reverting 29992a6a509 is not sufficient on its own: with default settings the standby sends feedback the walsender never reads, and closesocket() on a socket with unread input makes Windows send an RST. FD_CLOSE then carries a non-zero iErrorCode and every later recv() fails, so the peer notices without needing an event. wal_receiver_status_interval = 0 is what makes the close genuinely graceful. The sleep has to sit after the socket is associated with this event handle and before the sleep, not before WaitLatchOrSocket(): a FIN that arrives while no handle is registered is reported by WSAEventSelect() when one is registered, and the caller then notices the EOF. It is gated on PG_WALRCV_PRESLEEP_MS and on MyBackendType, so the build is unchanged unless that variable is set in the server's environment. To reproduce, against master: git revert -n 29992a6a509 git show a8458f508a7 -- src/backend/storage/ipc/latch.c \ | sed 's|.../latch\.c|.../waiteventset.c|g' | git apply -R -3 - git apply this-patch printf 'wal_receiver_status_interval = 0\n' > /tmp/tc.conf TEMP_CONFIG=/tmp/tc.conf PG_WALRCV_PRESLEEP_MS=300 \ meson test -C build commit_ts/002_standby a8458f508a7 moved from latch.c to waiteventset.c after it was committed, which is why it needs the path retargeted rather than a plain revert. TEMP_CONFIG is read by PostgreSQL::Test::Cluster, so the test itself is untouched. commit_ts/002_standby.pl then dies with "standby never caught up", which is what drongo, fairywren, jacana and tern reported in 2022. Leaving a8458f508a7 in place makes it pass again, and that is the only difference. Two cautions: - "standby never caught up" is not by itself proof of the hang. An IPC::Run newer than 0.116 makes binmode the default on Win32, which adds a stray CR to every safe_psql() result and kills this same test with the same message; and a pre-sleep large enough to throttle replication can time the test out while nothing is wedged. Both were mistaken for the hang during this work. The reliable check is that the walreceiver is wedged in the untimed wait and its log stops dead after "started streaming WAL from primary". - (B) and (C) are manufactured here. The 2022 animals got (C) for free at the default 10s status interval. The mechanism is faithful; the reproduction is engineered. Measured on a GitHub Actions windows-2022 runner, 300ms reproduces it and 500ms and 1000ms do too; without the sleep it does not reproduce there at all. On a slower buildfarm animal the window should open by itself, which is how this was found in the first place. --- src/backend/storage/ipc/waiteventset.c | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/src/backend/storage/ipc/waiteventset.c b/src/backend/storage/ipc/waiteventset.c index 5c807c3b274..d201cc45fc0 100644 --- a/src/backend/storage/ipc/waiteventset.c +++ b/src/backend/storage/ipc/waiteventset.c @@ -1687,6 +1687,26 @@ WaitEventSetWaitBlock(WaitEventSet *set, int cur_timeout, } } + /* + * REPRO ONLY, not for commit. Sleep while this handle is armed, so a peer + * that writes and then closes gracefully has time to do both: the FIN is + * latched here and consumed by the WSAEnumNetworkEvents() call below, + * leaving the next WaitLatchOrSocket()'s fresh handle nothing to wake it. + */ + if (MyBackendType == B_WAL_RECEIVER) + { + static int presleep_ms = -1; + + if (presleep_ms < 0) + { + const char *s = getenv("PG_WALRCV_PRESLEEP_MS"); + + presleep_ms = s ? atoi(s) : 0; + } + if (presleep_ms > 0) + pg_usleep(presleep_ms * 1000L); + } + /* * Sleep. * -- 2.56.0.windows.1