Re: pg_resetwal with replication slot (17.11)

From: Chao Li <li(dot)evan(dot)chao(at)gmail(dot)com>
To: Rıdvan Korkmaz <serkan(dot)ridvan(dot)korkmaz(at)gmail(dot)com>
Cc: pgsql-hackers(at)postgresql(dot)org
Subject: Re: pg_resetwal with replication slot (17.11)
Date: 2026-09-29 01:12:22
Message-ID: C8035643-1DD9-4E00-ADB6-EE612DA94C84@gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

> On Sep 21, 2026, at 20:21, Rıdvan Korkmaz <serkan(dot)ridvan(dot)korkmaz(at)gmail(dot)com> wrote:
>
> Hi Dear Experts,
> I hit a case seems odd. I wonder if I do something unexpected, or something is here I can't see.
>
> First, these are all on test environment. The version is PostgreSQL 17.11, uses homebrew installation on MacOS.
>
> I have a master - replica setup, both are on the same host.
> master's
> PGDATA = m
> port = 15432
> rs = rep slot for streaming replication used by instance "r", created on master (m instance)
> max_wal_size = 4GB
> min_wal_size = 2GB
> wal_level = replica
>
>
> replica's
> PGDATA = r
> port = 25432
> primary_conninfo = created by pg_basebackup
> primary_slot_name = rs
>
>
> Case: I have 16MB WAL files on master instance (so on replica). I want to utilize 1GB WAL files.
> Here are the steps I take.
>
> 1. setup master - replica run on same host in respective directories and on ports
> 2. verify streaming replication works
> 3. verify "rs" (replication slot), master ("m" instance), and replica ("r" instance) have SAME WAL lsn
> 4. stop master, keep replica online (simulation for actual case) (pg_ctl-17 stop -D m)
> 5. run "pg_resetwal-17 -D m --wal-segsize=1024" on master instance.
> 6. start master instance, started, fine. Replica complains about "ERROR: requested WAL segment 000000010000000000000000 has already been removed", no worry.
> 7. stop master again, cool, done. (last log lines: "checkpoint complete", "database system is shut down")
> 8. start master -> bamm, could not start. (pg_ctl-17 start -D m -l m.log)
>
> There is no read, write between after step 3 (after verification of WAL lsns)
>
> Final failure log (step 8):
> 2026-09-21 14:47:45.525 +03 [22817] LOG: starting PostgreSQL 17.11 (Homebrew) on aarch64-apple-darwin25.6.0, compiled by Apple clang version 21.0.0 (clang-2100.1.1.101), 64-bit
> 2026-09-21 14:47:45.525 +03 [22817] LOG: listening on IPv4 address "127.0.0.1", port 15432
> 2026-09-21 14:47:45.525 +03 [22817] LOG: listening on Unix socket "/tmp/.s.PGSQL.15432"
> 2026-09-21 14:47:45.528 +03 [22820] LOG: database system was shut down at 2026-09-21 14:46:42 +03
> 2026-09-21 14:47:45.528 +03 [22820] LOG: invalid checkpoint record
> 2026-09-21 14:47:45.528 +03 [22820] PANIC: could not locate a valid checkpoint record at 0/40000110
> 2026-09-21 14:47:45.528 +03 [22817] LOG: startup process (PID 22820) was terminated by signal 6: Abort trap: 6
> 2026-09-21 14:47:45.528 +03 [22817] LOG: terminating any other active server processes
> 2026-09-21 14:47:45.529 +03 [22817] LOG: shutting down due to startup process failure
> 2026-09-21 14:47:45.529 +03 [22817] LOG: database system is shut down
>
>
>
> After pg_resetwal, first start of master successful, but a second start fails.
> I guess this causes master to be lost.
>
> I'm able to spot the issue:
> The issue is replication slot. If I would have removed replication slot before second start (do it between 6 and 7), it succeeds.
>
> Questions:
> 1. Is this behavior is expected?
> 2. Should replication slot case mentioned in PostgreSQL documents? (I checked yet could not see)
> 3. Am I doing something out of order, unexpected?
> 4. Once I understood the case, I dropped replication slot and able to start master. Now I want to copy m/global/pg_control to replica and m/pg_wal to replica as well and complete wal segment size change. I wonder if this way is documented or supported. I can say "it works" but does not mean "supported or documented at all".
>
> Thank you in advance.
>
> Attachments:<1-master-replica-setup-info.txt><3-all-wal-lsn-same.txt>
>

To make the issue easier to reproduce, I created the attached repro_slot_wal_segsize.sh. It basically follows the procedure described by Rıdvan. I added "sleep 1" before the second server stop so that the log messages generated by that stop are easier to distinguish. I also added some temporary logging to show the problem more explicitly. See the attached temp_log.diff for those changes.

Here are the server logs from my reproduction:
```
2026-09-28 16:49:57.428 CST [76061] LOG: EVAN checkpoint cleanup after decrement: segment 0
2026-09-28 16:49:57.428 CST [76061] LOG: EVAN RemoveOldXlogFiles: remove through segment 0, boundary 000000000000000000000000, end segment 1, recycle through segment 128
2026-09-28 16:49:58.786 CST [76148] LOG: EVAN restoring replication slot "rs" from disk
2026-09-28 16:49:58.787 CST [76148] LOG: EVAN computed replication slot minimum LSN 0/0151FA80
2026-09-28 16:49:58.882 CST [76155] LOG: EVAN KeepLogSeg entry: end 0/400000E8, slot minimum 0/0151FA80, current segment 1, input segment 1
2026-09-28 16:49:58.882 CST [76155] LOG: EVAN KeepLogSeg mapped slot minimum to segment 0
2026-09-28 16:49:58.882 CST [76155] LOG: EVAN KeepLogSeg exit: candidate segment 0, output segment 0
```

The following messages are generated by the second server stop, as shown by their timestamps:
```
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN checkpoint cleanup before KeepLogSeg: redo 0/400000E8, end 0/40000170, segment 1
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN KeepLogSeg entry: end 0/40000170, slot minimum 0/0151FA80, current segment 1, input segment 1
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN KeepLogSeg mapped slot minimum to segment 0
```

Here is where the problem begins. The slot’s restart_lsn is 0/0151FA80. After wal_segment_size has been changed to 1 GB, this call in KeepLogSeg():
```
XLByteToSeg(keep, segno, wal_segment_size);
```

maps that LSN to segment zero. The slot’s restart_lsn refers to the old WAL history and is no longer meaningful after pg_resetwal has replaced that history:
```
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN KeepLogSeg exit: candidate segment 0, output segment 0
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN checkpoint cleanup after KeepLogSeg: segment 0
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN checkpoint cleanup before decrement: segment 0
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN checkpoint cleanup after decrement: segment 18446744073709551615
```

CreateCheckPoint() then makes the problem worse by unconditionally decrementing _logSegNo. This decrement is normally required because RemoveOldXlogFiles() removes files whose segment numbers are less than or equal to the supplied boundary. However, because _logSegNo is already 0 and XLogSegNo is unsigned, the decrement wraps to UINT64_MAX:
```
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN RemoveOldXlogFiles: remove through segment 18446744073709551615, boundary 00000000FFFFFFFF00000003, end segment 1, recycle through segment 2
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN RemoveOldXlogFiles: candidate 000000010000000000000001, next segment 1, recycle through segment 2
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN RemoveOldXlogFiles: removing WAL segment 000000010000000000000001
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN RemoveXlogFile: candidate 000000010000000000000001, next segment 1, recycle through segment 2
2026-09-28 16:50:00.899 CST [76146] LOG: EVAN InstallXLogFileSegment: renaming pg_wal/000000010000000000000001 to pg_wal/000000010000000000000002 as segment 2
```

Consequently, RemoveOldXlogFiles() incorrectly recycles 000000010000000000000001 as 000000010000000000000002. Segment 1 is no longer available under its expected name, even though it contains the checkpoint created after pg_resetwal. The next startup therefore cannot locate the checkpoint referenced by pg_control.

Running pg_resetwal again creates another checkpoint and allows the server to start again. However, this is only a temporary recovery. If the stale slot remains, a later checkpoint cleanup can reproduce the same failure.

Actually, the repro must be run with assertions disabled to reach the WAL recycling. I initially reproduced it that way because I had disabled assertions in my sandbox while load testing another patch. With assertions enabled, the second server stop fails earlier:
```
TRAP: failed Assert("!(possible_causes & RS_INVAL_WAL_REMOVED) || oldestSegno > 0"), File: "slot.c", Line: 2230, PID: 32858
0 postgres 0x0000000104d538a0 ExceptionalCondition + 216
1 postgres 0x0000000104a3b664 InvalidateObsoleteReplicationSlots + 148
2 postgres 0x0000000104546fb0 CreateCheckPoint + 3656
3 postgres 0x00000001045456a0 ShutdownXLOG + 424
4 postgres 0x00000001049c0328 CheckpointerMain + 2276
5 postgres 0x00000001049c5b98 postmaster_child_launch + 464
6 postgres 0x00000001049cabd4 StartChildProcess + 308
7 postgres 0x00000001049c9ae8 PostmasterMain + 6128
8 postgres 0x000000010483c158 main + 924
9 dyld 0x00000001827ac4e4 start + 6992
2026-09-28 14:05:39.589 CST [32855] LOG: checkpointer process (PID 32858) was terminated by signal 6: Abort trap: 6
```

This shows that passing segment 0 to InvalidateObsoleteReplicationSlots() violates the function’s existing precondition.

Changing wal_segment_size caused the stale restart_lsn in this repro to map to segment 0, which exposed the underflow. However, the real problem is that pg_resetwal discards the previous WAL history while leaving replication slot state unchanged. The stored restart_lsn values are no longer valid in the new WAL history. Even when the primary-startup failure does not occur, an existing standby cannot resume replication from the discarded history. Its slot must be recreated, and the standby must be rebuilt from a fresh base backup.

How to fix? My first thought was to prevent changing wal_segment_size when replication slots exist. In that case, users would have to drop all replication slots before changing wal_segment_size, otherwise, pg_resetwal would fail with a hint.

On second thought, I don't think we need to add this complexity to pg_resetwal. When pg_resetwal is used to recover a cluster with corrupted WAL or a corrupted control file, the doc already instructs users to immediately dump the data, run initdb, and restore into a new cluster. Replication slots do not play a meaningful role in that recovery procedure.

The relevant special case is using --wal-segsize to change the WAL segment size of an otherwise sound cluster without running initdb. In that case, existing replication slots retain restart_lsn values referring to the discarded WAL history. As this repro demonstrates, increasing the segment size can make such a value map to incorrect segment numbers and trigger the failure.

Since pg_resetwal is not frequently used, this problem is limited to this special use of --wal-segsize, and the server can be recovered by running pg_resetwal again and then immediately dropping the stale slots, I think enhancing the doc is sufficient. The doc for --wal-segsize should tell users to drop all replication slots before changing the segment size. It should also explain that existing standbys cannot resume replication afterward and must be rebuilt from a new base backup.

See the attached v1 patch for my proposed documentation change.

To answer Rıdvan's questions:

> Questions:
> 1. Is this behavior is expected?

No.

> 2. Should replication slot case mentioned in PostgreSQL documents? (I checked yet could not see)

I think so.

> 3. Am I doing something out of order, unexpected?

For a planned WAL segment-size change on a sound cluster, the standby should first be stopped and the old replication slot should be dropped before running pg_resetwal.

If the slot was not dropped, it may be possible to start the primary once and drop the slot before the next checkpoint or shutdown. The standby must be stopped first so that it does not reacquire the slot. If startup has already failed, another pg_resetwal may allow one more startup, but the stale slot must then be removed immediately.

> 4. Once I understood the case, I dropped replication slot and able to start master. Now I want to copy m/global/pg_control to replica and m/pg_wal to replica as well and complete wal segment size change. I wonder if this way is documented or supported. I can say "it works" but does not mean "supported or documented at all”.

AFAIK, no. Copying only global/pg_control and pg_wal does not produce a consistent standby and is not a supported procedure. After the primary’s WAL history has been reset, the standby should be recreated from a fresh base backup and attached using a newly created replication slot.

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/

Attachment Content-Type Size
temp_log.diff application/octet-stream 4.9 KB
v1-0001-Document-replication-slot-handling-with-pg_resetw.patch application/octet-stream 1.5 KB
repro_slot_wal_segsize.sh application/octet-stream 2.4 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Previous Message Zsolt Parragi 2026-09-29 00:38:56 Re: injection_points: canceled or terminated waiters leak their wait slots