Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation

From: Kirill Reshke <reshkekirill(at)gmail(dot)com>
To: nktpro(at)gmail(dot)com
Cc: pgsql-bugs(at)lists(dot)postgresql(dot)org
Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Date: 2026-09-27 10:33:39
Message-ID: CALdSSPjEM90Y_6RQGox7OFJC+oAurWy5YSuXaH8Mr7m-TMOyWA@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-bugs

On Sun, 27 Sept 2026 at 11:08, Jacky Nguyen <nktpro(at)gmail(dot)com> wrote:
>
> Hi PostgreSQL team,
>
> We encountered a standby startup PANIC after a switchover on PostgreSQL 18.6. A standalone reproducer and recovery TAP test reproduce it on 18.6, 17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10.
>
> Expected: a standby that has replayed a heap/visibility-map truncation restarts recovery successfully.
> Actual on restart:
>
> LOG: redo starts at 0/3059738
> WARNING: page 0 of relation base/5/16396_vm does not exist
> PANIC: WAL contains references to invalid pages
> LOG: startup process (PID 55286) was terminated by signal 6: Abort trap: 6
>
> Reproduction: Build PostgreSQL with --enable-tap-tests --enable-injection-points and install injection_points. Copy the attached t/058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run:
>
> make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl
>
> The test holds a standby restartpoint, promotes that standby, deletes and vacuums an all-visible table (truncating its heap and VM forks), rejoins the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs failed with this PANIC; on 18.4, 20/20 passed. A shell reproducer, exact steps, WAL excerpts, version matrix, and two candidate patches are included in the attachment.
>
> The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11 (083ac03). The runs were on macOS 15 / Darwin 24.6.0, arm64, clang 21.1.8, with debug and injection points enabled and without cassert. The original symptom was on Linux in a three-node streaming setup. The analysis in REPORT.md points to VM block reads without an FPI during redo after an already replayed VM truncation; that diagnosis and the candidate fixes are provided for review.
>
> Regards,
> Jacky Nguyen

Hi!
Thanks for report. Looks like this bug existed since first VM commit,
which was already missing invalid page guard from [0].
So, case from [0] reintroduces one we started to register VM pages.

Also I prefer candidate fix b from your email

[0] https://github.com/postgres/postgres/commit/defe93463c69f8e0bb717294a34d67c34ac0b03f

--
Best regards,
Kirill Reshke

In response to

Responses

Browse pgsql-bugs by date

  From Date Subject
Next Message Palak Chaturvedi 2026-09-27 15:08:08 Re: BUG #19701: GIN trigram index loses rows at similarity_threshold 0
Previous Message Jacky Nguyen 2026-09-27 06:07:49 PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation