| From: | Kirill Reshke <reshkekirill(at)gmail(dot)com> |
|---|---|
| To: | nktpro(at)gmail(dot)com |
| Cc: | pgsql-bugs(at)lists(dot)postgresql(dot)org |
| Subject: | Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation |
| Date: | 2026-09-27 10:33:39 |
| Message-ID: | CALdSSPjEM90Y_6RQGox7OFJC+oAurWy5YSuXaH8Mr7m-TMOyWA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-bugs |
On Sun, 27 Sept 2026 at 11:08, Jacky Nguyen <nktpro(at)gmail(dot)com> wrote:
>
> Hi PostgreSQL team,
>
> We encountered a standby startup PANIC after a switchover on PostgreSQL 18.6. A standalone reproducer and recovery TAP test reproduce it on 18.6, 17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10.
>
> Expected: a standby that has replayed a heap/visibility-map truncation restarts recovery successfully.
> Actual on restart:
>
> LOG: redo starts at 0/3059738
> WARNING: page 0 of relation base/5/16396_vm does not exist
> PANIC: WAL contains references to invalid pages
> LOG: startup process (PID 55286) was terminated by signal 6: Abort trap: 6
>
> Reproduction: Build PostgreSQL with --enable-tap-tests --enable-injection-points and install injection_points. Copy the attached t/058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run:
>
> make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl
>
> The test holds a standby restartpoint, promotes that standby, deletes and vacuums an all-visible table (truncating its heap and VM forks), rejoins the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs failed with this PANIC; on 18.4, 20/20 passed. A shell reproducer, exact steps, WAL excerpts, version matrix, and two candidate patches are included in the attachment.
>
> The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11 (083ac03). The runs were on macOS 15 / Darwin 24.6.0, arm64, clang 21.1.8, with debug and injection points enabled and without cassert. The original symptom was on Linux in a three-node streaming setup. The analysis in REPORT.md points to VM block reads without an FPI during redo after an already replayed VM truncation; that diagnosis and the candidate fixes are provided for review.
>
> Regards,
> Jacky Nguyen
Hi!
Thanks for report. Looks like this bug existed since first VM commit,
which was already missing invalid page guard from [0].
So, case from [0] reintroduces one we started to register VM pages.
Also I prefer candidate fix b from your email
[0] https://github.com/postgres/postgres/commit/defe93463c69f8e0bb717294a34d67c34ac0b03f
--
Best regards,
Kirill Reshke
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Palak Chaturvedi | 2026-09-27 15:08:08 | Re: BUG #19701: GIN trigram index loses rows at similarity_threshold 0 |
| Previous Message | Jacky Nguyen | 2026-09-27 06:07:49 | PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation |