Re: [PATCH] Clear FatalError earlier during crash restart

From: Muzzammil Sarwar <muzzammil(at)umaish(dot)com>
To: Michael Paquier <michael(at)paquier(dot)xyz>
Cc: Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org>, Noah Misch <noah(at)leadboat(dot)com>
Subject: Re: [PATCH] Clear FatalError earlier during crash restart
Date: 2026-10-10 21:50:20
Message-ID: CABUcsRSNPtOkAd41JKqztq1SyfH5GD_KHxpTJsz8M9Ad-bs6gg@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

Some data on the backpatch question, since the lack of reports was
mentioned.

I hit this hang outside of a test this week, on a server based on 18.6.
A backend crashed with SIGSEGV and my script ran "pg_ctl stop -m fast"
right after it. The stop never finished. The machine was busy, so the
restarted startup process was still syncing the data directory (about
3 seconds) when the shutdown request came in. The postmaster then kept
waiting for the checkpointer and the IO workers, as described upthread.
Only "-m immediate" brought it down.

It reproduces without a restore_command, with only a crash and the
timing. The attached script creates a cluster with about 9000 relation
files, so that the sync takes a few seconds, kills a backend with
SIGSEGV, and runs "pg_ctl stop -m fast" as soon as "reinitializing"
shows up in the log. What I got:

REL_17_STABLE (35c508af520) hangs (checkpointer)
REL_18_STABLE (3c6135f839d) hangs (checkpointer, IO workers)
REL_19_STABLE (28f836dacaf) hangs (checkpointer, IO workers)
master before 12f866cd5f7 (a32b8dc9865) hangs
master (f58ab18a904) stops in about a second
REL_18_STABLE + 12f866cd5f7 stops; 058 passes, and fails
without the fix
REL_19_STABLE + 12f866cd5f7 stops; 058 passes

12f866cd5f7 applies cleanly to REL_18_STABLE. On REL_19_STABLE the only
conflict is the test list in src/test/recovery/meson.build, where 059
is already listed.

I understand wanting it to settle on HEAD first. I mainly wanted to say
that it happens outside of tests: anything that stops the server right
after a backend crash can land in that window, and on a busy machine
the window is a few seconds long. Since 19 is not released yet, it may
be worth considering for it.

Thanks for the fix.

Regards,
Muzzammil Sarwar

On Wed, Oct 7, 2026 at 5:07 AM Michael Paquier <michael(at)paquier(dot)xyz> wrote:

> On Thu, Oct 01, 2026 at 10:54:55AM +0530, Ayush Tiwari wrote:
> > Thanks for this. The patch looks good to me.
> > I did multiple A/B testing and it works well.
>
> I've been going back and forth a bit in the postmaster code to try
> more if FatalError can be made again non-re-entrant for
> HandleFatalError() with an assertion. My first try was the
> introduction of a new flag to track if SIGQUIT was sent, but I was not
> completely sure that I got the semantics right. A second set of
> thoughts was pointing towards a more complicated scheme for the
> states, which felt not really exciting. At the end, I have just
> reused your solution, and applied it on HEAD. Let's let it brew first
> there. The lack of complaints on the matter (even if we had one back
> in 2023), and the v19 release being very close by don't argue in favor
> of a backpatch, at least it feels strongly so to me.
>
> Now down to the buildfarm to judge the stability of your test. The CI
> has cleared quite a few times.
> --
> Michael
>

Attachment Content-Type Size
repro_shutdown_hang.sh application/octet-stream 1.5 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Peter Geoghegan 2026-10-10 21:52:47 Re: index prefetching
Previous Message Zsolt Parragi 2026-10-10 21:30:29 Re: Proposal: Conflict log history table for Logical Replication