| From: | Muzzammil Sarwar <muzzammil(at)umaish(dot)com> |
|---|---|
| To: | Michael Paquier <michael(at)paquier(dot)xyz> |
| Cc: | Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org>, Noah Misch <noah(at)leadboat(dot)com> |
| Subject: | Re: [PATCH] Clear FatalError earlier during crash restart |
| Date: | 2026-10-10 21:50:20 |
| Message-ID: | CABUcsRSNPtOkAd41JKqztq1SyfH5GD_KHxpTJsz8M9Ad-bs6gg@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
Some data on the backpatch question, since the lack of reports was
mentioned.
I hit this hang outside of a test this week, on a server based on 18.6.
A backend crashed with SIGSEGV and my script ran "pg_ctl stop -m fast"
right after it. The stop never finished. The machine was busy, so the
restarted startup process was still syncing the data directory (about
3 seconds) when the shutdown request came in. The postmaster then kept
waiting for the checkpointer and the IO workers, as described upthread.
Only "-m immediate" brought it down.
It reproduces without a restore_command, with only a crash and the
timing. The attached script creates a cluster with about 9000 relation
files, so that the sync takes a few seconds, kills a backend with
SIGSEGV, and runs "pg_ctl stop -m fast" as soon as "reinitializing"
shows up in the log. What I got:
REL_17_STABLE (35c508af520) hangs (checkpointer)
REL_18_STABLE (3c6135f839d) hangs (checkpointer, IO workers)
REL_19_STABLE (28f836dacaf) hangs (checkpointer, IO workers)
master before 12f866cd5f7 (a32b8dc9865) hangs
master (f58ab18a904) stops in about a second
REL_18_STABLE + 12f866cd5f7 stops; 058 passes, and fails
without the fix
REL_19_STABLE + 12f866cd5f7 stops; 058 passes
12f866cd5f7 applies cleanly to REL_18_STABLE. On REL_19_STABLE the only
conflict is the test list in src/test/recovery/meson.build, where 059
is already listed.
I understand wanting it to settle on HEAD first. I mainly wanted to say
that it happens outside of tests: anything that stops the server right
after a backend crash can land in that window, and on a busy machine
the window is a few seconds long. Since 19 is not released yet, it may
be worth considering for it.
Thanks for the fix.
Regards,
Muzzammil Sarwar
On Wed, Oct 7, 2026 at 5:07 AM Michael Paquier <michael(at)paquier(dot)xyz> wrote:
> On Thu, Oct 01, 2026 at 10:54:55AM +0530, Ayush Tiwari wrote:
> > Thanks for this. The patch looks good to me.
> > I did multiple A/B testing and it works well.
>
> I've been going back and forth a bit in the postmaster code to try
> more if FatalError can be made again non-re-entrant for
> HandleFatalError() with an assertion. My first try was the
> introduction of a new flag to track if SIGQUIT was sent, but I was not
> completely sure that I got the semantics right. A second set of
> thoughts was pointing towards a more complicated scheme for the
> states, which felt not really exciting. At the end, I have just
> reused your solution, and applied it on HEAD. Let's let it brew first
> there. The lack of complaints on the matter (even if we had one back
> in 2023), and the v19 release being very close by don't argue in favor
> of a backpatch, at least it feels strongly so to me.
>
> Now down to the buildfarm to judge the stability of your test. The CI
> has cleared quite a few times.
> --
> Michael
>
| Attachment | Content-Type | Size |
|---|---|---|
| repro_shutdown_hang.sh | application/octet-stream | 1.5 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Peter Geoghegan | 2026-10-10 21:52:47 | Re: index prefetching |
| Previous Message | Zsolt Parragi | 2026-10-10 21:30:29 | Re: Proposal: Conflict log history table for Logical Replication |