Re: REPACK hits assertion failure on postmaster death exit

From: Manu <manuelreyesbravo(at)gmail(dot)com>
To: pgsql-hackers(at)lists(dot)postgresql(dot)org
Cc: Kirill Reshke <reshkekirill(at)gmail(dot)com>
Subject: Re: REPACK hits assertion failure on postmaster death exit
Date: 2026-10-08 13:00:21
Message-ID: 179146442179.602465.8633628195999085950@gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi Kirill,

> diff also contains some code churn that makes assert failure vanish,
> but I don't yet decide if fixing that way is OK. Looks like this just
> masks errors.

I reproduced it on master (5c74e122ef2) and REL_19_STABLE, and your
doubt looks right: part of the fix masks the problem.

After the postmaster dies the backend re-enters proc_exit more than
once. The FATAL -> LOG changes in parallel.c and repack.c handle two
of those paths cleanly (the "postmaster exited during a parallel
transaction" FATAL, and the repack decoding worker one). The pgstat
half is the one that masks: dlist_init(&pgStatPending) empties the list
but leaves each entry's pending data, so a different error fires
instead --
FATAL: releasing ref with pending data
-- which the TAP test does not catch, because it only greps for "TRAP:"
in the log. With the full fix the test passes, but REPACK's abort
never actually finishes (0 of 6 runs here).

The path that neither half covers is a direct proc_exit(1) from
WaitEventSetWait (WL_EXIT_ON_PM_DEATH, no ereport), reached from the
abort itself:
AbortTransaction -> smgrDoPendingDeletes -> DropRelationsAllBuffers
-> InvalidateBuffer -> WaitIO -> pgaio_io_wait
The abort drops REPACK's new relfilenode and then waits on an AIO that
never completes, because the IO workers died with the postmaster. With
the full fix applied, the abort only runs to the end under
io_method=sync; with the async method it is still cut short, and the
pgstat change just hides the resulting trap.

So the parallel and repack FATALs look real and worth the LOG
treatment, but the remaining problem seems to be the abort waiting on
IO that cannot complete after postmaster death, which the pgstat change
only hides. It may also be worth making the test fail on that pgstat
FATAL, not only on "TRAP:".

I can share the scripts and logs if useful.

Regards,
Manu

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Hayato Kuroda (Fujitsu) 2026-10-08 13:04:03 RE: Incorrect CONTEXT reported for errors from parallel apply worker in logical replication
Previous Message Ashutosh Bapat 2026-10-08 12:52:48 Re: Fix a wal_debug crash with the new shmem allocation API