Re: [PATCH] Fix vacuum_delay_point happening inside lock

From: Heikki Linnakangas <hlinnaka(at)iki(dot)fi>
To: Andres Freund <andres(at)anarazel(dot)de>, Kevin Rocker <me(at)kevinrocker(dot)com>
Cc: Greg Burd <greg(at)burd(dot)me>, pgsql-hackers(at)lists(dot)postgresql(dot)org, Tom Lane <tgl(at)sss(dot)pgh(dot)pa(dot)us>, Andrey Borodin <x4mmm(at)yandex-team(dot)ru>, Neil Chen <carpenter(dot)nail(dot)cz(at)gmail(dot)com>, rmt(at)lists(dot)postgresql(dot)org
Subject: Re: [PATCH] Fix vacuum_delay_point happening inside lock
Date: 2026-10-05 10:18:04
Message-ID: 13b5bdfa-d6a6-4d8f-90d2-dccb945b9927@iki.fi
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On 28/09/2026 18:24, Andres Freund wrote:
> Hi,
>
> On 2026-09-28 17:07:52 +0200, Kevin Rocker wrote:
>>> The reason this is more than a cancellation-latency issue on v19 is the
>>> buffer content-lock rewrite in fcb9c977aa5 tracks only one lock per
>>> buffer per backend (the single data.lockmode field). A second SHARE
>>> acquire on an already-share-locked buffer is now:
>>>
>>> - a hard Assert/crash in cassert builds (what unicorn shows), and
>>> - in a non-assert build, an asymmetric leak: BufferLockAttempt() adds
>>> a second BM_LOCK_VAL_SHARED to the shared state, data.lockmode
>>> records only one, and release subtracts one -- so that pg_class
>>> buffer is left permanently one shared-locker too high and can never
>>> again be locked exclusive. Any later exclusive waiter (VACUUM) on
>>> that buffer blocks for the life of the cluster.
>>>
>>> Pre-v19 the double SHARE was harmless (the held-lwlocks array could
>>> represent it), which is presumably why the call site survived so long.
>
> I think this was completely broken before 19 too. Acquiring a lock while
> holding the same lock just happened to be undiagnosed. Note that if you ever
> did this with an exclusive lock being involved, you'd just have ended up with
> an uninterruptible endless wait.
>
> It surely was never safe to call ProcessConfigFile(), or sane to sleep, while
> holding an lwlock. And calling vacuum_delay_point() with interrupts held, made
> it not actually properly work, due to not doing the CFI().

+1

I just arrived at this thread from this other thread discussing the same
issue:
https://www.postgresql.org/message-id/19628-c2b17d358181a1ea%40postgresql.org.

I committed and backported the fix to all stable branches. I included a
quick exit in vacuum_delay_point(), if !INTERRUPTS_CAN_BE_PROCESSED().

- Heikki

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Heikki Linnakangas 2026-10-05 10:39:50 Re: [PATCH v1] amcheck: Allow interrupting the child-level rightlink walk
Previous Message Hayato Kuroda (Fujitsu) 2026-10-05 10:13:03 RE: Bug in logical decoding with DDL and subtransactions