Re: Set 1s WaitLatch timeout if standby limit has expired in ResolveRecoveryConflictWithBufferPin

From: Andrey Borodin <x4mmm(at)yandex-team(dot)ru>
To: shihao zhong <zhong950419(at)gmail(dot)com>
Cc: Dmytro Astapov <dastapov(at)gmail(dot)com>, Anthony Hsu <erwaman(at)gmail(dot)com>, Álvaro Herrera <alvherre(at)kurilemu(dot)de>, pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: Set 1s WaitLatch timeout if standby limit has expired in ResolveRecoveryConflictWithBufferPin
Date: 2026-10-08 13:14:16
Message-ID: A65E3EDC-92B5-4C69-818F-0089F7AEB428@yandex-team.ru
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On 2 Oct 2026, Shihao Zhong wrote:
> It calls WaitLatch directly in standby.c, so proc.c and proc.h are
> not touched. That should be easier to backpatch.

We hit what looks like this bug today on a production standby running
our PostgreSQL 17.11 build. Both max_standby_*_delay settings were 30s,
but replay lag had already reached 5h41m. During investigation, startup
was waiting in BufferPin at an unchanged LSN. Following Dmytro's tip,
pg_log_backend_memory_contexts() woke it up. A buffer-pin recovery
conflict was logged, and the standby caught up without a restart.

I tested v3 on REL_17_STABLE, adapting the signal and wait-event enum
names. I used two cursors and SIGSTOP/SIGCONT on the first reader so
that the second pin was acquired after the cancellation broadcast.
Without the fix, the second reader survived cancellation of the first,
and replay needed the diagnostic wakeup. With v3, the second reader
was canceled in about a second, without an external wakeup.

I'll reopen the CF entry [0]. Please let me know if that's unnecessary
because this bug is already tracked elsewhere.

Thank you!

Best regards, Andrey Borodin.

[0] https://commitfest.postgresql.org/patch/6445/

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Álvaro Herrera 2026-10-08 13:29:00 Re: REPACK hits assertion failure on postmaster death exit
Previous Message Hayato Kuroda (Fujitsu) 2026-10-08 13:04:03 RE: Incorrect CONTEXT reported for errors from parallel apply worker in logical replication