| From: | Xuneng Zhou <xunengzhou(at)gmail(dot)com> |
|---|---|
| To: | Fujii Masao <masao(dot)fujii(at)gmail(dot)com> |
| Cc: | Vitaly Davydov <v(dot)davydov(at)postgrespro(dot)ru>, pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Subject: | Re: Deadlock detector fails to activate on a hot standby replica |
| Date: | 2026-10-02 03:00:58 |
| Message-ID: | CABPTF7UyBJ_Vrq0Mp+JKFMkRYU3BZZEg44wb8KARSDs_UGX9Uw@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
On Fri, Jun 5, 2026 at 7:58 PM Fujii Masao <masao(dot)fujii(at)gmail(dot)com> wrote:
>
> On Wed, May 27, 2026 at 12:01 AM Fujii Masao <masao(dot)fujii(at)gmail(dot)com> wrote:
> > + if (got_standby_delay_timeout)
> > + SendRecoveryConflictWithBufferPin(RECOVERY_CONFLICT_BUFFERPIN);
> > + else if (got_standby_deadlock_timeout)
> > + {
> >
> > Shouldn't we break out of the loop when either got_standby_delay_timeout or
> > got_standby_deadlock_timeout becomes true? Otherwise, the loop continues with
> > those flags still set, which could cause SendRecoveryConflictWithBufferPin() to
> > be called unnecessarily in the subsequent cycles.
> >
> >
> > + if (BufferGetRefCount(buffer) <= 1)
> >
> > Should this be "BufferGetRefCount(buffer) == 1" instead? I don't think
> > BufferGetRefCount(buffer) should ever return 0 here. If that's correct,
> > would it make sense to explicitly detect that case, for example:
> >
> > -----------------
> > uint32 refcount = BufferGetRefCount(buffer);
> >
> > Assert(refcount > 0);
> >
> > if (refcount == 0)
> > elog(ERROR, "buffer refcount dropped to zero while waiting for
> > cleanup lock");
> >
> > if (refcount == 1)
> > break;
> > -----------------
>
> I've updated the patch based on these comments.
> Attached is the latest version.
>
> I removed the TAP test from this patch for now. I'll consider adding
> a test for this separately later.
>
> BTW, I'm just wondering whether ResolveRecoveryConflictWithLock()
> might have the same issue. I need to investigate that further.
While working on something closely related, I revisited this thread.
ResolveRecoveryConflictWithLock seems to also suffer from the same
issue per investigation. I'll share some findings later.
--
Regards,
Xuneng Zhou
HighGo Software Co., Ltd.
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Umar Hayat | 2026-10-02 03:05:21 | Re: Exposing the cgroup memory limit to SQL? |
| Previous Message | shihao zhong | 2026-10-02 02:54:18 | Re: REPACK: warn about skipping foreign partitions |