| From: | Zhijie Hou <houzhijie22(at)gmail(dot)com> |
|---|---|
| To: | vignesh C <vignesh21(at)gmail(dot)com> |
| Cc: | "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com>, shveta malik <shveta(dot)malik(at)gmail(dot)com>, Tomas Vondra <tomas(at)vondra(dot)me>, Dilip Kumar <dilipbalaut(at)gmail(dot)com>, Andrei Lepikhov <lepihov(at)gmail(dot)com>, wenhui qiu <qiuwenhuifx(at)gmail(dot)com>, Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Peter Smith <smithpb2250(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com> |
| Subject: | Re: Parallel Apply |
| Date: | 2026-10-07 10:01:04 |
| Message-ID: | CAFvd2n8KmR0ZB9F18wqwhQLVMn2zeugQsMtKoMwrNBZC-gaVtA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
On Wed, Oct 7, 2026 at 2:48 PM vignesh C <vignesh21(at)gmail(dot)com> wrote:
>
> On Tue, 6 Oct 2026 at 12:07, Hayato Kuroda (Fujitsu)
> <kuroda(dot)hayato(at)fujitsu(dot)com> wrote:
> >
> > Hi Vignesh,
> >
> > > I was planning to review this patch set and noticed the patch does not
> > > apply anymore, could you post a rebased version for reviewing?
> >
> > Thanks for the contact. Here is an updated version.
>
> I ran into a reliability issue in pa_send_data()
> (src/backend/replication/logical/applyparallelworker.c).
> pa_send_data() retries shm_mq_send() for SHM_SEND_TIMEOUT_MS(9)
> seconds when the parallel apply worker is slow to consume its input.
> If it still cannot send the data, it returns false. All its callers in
> worker.c treat this as a hard error and terminate the apply worker:
> ereport(ERROR, ...,
> errmsg("could not send data to the logical replication "
> "parallel apply worker"));
Thanks for reporting.
This is a known behavior that has been there since the introduction of
streaming parallel mode. It was considered complicated to add a new
deadlock detector for this timeout case, so it wasn't deemed worth the effort.
And it would be rare to trigger this in production, and we haven't received any
reports of this kind of timeout yet, so I think this isn't a issue.
Since this is an existing behavior, and the new parallel apply is actually
harder to hit this because we're parallelizing non-streamed transactions which
tend to be small, I think we should consider it independently rather than doing
something in the current patches.
Best Regards,
Zhijie Hou
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Fujii Masao | 2026-10-07 10:18:14 | Re: remote_apply commit hangs when wal_receiver_status_interval = 0 |
| Previous Message | Zhijie Hou | 2026-10-07 09:55:49 | Re: Bug in logical decoding with DDL and subtransactions |