| From: | vignesh C <vignesh21(at)gmail(dot)com> |
|---|---|
| To: | "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com> |
| Cc: | shveta malik <shveta(dot)malik(at)gmail(dot)com>, Tomas Vondra <tomas(at)vondra(dot)me>, Dilip Kumar <dilipbalaut(at)gmail(dot)com>, Andrei Lepikhov <lepihov(at)gmail(dot)com>, wenhui qiu <qiuwenhuifx(at)gmail(dot)com>, Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Peter Smith <smithpb2250(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com> |
| Subject: | Re: Parallel Apply |
| Date: | 2026-10-07 06:47:26 |
| Message-ID: | CALDaNm11d=kVUbjg8Et9zPYN88EFZi27awb77ki22i54K3n0-Q@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Tue, 6 Oct 2026 at 12:07, Hayato Kuroda (Fujitsu)
<kuroda(dot)hayato(at)fujitsu(dot)com> wrote:
>
> Hi Vignesh,
>
> > I was planning to review this patch set and noticed the patch does not
> > apply anymore, could you post a rebased version for reviewing?
>
> Thanks for the contact. Here is an updated version.
I ran into a reliability issue in pa_send_data()
(src/backend/replication/logical/applyparallelworker.c).
pa_send_data() retries shm_mq_send() for SHM_SEND_TIMEOUT_MS(9)
seconds when the parallel apply worker is slow to consume its input.
If it still cannot send the data, it returns false. All its callers in
worker.c treat this as a hard error and terminate the apply worker:
ereport(ERROR, ...,
errmsg("could not send data to the logical replication "
"parallel apply worker"));
This can happen during a normal temporary delay in consuming the data.
For example, if the leader has more than 16 MB of changes to send
before the parallel worker catches up, pa_send_data() can time out and
bring down the apply worker. With disable_on_error, this can disable
the entire subscription due to a temporary delay in processing the
data. Without it, the apply worker is restarted and can hit the same
failure again. I felt that treating a slow parallel apply worker as an
error condition does not seem right, since this is an expected
situation during normal processing.
I have attached a patch with a test that reproduces the issue.
Regards,
Vignesh
| Attachment | Content-Type | Size |
|---|---|---|
| 0001-Test-to-show-a-slow-worker-brings-down-logical-repli.patch | application/octet-stream | 8.3 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Antonin Houska | 2026-10-07 07:34:58 | Re: REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes |
| Previous Message | Nisha Moond | 2026-10-07 06:31:34 | Re: Introduce XID age based replication slot invalidation |