| From: | vignesh C <vignesh21(at)gmail(dot)com> |
|---|---|
| To: | "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com> |
| Cc: | shveta malik <shveta(dot)malik(at)gmail(dot)com>, Tomas Vondra <tomas(at)vondra(dot)me>, Dilip Kumar <dilipbalaut(at)gmail(dot)com>, Andrei Lepikhov <lepihov(at)gmail(dot)com>, wenhui qiu <qiuwenhuifx(at)gmail(dot)com>, Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Peter Smith <smithpb2250(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com> |
| Subject: | Re: Parallel Apply |
| Date: | 2026-10-06 10:57:47 |
| Message-ID: | CALDaNm3iSKqzpOrO1ULzdhfw2=2wiAq3S21MbBTFnydoap3YUw@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Tue, 6 Oct 2026 at 12:07, Hayato Kuroda (Fujitsu)
<kuroda(dot)hayato(at)fujitsu(dot)com> wrote:
>
> Thanks for the contact. Here is an updated version.
Thanks for the updated version Kuroda-san.
I found an issue while reviewing the patch where an error relayed from
a parallel apply worker can get an unrelated CONTEXT line from the
leader process. ProcessParallelApplyMessage() in applyparallelworker.c
sets error_context_stack to the leader's apply_error_context_stack
before calling ereport(ERROR). As a result, when errfinish() processes
the new error, the leader's error context callback runs again and adds
the leader's current replication context.
This can be incorrect because the leader may be processing a
completely different transaction when it receives the error from the
parallel worker. For example, in my test:
- tab_a is processed by a parallel worker and fails due to a trigger.
- The leader is processing an unrelated insert into tab_b.
- The parallel worker's error is received while the leader is processing tab_b.
- The final error incorrectly contains context for both tab_a and tab_b.
The resulting error looks like:
ERROR: logical replication parallel apply worker exited due to error
CONTEXT: PL/pgSQL function public.tab_a_boom_fn() line 3 at RAISE
processing remote data ... relation "public.tab_a" ...
logical replication parallel apply worker
processing remote data ... relation "public.tab_b" ...
The tab_b context is unrelated to the actual failure and comes from
the leader's current activity. I have attached a regression test that
reproduces the issue. The test verifies that the relayed error
contains the context for tab_a but must not include the leader's
unrelated tab_b context. It currently fails on the second assertion.
I think the error handling in ProcessParallelApplyMessage() needs to
avoid running the leader's current error-context callbacks when
re-throwing the worker's error.
Thoughts?
Regards,
Vignesh
| Attachment | Content-Type | Size |
|---|---|---|
| 0001-Test-to-reproduce-parallel-apply-error-context-issue.patch | application/octet-stream | 8.3 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Vaibhav Dalvi | 2026-10-06 11:02:33 | Re: gist_trgm_ops '=' operator: planner picks it over btree, ~300x slower |
| Previous Message | Antonin Houska | 2026-10-06 10:44:12 | Re: REPACK (CONCURRENTLY) might keep dropped-column data |