Re: Parallel Apply

From: Zhijie Hou <houzhijie22(at)gmail(dot)com>
To: vignesh C <vignesh21(at)gmail(dot)com>
Cc: "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com>, shveta malik <shveta(dot)malik(at)gmail(dot)com>, Tomas Vondra <tomas(at)vondra(dot)me>, Dilip Kumar <dilipbalaut(at)gmail(dot)com>, Andrei Lepikhov <lepihov(at)gmail(dot)com>, wenhui qiu <qiuwenhuifx(at)gmail(dot)com>, Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Peter Smith <smithpb2250(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com>
Subject: Re: Parallel Apply
Date: 2026-10-06 11:27:27
Message-ID: CAFvd2n8i0KDMZJzpNoKOanK1T+rnzKXBk6ZrWZQ_PNK6VFFOpQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

On Tue, Oct 6, 2026 at 6:58 PM vignesh C <vignesh21(at)gmail(dot)com> wrote:
>
> On Tue, 6 Oct 2026 at 12:07, Hayato Kuroda (Fujitsu)
> <kuroda(dot)hayato(at)fujitsu(dot)com> wrote:
> >
> > Thanks for the contact. Here is an updated version.
>
> Thanks for the updated version Kuroda-san.
> I found an issue while reviewing the patch where an error relayed from
> a parallel apply worker can get an unrelated CONTEXT line from the
> leader process. ProcessParallelApplyMessage() in applyparallelworker.c
> sets error_context_stack to the leader's apply_error_context_stack
> before calling ereport(ERROR). As a result, when errfinish() processes
> the new error, the leader's error context callback runs again and adds
> the leader's current replication context.
>
> This can be incorrect because the leader may be processing a
> completely different transaction when it receives the error from the
> parallel worker. For example, in my test:
> - tab_a is processed by a parallel worker and fails due to a trigger.
> - The leader is processing an unrelated insert into tab_b.
> - The parallel worker's error is received while the leader is processing tab_b.
> - The final error incorrectly contains context for both tab_a and tab_b.
>
> The resulting error looks like:
> ERROR: logical replication parallel apply worker exited due to error
> CONTEXT: PL/pgSQL function public.tab_a_boom_fn() line 3 at RAISE
> processing remote data ... relation "public.tab_a" ...
> logical replication parallel apply worker
> processing remote data ... relation "public.tab_b" ...
>
> The tab_b context is unrelated to the actual failure and comes from
> the leader's current activity. I have attached a regression test that
> reproduces the issue. The test verifies that the relayed error
> contains the context for tab_a but must not include the leader's
> unrelated tab_b context. It currently fails on the second assertion.
>
> I think the error handling in ProcessParallelApplyMessage() needs to
> avoid running the leader's current error-context callbacks when
> re-throwing the worker's error.

IIUC, this is not an issue introduced by this patch, since the patch doesn't
change any error message handling here. It should have been there since the
streaming parallel mode was introduced. Additionally, streaming parallel mode
follows the same style as parallel query, so it looks possible that we have a
similar issue in that case as well. If so, I think we should handle this
separately in another thread, if it's really worth improving. Could you please
confirm once?

Best Regards,
Zhijie Hou

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Álvaro Rodríguez 2026-10-06 11:37:09 Re: tablecmds: fix bug where index rebuild loses replica identity on partitions
Previous Message Andrei Lepikhov 2026-10-06 11:26:48 Re: hashjoins vs. Bloom filters (yet again)