Re: Parallel Apply

From: vignesh C <vignesh21(at)gmail(dot)com>
To: "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com>
Cc: shveta malik <shveta(dot)malik(at)gmail(dot)com>, Tomas Vondra <tomas(at)vondra(dot)me>, Dilip Kumar <dilipbalaut(at)gmail(dot)com>, Andrei Lepikhov <lepihov(at)gmail(dot)com>, wenhui qiu <qiuwenhuifx(at)gmail(dot)com>, Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Peter Smith <smithpb2250(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com>
Subject: Re: Parallel Apply
Date: 2026-10-06 10:57:47
Message-ID: CALDaNm3iSKqzpOrO1ULzdhfw2=2wiAq3S21MbBTFnydoap3YUw@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, 6 Oct 2026 at 12:07, Hayato Kuroda (Fujitsu)
<kuroda(dot)hayato(at)fujitsu(dot)com> wrote:
>
> Thanks for the contact. Here is an updated version.

Thanks for the updated version Kuroda-san.
I found an issue while reviewing the patch where an error relayed from
a parallel apply worker can get an unrelated CONTEXT line from the
leader process. ProcessParallelApplyMessage() in applyparallelworker.c
sets error_context_stack to the leader's apply_error_context_stack
before calling ereport(ERROR). As a result, when errfinish() processes
the new error, the leader's error context callback runs again and adds
the leader's current replication context.

This can be incorrect because the leader may be processing a
completely different transaction when it receives the error from the
parallel worker. For example, in my test:
- tab_a is processed by a parallel worker and fails due to a trigger.
- The leader is processing an unrelated insert into tab_b.
- The parallel worker's error is received while the leader is processing tab_b.
- The final error incorrectly contains context for both tab_a and tab_b.

The resulting error looks like:
ERROR: logical replication parallel apply worker exited due to error
CONTEXT: PL/pgSQL function public.tab_a_boom_fn() line 3 at RAISE
processing remote data ... relation "public.tab_a" ...
logical replication parallel apply worker
processing remote data ... relation "public.tab_b" ...

The tab_b context is unrelated to the actual failure and comes from
the leader's current activity. I have attached a regression test that
reproduces the issue. The test verifies that the relayed error
contains the context for tab_a but must not include the leader's
unrelated tab_b context. It currently fails on the second assertion.

I think the error handling in ProcessParallelApplyMessage() needs to
avoid running the leader's current error-context callbacks when
re-throwing the worker's error.
Thoughts?

Regards,
Vignesh

Attachment Content-Type Size
0001-Test-to-reproduce-parallel-apply-error-context-issue.patch application/octet-stream 8.3 KB

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Vaibhav Dalvi 2026-10-06 11:02:33 Re: gist_trgm_ops '=' operator: planner picks it over btree, ~300x slower
Previous Message Antonin Houska 2026-10-06 10:44:12 Re: REPACK (CONCURRENTLY) might keep dropped-column data