| From: | "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com> |
|---|---|
| To: | 'vignesh C' <vignesh21(at)gmail(dot)com> |
| Cc: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | RE: Incorrect CONTEXT reported for errors from parallel apply worker in logical replication |
| Date: | 2026-10-08 07:53:50 |
| Message-ID: | OS7PR01MB183176B9D2D2B1ECDF12A30FDF5932@OS7PR01MB18317.jpnprd01.prod.outlook.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Dear Vignesh,
> While reviewing another thread at [1], I found an issue where an error
> relayed from a parallel apply worker can get an unrelated CONTEXT line
> from the leader process.
> ProcessParallelApplyMessage() in applyparallelworker.c sets
> error_context_stack to the leader's apply_error_context_stack before
> calling ereport(ERROR). As a result, when errfinish() processes the
> error, the leader's error context callback runs again and adds the
> leader's current replication context. This can be incorrect if the
> leader is processing a different transaction when it receives the
> error from the parallel worker.
Good catch and agreed your analysis.
I checked other examples in core, and parallel query seems to handle correctly.
It has an attribute ParallelContext::error_context_stack which preserves an error
context at that time, see CreateParallelContext(). When the parallel worker raises
an ERROR, the leader backend can consume its message and raise it again, and at
that time the error context is restored from the ParallelContext.
I think the easiest fix is to use apply_error_context_stack for preserving the
error context before entering the loop, attached patch implements the idea.
IIUC no contexts can be stacked at the begining of the worker, i.e.,
apply_error_context_stack can be NULL in any case. So the global variable can be
removed if we clean up more aggressively. I also felt the name can be changed,
but it's retained for now because the variable has been exposed.
BTW, I also checked the REPACK CONCURRENTLY, but it seemed to have the same issue
as the parallel apply. I locally found a reproducer and wrote a fix patch,
I can share here or another place.
Best regards,
Hayato Kuroda
FUJITSU LIMITED
| Attachment | Content-Type | Size |
|---|---|---|
| 0001-Test-to-reproduce-parallel-apply-error-context-issue.patch | application/octet-stream | 9.4 KB |
| 0002-Preserve-an-error-context-before-entering-an-apply-l.patch | application/octet-stream | 1.6 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Andrey Rachitskiy | 2026-10-08 08:02:55 | Re: Fix PGTYPESdate_fmt_asc overflow when a year does not fit "yyyy" |
| Previous Message | David Geier | 2026-10-08 07:52:27 | Re: Improving scalability of Parallel Bitmap Heap/Index Scan |