| From: | shihao zhong <zhong950419(at)gmail(dot)com> |
|---|---|
| To: | pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Two fixes for parallel query cleanup: DSM detach order and a stray statement timeout |
| Date: | 2026-09-30 04:12:15 |
| Message-ID: | CAGRkXqQ7d+B+XDCNR0oUmG55EJ421HPj3saWONU2Uc0VKqARWg@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
The REPACK decoding worker copied its shutdown sequence from
DestroyParallelContext(), and we recently found that the dsm_detach()
there ran too early [1]. So I went back to parallel.c and stress tested
cancellation of parallel queries, with Fable assistance. Two patches came
out of it.
0001 moves dsm_detach() in DestroyParallelContext() after
WaitForParallelWorkersToExit(). Today a worker that is still starting up
when the leader detaches fails to map the segment, or to attach to the
DSA or the SharedFileSet, and logs an error. Cancelling a 4 worker
parallel hash join 1200 times on master logged 407 such errors, with the
patch none. The error queues are still detached before the wait, so a
worker stuck in error reporting cannot block the leader. Cancelling
workers blocked on full tuple queues, parallel CREATE INDEX and parallel
VACUUM showed no hang and no leftover temp files.
0002 fixes something worse. A parallel hash join that spilled deletes
its temp files in dsm_detach() at the end of the query, with interrupts
held since 637668fb1d1. If statement_timeout fires in those few
milliseconds, disable_statement_timeout() skips the timer because it is
no longer active (22f6f2c1ccb) and the fired indicator stays set. The
next statement fails on its first CHECK_FOR_INTERRUPTS(). Master,
default settings, a 3M row self join (2 workers, 32 batches, about 560
ms), statement_timeout within 15% of that:
SELECT count(*) FROM big a JOIN big b USING (id); -- succeeds
SELECT 'marker';
ERROR: canceling statement due to statement timeout
The marker failed 5 times in 120 tries. In a BEGIN, INSERT, join, COMMIT
sequence the COMMIT failed 4 times in 242 tries and left the session in
an aborted transaction block. Tom described this case in [2]. 0002
clears the fired indicator when the timer is not active, as the code did
before 22f6f2c1ccb, and both counts drop to zero. Reproducer attached.
The patches are independent. 0002 is a regression in 13 and later, 0001
is only log noise.
[1]
https://postgr.es/m/CAGRkXqQwu8gFrpPXAkFSTOOK0bmf1MBSuLR6wENQH-JwcY4hRQ@mail.gmail.com
[2] https://postgr.es/m/1977913.1767894652@sss.pgh.pa.us
Regards,
Shihao Zhong
| Attachment | Content-Type | Size |
|---|---|---|
| v1-0001-Detach-the-parallel-DSM-segment-after-the-workers.patch | application/octet-stream | 2.2 KB |
| v1-0002-Forget-a-statement-timeout-that-fires-after-the-s.patch | application/octet-stream | 1.5 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Jobin Augustine | 2026-09-30 04:24:22 | Re: test: avoid redundant standby catchup in 049_wait_for_lsn |
| Previous Message | shveta malik | 2026-09-30 03:48:40 | Re: Proposal: Conflict log history table for Logical Replication |