| From: | shihao zhong <zhong950419(at)gmail(dot)com> |
|---|---|
| To: | pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Cc: | Antonin Houska <ah(at)cybertec(dot)at>, Álvaro Herrera <alvherre(at)kurilemu(dot)de> |
| Subject: | REPACK (CONCURRENTLY): do not block the table while waiting for the final lock |
| Date: | 2026-09-30 05:09:41 |
| Message-ID: | CAGRkXqQTRvmJJBjLmLa0boO1S78nz46s=B3htYMFTkSW-FS18Q@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Subject: REPACK (CONCURRENTLY): do not block the table while waiting for
the final lock
Hi hackers,
REPACK (CONCURRENTLY) requests AccessExclusiveLock at the end to swap
the files. While it is queued for that lock behind a long running
transaction, every new query on the table queues behind it. On master I
held AccessShareLock in one session for 15 seconds, and a single row
SELECT that arrived while REPACK was waiting took 14.7 seconds. The
usual defense, lock_timeout, makes it worse here. With lock_timeout = 3s
the REPACK fails after all the copying is done.
The attached patch makes REPACK stop queueing for that lock. It checks
whether the lock is available. While it is not, it applies the changes
that arrived meanwhile and checks again every 50 ms, the way
lazy_truncate_heap() does. With the patch the same SELECT took 0.9
seconds, and REPACK finished once the long transaction ended.
lock_timeout limits how long it keeps trying, so its meaning does not
change.
The cost is on the REPACK side. A conditional request only succeeds when
nobody holds a lock at that moment, so on a busy table it can take a
while. With 16 pgbench clients doing single row SELECTs on the table
(about 135k tps), REPACK needed 1 to 10 seconds to get the lock, against
0.5 seconds when queueing. After hours of copying I think that is fine.
There is a variant with a lock manager change that queues for
deadlock_timeout at a time and then leaves the queue. It gets the lock
in 0.5 seconds on that busy table, but it is 130 lines in lock.c and
proc.c for one caller. I can post it if there is interest.
One case needs the old behavior. If a session already waits behind the
lock REPACK holds, it may hold a lock on the table itself, and polling
would never end. So when LockHasWaitersRelation() reports a waiter,
REPACK waits for the lock the normal way. A real deadlock is then
reported as before, with REPACK as the victim.
Thanks,
Shihao
| Attachment | Content-Type | Size |
|---|---|---|
| v1-0001-Do-not-let-REPACK-CONCURRENTLY-block-the-table-wh.patch | application/octet-stream | 6.0 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Nisha Moond | 2026-09-30 05:34:49 | Re: Proposal: Conflict log history table for Logical Replication |
| Previous Message | Michael Paquier | 2026-09-30 05:07:33 | Re: Bug in logical decoding with DDL and subtransactions |