REPACK (CONCURRENTLY): do not block the table while waiting for the final lock

From: shihao zhong <zhong950419(at)gmail(dot)com>
To: pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Cc: Antonin Houska <ah(at)cybertec(dot)at>, Álvaro Herrera <alvherre(at)kurilemu(dot)de>
Subject: REPACK (CONCURRENTLY): do not block the table while waiting for the final lock
Date: 2026-09-30 05:09:41
Message-ID: CAGRkXqQTRvmJJBjLmLa0boO1S78nz46s=B3htYMFTkSW-FS18Q@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Subject: REPACK (CONCURRENTLY): do not block the table while waiting for
the final lock

Hi hackers,

REPACK (CONCURRENTLY) requests AccessExclusiveLock at the end to swap
the files. While it is queued for that lock behind a long running
transaction, every new query on the table queues behind it. On master I
held AccessShareLock in one session for 15 seconds, and a single row
SELECT that arrived while REPACK was waiting took 14.7 seconds. The
usual defense, lock_timeout, makes it worse here. With lock_timeout = 3s
the REPACK fails after all the copying is done.

The attached patch makes REPACK stop queueing for that lock. It checks
whether the lock is available. While it is not, it applies the changes
that arrived meanwhile and checks again every 50 ms, the way
lazy_truncate_heap() does. With the patch the same SELECT took 0.9
seconds, and REPACK finished once the long transaction ended.
lock_timeout limits how long it keeps trying, so its meaning does not
change.

The cost is on the REPACK side. A conditional request only succeeds when
nobody holds a lock at that moment, so on a busy table it can take a
while. With 16 pgbench clients doing single row SELECTs on the table
(about 135k tps), REPACK needed 1 to 10 seconds to get the lock, against
0.5 seconds when queueing. After hours of copying I think that is fine.

There is a variant with a lock manager change that queues for
deadlock_timeout at a time and then leaves the queue. It gets the lock
in 0.5 seconds on that busy table, but it is 130 lines in lock.c and
proc.c for one caller. I can post it if there is interest.

One case needs the old behavior. If a session already waits behind the
lock REPACK holds, it may hold a lock on the table itself, and polling
would never end. So when LockHasWaitersRelation() reports a waiter,
REPACK waits for the lock the normal way. A real deadlock is then
reported as before, with REPACK as the victim.

Thanks,
Shihao

Attachment Content-Type Size
v1-0001-Do-not-let-REPACK-CONCURRENTLY-block-the-table-wh.patch application/octet-stream 6.0 KB

Browse pgsql-hackers by date

  From Date Subject
Next Message Nisha Moond 2026-09-30 05:34:49 Re: Proposal: Conflict log history table for Logical Replication
Previous Message Michael Paquier 2026-09-30 05:07:33 Re: Bug in logical decoding with DDL and subtransactions