| From: | Andrew Dunstan <andrew(at)dunslane(dot)net> |
|---|---|
| To: | Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org> |
| Subject: | Re: Concurrent DROP TABLESPACE can miss a shared dependency |
| Date: | 2026-09-07 17:58:57 |
| Message-ID: | 0ec2d585-a5c2-4f43-aa27-749687a262cd@dunslane.net |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On 2026-08-30 Su 4:05 PM, Ayush Tiwari wrote:
> Hi,
>
> I found a race between DROP TABLESPACE and a concurrent command that adds a
> shared dependency on the tablespace. Both commands can succeed, leaving an
> object whose pg_class.reltablespace and pg_shdepend entries refer to a
> tablespace that no longer exists.
>
> One way to reproduce this is:
>
> 1. One session updates the pg_tablespace row and keeps the transaction open.
> 2. A second session runs DROP TABLESPACE. It finds no dependencies, then
> waits while deleting the pg_tablespace tuple.
> 3. A third session creates a partitioned table in the tablespace and commits.
> 4. The first session aborts, allowing DROP TABLESPACE to finish using the
> result of its earlier dependency check.
>
> ISTM the race is possible because shdepAddDependency() takes an
> AccessShareLock on the referenced shared object and rechecks that it still
> exists, but DropTableSpace() calls checkSharedDependencies() without first
> taking the corresponding conflicting lock. DropRole() appears to follow
> that protocol already.
>
> For a fix, my first thought was to have DROP TABLESPACE take an
> AccessExclusiveLock before checking its shared dependencies. However,
> doing only that seems to introduce a lock-order problem with commands that
> update pg_tablespace without first locking the tablespace. Such a command
> can hold the catalog tuple while DROP TABLESPACE holds the object lock, and
> then try to acquire the object lock itself when the transaction records a
> tablespace dependency.
>
> I went through the paths that update pg_tablespace, and I think ALTER
> TABLESPACE RENAME/SET, DROP OWNED, and REASSIGN OWNED need to acquire an
> AccessShareLock before updating the catalog tuple. Direct GRANT, REVOKE,
> and ALTER OWNER appear to acquire an object lock already. The attached
> 0002 contains those changes, separately from the DROP-side change in 0001.
> (They likely need to be squashed once the patch looks fine).
>
> Does this AccessExclusiveLock/AccessShareLock protocol seem like the right
> way to close the race? Also, have I missed another pg_tablespace update
> path that should participate in the same protocol?
>
>
Hi, I encountered this while working on cleaning up the ddl patches.
0002's deadlock analysis is correct. DROP
TABLESPACE takes the AccessExclusiveLock and then blocks in
CatalogTupleDelete on the uncommitted ALTER's row, while that same
transaction's later CREATE TABLE ... TABLESPACE blocks on the
AccessShareLock DROP already holds -- a real deadlock that 0001
introduces and 0002 closes by taking the AccessShareLock before
touching the catalog tuple in all four paths.
On your open question: I don't think you've missed a path. I tested
GRANT ON TABLESPACE and ALTER TABLESPACE ... OWNER TO the same way and
neither deadlocks -- DROP blocks, then correctly errors once the
transaction commits.
I think this should be applied as a single squashed commit,
(soon so I can rely on it for the fixes I mentioned).
I think it should be backpatched - all live branches have the same problem.
cheers
andrew
--
Andrew Dunstan
EDB: https://www.enterprisedb.com
| From | Date | Subject | |
|---|---|---|---|
| Previous Message | Ilia Evdokimov | 2026-09-07 17:39:55 | Re: Improve Hash/Merge Join estimate accuracy when all predicates are Hash/Merge clauses |