Re: Publication DDL can race with a concurrent UPDATE

From: Nikhil Sontakke <nikhil(at)planetscale(dot)com>
To: vignesh C <vignesh21(at)gmail(dot)com>
Cc: Zhijie Hou <houzhijie22(at)gmail(dot)com>, shveta malik <shveta(dot)malik(at)gmail(dot)com>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Publication DDL can race with a concurrent UPDATE
Date: 2026-10-09 09:46:10
Message-ID: CA+UBoq0Evs3uJ9w+bayPCk+Li2FLUGQW9oOOqu6B4oCMyhV9BA@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi Vignesh,

>
> > The root issue discussed here is that the relcache is stale due to
> insufficient
> > locking, and the standard approach to fixing such issues is to take a
> strong
> > lock during DDL rather than DML. The publication DDL case is special
> since it
> > needs to lock many tables, but I think we should build on that and
> improve
> > gradually rather than immediately strengthening the lock on the DML
> side. My
> > point is we should first put more effort into engineering the DDL side
> to solve
> > the issue, and if that turns out not to work, then we might consider the
> > DML-side approach as a last resort.
> >
> > Of course, testing is also worthwhile, but I think even if the results
> show
> > nothing, we still need broad consensus if we want to choose that
> approach.
>
> Thanks, let's proceed with this approach.
> I found a few improvements that could be done:
>

Moving toward taking a stronger lock during DDL without penalizing DML
sounds attractive.

I tried v3 locally and found a lock-upgrade deadlock involving
pg_publication, even when the publications cover disjoint schemas.

CreatePublication() already holds RowExclusiveLock on pg_publication when
LockTablesWithoutReplicaIdentity() requests ShareLock on the same catalog.
Although two ShareLock requests are compatible, each conflicts with the
other transaction's existing RowExclusiveLock.

Here is a reproducer (the spec file is attached with this email as well):

-- Setup
CREATE SCHEMA review_a;
CREATE TABLE review_a.t (id int PRIMARY KEY, val int);

CREATE SCHEMA review_b;
CREATE TABLE review_b.t (id int PRIMARY KEY, val int);

CREATE TABLE review_outside (id int, val int);
INSERT INTO review_outside VALUES (1, 1);

-- Session 1
BEGIN;
UPDATE review_outside SET val = 2;

Although review_outside is not published, its descriptor build acquires
RowExclusiveLock on pg_publication because it lacks replica identity.
That lock is retained until transaction end.

-- Session 2
CREATE PUBLICATION review_pb FOR TABLES IN SCHEMA review_b;

Session 2 blocks requesting ShareLock on pg_publication while also holding
the RowExclusiveLock acquired by CreatePublication().

-- Session 1
CREATE PUBLICATION review_pa FOR TABLES IN SCHEMA review_a;

This deadlocks: Session 1 requests ShareLock while Session 2 holds
RowExclusiveLock, and Session 2 is waiting for Session 1's RowExclusiveLock.

The server reports both sessions waiting for ShareLock on relation
(pg_publication), each blocked by the other. The same sequence completes
without blocking on the unpatched build.

This cycle involves only the publication catalog locks, so sorting the
user-table locks does not address it.

A dedicated synchronization lock would avoid colliding with ordinary catalog
write locks and remove this additional lock-upgrade cycle. However, that would
not address the table/interlock deadlocks already acknowledged in the function
comment.

Regards,
Nikhil
---
Nikhil Sontakke
PlanetScale Postgres Core Team

Attachment Content-Type Size
review-catalog-upgrade.spec application/octet-stream 899 bytes

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Pavel Borisov 2026-10-09 09:58:00 Re: Looking for a good first patch to author
Previous Message shveta malik 2026-10-09 09:43:30 Re: Proposal: Conflict log history table for Logical Replication