| From: | Nikhil Sontakke <nikhil(at)planetscale(dot)com> |
|---|---|
| To: | vignesh C <vignesh21(at)gmail(dot)com> |
| Cc: | Zhijie Hou <houzhijie22(at)gmail(dot)com>, shveta malik <shveta(dot)malik(at)gmail(dot)com>, "Zhijie Hou (Fujitsu)" <houzj(dot)fnst(at)fujitsu(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Publication DDL can race with a concurrent UPDATE |
| Date: | 2026-10-09 09:46:10 |
| Message-ID: | CA+UBoq0Evs3uJ9w+bayPCk+Li2FLUGQW9oOOqu6B4oCMyhV9BA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi Vignesh,
>
> > The root issue discussed here is that the relcache is stale due to
> insufficient
> > locking, and the standard approach to fixing such issues is to take a
> strong
> > lock during DDL rather than DML. The publication DDL case is special
> since it
> > needs to lock many tables, but I think we should build on that and
> improve
> > gradually rather than immediately strengthening the lock on the DML
> side. My
> > point is we should first put more effort into engineering the DDL side
> to solve
> > the issue, and if that turns out not to work, then we might consider the
> > DML-side approach as a last resort.
> >
> > Of course, testing is also worthwhile, but I think even if the results
> show
> > nothing, we still need broad consensus if we want to choose that
> approach.
>
> Thanks, let's proceed with this approach.
> I found a few improvements that could be done:
>
Moving toward taking a stronger lock during DDL without penalizing DML
sounds attractive.
I tried v3 locally and found a lock-upgrade deadlock involving
pg_publication, even when the publications cover disjoint schemas.
CreatePublication() already holds RowExclusiveLock on pg_publication when
LockTablesWithoutReplicaIdentity() requests ShareLock on the same catalog.
Although two ShareLock requests are compatible, each conflicts with the
other transaction's existing RowExclusiveLock.
Here is a reproducer (the spec file is attached with this email as well):
-- Setup
CREATE SCHEMA review_a;
CREATE TABLE review_a.t (id int PRIMARY KEY, val int);
CREATE SCHEMA review_b;
CREATE TABLE review_b.t (id int PRIMARY KEY, val int);
CREATE TABLE review_outside (id int, val int);
INSERT INTO review_outside VALUES (1, 1);
-- Session 1
BEGIN;
UPDATE review_outside SET val = 2;
Although review_outside is not published, its descriptor build acquires
RowExclusiveLock on pg_publication because it lacks replica identity.
That lock is retained until transaction end.
-- Session 2
CREATE PUBLICATION review_pb FOR TABLES IN SCHEMA review_b;
Session 2 blocks requesting ShareLock on pg_publication while also holding
the RowExclusiveLock acquired by CreatePublication().
-- Session 1
CREATE PUBLICATION review_pa FOR TABLES IN SCHEMA review_a;
This deadlocks: Session 1 requests ShareLock while Session 2 holds
RowExclusiveLock, and Session 2 is waiting for Session 1's RowExclusiveLock.
The server reports both sessions waiting for ShareLock on relation
(pg_publication), each blocked by the other. The same sequence completes
without blocking on the unpatched build.
This cycle involves only the publication catalog locks, so sorting the
user-table locks does not address it.
A dedicated synchronization lock would avoid colliding with ordinary catalog
write locks and remove this additional lock-upgrade cycle. However, that would
not address the table/interlock deadlocks already acknowledged in the function
comment.
Regards,
Nikhil
---
Nikhil Sontakke
PlanetScale Postgres Core Team
| Attachment | Content-Type | Size |
|---|---|---|
| review-catalog-upgrade.spec | application/octet-stream | 899 bytes |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Pavel Borisov | 2026-10-09 09:58:00 | Re: Looking for a good first patch to author |
| Previous Message | shveta malik | 2026-10-09 09:43:30 | Re: Proposal: Conflict log history table for Logical Replication |