| From: | Nisha Moond <nisha(dot)moond412(at)gmail(dot)com> |
|---|---|
| To: | Bharath Rupireddy <bharath(dot)rupireddyforpostgres(at)gmail(dot)com> |
| Cc: | Amit Kapila <amit(dot)kapila16(at)gmail(dot)com>, Masahiko Sawada <sawada(dot)mshk(at)gmail(dot)com>, Srinath Reddy Sadipiralla <srinath2133(at)gmail(dot)com>, SATYANARAYANA NARLAPURAM <satyanarlapuram(at)gmail(dot)com>, "Hayato Kuroda (Fujitsu)" <kuroda(dot)hayato(at)fujitsu(dot)com>, John H <johnhyvr(at)gmail(dot)com>, PostgreSQL-development <pgsql-hackers(at)postgresql(dot)org> |
| Subject: | Re: Introduce XID age based replication slot invalidation |
| Date: | 2026-10-07 06:31:34 |
| Message-ID: | CABdArM6RtWPFHiZZZV71Pvu_b7_yApyT7YGCRgCubm-95Sc74A@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Tue, Oct 6, 2026 at 5:18 PM Nisha Moond <nisha(dot)moond412(at)gmail(dot)com> wrote:
>
> On Wed, Sep 23, 2026 at 11:44 PM Bharath Rupireddy
> <bharath(dot)rupireddyforpostgres(at)gmail(dot)com> wrote:
> >
> >
> > Thanks for reading this far. Please have a look at the attached v15 patches.
> >
>
> Thanks for the updated patches.
> - v15-0001 fails to apply on HEAD and needs to be rebased after commit 46024c5.
>
Found one issue: a long-running transaction on the publisher can get
every healthy logical slot invalidated, without vacuum gaining
anything.
A logical slot's catalog_xmin can't advance past the oldest running
XID (oldestRunningXid in xl_running_xacts, see
SnapBuildProcessRunningXacts()). So a old txn that has an XID pins the
catalog_xmin of all logical slots, even when their subscribers are
active and caught up. Once that txn is older than max_slot_xid_age,
the checkpoint invalidates all of these slots, terminates their
walsenders, and the subscriptions break.
I reproduced this with max_slot_xid_age = 100. I left an open
transaction in one session on the publisher:
postgres=# BEGIN;
postgres=*# SELECT pg_current_xact_id();
pg_current_xact_id
--------------------
674
Then for the test, I created a dead tuple in pg_class (CREATE TABLE
t2; DROP TABLE t2;), which vacuum can't remove because of that
transaction:
postgres=# VACUUM (VERBOSE) pg_class;
tuples: 0 removed, 428 remain, 1 are dead but not yet removable
removable cutoff: 674, which was 3 XIDs old when operation ended
After burning ~150 XIDs to exceed max_slot_xid_age, slot s1 is active
and caught up, but its catalog_xmin is stuck at 674:
slot_name | active | active_pid | catalog_xmin | cxmin_age |
confirmed_flush_lsn | invalidation_reason
-----------+--------+------------+--------------+-----------+---------------------+---------------------
s1 | t | 70432 | 674 | 154 |
0/017E6A90 |
The checkpoint invalidates the slot, but vacuum is still blocked at
the same cutoff:
postgres=# CHECKPOINT;
slot_name | active | catalog_xmin | cxmin_age | confirmed_flush_lsn |
invalidation_reason
-----------+--------+--------------+-----------+---------------------+---------------------
s1 | f | 674 | 154 | 0/017E6A90 | xid_aged
postgres=# VACUUM (VERBOSE) pg_class;
tuples: 0 removed, 428 remain, 1 are dead but not yet removable
removable cutoff: 674, which was 154 XIDs old when operation ended
The dead row is removed only after the long open txn ends, and by then
subscription s1 is already lost.
This is different from wal_removed (max_slot_wal_keep_size). There, a
long running txn that has written WAL can also hold back a slot's
restart_lsn and get it invalidated, but it's the slot, not the txn,
that retains the WAL, so invalidation does free disk space. Here the
txn itself holds the same horizon, so invalidating the slot frees
nothing.
I tested only the checkpoint path, but I think the same applies to the
vacuum path (002) too.
IMO, a slot shouldn't be invalidated for XID age when it isn't what's
holding the horizon back, i.e. invalidate a slot only if its aged
xmin/catalog_xmin precedes the oldest running XID (e.g.
GetOldestActiveTransactionId()).
Thoughts?
--
Thanks,
Nisha
| From | Date | Subject | |
|---|---|---|---|
| Previous Message | David Geier | 2026-10-07 06:15:29 | Re: Improving scalability of Parallel Bitmap Heap/Index Scan |