REPACK (CONCURRENTLY) might keep dropped-column data

From: Radim Marek <radim(at)boringsql(dot)com>
To: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Cc: "ah(at)cybertec(dot)at" <ah(at)cybertec(dot)at>
Subject: REPACK (CONCURRENTLY) might keep dropped-column data
Date: 2026-09-30 06:27:08
Message-ID: CAJgoLkK2UBzB1J9buCsSUbjf7bOqz-o_0=CeiTruU789BhBw-Q@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hello,

last night I found my small issue as one of REPACK (CONCURRENTLY) testing.
It's similar to an old problem with pg_squeeze reported to Antonin some
time ago.

I played with the idea how it might cope under the logical replication (on
subscriber) and given the previous experience with dropped columns, I
managed to hit scenario where it leaves old data behind.

While the initial copy removes the dropped column data as expected, any
changes that don't follow regular UPDATE path seems to retain the old value.

The table sizes shows the problem nicely

before 2424 kB
REPACK 224 kB
CONCURRENTLY, no replicated updates 288 kB
CONCURRENTLY, rows updated 2616 kB

The local apply worker seems to build the data from the original tuple,
without setting dropped value to NULL.

How to replicate:
1. On publisher create table demo(id int primary key, a text)
2. On subscriber set the table demo(id int primary key, a text, b text)
3. Populate values on publishers, set random data in b on subscriber
4. Drop the column 'b' on subscriber
5. Keep updating data on publisher
6. run REPACK (CONCURRENTLY) on subscriber table

Hope this helps. I will try to look more into the source of the problem
once I have time (if needed).

Radim

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Narayanan Venkateswaran 2026-09-30 06:57:38 Re: Proposal: Conflict log history table for Logical Replication
Previous Message Shubhra Jain 2026-09-30 06:24:21 Re: docs: Include database collation check on SQL from alter_collation.sgml