Re: REPACK (CONCURRENTLY) might keep dropped-column data

From: Radim Marek <radim(at)boringsql(dot)com>
To: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Cc: "ah(at)cybertec(dot)at" <ah(at)cybertec(dot)at>
Subject: Re: REPACK (CONCURRENTLY) might keep dropped-column data
Date: 2026-09-30 07:37:20
Message-ID: CAJgoLk+aoiyMA5batXhKn9pWg_TQHB9er6U11jN1Wa63CtAgPg@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Aha, so on my way to office I started thinking and got more silly ideas,
and now can confirm this is more widespread than logical subscriber use
case.

There's a case for BEFORE UPDATE trigger that might RETURN OLD. I.e. the
cases

-- 1
BEGIN RETURN OLD; END

-- 2
BEGIN NEW := OLD; RETURN NEW; END

-- 3
BEGIN OLD.a := NEW.a; RETURN OLD; END

are all affected by the same problem

CREATE TRIGGER t BEFORE UPDATE ON demo
FOR EACH ROW EXECUTE FUNCTION trg_return_old();

Confirmed by a single run

before 2288 kB
REPACK 72 kB
CONCURRENTLY, no updates 128 kB
CONCURRENTLY, rows updated (trigger) 2360 kB

How to replicate:
1. create table demo(id int primary key, a text, b text), fill it with
random data in b
2. add a BEFORE UPDATE trigger that does RETURN OLD
3. drop column b
4. keep updating rows while running REPACK (CONCURRENTLY) demo

Radim

On Wed, 30 Sept 2026 at 08:27, Radim Marek <radim(at)boringsql(dot)com> wrote:

> Hello,
>
> last night I found my small issue as one of REPACK (CONCURRENTLY) testing.
> It's similar to an old problem with pg_squeeze reported to Antonin some
> time ago.
>
> I played with the idea how it might cope under the logical replication (on
> subscriber) and given the previous experience with dropped columns, I
> managed to hit scenario where it leaves old data behind.
>
> While the initial copy removes the dropped column data as expected, any
> changes that don't follow regular UPDATE path seems to retain the old value.
>
> The table sizes shows the problem nicely
>
> before 2424 kB
> REPACK 224 kB
> CONCURRENTLY, no replicated updates 288 kB
> CONCURRENTLY, rows updated 2616 kB
>
> The local apply worker seems to build the data from the original tuple,
> without setting dropped value to NULL.
>
> How to replicate:
> 1. On publisher create table demo(id int primary key, a text)
> 2. On subscriber set the table demo(id int primary key, a text, b text)
> 3. Populate values on publishers, set random data in b on subscriber
> 4. Drop the column 'b' on subscriber
> 5. Keep updating data on publisher
> 6. run REPACK (CONCURRENTLY) on subscriber table
>
> Hope this helps. I will try to look more into the source of the problem
> once I have time (if needed).
>
> Radim
>

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Andrey Borodin 2026-09-30 07:53:43 Re: Protocol Compression (fourth attempt)
Previous Message Xuneng Zhou 2026-09-30 07:32:03 Re: test: avoid redundant standby catchup in 049_wait_for_lsn