| From: | Radim Marek <radim(at)boringsql(dot)com> |
|---|---|
| To: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Cc: | "ah(at)cybertec(dot)at" <ah(at)cybertec(dot)at> |
| Subject: | Re: REPACK (CONCURRENTLY) might keep dropped-column data |
| Date: | 2026-09-30 07:37:20 |
| Message-ID: | CAJgoLk+aoiyMA5batXhKn9pWg_TQHB9er6U11jN1Wa63CtAgPg@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Aha, so on my way to office I started thinking and got more silly ideas,
and now can confirm this is more widespread than logical subscriber use
case.
There's a case for BEFORE UPDATE trigger that might RETURN OLD. I.e. the
cases
-- 1
BEGIN RETURN OLD; END
-- 2
BEGIN NEW := OLD; RETURN NEW; END
-- 3
BEGIN OLD.a := NEW.a; RETURN OLD; END
are all affected by the same problem
CREATE TRIGGER t BEFORE UPDATE ON demo
FOR EACH ROW EXECUTE FUNCTION trg_return_old();
Confirmed by a single run
before 2288 kB
REPACK 72 kB
CONCURRENTLY, no updates 128 kB
CONCURRENTLY, rows updated (trigger) 2360 kB
How to replicate:
1. create table demo(id int primary key, a text, b text), fill it with
random data in b
2. add a BEFORE UPDATE trigger that does RETURN OLD
3. drop column b
4. keep updating rows while running REPACK (CONCURRENTLY) demo
Radim
On Wed, 30 Sept 2026 at 08:27, Radim Marek <radim(at)boringsql(dot)com> wrote:
> Hello,
>
> last night I found my small issue as one of REPACK (CONCURRENTLY) testing.
> It's similar to an old problem with pg_squeeze reported to Antonin some
> time ago.
>
> I played with the idea how it might cope under the logical replication (on
> subscriber) and given the previous experience with dropped columns, I
> managed to hit scenario where it leaves old data behind.
>
> While the initial copy removes the dropped column data as expected, any
> changes that don't follow regular UPDATE path seems to retain the old value.
>
> The table sizes shows the problem nicely
>
> before 2424 kB
> REPACK 224 kB
> CONCURRENTLY, no replicated updates 288 kB
> CONCURRENTLY, rows updated 2616 kB
>
> The local apply worker seems to build the data from the original tuple,
> without setting dropped value to NULL.
>
> How to replicate:
> 1. On publisher create table demo(id int primary key, a text)
> 2. On subscriber set the table demo(id int primary key, a text, b text)
> 3. Populate values on publishers, set random data in b on subscriber
> 4. Drop the column 'b' on subscriber
> 5. Keep updating data on publisher
> 6. run REPACK (CONCURRENTLY) on subscriber table
>
> Hope this helps. I will try to look more into the source of the problem
> once I have time (if needed).
>
> Radim
>
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Andrey Borodin | 2026-09-30 07:53:43 | Re: Protocol Compression (fourth attempt) |
| Previous Message | Xuneng Zhou | 2026-09-30 07:32:03 | Re: test: avoid redundant standby catchup in 049_wait_for_lsn |