| From: | Radim Marek <radim(at)boringsql(dot)com> |
|---|---|
| To: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Cc: | "ah(at)cybertec(dot)at" <ah(at)cybertec(dot)at> |
| Subject: | REPACK (CONCURRENTLY) might keep dropped-column data |
| Date: | 2026-09-30 06:27:08 |
| Message-ID: | CAJgoLkK2UBzB1J9buCsSUbjf7bOqz-o_0=CeiTruU789BhBw-Q@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hello,
last night I found my small issue as one of REPACK (CONCURRENTLY) testing.
It's similar to an old problem with pg_squeeze reported to Antonin some
time ago.
I played with the idea how it might cope under the logical replication (on
subscriber) and given the previous experience with dropped columns, I
managed to hit scenario where it leaves old data behind.
While the initial copy removes the dropped column data as expected, any
changes that don't follow regular UPDATE path seems to retain the old value.
The table sizes shows the problem nicely
before 2424 kB
REPACK 224 kB
CONCURRENTLY, no replicated updates 288 kB
CONCURRENTLY, rows updated 2616 kB
The local apply worker seems to build the data from the original tuple,
without setting dropped value to NULL.
How to replicate:
1. On publisher create table demo(id int primary key, a text)
2. On subscriber set the table demo(id int primary key, a text, b text)
3. Populate values on publishers, set random data in b on subscriber
4. Drop the column 'b' on subscriber
5. Keep updating data on publisher
6. run REPACK (CONCURRENTLY) on subscriber table
Hope this helps. I will try to look more into the source of the problem
once I have time (if needed).
Radim
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Narayanan Venkateswaran | 2026-09-30 06:57:38 | Re: Proposal: Conflict log history table for Logical Replication |
| Previous Message | Shubhra Jain | 2026-09-30 06:24:21 | Re: docs: Include database collation check on SQL from alter_collation.sgml |