Re: REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes

From: Alvaro Herrera <alvherre(at)kurilemu(dot)de>
To: Radim Marek <radim(at)boringsql(dot)com>
Cc: shihao zhong <zhong950419(at)gmail(dot)com>, Nathan Bossart <nathandbossart(at)gmail(dot)com>, Antonin Houska <ah(at)cybertec(dot)at>, pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes
Date: 2026-10-08 12:34:50
Message-ID: aseKlWSW6ScgT2sX@alvherre.pgsql
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

On 2026-Oct-08, Radim Marek wrote:

> As a medium-heavy repack user over the last decade, I'd say this limit is
> acceptable for now.

Agreed -- I don't think we need any localized hacking for this,
considering that we have a patch in the queue for 20 that will
completely remove the need for this. People that can't tolerate 105
million rows concurrently updated during REPACK (CONCURRENTLY) can do
whatever they've been doing so far to cope with such a scenario. Maybe
they run pg_repack, or maybe they run VACUUM FULL, I don't really care.
(What I actually think is that such people don't actually exist.)

IMO this is not a regression and we don't *need* to cope with it.

> Online repacking of a large table is rarely done
> without planning, and what decides whether you hit the limit is storage
> throughput (how long the run takes) times the rate of change on the table.
>
> I put numbers on that in [1]. The budget is roughly 29k changed rows/s
> divided by the run's duration in hours.
>
> [1] https://boringsql.com/posts/repack-concurrently-costs/#when-the-run-is-long

Right, I found this article a couple of days ago while preparing the doc
patch. It's a great read, thanks for putting it together.

> Generous for a 250 GB table, tight (circa 1.5k rows/s) for 5 TB on a
> modest cloud VM.

I would argue that nobody would run a database containing a single 5 TB
table on "a modest cloud VM".

> My point in starting this thread was to make sure people affected on v19
> know the limit exists, and the fact it might attract new audience that
> previously was not using repacking heavily. The worst case is losing all
> that work at the very last step, quite possibly while you're already
> dealing with something close to an incident. If 19 ships with it, I think
> it deserves a sentence in the REPACK docs, next to the disk space note.

Yeah, the note has been added, you can read it here:
https://www.postgresql.org/docs/devel/sql-repack.html
If you have further feedback, I'm glad to do further changes.

> PS: I've been using pg_squeeze for exactly those scenarios as pg_repack has
> some limitations (ironically one of them was FOR ALL TABLES without
> exception when logical replication is already used). For me
> personally REPACK (CONCURRENTLY) as now planned for v19 is improvement in
> all operational aspects.

That's great to hear, thank you very much.

--
Álvaro Herrera 48°01'N 7°57'E — https://www.EnterpriseDB.com/
"Find a bug in a program, and fix it, and the program will work today.
Show the program how to find and fix a bug, and the program
will work forever" (Oliver Silfridge)

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Ashutosh Bapat 2026-10-08 12:52:48 Re: Fix a wal_debug crash with the new shmem allocation API
Previous Message Manu 2026-10-08 11:48:02 Re: [PATCH] Extensible ReadyForQuery wire protocol message and C hook, for connection pools and WAIT FOR LSN