Re: Direct TOAST v2, faster, smaller and no migration needed

From: Michael Paquier <michael(at)paquier(dot)xyz>
To: Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>
Cc: Hannu Krosing <hannuk(at)google(dot)com>, pgsql-hackers <pgsql-hackers(at)postgresql(dot)org>, Dilip Kumar <dilipkumarb(at)google(dot)com>, Yugo Nagata <nagata(at)sraoss(dot)co(dot)jp>
Subject: Re: Direct TOAST v2, faster, smaller and no migration needed
Date: 2026-09-30 02:56:23
Message-ID: arx6V8dC3EhZrDBq@paquier.xyz
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, Sep 29, 2026 at 02:18:41PM +0200, Matthias van de Meent wrote:
> But the independent repack-ability of toast tables is a very valuable
> feature. Main tables have indexes that may be very expensive to
> rebuild. Yes, CONCURRENTLY could help, but it'll still take a
> possibly huge amount of resources, in time, CPU, memory, disk, IO.
> Repacking the toast table separately solves bloat issues in the toast
> table, and does not require the main table to be rebuilt, avoiding the
> related resource consumptions of index builds etc.

Note that TOAST-only nightly VACUUMs can be quite common for some
workloads, a reason why VACUUM (PROCESS_TOAST) exists so as we can
clean up only the underlying TOAST table. I do not see enough
arguments in favor of dropping this kind of properties that many
people rely on, based on only the claims posted on this thread.
That's just too much risk compared to the potential reward.

> No, we don't recheck the valueid in the heaptuple. However, we do
> check that the chunk counter matches the expected values, and that's
> something that won't be possible in the direct toast design -- there
> is nothing in the toast pointer nor the toast tuple that links it to
> its own identity, especially not outside the scope of MVCC snapshots.

Yes, we've relied on the concept of TOAST tuple "identity" heavily for
20-ish years.

You could introduce a new behavior based on tids as an opt-in, I
guess, leaving the default 4-bytes be, giving access to tid-based
access tables for newly-created TOAST tables.

Anyway, claiming that this proposal should enforce a new behavior
across the board for everything is not something I can vouch for. I
get the cheap rewrite argument with the addition of the partial index,
but no solution has magical properties; there are always two sides of
the coin, and we have to think about both sides. Using tids for
lookups has benefits in some cases, clearly, just claiming that it's
always beneficial for every kind of workload is something I find very
hard to believe. The decision to make impossible TOAST-based VACUUMs
is one downside that I do not and will not underestimate, for one. I
cannot speak for other committers, but at least that's one committer
opinion. My opinion can of course be overruled, it happens a lot.

Another concept Greg Burd has been mentioning to me is overflow pages,
used by some other systems: keep all the chunks in heap, with tid
redirections. That kind of mimics some of the builds out there where
folks compile the code with larger base page sizes to avoid the TOAST
lookups, for workloads where the blobs are large enough to be in TOAST
but could fit on a page if made larger. It has its own set of
advantages and disadvantages as well, like any other approach. Back
to the two sides of the coin.
--
Michael

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Michael Paquier 2026-09-30 03:05:09 Re: Bug in logical decoding with DDL and subtransactions
Previous Message Richard Guo 2026-09-30 02:47:57 Re: Assert failure in get_baserel_parampathinfo with lateral UNION ALL