| From: | Hannu Krosing <hannuk(at)google(dot)com> |
|---|---|
| To: | pgsql-hackers <pgsql-hackers(at)postgresql(dot)org>, Michael Paquier <michael(at)paquier(dot)xyz>, Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>, Hannu Krosing <hannuk(at)google(dot)com>, Dilip Kumar <dilipkumarb(at)google(dot)com>, Yugo Nagata <nagata(at)sraoss(dot)co(dot)jp>, Nikita Malakhov <hukutoc(at)gmail(dot)com> |
| Subject: | Re: Direct TOAST v2, faster, smaller and no migration needed |
| Date: | 2026-09-30 09:07:46 |
| Message-ID: | CAMT0RQTFsMvsisWVS8L_SP-BzEdkQGhi+cAxh=Jnkv7hJ5EcZw@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Adding back the list and other recipients again
On Wed, Sep 30, 2026 at 4:56 AM Michael Paquier <michael(at)paquier(dot)xyz> wrote:
>
> On Tue, Sep 29, 2026 at 02:18:41PM +0200, Matthias van de Meent wrote:
> > But the independent repack-ability of toast tables is a very valuable
> > feature. Main tables have indexes that may be very expensive to
> > rebuild. Yes, CONCURRENTLY could help, but it'll still take a
> > possibly huge amount of resources, in time, CPU, memory, disk, IO.
> > Repacking the toast table separately solves bloat issues in the toast
> > table, and does not require the main table to be rebuilt, avoiding the
> > related resource consumptions of index builds etc.
>
> Note that TOAST-only nightly VACUUMs can be quite common for some
> workloads, a reason why VACUUM (PROCESS_TOAST) exists so as we can
> clean up only the underlying TOAST table. I do not see enough
> arguments in favor of dropping this kind of properties that many
> people rely on, based on only the claims posted on this thread.
> That's just too much risk compared to the potential reward.
I have a patch to treat deleted Direct TOAST chunks similar to HOT pages and
prune them to LP_UNUSED on access, together with page map updates, so
you can get to a steady state without running even autovacuum.
This is now possible because there is no index cleanup needed.
> > No, we don't recheck the valueid in the heaptuple. However, we do
> > check that the chunk counter matches the expected values, and that's
> > something that won't be possible in the direct toast design -- there
> > is nothing in the toast pointer nor the toast tuple that links it to
> > its own identity, especially not outside the scope of MVCC snapshots.
>
> Yes, we've relied on the concept of TOAST tuple "identity" heavily for
> 20-ish years.
>
> You could introduce a new behavior based on tids as an opt-in, I
> guess, leaving the default 4-bytes be, giving access to tid-based
> access tables for newly-created TOAST tables.
The latest patch series does this for both oid and oi8 TOAST.
And it has been something you can turn on and off via ALTER TABLE ...
(toast_flavour = direct)
It can be turned on and off any time you like, as both types of toas
can live together.
> Anyway, claiming that this proposal should enforce a new behavior
> across the board for everything is not something I can vouch for.
It does not force anything, is controlled by a dynamic table storage
parameter toast_flavour
And we probably should also allow toast_flavor :)
Patch 2: Add Direct TOAST catalog, GUC, and reloptions infrastructure
Defines struct varatt_direct and VARTAG_DIRECT in varatt.h. Adds the
default_toast_flavour GUC ('plain' | 'direct', default 'plain') and the
table storage parameter toast_flavour. Extends create_toast_table() to add
chunk_tids (tid[]) and chunk_tid_offsets (int8[]) to TOAST relations and
creates the partial index predicate (WHERE chunk_id IS NOT NULL).
> I get the cheap rewrite argument with the addition of the partial index,
> but no solution has magical properties; there are always two sides of
> the coin, and we have to think about both sides. Using tids for
> lookups has benefits in some cases, clearly, just claiming that it's
> always beneficial for every kind of workload is something I find very
> hard to believe.
That's why I have asked you to find a counterexample :)
> The decision to make impossible TOAST-based VACUUMs
> is one downside that I do not and will not underestimate, for one. I
> cannot speak for other committers, but at least that's one committer
> opinion. My opinion can of course be overruled, it happens a lot.
For a VACUUM FULL equivalent you can still do REPACK on main table
which has the upside of clustering the toast values in table order
And the normal vacuum is made drastically faster as it can skip the
index scan and prune tuples immediately.
> Another concept Greg Burd has been mentioning to me is overflow pages,
> used by some other systems: keep all the chunks in heap, with tid
> redirections.
Yes, this was actually the original Ingres design from 1976, see page 204
in https://www.engineering.upenn.edu/~zives/cis650/papers/INGRES.PDF
The downside is that it would updates more expensive when the toasted
values are not updated.
> That kind of mimics some of the builds out there where
> folks compile the code with larger base page sizes to avoid the TOAST
> lookups, for workloads where the blobs are large enough to be in TOAST
> but could fit on a page if made larger.
The other page sizes seem to be largely neglected.
Currently some tests fail and someone told me that the b-tree still
would fill the larger pages only up to 8k :(
I made a weak attempt to set up a build farm animal to test 16k ad 32k
pages but never got my buildfarm key back from the registraion form.
> It has its own set of
> advantages and disadvantages as well, like any other approach. Back
> to the two sides of the coin.
Agreed, without more invasive changes larger pages will cause at least
more read and write amplification for sparse accesses to just mention
one downside.
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Bingshuai Li | 2026-09-30 09:12:31 | RE: Bug in logical decoding with DDL and subtransactions |
| Previous Message | Anthonin Bonnefoy | 2026-09-30 08:48:10 | Re: Protocol Compression (fourth attempt) |