Re: UNDO with constant time recovery (CTR)

From: Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>
To: Greg Burd <greg(at)burd(dot)me>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: UNDO with constant time recovery (CTR)
Date: 2026-10-05 10:59:32
Message-ID: CAEze2WgeRpQyMaH-WJXPePOZ=AdTCt-sV80nMcue8k3R8=XXnA@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, 29 Sept 2026, 00:29 Greg Burd, <greg(at)burd(dot)me> wrote:
>
> Hello hackers,
>

Hi Greg,

A few things popped out in this thread, so here's a few comments.

[...]

> All these authors and more deserve credit for their hard work, thank
> you. In the case of ZHeap and Zedstore the ultimate goal was a new table
> AM. I'm not going to boil that ocean up front, here is where I'm going
> to diverge from those past projects.

I'd prefer if you didn't boil any oceans in the process of making
contributions. Not up front is a good start, but please don't do that
later on, either.

[...]

> If a transaction is in flight and creates a new file and then crashes to
> me it makes sense that after recovery that new file is gone and the
> system is consistent again.

That very much depends on the type of file created. E.g. WAL files,
SLRU files, and new segments of existing relations should not be
removed just because they were created by a transaction that crashed
or rolled back -- other backends may have have written to those files.

Also note that we may want to have UNDO outside transactional
boundaries: after all, not all DDL is transactional, and I think
there's some gain to be had for e.g. REINDEX CONCURRENTLY; it allows
us to truncate the !indisready index files it leaves behind after a
crash, reducing disk bloat in those cases.

Side note: I'd prefer if this change does _not_ mean that unlogged
relations get to WAL-log proportional to data operations. Undo
logging can be useful, but should not be predicated on WAL if a
tableAM requires UNDO.

[...]

> Also, I realize that this kind of change takes a lot of time to gain
> traction and adoption into core, if at all. I'm ready for that, sure
> we're working on v20 now and v19 is inching out the door maybe this
> merges into v25 or maybe it gets shelved along the way for good reason.
> Who knows, but I do know that it's worth the effort to advocate and to
> have a durable record for others interested in it even if it doesn't get
> merged in this time. I look forward to seeing what happens. :)

The title indicates constant-time recovery, but I don't see anything
in this thread that supports this claim. Could you expand on the
mechanisms you're using to guarantee this?

I also don't see how your claimed constant-time rollback can work,
given that the size of a transaction's modified working set is bounded
only by time and the space available to store the database, and that
undoing changes in files can't really be done faster than the
bandwidth of your CPU. Undoing a set of transaction operations
therefore can't really be constant-time unless you're limiting the
number of undo-able transaction operations, and a limited transaction
size is not really something current users have to consider (and thus,
don't expect).

Kind regards,

Matthias van de Meent
Databricks (https://www.databricks.com)

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Kirill Reshke 2026-10-05 11:20:04 Re: pg_dump/restore failure (dependency?) on BF serinus
Previous Message Heikki Linnakangas 2026-10-05 10:39:50 Re: [PATCH v1] amcheck: Allow interrupting the child-level rightlink walk