Re: Offline data checksum changes can cause incorrect checksum state on standbys

From: Bertrand Drouvot <bertranddrouvot(dot)pg(at)gmail(dot)com>
To: Daniel Gustafsson <daniel(at)yesql(dot)se>
Cc: Zsolt Parragi <zsolt(dot)parragi(at)percona(dot)com>, pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys
Date: 2026-08-31 08:01:00
Message-ID: apU0vKrffQDmmv68@bdtpg
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

On Mon, Aug 31, 2026 at 09:16:52AM +0200, Daniel Gustafsson wrote:
> > On 31 Aug 2026, at 07:06, Bertrand Drouvot <bertranddrouvot(dot)pg(at)gmail(dot)com> wrote:
> >
> > Fixing this would probably require recording additional ordering information for
> > offline changes, adding even more complexity to v5. Another option would be to
> > document that the standby must be fully caught up before both nodes are stopped
> > for the offline operation.
>
> I think we really need to think about documenting a lot of this, potentially
> even to the point of saying that offline and online changes should not be mixed

Yeah, that could make sense, at least to make the limitation explicit and warn
users about these cases.

> as they work with completely different durability models.

Yeap.

> The more I think about this the less excited I am about contorting the logic of
> a feature which does proper WAL logging to cope with a tool that doesn't,
> including misuses like creating mismatched clusters. We should probably start
> to look at improving pg_checksums such that transitions are WAL logged rather
> than shoehorning in such changes with a WAL logged flow. pg_checksums rewrites
> the datadirectory without the postmaster given any information that any change
> was made, which in itself should be a red flag.

I agree.

> Making sure that StartupXLOG
> can detect the offline change (or something along those lines) and properly log
> it seems like a better starting point.

That sounds worth exploring. Would the idea be for pg_checksums to leave a
marker in pg_control which StartupXLOG would turn into a WAL logged transition?

> >>> Do you see the control version change as a concern?
> >>
> >> Yes, it is another non-trivial change in an already complex patch,
> >> really close to RC1. It's also not an area where we could easily
> >> implement bug fixes in a minor version, if we discover something
> >> later.
> >
> > Yeah, and I think the case above reinforces that concern.
>
> I don't think the above reinforces not wanting to do a pg_control change at
> this point. I think it reinforces that changing datafiles without WAL logging
> is a fairly slippery slope.

Yeah, I see your point that the underlying issue is the non WAL logged change.
My concern was not about a pg_control change in itself, but that addressing this
case might require further changes to the current pg_control design, which adds
risk this close to RC1.

Regards,

--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Zsolt Parragi 2026-08-31 08:01:14 Re: Offline data checksum changes can cause incorrect checksum state on standbys
Previous Message Michael Paquier 2026-08-31 07:57:36 Re: pg_restore_attribute_stats() accepts non-finite values