| From: | Rogers Wang <rogers(dot)ww(at)qq(dot)com> |
|---|---|
| To: | pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Cc: | Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>, Álvaro Herrera <alvherre(at)kurilemu(dot)de>, Michael Paquier <michael(at)paquier(dot)xyz> |
| Subject: | Re: WAL_LOG CREATE DATABASE strategy broken for non-standard page layouts |
| Date: | 2026-08-23 17:04:44 |
| Message-ID: | tencent_2E870046716FD94285045E96505A2D4E2908@qq.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
Reviving this thread: the bug reported here in 2024 is still present
on master, and has become more harmful since.
To recap the root cause: RelationCopyStorageUsingBuffer() copies every
fork page by page and WAL-logs each copied page with
log_newpage_buffer(buf, page_std = true). page_std=true punches the
[pd_lower, pd_upper) range out of the full-page image as a hole, to
save WAL. That works for heap and index pages, but VM and FSM pages
never maintain pd_lower/pd_upper - they keep whatever PageInit() left
(pd_lower = SizeOfPageHeaderData, pd_upper = BLCKSZ). So for those
forks the "hole" covers the whole content area, the FPI carries
nothing but the header, and replay zero-fills it. The primary lands
the real VM/FSM contents on disk, while the standby replays all-zero
pages. As far as I can see, this call site is the only one that logs
VM/FSM pages in standard mode - everywhere else they are logged with
page_std=false.
What changed since the last discussion
--------------------------------------
The divergence was assessed back then as "not actively harmful". That
no longer holds in v19: commit add323da40a added
Assert(BufferIsDirty(vmbuffer)) to heap_xlog_prune_freeze(), which
makes the divergence fatal on assert builds after a switchover:
1. The promoted standby has a zeroed VM, so VACUUM FREEZE sets the
bits again, emitting PRUNE records with no FPI for the VM buffer.
2. The old primary, rejoined as standby, still has the bits set with
an older page LSN, so replay takes the redo branch - but
visibilitymap_set() is a no-op there (bits already set) and the
buffer stays clean.
3. Assert(BufferIsDirty(vmbuffer)) fires, and the node crash-loops on
that record.
And even without assertions, the new database on the standby silently
loses all all-visible/all-frozen bits and all free space data of every
copied relation.
Repro
-----
With a streaming standby attached:
-- on the primary
CREATE DATABASE src_db;
\c src_db
CREATE EXTENSION pageinspect;
CREATE TABLE t(a int);
INSERT INTO t SELECT generate_series(1, 10000);
VACUUM FREEZE t;
-- back in postgres
CREATE DATABASE dst_db TEMPLATE src_db STRATEGY WAL_LOG;
-- then on both nodes, in dst_db:
SELECT encode(get_raw_page('t', 'vm', 0), 'hex');
SELECT encode(get_raw_page('t', 'fsm', 0), 'hex');
-- primary shows real content, standby is all zeros
For the assert crash: run with full_page_writes=off so the PRUNE
record carries no FPI for the VM page, promote the standby, rejoin the
old primary as its standby, then VACUUM (FREEZE, DISABLE_PAGE_SKIPPING)
t on the new primary. The old primary dies on the assert while
replaying the PRUNE records.
0002 adds a TAP test to src/test/recovery/t/034_create_database.pl
which automates the scenario above: all new checks fail without 0001
and pass with it (verified on REL_19_STABLE with --enable-cassert).
Fix
---
The v1 patch logged all copied pages with page_std=false, which was
objected to because of the WAL volume increase on the main fork: holes
would no longer be punched from heap/index pages. Attached v2 instead
makes the decision per fork, along the lines Michael suggested. The
VM and FSM forks are the only ones whose pages never maintain a
standard page layout, so only their pages are logged whole:
- /* WAL-log the copied page. */
+ /*
+ * WAL-log the copied page. VM/FSM pages never maintain
+ * pd_lower/pd_upper, so hole punching would omit their entire
+ * content from the FPI.
+ */
if (use_wal)
- log_newpage_buffer(dstBuf, true);
+ log_newpage_buffer(dstBuf,
+ forkNum != VISIBILITYMAP_FORKNUM &&
+ forkNum != FSM_FORKNUM);
Main and init forks keep hole punching, so their WAL volume is
unchanged; the extra WAL is limited to the VM/FSM forks, which are
small. The fork number is available locally in
RelationCopyStorageUsingBuffer() (it already branches on it for
use_wal), so no plumbing through dbcommands.c is needed.
The bug goes back to v15, where STRATEGY WAL_LOG was added
(9c08aea6a30); the assertion crash additionally requires v19
(add323da40a), so older branches see only the VM/FSM divergence.
This should be backpatched to all supported versions.
Regards,
Rogers Wang
| Attachment | Content-Type | Size |
|---|---|---|
| 0001-Fix-hole-punching-of-VM-FSM-pages-in-CREATE-DATABASE.patch | text/plain | 754 bytes |
| 0002-Add-test-for-VM-FSM-divergence-after-CREATE-DATABASE.patch | text/plain | 6.3 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Tomas Vondra | 2026-08-23 17:30:32 | Re: Changing the state of data checksums in a running cluster |
| Previous Message | Alexander Lakhin | 2026-08-23 17:00:01 | Re: Changing the state of data checksums in a running cluster |