| From: | Nazir Bilal Yavuz <byavuz81(at)gmail(dot)com> |
|---|---|
| To: | Alexander Lakhin <exclusion(at)gmail(dot)com> |
| Cc: | Daniel Gustafsson <daniel(at)yesql(dot)se>, Bertrand Drouvot <bertranddrouvot(dot)pg(at)gmail(dot)com>, Zsolt Parragi <zsolt(dot)parragi(at)percona(dot)com>, Heikki Linnakangas <hlinnaka(at)iki(dot)fi>, pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Subject: | Re: Offline data checksum changes can cause incorrect checksum state on standbys |
| Date: | 2026-09-28 06:58:05 |
| Message-ID: | CAN55FZ2Hvq9DKj9+pG8-R4EFQcdXO+fshb1vvN0LqjHeSiathg@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
On Sat, 26 Sept 2026 at 08:00, Alexander Lakhin <exclusion(at)gmail(dot)com> wrote:
>
> 14.09.2026 16:48, Daniel Gustafsson wrote:
>
> BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail
> a test added in b52a1c2c8:
> [22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint
> [22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint
> [22:51:29.088](0.001s)
> [22:51:29.088](0.000s) # Failed test 'last common checkpoint is a shutdown checkpoint'
> # at t/013_rewind.pl line 161.
> [22:51:29.089](0.001s) # ''
> # doesn't match '(?^:CHECKPOINT_SHUTDOWN)'
> [22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file
> ...
> Test Summary Report
> -------------------
> t/013_rewind.pl (Wstat: 256 (exited 1) Tests: 13 Failed: 1)
>
> I've reproduced this failure with the following modification:
> --- a/src/backend/access/transam/xlog.c
> +++ b/src/backend/access/transam/xlog.c
> @@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,
> */
> +pg_usleep(100000);
> installed_segno = logsegno;
>
> which makes all-zero 000000020000000000000005 appear inside
> $node_a->data_dir . '/pg_wal':
> tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005 | wc
> 0 0 0
>
> and then if readdir() happens to return this file first (I'm observing
> this on ext4):
> perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \
> src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal
> 000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003
>
> pg_waldump with no explicit segment specification fails:
> .../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000
> pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes)
> pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB.
Thanks for the report! I think your analysis correct. I encountered
the same problem on my local, there is another thread for fixing this
problem [1].
I can reproduce the problem with your reproducer, though not on the
first try; I needed to run the test a couple of times. Then, I confirm
that the patch in [1] fixes the problem, I run 013_rewind test 100
times and there was no failure.
[1] https://postgr.es/m/CAN55FZ1Yak_xBqMaDQsD7atpBkGLEkF-DKXcs3nLHM1Uq4YRew%40mail.gmail.com
--
Regards,
Nazir Bilal Yavuz
Microsoft
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Ayush Tiwari | 2026-09-28 07:14:38 | Re: [PATCH] Two remaining shmem attachment issues in single-user mode |
| Previous Message | David Steele | 2026-09-28 06:03:21 | Re: Return pg_control from pg_backup_stop(). |