| From: | shihao zhong <zhong950419(at)gmail(dot)com> |
|---|---|
| To: | Andrey Borodin <x4mmm(at)yandex-team(dot)ru> |
| Cc: | Melanie Plageman <melanieplageman(at)gmail(dot)com>, Chao Li <li(dot)evan(dot)chao(at)gmail(dot)com>, Nazir Bilal Yavuz <byavuz81(at)gmail(dot)com>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, Andres Freund <andres(at)anarazel(dot)de>, tristan(dot)yim(at)gmail(dot)com |
| Subject: | Re: Checkpointer write combining |
| Date: | 2026-10-04 05:07:07 |
| Message-ID: | CAGRkXqRqJq+_wmyzbD5+HOQ4C4DPQ6f68d00qaVJzXHq5+iEbQ@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi Melanie,
Thanks for working on this. I tried v16 and have a few notes.
1. Recovery from a base backup crashes for me on an assert build.
On master I get
WARNING: xlog min recovery request 0/045E5260 is past current
point 0/0210CF08
CONTEXT: writing block 0 of relation "base/5/16384_vm"
and recovery finishes. With v16 the startup process then hits
TRAP: failed Assert("... || !XLogNeedsFlush(BufferGetLSN(bufhdr))")
I think it is because XLogFlush() only warns in this case, so the
new Assert in WriteBuffers() does not hold in recovery. It needs
checksums off.
2. I see the same pacing problem as Andrey. A spread checkpoint
takes 26.3s on master and 11.2s with v16. With the attached patch it
is 26.7s.
3. In 0018 the neighbors count against bgwriter_lru_maxpages. With
pgbench at scale 100 and 256MB shared_buffers, the bgwriter hits the
limit in most rounds on both. In 60s it cleaned 28000 clock hand
buffers on master and 20218 with v16. tps was about the same. Is
that intended? I also wonder if the bgwriter falls behind the clock sweep
more
easily now, especially with direct I/O. It spends more time on each
write, so the clock tick may pass it in the middle of a round. I did not
test this yet but I think worth to discuss.
4. A failed combined write sets BM_IO_ERROR on all buffers in the
batch. With EIO on block 690, the retry also logs "write error might
be permanent" for 691 and 692. I think technically we can do better but
given IO_ERROR is rare, maybe it is not worth it in the first path.
5. 0004 renames buffer-sync-written, but monitoring.sgml still has
the old name. I think that should go with 0004. The docs for the new
probes can wait until the functional parts are done.
Thanks,
Shihao
| Attachment | Content-Type | Size |
|---|---|---|
| nocfbot-repro-writebuffers-assert.sh | text/x-sh | 1.4 KB |
| nocfbot-standby-full-v16.log | application/octet-stream | 3.5 KB |
| From | Date | Subject | |
|---|---|---|---|
| Previous Message | shihao zhong | 2026-10-04 04:22:21 | Re: BUG #19686: Rolling back SET TABLESPACE |