| From: | Heikki Linnakangas <hlinnaka(at)iki(dot)fi> |
|---|---|
| To: | Jakub Wartak <jakub(dot)wartak(at)enterprisedb(dot)com>, Gustavo William <gustavowilliam0805(at)gmail(dot)com> |
| Cc: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O) |
| Date: | 2026-10-05 16:24:13 |
| Message-ID: | dd7a21be-3c16-48de-bfad-58e917cdf40a@iki.fi |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
I started to look at these, starting from this patch:
> 0004 increases SINK_BUFFER_LENGTH from its current 32 kB. Profiling the
> server side shows a huge number of small pread() calls coming from
> basebackup_read_file(), which is surprising given that elsewhere in the tree
> (backend/storage/buffer/README) we already mention that a 256 kB ring for seq.
> scans scans is used because it fits comfortably in L2 cache. On my laptop,
> with no SSL, no checksum generation or verification, a hot filesystem cache,
> and a client blackhole to eliminate client I/O, a 10 GB backup over loopback
> runs at:
> 3.1 GB/s with the 32 kB buffer,
> 5.1 GB/s at 128 kB,
> 5.5 GB/s at 256 kB,
> 6.0 GB/s at 1 MB
> (those are average of five runs).
Cool
> The win comes from letting pread()
> swallow much larger chunks of each segment in one go: per-core L2 caches are
> 1-2 MB even on laptops these days, syscalls have become more expensive since
> the Spectre/Meltdown mitigations, and the kernel's default readahead is
> already in the 128-512 kB range, so there's little reason to trickle the data
> through 32 kB at a time.
While we're at it, we should probably allocate the buffer with
palloc_aligned(PG_IO_ALIGN_SIZE); a pread() into an aligned buffer
should be a little faster.
I wonder about this in read_file_data_into_buffer():
> cnt = basebackup_read_file(fd, sink->bbs_buffer,
> Min(sink->bbs_buffer_length, length),
> offset, readfilename, true);
>
> /* Can't verify checksums if read length is not a multiple of BLCKSZ. */
> if (!verify_checksum || (cnt % BLCKSZ) != 0)
> return cnt;
It doesn't seem right to assume that the count must be a multiple of
BLCKSZ, especially if we use a larger buffer. Should we do something
about that?
- Heikki
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Tomas Vondra | 2026-10-05 16:24:20 | Re: hashjoins vs. Bloom filters (yet again) |
| Previous Message | Masahiko Sawada | 2026-10-05 16:10:56 | Re: parallel autovacuum: Propagate track_cost_delay_timing to parallel workers |