Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)

From: Heikki Linnakangas <hlinnaka(at)iki(dot)fi>
To: Jakub Wartak <jakub(dot)wartak(at)enterprisedb(dot)com>, Gustavo William <gustavowilliam0805(at)gmail(dot)com>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)
Date: 2026-10-05 16:24:13
Message-ID: dd7a21be-3c16-48de-bfad-58e917cdf40a@iki.fi
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

I started to look at these, starting from this patch:

> 0004 increases SINK_BUFFER_LENGTH from its current 32 kB. Profiling the
> server side shows a huge number of small pread() calls coming from
> basebackup_read_file(), which is surprising given that elsewhere in the tree
> (backend/storage/buffer/README) we already mention that a 256 kB ring for seq.
> scans scans is used because it fits comfortably in L2 cache. On my laptop,
> with no SSL, no checksum generation or verification, a hot filesystem cache,
> and a client blackhole to eliminate client I/O, a 10 GB backup over loopback
> runs at:
> 3.1 GB/s with the 32 kB buffer,
> 5.1 GB/s at 128 kB,
> 5.5 GB/s at 256 kB,
> 6.0 GB/s at 1 MB
> (those are average of five runs).

Cool

> The win comes from letting pread()
> swallow much larger chunks of each segment in one go: per-core L2 caches are
> 1-2 MB even on laptops these days, syscalls have become more expensive since
> the Spectre/Meltdown mitigations, and the kernel's default readahead is
> already in the 128-512 kB range, so there's little reason to trickle the data
> through 32 kB at a time.

While we're at it, we should probably allocate the buffer with
palloc_aligned(PG_IO_ALIGN_SIZE); a pread() into an aligned buffer
should be a little faster.

I wonder about this in read_file_data_into_buffer():

> cnt = basebackup_read_file(fd, sink->bbs_buffer,
> Min(sink->bbs_buffer_length, length),
> offset, readfilename, true);
>
> /* Can't verify checksums if read length is not a multiple of BLCKSZ. */
> if (!verify_checksum || (cnt % BLCKSZ) != 0)
> return cnt;

It doesn't seem right to assume that the count must be a multiple of
BLCKSZ, especially if we use a larger buffer. Should we do something
about that?

- Heikki

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Tomas Vondra 2026-10-05 16:24:20 Re: hashjoins vs. Bloom filters (yet again)
Previous Message Masahiko Sawada 2026-10-05 16:10:56 Re: parallel autovacuum: Propagate track_cost_delay_timing to parallel workers