Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)

From: Jakub Wartak <jakub(dot)wartak(at)enterprisedb(dot)com>
To: Gustavo William <gustavowilliam0805(at)gmail(dot)com>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)
Date: 2026-09-28 08:58:58
Message-ID: CAKZiRmxEha69zr_prZbbSi=uzzu9pd9o1kHBfyY872HcNGMuog@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, Sep 22, 2026 at 9:40 PM Gustavo William
<gustavowilliam0805(at)gmail(dot)com> wrote:
>
> Hi Jakub,
>
> Thanks for reviewing it. I'm resending the chain of patches here, with 0005 already merged + your suggestions applied.
[..]

Hi Gustavo,

TLDR; I've verified ~9.4x improvement with 0005-incremental-backup just for
the read path on the NVMe (by avoiding stalled pread()). Good job :)

Testing method:
initdb # summarize_wal=on , maintenance_io_concurrency=2, ...
pgbench -i -s 500 --fillfactor=100 # ~7.5GB base/
pg_basebackup -D /tmp/full -c fast -v
pgbench -N -c 8 -j 8 -t 10000 postgres # -N=updates
psql -c "CHECKPOINT" postgres
sudo /usr/local/bin/drop_fs_cache.sh
time pg_basebackup -i /tmp/full/backup_manifest --target=server-blackhole \
-Xnone -c fast --no-sync --no-verify-checksums

Avg of 3 runs (drop_caches + time pg_basebackup -i)
Without 0005: 7.53s
With 0005: 0.80s (note it's just reading due to blackhole, so no writes!)

I don't have access to SATA drive, but I suspect it would be way better
there.

I was kind of sceptical of the result at first, but explanation of the
phehomena seems to be like this: basebackup_incremnetal.c has
GetFileBackupMethod()->qsort() so it sorts the blocks to be read. Without 0005
patch, preads() of increasing block numbers (offsets) all cause hit page cache
misses, and apparently - at least here on kernel 6.17.x, but earlier for sure
too - this seem to kick off the readahead (because kernel's readahead seems to
be implemented to only kick in by by page-cache misses , it won't even start
if there are hits, kind of makes sense), so pure preads() tend to overread a
lot (like reading N-times more data from a a 1GB segment than we need).

I think the kernel's assumption is that reading sequentially is cheaper than
reading one by one, but depending on density of changes we often - at least
in this cenario, that is in incremetanl basebackup mode - we read much less
than the kernel reads with readahead (simply said: it overreads a lot!).

So using fadvise(FADV_WILLNEED) here caused kind of disarms such aggressive
readahead completley (as it results in 100% hit to pagecaches, so readahead is
not even activated), and this causes way faster incremental times, so fast
that I'm had to actually reserach why and how it happens :o

Issues with the patch
- it seems maintenance_io_concurrency=1 disables prefetch? (it's not queue
depth, it's more of how many to prefetch, so "1" means IMHO pread() +
one posix_fadvise() or am I wrong?)
- pgindent might be needed.

-J.

In response to

Browse pgsql-hackers by date

  From Date Subject
Previous Message Cagri Biroglu 2026-09-28 08:54:35 Re: Per-table resync for logical replication subscriptions