Re: Recovery does not honor io_combine_limit, causing IOPS saturation

From: Mats Kindahl <mats(dot)kindahl(at)gmail(dot)com>
To: Andrey Borodin <x4mmm(at)yandex-team(dot)ru>, Markos Fountoulakis <markos(at)planetscale(dot)com>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Recovery does not honor io_combine_limit, causing IOPS saturation
Date: 2026-09-30 17:46:25
Message-ID: 39bc9cec-009d-4712-8212-aa17a9facb0f@gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi Andrey, Markos,

Thank you for reviewing this. Here is a new patch fixing the bug that
Andrey found. I need to get back to you regarding the benchmarks.

Best wishes,
Mats Kindahl

On 9/14/26 20:44, Andrey Borodin wrote:
> Hi Mats, Markos,
>
> Mats, in the original gp3 setup, did you also measure fewer device-level
> reads and shorter recovery time with the normal readahead settings?
>
> I'd also be interested in replay performance with WAL and relation pages
> already cached. Fewer syscalls could help even when the kernel already
> combines disk reads, but does that saving outweigh the extra per-page
> memcpy in this implementation?
>
> Looking at v1, I think there is a cache-validity issue with streaming:
>
>> + if (readSource == XLOG_FROM_STREAM)
>> + wanted = Min(wanted, readLen);
>> ...
>> + wanted = Max(wanted, XLOG_BLCKSZ);
> readLen is the valid prefix of the requested page, not the amount of
> available WAL from that page onwards. It is at most XLOG_BLCKSZ, so
> these bounds also leave streaming reads at one page regardless of
> io_combine_limit.
>
> More importantly, when only part of the page has arrived, pread() still
> reads the whole page and readAheadLen records the full result. Once
> more WAL arrives, XLogPageRead() can be called again for the same page
> with a larger reqLen. The new cache then hits and returns the old copy,
> including bytes that were not valid when it was filled.
>
> Should the cache track how many bytes were valid at the time of the
> read, and refill when reqLen exceeds that prefix? The available WAL
> distance from targetPagePtr to flushedUpto could separately bound the
> combined read size.
>
> Thanks!
>
>
> Best regards, Andrey Borodin.

Attachment Content-Type Size
v2.0001-Combine-WAL-segment-reads-during-recovery-into-large.patch text/x-patch 50.3 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Andrey Borodin 2026-09-30 18:00:26 Re: injection_points: canceled or terminated waiters leak their wait slots
Previous Message Andrey Borodin 2026-09-30 17:41:40 Re: Open SSI correctness issues