Re: aio: Async fsyncs for crash recovery and checkpointer

From: Nazir Bilal Yavuz <byavuz81(at)gmail(dot)com>
To: Nitin Jadhav <nitinjadhavpostgres(at)gmail(dot)com>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, Andres Freund <andres(at)anarazel(dot)de>, iamqyh(at)gmail(dot)com
Subject: Re: aio: Async fsyncs for crash recovery and checkpointer
Date: 2026-09-30 08:43:42
Message-ID: CAN55FZ2i-=87NWg8v66hbbm18o+kBmCjyLyX4fE7s49c9JbEBw@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

On Tue, 22 Sept 2026 at 18:21, Nitin Jadhav
<nitinjadhavpostgres(at)gmail(dot)com> wrote:
>
> I found two further points, and the worker-side SLRU test-coverage
> question from v1 still appears applicable.
>
> An I/O worker can skip a required fsync after fsync is enabled. As
> Yuhang noted, there appears to be a race when reloading fsync from off
> to on. pgaio_io_start_fsync() makes the dispatch decision using the
> issuing process's enableFsync.

You are right, that is fixed in v3 [1]. I used the issuing process's enableFsync

> Worker SMGR cleanup appears dependent on the worker becoming idle.
> Patch 0002 calls smgrdestroyall() only after the worker finds that no
> request is available. This means a continuously busy worker may never
> perform the cleanup. Worker-side relation reopening calls smgropen(),
> and those unpinned SMgrRelation objects remain in the worker's SMGR
> hash until smgrdestroyall() is called. With sustained I/O over many
> distinct relations, a worker whose queue never becomes empty could
> therefore retain an increasing number of SMGR entries, including
> entries for relations that have since been dropped. Would it be safer
> to check FirstCallSinceLastCheckpoint() at a safe point after
> completing each request, before consuming the next request, rather
> than only on the idle path? At that point any descriptor reopened for
> the completed operation has already been released.

We have created another thread for this. Now, we have a dedicated SMGR
entry limits, and the worker can't surpass this limit. For more
information you can look [2], but briefly, SMGR entries can't surpass
this limit when the worker is busy; they are cleaned when we reach
this limit. Also, SMGR entries are cleaned before the worker goes to
idle. These should solve the problem.

> The worker-side SLRU path does not appear to have targeted test
> coverage. The test_slru_page_sync() still registers the test SLRU with
> SYNC_HANDLER_NONE. Consequently, SlruSyncFileTag() selects
> PGAIO_TID_SYNC, which has no reopen callback and is executed
> synchronously in the submitting process under io_method=worker.

Yes, that is not tested. I will add that.

[1] https://postgr.es/m/CAN55FZ2NT6529QhB_dcrZ5LaSL20aayhP33ZisZVFSg2hOqLXg%40mail.gmail.com
[2] https://postgr.es/m/CAN55FZ2BesKUnajdgpw1fPSe3S6_CHOugryaUEtD7vdP%3DdRKEQ%40mail.gmail.com

--
Regards,
Nazir Bilal Yavuz
Microsoft

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Anthonin Bonnefoy 2026-09-30 08:48:10 Re: Protocol Compression (fourth attempt)
Previous Message Nazir Bilal Yavuz 2026-09-30 08:42:11 Re: aio: Async fsyncs for crash recovery and checkpointer