Re: Use instr_time for pg_stat_database block read/write time counters

From: ahmed <gouda0x(at)gmail(dot)com>
To: Andres Freund <andres(at)anarazel(dot)de>
Cc: Michael Paquier <michael(at)paquier(dot)xyz>, Melanie Plageman <melanieplageman(at)gmail(dot)com>, Lukas Fittl <lukas(at)fittl(dot)com>, Bernd Reiß <bd_reiss(at)gmx(dot)at>, pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: Use instr_time for pg_stat_database block read/write time counters
Date: 2026-10-02 17:03:24
Message-ID: CAFTkQVLbbNJoasheK3k0LbeT_ONWE4Sn_W0iP68t7xan75eXjQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi Andres,

Thanks a lot for the review, and first, please ignore the previous email;
it has bad formatting = Gmail tricked me :(.

> Of course it'd be even better if would be to stop counting the same stuff
in
> pgBufferUsage.{shared,local}_blk_{read,write}_time and
> pgStatBlock{Read,Write}Time. That's pretty darn silly.
>
> pgstat_update_dbstats() should just keep a pgBufferUsage snapshot from the
> last report and add up the relevant pgBufferUsage fields. And vacuum &
analyze
> already diff, so they just would need to add the fields to get the same
> results.

We tried to implement this and directly use `pgBufferUsage` in
`pgstat_update_dbstats()` but we came across some issues that caused
pg_stat_database to overcount, specifically when having parallel workers in
the plan, each worker will acuumulate its own statistics in its own
`pgBufferUsage` and will report them at the end and update the db stats
entry, but the problem is after that the parent backend will also acuumulate
all the statistics from its parallel workers and add them to its
`pgBufferUsage` and update the db stats entry again hence causing each
worker's I/O write/read times to be counted twice.

Meanwhile using `pgStatBlock{Read,Write}Time` in
`pg_stat_count_io_op_time()` will count every I/O timing a single time,
without having to introduce changes in the parallel worker infrasturcture
to handle the above case? Maybe thats why `pgStatBlock{Read,Write}Time` was
introduced in the beginning?

Or else, would it be the right approach to look into changing the existing
infastructure for I/O time counting in the context of parallel workers in
your opinion?

What do you think? Are we missing something?

Best Regards,
Ahmed Gouda and Bernd Reiß

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Kirill Reshke 2026-10-02 17:03:39 Fix reindexdb with parallel index-level conrurrent run
Previous Message Alexandre Felipe 2026-10-02 17:01:59 Re: Throwing away unnecessary spin-locks