Re: Improving scalability of Parallel Bitmap Heap/Index Scan

From: John Naylor <johncnaylorls(at)gmail(dot)com>
To: David Geier <geidav(dot)pg(at)gmail(dot)com>
Cc: PostgreSQL Developers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Improving scalability of Parallel Bitmap Heap/Index Scan
Date: 2026-10-08 11:21:53
Message-ID: CANWCAZayXXZfcDGXLhZGw6Pa4Qh88LvthhK5JiPbx5vpkaidpQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Thu, Oct 8, 2026 at 2:52 PM David Geier <geidav(dot)pg(at)gmail(dot)com> wrote:
> > Okay, so worst case N workers will use N times more memory. (not
> > counting transient use during repartitioning)
>
> Yes, but only temporarily and only in the BIS. The per-worker memory
> consumption also stays the same. It's just that there are more workers
> now that each can encounter the same TIDs / blocks. And once the
> repartitioning step is complete, the old TIDBitmaps are freed and the
> memory consumption is as before because now every participant holds a
> TIDBitmap containing only 1/nth of the total set of TIDs / blocks.

Okay.

> > One mitigation could be to use hash_mem_multiplier to expand work_mem,
> > and divide that by the number of workers.
>
> I thought work_mem is per-node, per-worker. Given that the per-node,
> per-worker memory consumption doesn't change, things should be fine?

Right, although I'm anticipating complaints about more memory used. I
believe parallel hash join started with per-worker private hash
tables, so it's probably defensible to follow that precedent. My idea
didn't really fit with how things normally work...

> > (I'm not sure why we don't
> > do that already; seems like a natural use for it). This will need
> > testing.
>
> Using hash_mem_multiplier in TIDBitmap makes sense, given that TIDBitmap
> is hash-based. I'll add that.

I still think it makes sense in principle, but if we're keeping
work_mem per node as is customary, having two mechanisms to possibly
increase memory usage is a step two far, especially in the same patch.

> Additionally, we could improve on the increased memory consumption by
> having the participants feed the found TIDs into lock-free ring buffers,
> one per partition. The participants would alternate between reading from
> the index to feed the ring buffers and pulling data out from their
> assigned ring buffer to store it in their local partition TIDBitmap.
>
> That's for sure less work than making TIDBitmap parallel-safe for insert
> (which includes lossification), should be on-par performance-wise and
> avoids the increase in memory consumption.

That sounds difficult to review.

> >> For that to pay off, the increased cost for inserting must be reasonably
> >> low. Is that the case?
> >
> > Actually, no, I just reminded myself that in shared memory, the node
> > pointers use DSA pointers, and dereferencing them requires a function
> > call.
>
> That makes sense: for what I understood, in parallel VACUUM the TidStore
> is filled by a single leader but then read by multiple per-index workers.

Right, and two total workers only shows a modest (still noticeable)
improvement over one because of the increased overhead.

> > Parallel index scan with shared partitioned hash table has been tried
> > before without success, I believe.
> Do you have any pointers to mailing list discussions or similar?

I found this one, not sure if it's the only one:

https://www.postgresql.org/message-id/CAFiTN-t4NtRzafw94x%2BUb_gUqQsv%3Du9nK%3DOaJmS612M_bZv6%2BQ%40mail.gmail.com

--
John Naylor
Amazon Web Services

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Kirill Reshke 2026-10-08 11:38:59 REPACK hits assertion failure on postmaster death exit
Previous Message Nazir Bilal Yavuz 2026-10-08 11:00:46 Re: Adding init-po and update-po targets to the meson build system