| From: | John Naylor <johncnaylorls(at)gmail(dot)com> |
|---|---|
| To: | David Geier <geidav(dot)pg(at)gmail(dot)com> |
| Cc: | PostgreSQL Developers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Improving scalability of Parallel Bitmap Heap/Index Scan |
| Date: | 2026-10-08 11:21:53 |
| Message-ID: | CANWCAZayXXZfcDGXLhZGw6Pa4Qh88LvthhK5JiPbx5vpkaidpQ@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Thu, Oct 8, 2026 at 2:52 PM David Geier <geidav(dot)pg(at)gmail(dot)com> wrote:
> > Okay, so worst case N workers will use N times more memory. (not
> > counting transient use during repartitioning)
>
> Yes, but only temporarily and only in the BIS. The per-worker memory
> consumption also stays the same. It's just that there are more workers
> now that each can encounter the same TIDs / blocks. And once the
> repartitioning step is complete, the old TIDBitmaps are freed and the
> memory consumption is as before because now every participant holds a
> TIDBitmap containing only 1/nth of the total set of TIDs / blocks.
Okay.
> > One mitigation could be to use hash_mem_multiplier to expand work_mem,
> > and divide that by the number of workers.
>
> I thought work_mem is per-node, per-worker. Given that the per-node,
> per-worker memory consumption doesn't change, things should be fine?
Right, although I'm anticipating complaints about more memory used. I
believe parallel hash join started with per-worker private hash
tables, so it's probably defensible to follow that precedent. My idea
didn't really fit with how things normally work...
> > (I'm not sure why we don't
> > do that already; seems like a natural use for it). This will need
> > testing.
>
> Using hash_mem_multiplier in TIDBitmap makes sense, given that TIDBitmap
> is hash-based. I'll add that.
I still think it makes sense in principle, but if we're keeping
work_mem per node as is customary, having two mechanisms to possibly
increase memory usage is a step two far, especially in the same patch.
> Additionally, we could improve on the increased memory consumption by
> having the participants feed the found TIDs into lock-free ring buffers,
> one per partition. The participants would alternate between reading from
> the index to feed the ring buffers and pulling data out from their
> assigned ring buffer to store it in their local partition TIDBitmap.
>
> That's for sure less work than making TIDBitmap parallel-safe for insert
> (which includes lossification), should be on-par performance-wise and
> avoids the increase in memory consumption.
That sounds difficult to review.
> >> For that to pay off, the increased cost for inserting must be reasonably
> >> low. Is that the case?
> >
> > Actually, no, I just reminded myself that in shared memory, the node
> > pointers use DSA pointers, and dereferencing them requires a function
> > call.
>
> That makes sense: for what I understood, in parallel VACUUM the TidStore
> is filled by a single leader but then read by multiple per-index workers.
Right, and two total workers only shows a modest (still noticeable)
improvement over one because of the increased overhead.
> > Parallel index scan with shared partitioned hash table has been tried
> > before without success, I believe.
> Do you have any pointers to mailing list discussions or similar?
I found this one, not sure if it's the only one:
--
John Naylor
Amazon Web Services
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Kirill Reshke | 2026-10-08 11:38:59 | REPACK hits assertion failure on postmaster death exit |
| Previous Message | Nazir Bilal Yavuz | 2026-10-08 11:00:46 | Re: Adding init-po and update-po targets to the meson build system |