Re: Costing for parallel scans with few/single row produced in the outer side

From: Robert Haas <robertmhaas(at)gmail(dot)com>
To: Matthias van de Meent <boekewurm(at)gmail(dot)com>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Costing for parallel scans with few/single row produced in the outer side
Date: 2026-09-29 14:50:13
Message-ID: CA+TgmobUaa0OAkUBOy0E-BxkG1k+0FZacFzp93_sPJPDQgx3Tw@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Mon, Sep 28, 2026 at 6:06 PM Matthias van de Meent
<boekewurm(at)gmail(dot)com> wrote:
> At Databricks, we noticed a weird plan in one of our TPCC benchmarks
> on pg18 (and reproduced on master), which was basically as follows:
>
> Gather
> NestedLoop
> NestedLoop
> Parallel Seq Scan (Filter: <primary key = constant>)
> BitmapScan
> BitmapScan
>
> At face value that plan looks OK, but when you look at it in detail
> it's quite a strange plan: The outermost plan is expected to only
> produce one row. Joins can generally only produce tuples once rows are
> produced on the outer side, so the paralellism is wasted: only a
> single worker will find a row, and thus only a single worker will
> execute the joins. Given that cap on parallelism, an index scan on
> the primary key would've been much cheaper.

I went looking for other reports of this issue. The only really clear
example I found was
http://postgr.es/m/f16e6fd6-7b7e-4c09-b26e-d7979ef86d2e@gmail.com --
in that email, Mark Kirkwood complains of what looks like the exact
same problem. In
http://postgr.es/m/152840735359.22458.3303333403164396853@wrigleys.postgresql.org
there's an interesting case where the driving table returns only 23
rows, but needs to be joined to a large sequential scan. Note that the
*estimate* for the area table is just 1 row, so this is very close to
being a case that your patch would affect, but I think that it isn't,
quite. http://postgr.es/m/872bffe7-82d0-86db-e3d6-2e20b1a72c4b@codata.eu
is a sort of opposite case: the estimates are high, the actual row
counts are low, and parallelism loses, but the issue there may have
more to do with worker startup and shutdown being expensive on that
machine than with the work distribution being uneven.

As far as the approach taken by the patch, I'm not sure that making
Path bigger for this is a good idea. It might be fine if we had lots
of reports of this being a problem, but it seems expensive as things
are. Still, that's not necessarily to say I think we should do
nothing. The approach makes me a little nervous in that it treats very
small number of rows as a very special case in need of very special
handling, but there's some argument to be made that this is actually
the case. I mean, small LIMIT values have similar problems, and we get
those wrong frequently, arguably because we don't treat that as a
sufficiently special case. Still, your patch takes the idea further:
the correction drops to zero as soon as #rows >= #workers, but lumpy
work distribution could still be an issue past that point (e.g. 5
rows, 4 workers). I'm not really sure what's best here.

--
Robert Haas
Databricks

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Álvaro Herrera 2026-09-29 14:58:12 Re: ATTACH PARTITION cost grows linearly with pg_constraint size (seqscan in CloneFkReferenced), much worse since not-null constraints are in pg_constraint (PG 18)
Previous Message Andrew Dunstan 2026-09-29 14:24:23 Re: [PATCH] Fix TAP tests with recent IPC::Run on Windows