Re: PG18: use-after-free in exec partition pruning after an EPQ recheck in LockRows

From: Vladimir Savin <vladimir(at)encord(dot)com>
To: Andrey Rachitskiy <pl0h0yp1(at)gmail(dot)com>
Cc: pgsql-bugs(at)lists(dot)postgresql(dot)org
Subject: Re: PG18: use-after-free in exec partition pruning after an EPQ recheck in LockRows
Date: 2026-10-01 19:36:08
Message-ID: CAP2G_gEgkKHs-_0GHLBRDso+HdagH+DaRhVCH9yT88=0LoCx1w@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-bugs

Hi Andrey,

Thank you for looking into this!

Is there anything else I can do to help, or are we all set?

Regards,
Vladimir

On Thu, 1 Oct 2026 at 17:15, Andrey Rachitskiy <pl0h0yp1(at)gmail(dot)com> wrote:

> This is the first time you're receiving an email from this person. Make
> sure you check the email address to confirm their identity before
> interacting with the email.
>
>
>
> чт, 1 окт. 2026 г. в 19:28, Vladimir Savin <vladimir(at)encord(dot)com>:
>
>> Hi,
>>
>> A backend process is terminated by SIGSEGV in run-time partition pruning
>> when a statement contains a FOR UPDATE CTE that has to recheck a
>> concurrently
>> updated row (EvalPlanQual) and a separate CTE that updates a partitioned
>> table
>> through an Append with run-time pruning. It reproduces on REL_18_6 and on
>> master.
>>
>> == Steps to reproduce
>>
>> Attached epq-prune-repro.sql, run against a freshly initdb'd cluster:
>>
>> psql -X -d postgres -f epq-prune-repro.sql
>>
>> It needs contrib/dblink, which it uses to open the second session.
>>
>> The enable_* settings only force the nested loop plan on these small
>> tables;
>> we first hit this with a plan the planner picked on its own (Update ->
>> Nested
>> Loop -> CTE Scan / Append with run-time pruning).
>>
>> == Output I got (REL_18_6, built with --enable-cassert; master gives the
>> same)
>>
>> psql:
>>
>> COMMIT
>> psql:epq-prune-repro.sql:44: ERROR: 08006: server closed the
>> connection unexpectedly
>> This probably means the server terminated abnormally
>> before or while processing the request.
>> invalid socket
>> CONTEXT: while executing query on dblink connection named "s2"
>> LOCATION: dblink_res_error, dblink.c:2863
>>
>> Server log (log_error_verbosity = verbose):
>>
>> 2026-10-01 14:07:00.975 UTC [13783] LOG: 00000: client backend (PID
>> 13824) was terminated by signal 11: Segmentation fault
>> 2026-10-01 14:07:00.975 UTC [13783] DETAIL: Failed process was running:
>> WITH inputs AS MATERIALIZED (
>> ... (the session 2 query above)
>>
>> Backtrace of the backend core (REL_18_6):
>>
>> #0 ExecEvalExprSwitchContext (state=0x7f7f7f7f7f7f7f7f,
>> econtext=0xaaaaf14182f0, isNull=0xffffe735e23f) at
>> ../../../src/include/executor/executor.h:440
>> #1 partkey_datum_from_expr (context=0xaaaaf1418568,
>> expr=0xaaaaf13fb440, stateidx=0, value=0xffffe735e240,
>> isnull=0xffffe735e23f) at partprune.c:3823
>> #2 perform_pruning_base_step (context=0xaaaaf1418568,
>> opstep=0xaaaaf13fb3f0) at partprune.c:3503
>> #3 get_matching_partitions (context=0xaaaaf1418568,
>> pruning_steps=0xaaaaf13fb4c0) at partprune.c:881
>> #4 find_matching_subplans_recurse (prunedata=0xaaaaf14184d0,
>> pprune=0xaaaaf14184d8, initial_prune=false, validsubplans=0xffffe735e488,
>> validsubplan_rtis=0x0)
>> #5 ExecFindMatchingSubPlans (prunestate=0xaaaaf1418480,
>> initial_prune=false, validsubplan_rtis=0x0) at execPartition.c:2535
>> #6 choose_next_subplan_locally (node=0xaaaaf143fe58) at
>> nodeAppend.c:584
>> #7 ExecAppend (pstate=0xaaaaf143fe58) at nodeAppend.c:330
>> #8 ExecProcNode (node=0xaaaaf143fe58) at
>> ../../../src/include/executor/executor.h:315
>> #9 ExecNestLoop (pstate=0xaaaaf142d410) at nodeNestloop.c:159
>> #10 ExecProcNode (node=0xaaaaf142d410) at
>> ../../../src/include/executor/executor.h:315
>> #11 ExecModifyTable (pstate=0xaaaaf142c260) at nodeModifyTable.c:4280
>> #12 ExecProcNode (node=0xaaaaf142c260) at
>> ../../../src/include/executor/executor.h:315
>> ...
>>
>> On master the frames are the same (partprune.c:3841, execPartition.c:2642,
>> nodeModifyTable.c:4497), also with state=0x7f7f7f7f7f7f7f7f.
>>
>>
>> == Output I expected
>>
>> The session 2 statement completes and dblink_get_result returns updated =
>> 8.
>> That is what the patched build below returns.
>>
>>
>> == Builds and settings
>>
>> - REL_18_6, commit 724edf9bde9d356724ad384a2e196edc3c9f80f7:
>> SELECT version() = PostgreSQL 18.6 on aarch64-unknown-linux-gnu,
>> compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit
>> - master, commit 42e96cf2fe095c822f76bc94669ea875cb1e3351:
>> SELECT version() = PostgreSQL 20devel on aarch64-unknown-linux-gnu,
>> compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit
>> - configure: '--prefix=/pg' '--enable-cassert' '--enable-debug'
>> '--without-icu' 'CFLAGS=-O0 -g'
>> plus `make -C contrib/dblink install`. No other deviations from the
>> installation instructions.
>> - Cluster: `initdb -A trust`, started with pg_ctl. The only
>> postgresql.conf change from the
>> initdb defaults is log_error_verbosity = verbose. No environment
>> variables set.
>>
>> Platform: Linux 7.0.12-linuxkit #1 SMP PREEMPT aarch64 (Docker Desktop on
>> macOS,
>> Apple silicon), Debian GNU/Linux 13 (trixie), glibc 2.41-12+deb13u4,
>> gcc 14.2.0-19, 14 CPUs, 23 GiB RAM.
>>
>> Release builds: with these exact steps the stock postgres:18.6 Docker
>> image
>> (Debian 18.6-1.pgdg13+2, aarch64) returns 8 and does not crash (3 of 3
>> runs).
>>
>> Release builds do crash on the same query shape under concurrent load
>> from the
>> Hatchet workflow engine, whose UpdateDurableEventLogEntriesSatisfied
>> query has
>> this structure:
>> - postgres:18.6 (18.6-1.pgdg13+2) under a synthetic load test: a segfault
>> every few minutes. A core taken with the dbgsym package installed has
>> the
>> same frames from ExecModifyTable (called from CteScanNext) up to
>> ExecEvalExprSwitchContext, with frame #0 at 0x000004a00000000c.
>> - Google Cloud SQL for PostgreSQL 18, production: "terminated by signal
>> 11"
>> in that query roughly once a day.
>> - postgres:17.11 under the same synthetic load: no crashes in 9 runs.
>>
>> Thanks,
>> Vladimir Savin
>>
>> *This email and any files transmitted with it are confidential and
>> intended solely for the use of the individual or entity to whom they are
>> addressed. If you have received this email in error, please notify the
>> sender immediately by replying to this email and delete it from your
>> system. Please do not copy, distribute, or take action based on the
>> contents of this email if you are not the intended recipient. Any
>> unauthorized use or disclosure of this email's contents is strictly
>> prohibited.*
>>
>
> Hi, Vladimir!
>
> Thanks for the report.
>
> I can reproduce it on the master branch.
>
> EvalPlanQualStart() shares the parent's PartitionPruneState so that EPQ
> initializes the same Append and MergeAppend subplans (8741e48e5dd).
> ExecInit of those nodes called InitExecPartitionPruneContexts() again.
> That replaced exec prune ExprStates with ones allocated in the EPQ
> memory context. EvalPlanQualEnd() frees that context.
>
> A later CTE that updates a partitioned table through Nested Loop plus
> Append with run-time pruning then used the dangling ExprStates and
> could SIGSEGV in partkey_datum_from_expr().
>
> Skip InitExecPartitionPruneContexts() when es_epq_active is set. A
> case is added to eval-plan-qual.
>
>
> --
> Regards,
> Rachitskiy Andrey
>

--
*This email and any files transmitted with it are confidential and intended
solely for the use of the individual or entity to whom they are addressed.
If you have received this email in error, please notify the sender
immediately by replying to this email and delete it from your system.
Please do not copy, distribute, or take action based on the contents of
this email if you are not the intended recipient. Any unauthorized use or
disclosure of this email's contents is strictly prohibited.*

In response to

Responses

Browse pgsql-bugs by date

  From Date Subject
Next Message shihao zhong 2026-10-02 03:17:26 Re: BUG #19730: PostgreSQL `pg_backup_start` accepts newlines in the backup label, producing an unrestorable `backup
Previous Message Masahiko Sawada 2026-10-01 18:31:17 Re: autovacuum: automatically propagate updated parameters