Re: BUG #19728: A logical replication apply worker segfaults dereferencing a NULL `MyLogicalRepWorker->stream_filese

From: Ajin Cherian <itsajin(at)gmail(dot)com>
To: bob(at)emrge(dot)ai, pgsql-bugs(at)lists(dot)postgresql(dot)org
Subject: Re: BUG #19728: A logical replication apply worker segfaults dereferencing a NULL `MyLogicalRepWorker->stream_filese
Date: 2026-10-01 03:09:19
Message-ID: CAFPTHDY6vZW7aww9NOg4VP5FpSRBW33CELVMVzHFHjxVjdzXgg@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-bugs

On Thu, Oct 1, 2026 at 12:58 PM PG Bug reporting form
<noreply(at)postgresql(dot)org> wrote:
>
> The following bug has been logged on the website:
>
> Bug reference: 19728
> Logged by: Robert Schmitt
> Email address: bob(at)emrge(dot)ai
> PostgreSQL version: 18.6
> Operating system: Linux Kaitain 7.0.0-1019-nvidia #19~24.04.2-Ubuntu
> Description:
>
> ================================================================================
> SUBMIT VIA: https://www.postgresql.org/account/submitbug/
> (or email the body below to pgsql-bugs(at)lists(dot)postgresql(dot)org)
>
> Form fields
> -----------
> PostgreSQL version: 18.6
> Operating system: Ubuntu 24.04.5 LTS, aarch64 (subscriber) / macOS 26.6.2,
> arm64 (publisher)
> Short description: Logical replication apply worker segfaults on a STREAM
> ABORT for a
> transaction that was never streamed (NULL
> stream_fileset)
> ================================================================================
>
>
> SUMMARY
> -------
> A logical replication apply worker segfaults dereferencing a NULL
> `MyLogicalRepWorker->stream_fileset` in subxact_info_read(), reached from
> apply_handle_stream_abort().
>
> The subscription has `streaming = off`, and I have confirmed from the
> publisher's
> pg_stat_activity that the START_REPLICATION command carries no `streaming`
> option at all.
> The publisher nevertheless sends a STREAM ABORT message. The apply worker
> has no streaming
> state for that xid, so the fileset is NULL and it crashes.
>
> Because the postmaster reinitialises the whole cluster when a background
> worker dies on
> SIGSEGV, this takes down every database on the subscriber, not just
> replication. And because
> the replication origin cannot advance past the offending record, the same
> message replays on
> every restart. In our case that was 484 cluster restarts over 8.5 hours,
> roughly one a minute,
> until we intervened.
>
>
> ENVIRONMENT
> -----------
> Publisher: PostgreSQL 18.6 (Homebrew) on aarch64-apple-darwin25.6.0, macOS
> 26.6.2
> Subscriber: PostgreSQL 18.6 (Ubuntu 18.6-1.pgdg24.04+2) on
> aarch64-unknown-linux-gnu,
> Ubuntu 24.04.5 LTS
> Both nodes are 18.6; there is no version skew.
>
> Output plugin: pgoutput. Slot two_phase = false, failover = false.
> Subscription: substream = 'f' (streaming = off), verified in
> pg_subscription.
> Publisher logical_decoding_work_mem was 64MB (the default) when this
> occurred.
>
> Disclosure of non-vanilla elements, since you will ask:
> - Both nodes have `timescaledb` in shared_preload_libraries.
> - However, the affected databases do NOT have the extension installed. The
> publisher-side
> source database (traydur_development) has only: plpgsql, pgcrypto,
> vector, postgres_fdw,
> pg_trgm, pg_stat_statements, btree_gist. The subscriber-side target has
> only `vector`.
> - I have not attempted a minimal reproduction on a build without
> timescaledb preloaded.
> See OPEN QUESTIONS below.
>
>
> BACKTRACE
> ---------
> Captured twice, from two separate crashes hours apart, with
> postgresql-18-dbgsym matching
> the running binary exactly. Identical both times, including the xid.
>
> #0 ChooseTablespace (name=0x... "14315996-615875138.subxacts.0",
> fileset=0x0)
> at src/backend/storage/file/fileset.c:190
> #1 FilePath (path=..., fileset=0x0, name=...
> "14315996-615875138.subxacts.0")
> at src/backend/storage/file/fileset.c:201
> #2 FileSetOpen (mode=0, name=... "14315996-615875138.subxacts.0",
> fileset=0x0)
> at src/backend/storage/file/fileset.c:119
> #3 BufFileOpenFileSet (fileset=0x0, name=...
> "14315996-615875138.subxacts", mode=0,
> missing_ok=true) at src/backend/storage/file/buffile.c:316
> #4 subxact_info_read (subid=<optimized out>, xid=615875138)
> at src/backend/replication/logical/worker.c:4185
> #5 stream_abort_internal (xid=615875138, subxid=615875139)
> at src/backend/replication/logical/worker.c:1789
> #6 apply_handle_stream_abort (s=0x...)
> at src/backend/replication/logical/worker.c:1876
> #7 apply_dispatch (s=0x...) at
> src/backend/replication/logical/worker.c:3452
> #8 LogicalRepApplyLoop (last_received=6035351251824)
> at src/backend/replication/logical/worker.c:3698
> #9 start_apply (origin_startpos=6035349339912)
> at src/backend/replication/logical/worker.c:4525
> #10 run_apply_worker () at src/backend/replication/logical/worker.c:4663
> #11 ApplyWorkerMain (main_arg=<optimized out>)
> at src/backend/replication/logical/worker.c:4839
>

This looks similar to the issue as bug #19616 [1]. It was fixed in
[2] and backpatched to REL_18_STABLE in [3]. The fix went in after
18.6 was released, so it will be in the next minor release (18.7).

[1] https://www.postgresql.org/message-id/19616-f6153af509910853@postgresql.org
[2] https://github.com/postgres/postgres/commit/aa4c52b808f76870f80189757af7217358544d60
[3] https://github.com/postgres/postgres/commit/caa463e3f8d2b9c1441e6a59c50f90b8598439f3

regards,
Ajin Cherian
Fujitsu Australia

In response to

Browse pgsql-bugs by date

  From Date Subject
Next Message Grigorev Jurij 2026-10-01 03:18:11 Re: BUG #19599: RestoreBlockImage: the decode cross-checks never bound hole_offset + hole_length against BLCKSZ
Previous Message shihao zhong 2026-10-01 03:04:19 Re: BUG #19705: One NaN box makes a BRIN box_inclusion_ops index omit unrelated rows