Re: injection_points: canceled or terminated waiters leak their wait slots

From: Zsolt Parragi <zsolt(dot)parragi(at)percona(dot)com>
To: Michael Paquier <michael(at)paquier(dot)xyz>
Cc: Nikolay Samokhvalov <nik(at)postgres(dot)ai>, pgsql-hackers(at)lists(dot)postgresql(dot)org, Andrey Borodin <x4mmm(at)yandex-team(dot)ru>
Subject: Re: injection_points: canceled or terminated waiters leak their wait slots
Date: 2026-09-29 00:38:56
Message-ID: CAN4CZFM4iAESOc08pNiA87nboNP35Nj_7z7eXvfS1-g=08KZ7A@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

> How exactly?

If we assume the linked thread about windows skipping the last message
is correct, we can simulate that by adding an option not to send the
fatal message to the client:

diff --git a/src/backend/utils/error/elog.c b/src/backend/utils/error/elog.c
index b9d2c96b97a..e043717b943 100644
--- a/src/backend/utils/error/elog.c
+++ b/src/backend/utils/error/elog.c
@@ -1924,7 +1924,8 @@ EmitErrorReport(void)
send_message_to_server_log(edata);

/* Send to client, if enabled */
- if (edata->output_to_client)
+ if (edata->output_to_client &&
+ !(edata->elevel == FATAL && getenv("PG_TEST_DROP_FATAL")))
send_message_to_frontend(edata);

MemoryContextSwitchTo(oldcontext);

On windows, it seems it can randomly happen, but it's very unlikely

> TBH, I am not completely sure what you are proposing here. v1-0001
> from Andrey and AI-generated-not-reviewed 0003 from Nikolay step on
> each other

0001+0002 is enough to fix the existing reports. I think
temp-schema-cleanup could also use an alternative output, as it uses
the same construction and fails with the above modification.

0003+0004 seems to fix a blocking issue Nikolay mentioned, but I don't
think that's currently required by any of the existing test cases, it
seems more like an AI-discovered scenario.

One thing I realized after sending that message is that if we accept
the limitation in this comment:

+ /*
+ * Save the error before PQgetResult() adds another complaint about
+ * attempting to read from the dead socket.
+ */
+ connection_error = pg_strdup(PQerrorMessage(conn));

Which means an additional "invalid socket" message in the output,
0001+0003 together becomes 3 simple condition changes. See the
attached patch, the only difference is that there's one more extra
line in the alternative test outputs, but now the code change is
simpler, so this might be actually better. (I also validated Nikolay's
0004 test case with this, but I didn't include it in the patch)

This still need actual windows testing, I didn't do that part yet.

Attachment Content-Type Size
v2-0001-isolationtester-Report-lost-connection-as-step-re.patch application/octet-stream 12.2 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Previous Message Amit Kapila 2026-09-29 00:35:06 Re: Fix "unexpected logical decoding status change" error; from concurrent logical decoding activation