How to reproduce the observations on the ReadyForQuery 'rich' patch (v1-0001 / v2-0001; the behaviour is identical) =================================================================== All four scripts are attached as .txt so the commitfest bot ignores them; drop the "nocfbot-" prefix and the ".txt" suffix to use them. rfq-wire-client.py raw protocol v3 client, no libpq rfq-old-libpq-client.c libpq client to test an unpatched libpq rfq-pgbench.sh pgbench -S, three variants interleaved rfq-instructions.sh perf instructions:u per ReadyForQuery 1. Builds and servers --------------------- Two installs from the same master commit (5c74e122ef2), one with v1-0001 applied (it applies cleanly). Both built with ./configure --prefix= --enable-tap-tests --enable-depend \ CFLAGS="-O2 -fno-omit-frame-pointer" (no --enable-cassert). Then two servers on one Unix socket directory, trust auth, no TCP: SOCK=/tmp/rfq-sock; mkdir -p $SOCK BASE=; PATCHED= $BASE/bin/initdb -D /tmp/rfq-data/base -U postgres -A trust --no-sync $PATCHED/bin/initdb -D /tmp/rfq-data/patched -U postgres -A trust --no-sync OPTS="-k $SOCK -c listen_addresses='' -c max_prepared_transactions=5 -c shared_buffers=256MB" $BASE/bin/pg_ctl -D /tmp/rfq-data/base -l /tmp/rfq-data/base.log -o "-p 54501 $OPTS" start $PATCHED/bin/pg_ctl -D /tmp/rfq-data/patched -l /tmp/rfq-data/patched.log -o "-p 54502 $OPTS" start The scripts default to SOCK=/tmp/rfq-sock, BASE_PORT=54501 and PATCHED_PORT=54502. 2. Wire bytes in plain mode (rfq-wire-client.py capture) -------------------------------------------------------- python3 rfq-wire-client.py capture $SOCK 54501 > base.json python3 rfq-wire-client.py capture $SOCK 54502 > patched-plain.json Compare the hex of each message (BackendKeyData differs on every connection). What we saw: the 'Z' after startup, after a simple Query and after Sync is 5a 00000005 49 on both. The startup sequence of the patched server has one more ParameterStatus, ready_for_query_message=plain, because the GUC is GUC_REPORT. 3. What T, H, P and L report (rfq-wire-client.py script) --------------------------------------------------------- Each statement is sent as a simple Query; the output has the 'Z' that closes it, decoded. python3 rfq-wire-client.py script $SOCK 54502 '-c ready_for_query_message=rich' -- \ 'CREATE TEMP TABLE t(a int)' 'DROP TABLE t' 'DISCARD TEMP' 'DISCARD ALL' \ 'SELECT count(*) FROM pg_class WHERE relpersistence = $$t$$' T: it goes to 1 on CREATE TEMP TABLE and stays 1 after DROP TABLE, DISCARD TEMP and DISCARD ALL, with no temporary relation left. P: from a second session, prepare a transaction, then query from the first one: $PATCHED/bin/psql -h $SOCK -p 54502 -U postgres \ -c "BEGIN" -c "CREATE TABLE p(a int)" -c "PREPARE TRANSACTION 'x'" python3 rfq-wire-client.py script $SOCK 54502 '-c ready_for_query_message=rich' -- 'SELECT 1' $PATCHED/bin/psql -h $SOCK -p 54502 -U postgres -c "COMMIT PREPARED 'x'" The session that never prepared anything gets P=1 until the COMMIT PREPARED. L: compare it with the server function in the same session: python3 rfq-wire-client.py script $SOCK 54502 '-c ready_for_query_message=rich' -- \ 'CREATE TABLE l(a int)' 'SELECT pg_current_wal_insert_lsn()' The values are equal as LSNs, but the text differs: L is printed with %X/%X (0/17F5830) and pg_lsn with %X/%08X (0/017F5830). This only shows while the low half has fewer than 8 hex digits, i.e. on a fresh cluster. 4. An unpatched libpq against rich mode (rfq-old-libpq-client.c) ---------------------------------------------------------------- gcc -o old_client rfq-old-libpq-client.c -I$BASE/include -L$BASE/lib -lpq for lib in $BASE/lib $PATCHED/lib; do LD_LIBRARY_PATH=$lib ./old_client \ "host=$SOCK port=54502 user=postgres dbname=postgres options='-c ready_for_query_message=rich'" \ trace-$(basename $(dirname $lib)).txt done With the unpatched libpq the connection fails with: message contents do not agree with length in message type "Z". We saw the same with a stock PostgreSQL 18 libpq. With options set to plain, all steps pass with any libpq. Switching after connecting, since the GUC is PGC_USERSET: LD_LIBRARY_PATH=$BASE/lib timeout 8 $BASE/bin/psql -X -h $SOCK -p 54502 -U postgres \ -c 'select 1' -c 'set ready_for_query_message = rich' -c 'select 2' The SET is answered, then the same error appears and psql does not return (exit 124 from timeout). With the patched libpq the trace file shows "mismatched message length: consumed 5, expected 23" on every rich 'Z' (fe-trace.c is not changed by the patch); this is also why the "trace match" checks of libpq_pipeline fail when it runs with PGOPTIONS='-c ready_for_query_message=rich'. 5. Throughput (rfq-pgbench.sh) ------------------------------ PATCHED_BIN=$PATCHED/bin CPUS= bash rfq-pgbench.sh 10 10 8 pgbench -S -M prepared, 8 clients, 10 s per run, 10 rounds, the three variants rotated in every round, always with the patched pgbench. Prints the median tps per variant and the relative differences. 6. Instructions per ReadyForQuery (rfq-instructions.sh) ------------------------------------------------------- PATCHED_BIN=$PATCHED/bin bash rfq-instructions.sh 5 Counts instructions:u of a single backend for 10000 and for 30000 'SELECT 1', and divides the difference by 20000. Cases: a clean session, and 100 or 1000 cursors without HOLD open in the transaction. Our numbers (identical over two separate runs): clean plain 29114 rich 30828 (+1714, +5.9%) 100 cursors plain 26115 rich 31703 (+5588, +21.4%) 1000 cursors plain 26112 rich 67815 (+41703, +159.7%) Plain minus master is +2 instructions per ReadyForQuery. The rich cost grows with the number of open portals, about 40 instructions per portal, which matches HasActiveWithHoldCursors() walking the whole portal hash table on every ReadyForQuery.