| From: | Anthonin Bonnefoy <anthonin(dot)bonnefoy(at)datadoghq(dot)com> |
|---|---|
| To: | PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org> |
| Subject: | Protocol Compression (fourth attempt) |
| Date: | 2026-09-29 09:30:51 |
| Message-ID: | CAO6_Xqr-kr5=9nh_PvyKHzw5f82M8C-2_NJ84P+s4uLDgyKZQQ@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi all,
Here's a fourth attempt at implementing protocol compression (past
attempts: [1], [2], [3]). I've read through the past threads and tried
to address the concerns, however, the threads are long so it's very
likely I've missed or forgot about some details in the discussion.
I've also taken a different approach for the implementation, so I
haven't reused code from the previous attempts.
* Backend implementation (src/backend/libpq/pqcomm_compress{,_zstd,_lz4}) *
A new protocol message, CompressedMessages, is introduced and contains:
- The compression algorithm used (lz4 or zstd).
- The message types contained in the payload. For example, 'TDD' means
the CompressedMessages contains 1 RowDescription and 2 DataRows
- The compressed payload containing 0 to n messages
I've tried to fit the compression logic as much as possible in the
available interfaces, and PQcommMethods is a surprisingly good fit to
handle streaming compression. On a very high level, we have:
PqCommMethods->putmessage() adds the new message to the ongoing
compression buffer with 'ZSTD_compressStream2(..., ZSTD_e_continue)'
PqCommMethods->flush() calls 'ZSTD_compressStream2(...,
ZSTD_e_flush)' and sends the compressed payload in a new
CompressedMessages.
Enabling compression can be done by modifying PqCommMethods from:
PqCommMethods -> PqCommSocketMethods
To
PqCommMethods -> PqCompressMethods -> PqCommSocketMethods
The previous PqCommMethods (which is likely PqCommSocketMethods) is
saved and used to send the compressed (or uncompressed) messages.
* Backend memory footprint *
As streaming is used, the memory usage on the backend is capped.
- cctx: up to 1.3MB with default compression level
- outBuf: 131KB (set to ZSTD_CStreamOutSize())
When the output buffer is full, the content is sent in a new
CompressedMessages, allowing to clear the output buffer and continue
compression. In this case, the last message was only partially
compressed and will need additional payloads to be rebuilt.
* Compression Activation *
Compression is controlled through multiple GUCs:
protocol_backend_compression_allowed_algorithms: List of compression
algorithms allowed (and supported) by the server.
protocol_backend_compression: Compression algorithm to use
protocol_backend_compression_threshold: Minimum threshold needed to
trigger compression, currently set at 100 bytes
protocol_backend_compression_number_messages: Number of messages per
CompressedMessages before sending it. Currently 0 (CompressedMessages
is sent if the buffer is full or a flush is called).
It's possible for a client to freely enable/disable/switch compression
algorithms within a session. The server can forbid compression by
setting protocol_backend_compression_allowed_algorithms='none'. As
compression only starts after GUCs are processed, startup and
authentication packets are not compressed.
For example, a session can be started with compression enabled and
protocol tracing with:
PQTRACE="-" psql "options='-cprotocol_backend_compression=zstd
-cprotocol_backend_compression_threshold=0'" -c"select 1"
It will yield the following protocol trace:
F 14 Query "SELECT 1;"
B* 33 RowDescription 1 "?column?" 0 0 23 4 -1 0
B* 11 DataRow 1 1 '1'
B 63 CompressedMessages \x03 "TD" '\x28\xb5\x2f\xfd\x00...'
B 13 CommandComplete "SELECT 1"
B 5 ReadyForQuery I
Protocol tracing was modified to prefix messages nested in a
CompressedMessages with B*. Here, RowDescription and DataRow are both
part of the compressed payload. The CompressedMessages itself is
displayed and consumed only after all nested messages are processed.
* Proxy compatibility *
From checking pgbouncer's code, it only needs to inspect the content
of a subset of messages: ReadyForQuery, ErrorResponse,
ParameterStatus, CommandComplete. For other messages, just knowing
their types seems to be enough. Thus, the messages that need to be
inspected are never compressed (plus they don't benefit much from
compression given how small they are). With that, it should be
possible for pgbouncer to forward CompressedMessages without having to
decompress them (assuming the client supports the compression
algorithm).
This does require compression frames to be independent between
transactions, otherwise the client won't be able to decompress the
payload if some part of the compression frame was sent to a different
client. This can be achieved with
'protocol_backend_compression_transaction_frame'. With this GUC
enabled, the compression frame will be closed with ZSTD_e_end (or the
LZ4 equivalent) when a transaction ends.
With protocol_backend_compression_transaction_frame enabled, the
previous protocol trace becomes:
PQTRACE="-" psql "options='-cprotocol_backend_compression=zstd
-cprotocol_backend_compression_threshold=0
-cprotocol_backend_compression_transaction_frame=true'" -c"select 1"
F 13 Query "select 1"
B* 33 RowDescription 1 "?column?" 0 0 23 4 -1 0
B* 11 DataRow 1 1 '1'
B 63 CompressedMessages \x03 "TD" '\x28\xb5\x2f\xfd\x00...'
B 13 CommandComplete "SELECT 1"
B 9 CompressedMessages \x03 "" '\x01\x00\x00'
B 5 ReadyForQuery I
There's an additional CompressedMessages with no messages and a
'\x01\x00\x00' content: this is the end of a zstd frame. With this,
each transaction has a dedicated compression frame, making them
independent of each other.
* Frontend Implementation src/interfaces/libpq/fe-compress{,-zstd,-lz4}.c *
This time, I couldn't fit well in the available code as the existing
code consumes from PGConn's inBuffer in three different places:
pqParseInput3, getCopyMessageData, pqFunctionCall3 and all pqGet*
functions. To reduce the amount of direct access to inBuffer, I've
introduced a new msg_buffer struct that's passed to the pqGet*
functions.
To handle CompressedMessages in the frontend, a new decompressBuffer
msg_buffer was added to the connection which will be used for
decompression. pqParseInput3/getCopyMessageData/pqFunctionCall3 will
use decompressBuffer if there's at least one full message available,
otherwise, inBuffer is used.
* Protocol negotiation *
While the server can advertise the allowed and supported algorithms
through protocol_backend_compression_allowed_algorithms, the server
should only send compressed messages if the client actually supports
it (or at least claims to support it). This is done with the client
sending the supported algorithms through _pq_.supported_compressions
startup option. The server saves the information
port->supported_compress_{zstd,lz4} and will only accept the client
setting protocol_backend_compression to the algorithm it supports.
Given that this is more a protocol extension, the backend only sends
the new CompressedMessages if the client sent the
_pq_.supported_compressions list, I don't think a protocol version
bump is needed (though I'm not sure what are the rules here).
* Tests *
Protocol tracing provided by fe-trace is incredibly valuable to check
the new CompressedMessages. However, there's currently only one module
that relies on it, libpq_pipeline, using dedicated code to enable
protocol tracing. I've built a new trace test harness that relies on
an ongoing PQTRACE patch[4] which is very similar to regress: SQL
scripts are executed, and protocol trace is captured along the output,
and checked against the expected values. This allows tests to check
things like if a message is compressed as expected or if a compression
frame is correctly closed.
* Replication Compression *
Since WAL replication is sent through backend messages, protocol
compression can be applied to the WAL stream by modifying conninfo:
primary_conninfo = '... options=''-cprotocol_backend_compression=zstd'' ... '
* Some Benchmarks *
I've tested running 3 different queries while capturing the network
footprint with tcpdump:
- select_all: SELECT * FROM pgbench_accounts (100 times)
- select_100: SELECT * FROM pgbench_accounts LIMIT 100 (20000 times)
- select_random: SELECT gen_random_uuid() FROM generate_series(1,
50000) (100 times)
With different compression parameters:
- No compression/zstd/lz4
- tx frame enabled/disabled
- 1,2,3,2000 and default number messages per CompressedMessages
The results are available in benchmarks.csv. I've also attached graphs
showing the latency and network size specifically for the select_all
query.
All 3 queries have similar results, but looking specifically at select_all test:
- There's an increased latency (at least for zstd, 55ms -> 70ms) when
using the default options (CompressedMessages is only sent when the
output buffer is full or during a flush). Since the rows are heavily
compressible, everything is sent in a single CompressedMessage and
pcap shows the server sends the first TCP packet after 41ms. Meaning
the client is idle during those first 41ms and then has to process
everything.
- Limiting the number of messages per CompressedMessage can mitigate
this behaviour. At 2K messages/batch, zstd even has a similar latency
to uncompressed.
- Batching messages helps a lot with compression and latency. zstd
compresses from 1GB to 18MB with >2K messages/batch. Given that a
ZSTD_e_flush closes a zstd block, having only 1 msg/batch prevents
zstd from finding patterns between rows (like the same bid and
abalance values).
- It doesn't seem like there's much difference between enabling or
disabling Transaction Frame. So maybe it would be better to make this
the default and remove the option? Though there might be some use
cases of keeping the frame opened with a high level of compression and
long distance enabled.
With the random query, sending only 1 message per message leads to an
increase in wire size with the additional CompressedMessages bytes and
compression frame overhead. With >2K msgs/batch, we still manage to
reduce the size from 227MB to 105MB.
Current limitations:
- The patch only implements lz4 and zstd compression. gzip can be added later
- Only backend messages can be compressed for now. There's nothing
preventing the use of CompressedMessages for the frontend, but I
wanted to keep the patch focused, and limiting it to backend
compression sounded like a good stopping point.
- I haven't written the related doc yet
- protocol_backend_compression_allowed_algorithms only limit the
algorithms used. Maybe we want to also limit the compression level?
The patchset has the following files:
001 to 006: Those are from the PQTRACE[4] patch. I've joined them
since they are needed to run the tests.
007: A small change to libpq to hide notifications so LISTEN can be
used in regression tests
008: This creates a new trace test harness. It's very similar to the
regress harness, except it also captures and compare the protocol
traces using PQTRACE
009: Make parse_compress_options available to the backend. The goal is
to have the protocol_backend_compression GUC reuse the same
METHOD:DETAIL format.
010: Add backend support for CompressedMessages. This creates the new
PQcommMethods with the zstd and lz4 implementations.
011: Create a new msg_buffer struct to store PGconn's inBuffer and
pass it to pqGet* functions.
012: Frontend support for decompressing CompressedMessages.
013: Add CompressedMessages to protocol tracing and tag compressed
message with 'B*'
014: This adds the protocol negotiation with the client sending
_pq_.supported_compressions and the backend parsing it and storing the
zstd/lz4 support.
015 to 016: The regression and trace tests
[1]: https://www.postgresql.org/message-id/aad16e41-b3f9-e89d-fa57-fb4c694bec25@postgrespro.ru
[2]: https://www.postgresql.org/message-id/flat/ABAA09C6-BB95-47A5-890D-90353533F9AC%40yandex-team.ru
[3]: https://www.postgresql.org/message-id/flat/CACzsqT4cJG0kaCbz24Sd%3DGAEgiQDpzU8yuD6vF25zo870%2B3M6g%40mail.gmail.com
[4]: https://www.postgresql.org/message-id/CAO6_XqrOuYe8J1vOuErx0XarkF3e96fCO4fAHWOK-6amJ_UCew@mail.gmail.com
| From | Date | Subject | |
|---|---|---|---|
| Next Message | wenhui qiu | 2026-09-29 09:39:02 | Re: [PATCH] Reduce LWLockWaitListLock() cache-line contention with adaptive spin reads |
| Previous Message | Chao Li | 2026-09-29 09:18:10 | Re: Silence -fsanitize=function where we cast function pointers on purpose |