Re: Protocol Compression (fourth attempt)

From: Anthonin Bonnefoy <anthonin(dot)bonnefoy(at)datadoghq(dot)com>
To: Andrey Borodin <x4mmm(at)yandex-team(dot)ru>
Cc: PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org>
Subject: Re: Protocol Compression (fourth attempt)
Date: 2026-09-29 15:37:46
Message-ID: CAO6_XqrJOSkmWakAP0dHS071NULFFteFcytOnbh=tLqNeGAy-g@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, Sep 29, 2026 at 4:06 PM Andrey Borodin <x4mmm(at)yandex-team(dot)ru> wrote:
> Could we combine our efforts on this?

Yeah definitely. Sorry I've missed your proposal.

> Thanks, Michael, for linking my proposal. My goal is a minimal,
> future-proof design: a small useful protocol and configuration
> interface, with a clear way to negotiate extensions later. My takeaway
> from the previous attempts is that expanding scope kept us from
> agreeing on that core.

I had the same take.

> The separate msg_buffer avoids copying decompressed messages
> back into the input buffer, as my prototype does.

I had a similar approach initially, but that definitely required a lot
of memmove. I'm still not 100% happy about the frontend and feel like
there's still some possible simplification that could be done.
You do have something in your prototype that I'm not doing: shrinking
the buffer once a large result has been processed.

> I would also like to keep your trace-based tests for compressed
> messages and frame boundaries.

+1

> There are several choices I would simplify for v1:
>
> - Multiple codecs, levels and long-distance matching. I would start
> with Zstandard at its default level.

Sounds reasonable. I've mostly added LZ4 to validate that the
interface works fine with another codec.

> - Switching codecs within a session and identifying them in every
> wrapper. Choosing once at startup avoids those state transitions.

Yeah, codec switching definitely adds a fair share of complexity and
tricky edge cases. I would still keep the capacity of enabling and
disabling compression freely in a session, and a session will be
locked to the codec used the first time. The use of the PqCommMethods
layer makes enabling and disabling compression straightforward.

> - Compressing additional message types and listing their types in the
> wrapper. Starting with DataRow and CopyData leaves other messages
> visible to poolers without that extra metadata.

The message types were definitely for the poolers, I didn't use them
in libpq. If we only compress messages that can be ignored by poolers,
then they can definitely be removed.

> - GUCs for thresholds and batch sizes. We can improve the buffering
> policy while keeping these choices internal.

The GUCs were there as a way to be able to test different values since
I had no idea what would be a good value. Also, it's likely the
defaults may be inappropriate for some workloads, so I feel like
leaving the option to tune them would help.

> - Configurable frame lifetime. I agree with your suggestion to remove
> that option and settle on a mandatory reset rule for poolers.

+1

> I am open to using your patchset as the base and reducing its scope, or
> taking your implementation ideas into mine. We can choose the base
> once we agree on the common design.

Given you sent your proposal earlier, you have precedence, so I leave
you the choice :). I could start by implementing the simplifications
you've mentioned to get a v2 on my patch to make it easier to agree
and merge the approach.

Thank you!
Anthonin Bonnefoy

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Bharath Rupireddy 2026-09-29 15:45:17 Re: parallel autovacuum: Propagate track_cost_delay_timing to parallel workers
Previous Message Vik Fearing 2026-09-29 15:20:11 Re: Logical Implication