Re: Protocol Compression (fourth attempt)

From: Anthonin Bonnefoy <anthonin(dot)bonnefoy(at)datadoghq(dot)com>
To: Andrey Borodin <x4mmm(at)yandex-team(dot)ru>
Cc: PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org>
Subject: Re: Protocol Compression (fourth attempt)
Date: 2026-09-30 08:48:10
Message-ID: CAO6_XqrhOnpon4pg+aceN_-Z+qOAW_Xp1DsmgDUQz6+wiPMwBA@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Wed, Sep 30, 2026 at 9:53 AM Andrey Borodin <x4mmm(at)yandex-team(dot)ru> wrote:
> Would you mind if I prepared that revision of your patchset? It would
> help me understand your implementation better and give us a concrete
> starting point for combining our work. I'd focus on the simplifications
> we agree on.
>
> If you've already started on v2, let me know so we can split the work.

Sounds good to me, I haven't started the v2 yet.

> On enabling and disabling compression within a session, I agree that
> PQcommMethods makes the switch simple. But a USERSET GUC must handle
> SET LOCAL and rollback while output is buffered or a frame is open.
> Could we keep the choice at connection startup for v1? Ordinary messages
> would still be allowed on a compressed connection, so we could add
> session-level control later without changing the wire format. Is there a
> use case where choosing at startup would not be enough?

That's currently handled by flushing + pq_send_messages which closes
the frame and sends the buffered messages. I imagine that on some
workloads (table with mostly random data), compression would have
mostly negative effects and users may want to disable it for specific
queries. For a v1, that's fine to leave this out.

> On thresholds and batch sizes, I agree that the knobs are useful for
> experiments. Your first-packet timings show that the flush policy needs
> work. I'd first try to address that internally, for example by bounding
> the amount of uncompressed input processed before flushing output.
> Publishing these thresholds as GUCs would mean supporting their
> semantics as we change the buffering policy. Could we keep them in a
> benchmarking patch for now, and add public controls if measurements
> show a trade-off that users need to choose themselves?

Sounds reasonable to me. Using the input size to trigger flush was
something I had in mind, but that wouldn't be useful if the messages
are incomplete as the client won't start processing the message until
it is fully available. Maybe a combination of both (limited input size
with at least x full messages) could work? Adding it and benchmarking
it would definitely help to see the impact.

Regards,
Anthonin Bonnefoy

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Hannu Krosing 2026-09-30 09:07:46 Re: Direct TOAST v2, faster, smaller and no migration needed
Previous Message Nazir Bilal Yavuz 2026-09-30 08:43:42 Re: aio: Async fsyncs for crash recovery and checkpointer