Re: doc: Document Linux cgroup memory limits

From: Jakub Wartak <jakub(dot)wartak(at)enterprisedb(dot)com>
To: Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com>
Cc: Manu <manuelreyesbravo(at)gmail(dot)com>, pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: doc: Document Linux cgroup memory limits
Date: 2026-10-05 06:37:28
Message-ID: CAKZiRmxJRL_UKF4fxtj8u0kem5xoKpcF=Mv2v8++_GbaUpzpcQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Mon, Oct 5, 2026 at 4:17 AM Joao Detomini
<joao(dot)detomini(at)enterprisedb(dot)com> wrote:
>
> Hi Manu,
>
> Thanks for the quick response. Attached v4 with both points: the systemd
> OOMPolicy note, and the oom_score_adj -1000 exemption for the group
> kill, with a pointer to the overcommit section. I also resolved the ambiguity
> that arose after this new insertion.

+1 to the documenting this, but it's more than just set couple of GUCs
to avoid OOM:

We are kind of weak on self-limiting memory usage (there's new leak every
now and then, and people are complaining about work_mem not working). IMHO
running without vm.overcommit_memory=2 and without rotating connections
every couple hours/days [to release leaks/fat from free() caching by e.g.
glibc [1][2]] is basically prolonged outage sooner or later before one hits
suchleaks, dunno what people do with cgroups2 do to avoid this. So to me it's
some kind of risky to run PostgreSQL on cgroups2 in the first place. IMHO
kernel should be enhanced to implement overcommit_memory=2 within cgroups,
but I understand that stuff that some stuff is implemented during page-fault
handling (and not memory allocation - the brk() itself), so it is impossible?

The patch mentions some GUC, but do in my experience, that's not the whole
truth, there's plenty of other things: number of relations involved, libc/
malloc implementation, and so on.

Two years ago I was thinking we could simply - in theory - issue
setrlimit(RLIMIT_DATA) (per process to limit each backend to strict X MB),
but AFAIR(?) Andres told me it is not really usable there as it may cause
SIGSEGVs in case of hitting the limit when growing stack and not just
overallocating memory (if I recall correctly). Or maybe I misunderstood and
I couldn't reproducce this on modern kernels.

-J.

[1] - https://www.postgresql.org/message-id/CA%2BT%3D_GV-igvf8gLRK4w22aN1qB6F1zgHD9dxnEHtK%3DXEoTTPEA%40mail.gmail.com
[2] - https://www.postgresql.org/message-id/flat/3424675.QJadu78ljV%40aivenlaptop

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Michael Paquier 2026-10-05 06:49:16 Re: pgstat: allow a stats kind to use its own dedicated dsa/dshash
Previous Message Andrey Borodin 2026-10-05 06:32:09 Re: Check for non-deterministic FK collations before upgrade