| From: | Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com> |
|---|---|
| To: | Jakub Wartak <jakub(dot)wartak(at)enterprisedb(dot)com> |
| Cc: | Manu <manuelreyesbravo(at)gmail(dot)com>, pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Subject: | Re: doc: Document Linux cgroup memory limits |
| Date: | 2026-10-05 12:18:56 |
| Message-ID: | CABH8dKwvPwBtUAdHrCPJKvw05Yk25JAY=NC7xU_9BrqSRpJzsA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi Jakub,
Thanks for looking at it. You're right that the last paragraph read
as if choosing those settings is enough to stay under the limit, and
that isn't true: the total can be many times work_mem. Attached v5
softens that sentence and recommends leaving a safety margin below
the limit, the same way the work_mem docs put it.
Overcommit inside cgroups and RLIMIT_DATA seem like a separate
discussion from this docs patch, so I left them out.
Changes since v4: only the last paragraph of the section.
Thanks,
João Marcelo
Em seg., 5 de out. de 2026 às 03:37, Jakub Wartak <
jakub(dot)wartak(at)enterprisedb(dot)com> escreveu:
> On Mon, Oct 5, 2026 at 4:17 AM Joao Detomini
> <joao(dot)detomini(at)enterprisedb(dot)com> wrote:
> >
> > Hi Manu,
> >
> > Thanks for the quick response. Attached v4 with both points: the systemd
> > OOMPolicy note, and the oom_score_adj -1000 exemption for the group
> > kill, with a pointer to the overcommit section. I also resolved the
> ambiguity
> > that arose after this new insertion.
>
> +1 to the documenting this, but it's more than just set couple of GUCs
> to avoid OOM:
>
> We are kind of weak on self-limiting memory usage (there's new leak every
> now and then, and people are complaining about work_mem not working). IMHO
> running without vm.overcommit_memory=2 and without rotating connections
> every couple hours/days [to release leaks/fat from free() caching by e.g.
> glibc [1][2]] is basically prolonged outage sooner or later before one hits
> suchleaks, dunno what people do with cgroups2 do to avoid this. So to me
> it's
> some kind of risky to run PostgreSQL on cgroups2 in the first place. IMHO
> kernel should be enhanced to implement overcommit_memory=2 within cgroups,
> but I understand that stuff that some stuff is implemented during
> page-fault
> handling (and not memory allocation - the brk() itself), so it is
> impossible?
>
> The patch mentions some GUC, but do in my experience, that's not the whole
> truth, there's plenty of other things: number of relations involved, libc/
> malloc implementation, and so on.
>
> Two years ago I was thinking we could simply - in theory - issue
> setrlimit(RLIMIT_DATA) (per process to limit each backend to strict X MB),
> but AFAIR(?) Andres told me it is not really usable there as it may cause
> SIGSEGVs in case of hitting the limit when growing stack and not just
> overallocating memory (if I recall correctly). Or maybe I misunderstood and
> I couldn't reproducce this on modern kernels.
>
> -J.
>
> [1] -
> https://www.postgresql.org/message-id/CA%2BT%3D_GV-igvf8gLRK4w22aN1qB6F1zgHD9dxnEHtK%3DXEoTTPEA%40mail.gmail.com
> [2] -
> https://www.postgresql.org/message-id/flat/3424675.QJadu78ljV%40aivenlaptop
>
| Attachment | Content-Type | Size |
|---|---|---|
| v5-0001-doc-Document-Linux-cgroup-memory-limits.patch | application/octet-stream | 4.6 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | David Rowley | 2026-10-05 12:23:43 | Re: tuplesort_putdatum() does not account for tuple memory |
| Previous Message | Andrei Lepikhov | 2026-10-05 12:04:51 | Re: hashjoins vs. Bloom filters (yet again) |