| From: | Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com> |
|---|---|
| To: | Manu <manuelreyesbravo(at)gmail(dot)com> |
| Cc: | pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Subject: | Re: doc: Document Linux cgroup memory limits |
| Date: | 2026-10-04 21:15:59 |
| Message-ID: | CABH8dKzcgQtrP=z6WdGY-JWOVP6vzK1-Bkk+K1pUCfq6xOD48Q@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
Here is v2, with these changes based on Manu's testing:
- the huge pages accounting now depends on memory_hugetlb_accounting
instead of being described as off by default
- a new paragraph covers memory.oom.group, where the whole cgroup
(postmaster included) is killed and nothing is logged
Thanks again for the review.
Regards,
João Marcelo
Em sáb., 3 de out. de 2026 às 00:08, Manu <manuelreyesbravo(at)gmail(dot)com>
escreveu:
> Hi João,
>
> Thanks for writing this up. I tried it on a cgroup v2 host (kernel
> 7.2, systemd 262), running master in a systemd scope with
> MemoryMax=400M and no swap.
>
> The main behaviour checks out. With four sessions sorting with
> work_mem = 300MB, one process was SIGKILLed, the postmaster
> reinitialized, every session lost its connection, and nobody got an
> out-of-memory error. Reading a 638 MB table twice took usage up to
> memory.max without any kill, since the page cache was reclaimed.
>
> Two things I would change:
>
> > Memory allocated from the huge page pool (see huge_pages) is by
> > default not counted towards the memory usage of the cgroup.
>
> That is the kernel default, but systemd 259 and later mount cgroup2
> with memory_hugetlb_accounting [1], and then it is counted. Here,
> with shared_buffers = 256MB and huge_pages = on, memory.stat showed
> 290MB under hugetlb. Perhaps say that whether it counts depends on
> that mount option.
>
> > If that is a server process other than the postmaster, the
> > postmaster treats it as a crash
>
> With memory.oom.group set to 1, the OOM killer kills every process in
> the cgroup, the postmaster included. There is no crash recovery then,
> and nothing reaches the server log: the last line was "database system
> is ready to accept connections". Kubernetes sets memory.oom.group for
> containers on cgroup v2 unless singleProcessOOMKill is enabled [2], so
> this is probably the case most container users will see. A sentence
> about it would help.
>
> One observation, maybe not for the docs: the first process killed was
> not one of the sorting backends but an io worker, with about 90MB of
> shared buffers mapped and almost no private memory. The OOM killer
> counts mapped shared memory, so the server log named the io worker
> rather than the sessions that used the memory.
>
> [1] https://github.com/systemd/systemd/blob/main/NEWS (CHANGES WITH 259)
> [2]
> https://kubernetes.io/docs/reference/config-api/kubelet-config.v1beta1/
> (singleProcessOOMKill)
>
> Regards,
> Manu
>
| Attachment | Content-Type | Size |
|---|---|---|
| v2-0001-doc-Document-Linux-cgroup-memory-limits.patch | application/octet-stream | 4.1 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Joao Detomini | 2026-10-04 21:22:20 | Re: Exposing the cgroup memory limit to SQL? |
| Previous Message | Manu | 2026-10-04 20:54:51 | Re: Planning time quadratic in the IN-list length for "c = X AND (a, b) IN (...)" with BitmapOr |