From e26973abdd64af8f5044ec18dd7d6ff3aa6d62ef Mon Sep 17 00:00:00 2001 From: Joao Detomini Date: Thu, 1 Oct 2026 19:14:27 -0300 Subject: [PATCH v5] doc: Document Linux cgroup memory limits The section on Linux memory overcommit does not cover memory limits imposed by control groups, which are commonly used for containers. Such limits are enforced independently of vm.overcommit_memory, and the OOM killer is invoked within the cgroup instead of allocations failing. Document this behavior and its consequences for the postmaster, including the case where memory.oom.group is set and the whole cgroup, postmaster included, is killed without crash recovery. Also document what counts towards the cgroup's memory usage, and that the accounting of huge pages depends on the memory_hugetlb_accounting mount option. --- doc/src/sgml/runtime.sgml | 69 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 69 insertions(+) diff --git a/doc/src/sgml/runtime.sgml b/doc/src/sgml/runtime.sgml index d9984910cc..41796b2238 100644 --- a/doc/src/sgml/runtime.sgml +++ b/doc/src/sgml/runtime.sgml @@ -1438,6 +1438,75 @@ export PG_OOM_ADJUST_VALUE=0 + + Linux Control Group Memory Limits + + + control group + + + + cgroup + + + + On Linux, the memory used by a group of processes can be limited using + control groups (cgroups), a mechanism commonly + used to limit the memory of containers. Such a limit is enforced + regardless of the amount of memory available on the host and of + host-wide settings such as vm.overcommit_memory (see + ). + + + + With cgroup version 2, the limit is set in the + memory.max file. (Version 1 uses different files and + is not covered here.) When memory usage reaches the limit and cannot be + reduced by reclaiming memory, the kernel's OOM killer is invoked within + the cgroup. In this case memory allocations usually do not fail with an + out-of-memory error; instead, the OOM killer terminates a process of the + cgroup. If that is a server process other than the postmaster, the + postmaster treats it as a crash and, unless + is turned off, restarts all + server processes, disconnecting all sessions. If the server runs as a + systemd service, the service's + OOMPolicy setting also applies; its default, + stop, stops the whole service instead. + + + + If memory.oom.group is set for the cgroup, which + container runtimes often do (recent versions of Kubernetes do so by + default), the OOM killer instead terminates all processes of the cgroup + except those with an oom_score_adj of -1000, so the + postmaster too, unless it is protected as described in + . If the postmaster is + terminated, the server is not restarted automatically, and it does not log + the termination. For details see the kernel documentation file + . + + + + The shared memory used by PostgreSQL + (including the shared buffers) and the private memory of each server + process count towards the memory usage of the cgroup. Therefore, + , plus the memory that concurrent + sessions might use (see , + , + , and + ), should be chosen with the limit + in mind. Since the total memory used can be many times the value of + work_mem, it is advisable to leave a safety margin + below the limit. Whether memory allocated from the huge + page pool (see ) counts towards the memory + usage of the cgroup depends on the + memory_hugetlb_accounting mount option of the cgroup2 + file system. The kernel does not enable it by default, but recent versions + of systemd do. + + + + Linux Huge Pages -- 2.50.1 (Apple Git-155)