Re: doc: Document Linux cgroup memory limits

From: Manu <manuelreyesbravo(at)gmail(dot)com>
To: Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com>
Cc: pgsql-hackers(at)lists(dot)postgresql(dot)org
Subject: Re: doc: Document Linux cgroup memory limits
Date: 2026-10-05 01:19:22
Message-ID: 179116316208.1696871.11500242683868085248@gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi João,

Thanks for v3. The new sentence is right, and the docs build. Going
through the whole section once more, two cases behave differently
from what it says; I should have raised them with v2.

> If that is a server process other than the postmaster, the
> postmaster treats it as a crash and, unless restart_after_crash is
> turned off, restarts all server processes, disconnecting all sessions.

Not when the server runs as a systemd service, which is a common way
to set a memory limit outside containers. systemd's OOMPolicy applies
too, and its default, stop, stops the whole service after an OOM kill.
With MemoryMax=400M and no OOMPolicy set, two of three runs ended with
"received smart shutdown request" and the unit in the oom-kill failed
state, and the third was already deactivating; with
OOMPolicy=continue, the postmaster restarted all three times. Perhaps
add:

If the server runs as a systemd service, systemd's OOMPolicy setting
also applies; its default, stop, stops the whole service instead.

> If memory.oom.group is set for the cgroup, [...] the OOM killer
> instead terminates all processes of the cgroup, including the
> postmaster.

The kernel exempts processes whose oom_score_adj is -1000 from the
group kill, and that is what linux-memory-overcommit recommends for
the postmaster. With the postmaster at -1000 and memory.oom.group set
to 1, the group kill happened but the postmaster survived and
restarted the other processes, in three runs out of three. Inside a
container the postmaster usually cannot lower its own score, so the
Kubernetes case stays as you describe. Perhaps:

... the OOM killer instead terminates all processes of the cgroup
except those with an oom_score_adj of -1000, so the postmaster too,
unless it is protected as described in
<xref linkend="linux-memory-overcommit"/>.

Regards,
Manu

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Scott Ray 2026-10-05 01:41:56 Re: pg_xmin_horizon: a system view of everything pinning the xmin horizon
Previous Message David Rowley 2026-10-05 00:46:26 Re: Material node can report incorrect "Maximum Storage" in EXPLAIN