Memory leak in statext_ndistinct_build() during ANALYZE

From: Ilia Evdokimov <ilya(dot)evdokimov(at)tantorlabs(dot)com>
To: pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Memory leak in statext_ndistinct_build() during ANALYZE
Date: 2026-10-08 19:06:20
Message-ID: 20d6689e-f490-4ca7-a92f-e92f691ba07e@tantorlabs.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi hackers,

While testing ANALYZE with large statistics targets, I found that
building ndistinct extended statistics leaks memory proportional to the
sample size for every combination of columns.

Example:
postgresql.conf settings
```
default_statistics_target = 1000
```
Commands
```
SET client_min_messages = log;
SET log_statement_stats = on;
DROP TABLE IF EXISTS ndist_leak;
CREATE TABLE ndist_leak AS
SELECT (random() * 3000)::int AS a,
              (random() * 3000)::int AS b,
              (random() * 50)::int   AS c,
              (random() * 20)::int   AS d,
              (random() * 10)::int   AS e,
              (random() * 5)::int    AS f
FROM generate_series(1, 1000000);

CREATE STATISTICS ndist_leak_nd (ndistinct) ON a, b, c, d, e, f FROM
ndist_leak;

ANALYZE ndist_leak;
```

*Peak memory of the backend:*
Before patch - 971076 kB max resident size
After    patch - 112780 kB max resident size

ndistinct_for_combination() allocates the items, values and isnull
arrays for the whole sample, plus the MultiSortSupport, and never frees
them. statext_ndistinct_build() calls it once per combination of
columns, so these allocations accumulate until the whole ndistinct
object is built.

The attached patch changes statext_ndistinct_build() to call
ndistinct_for_combination() in a separate memory context and reset it
after each combination, the same way as is already done for
dependency_degree() in statext_dependencies_build().

Similar memory issues in extended statistics were fixed in [0]. The
original report there actually used ndistinct objects, but the fixes
addressed memory accumulated across statistics objects and in
dependency_degree(); the per-combination allocations in
ndistinct_for_combination() were not discussed there.

Any suggestions?

[0]:
https://www.postgresql.org/message-id/20210915200928.GP831%40telsasoft.com

--
Best regards,
Ilia Evdokimov,
Tantor Labs LLC,
https://tantorlabs.com/

Attachment Content-Type Size
v1-0001-Release-memory-allocated-by-ndistinct_for_combina.patch text/x-patch 3.0 KB

Browse pgsql-hackers by date

  From Date Subject
Next Message Tom Lane 2026-10-08 19:10:56 Re: Wrong results from a parameterized Append
Previous Message Yash Jadhav 2026-10-08 19:00:04 Re: [PATCH] Add memory/disk usage for Function Scan nodes in EXPLAIN