| From: | Ilia Evdokimov <ilya(dot)evdokimov(at)tantorlabs(dot)com> |
|---|---|
| To: | pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Memory leak in statext_ndistinct_build() during ANALYZE |
| Date: | 2026-10-08 19:06:20 |
| Message-ID: | 20d6689e-f490-4ca7-a92f-e92f691ba07e@tantorlabs.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi hackers,
While testing ANALYZE with large statistics targets, I found that
building ndistinct extended statistics leaks memory proportional to the
sample size for every combination of columns.
Example:
postgresql.conf settings
```
default_statistics_target = 1000
```
Commands
```
SET client_min_messages = log;
SET log_statement_stats = on;
DROP TABLE IF EXISTS ndist_leak;
CREATE TABLE ndist_leak AS
SELECT (random() * 3000)::int AS a,
(random() * 3000)::int AS b,
(random() * 50)::int AS c,
(random() * 20)::int AS d,
(random() * 10)::int AS e,
(random() * 5)::int AS f
FROM generate_series(1, 1000000);
CREATE STATISTICS ndist_leak_nd (ndistinct) ON a, b, c, d, e, f FROM
ndist_leak;
ANALYZE ndist_leak;
```
*Peak memory of the backend:*
Before patch - 971076 kB max resident size
After patch - 112780 kB max resident size
ndistinct_for_combination() allocates the items, values and isnull
arrays for the whole sample, plus the MultiSortSupport, and never frees
them. statext_ndistinct_build() calls it once per combination of
columns, so these allocations accumulate until the whole ndistinct
object is built.
The attached patch changes statext_ndistinct_build() to call
ndistinct_for_combination() in a separate memory context and reset it
after each combination, the same way as is already done for
dependency_degree() in statext_dependencies_build().
Similar memory issues in extended statistics were fixed in [0]. The
original report there actually used ndistinct objects, but the fixes
addressed memory accumulated across statistics objects and in
dependency_degree(); the per-combination allocations in
ndistinct_for_combination() were not discussed there.
Any suggestions?
[0]:
https://www.postgresql.org/message-id/20210915200928.GP831%40telsasoft.com
--
Best regards,
Ilia Evdokimov,
Tantor Labs LLC,
https://tantorlabs.com/
| Attachment | Content-Type | Size |
|---|---|---|
| v1-0001-Release-memory-allocated-by-ndistinct_for_combina.patch | text/x-patch | 3.0 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Tom Lane | 2026-10-08 19:10:56 | Re: Wrong results from a parameterized Append |
| Previous Message | Yash Jadhav | 2026-10-08 19:00:04 | Re: [PATCH] Add memory/disk usage for Function Scan nodes in EXPLAIN |