[PATCH] numeric: canonicalize wide digit accumulation

From: zengxx <xiangxin_zeng(at)qq(dot)com>
To: pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: [PATCH] numeric: canonicalize wide digit accumulation
Date: 2026-10-10 16:16:09
Message-ID: tencent_9DA409B9B2CAA6FFF985C638A28D35C73705@qq.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

While looking at wide numeric aggregates, I found that the digit
accumulation loop in accum_sum_add() did not auto-vectorize. It advances
two indexes together:

for (val_i = 0; val_i < val_ndigits; val_i++)
{
accum_digits[i] += (int32) val_digits[val_i];
i++;
}

The attached patch advances the destination pointer once before the loop:

accum_digits += accum->weight - val->weight;
for (val_i = 0; val_i < val_ndigits; val_i++)
accum_digits[val_i] += (int32) val_digits[val_i];

This follows the same general loop-canonicalization approach already used
for the hot numeric multiplication and division loops. It is a semantic
no-op and adds no SIMD intrinsics, simd.h changes, ISA-specific code, or
runtime CPU dispatch.

There is prior precedent for this kind of change in numeric.c, and those
patches were useful references:

* commit 8870917623 enabled -ftree-vectorize for numeric.c and adjusted
the mul_var() inner loop so GCC could auto-vectorize it [1].
* commit 9c79e646c6 adjusted the same loop form so clang could
vectorize it too [2].
* commit e3d41d08a1 used div_qi = &div[qi] in div_var_fast() to
canonicalize the destination array reference for auto-vectorization
[3].

The last one is especially close to this patch: it replaces an indexed
destination with a pointer to the current offset before the hot loop.

With GCC 13.3.0/x86_64 and PostgreSQL's normal -O2 -ftree-vectorize
flags, the new loop is vectorized with runtime alias versioning:

numeric.c:11946: optimized: loop vectorized using 16 byte vectors
numeric.c:11946: optimized: loop versioned for vectorization because
of possible aliasing
numeric.c:11946: optimized: loop vectorized using 8 byte vectors

The generated code contains movdqu/punpckwd/paddd. The base report for
the original loop reports that it could not be vectorized because of a
possible alias involving gather/scatter.

Performance below is from one local machine, so it is motivation rather
than a universal claim. The machine is WSL2/Linux on an i5-14400F using
GCC 13.3.0. Each number is the median of 11 runs after 2 warmups; base
and feature used identical task files, settings, row counts, and
disposable clusters.

1. Single-column wide numeric SUM
- 200000 rows, 1020 decimal digits per value
- serial aggregate
- round 1: 43.594 ms -> 26.594 ms (1.639x)
- round 2: 42.714 ms -> 26.215 ms (1.629x)

2. Width sweep over a 200000-row table
Each query scans the same table, so fixed executor/scan overhead
dilutes the accumulator gain. Values are geomeans of two rounds:

- 32 digits: 1.026x
- 64 digits: 1.028x
- 128 digits: 1.073x
- 256 digits: 1.161x
- 512 digits: 1.154x
- 1024 digits: 1.230x

At 1024 decimal digits:
- serial AVG: 1.227x
- forced-parallel SUM with 4 workers: 1.264x

The narrow cases are within the noise range and I would not claim a win
there. The benefit becomes clearer as the numeric value gets wider and
accumulation is a larger fraction of total runtime.

Correctness:

* The full core regression suite passes (243/243).
* Differential SQL covers 32, 64, 128, 256, 512, and 1024 decimal-digit
values with both positive and negative inputs, and checks exact SUM and
AVG results.
* A separate check compares exact serial and forced-parallel SUM results
over 100000 wide values.

I do not think this promises vectorization on every compiler or
architecture. It only puts the loop in the canonical form that GCC can
vectorize here; other compilers can still emit the scalar form. With
that caveat, it is a small local change and gives a useful improvement
for wide numeric aggregates on the tested configuration.

Patch:
0001-numeric-canonicalize-wide-digit-accumulation.patch

[1]
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commitdiff;h=88709176236caf3cb9655acda6bad2df0323ac8f
[2]
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commitdiff;h=9c79e646c6f0f8df06d966c536d0c6aa33bf1b06
[3]
https://git.postgresql.org/gitweb/?p=postgresql.git;a=commitdiff;h=e3d41d08a17549fdc60a8b9450c0511c11d666d7

Thanks,
Xiangxin Zeng

Attachment Content-Type Size
0001-numeric-canonicalize-wide-digit-accumulation.patch application/octet-stream 1.4 KB

Browse pgsql-hackers by date

  From Date Subject
Next Message David E. Wheeler 2026-10-10 16:51:48 Policy for Abandoned Extensions
Previous Message Greg Burd 2026-10-10 16:14:37 Re: Tepid: selective index updates for heap relations