| From: | ChenhuiMo <chenhuimo(dot)mch(at)qq(dot)com> |
|---|---|
| To: | Jan Nidzwetzki <jan(at)planetscale(dot)com>, pgsql-hackers <pgsql-hackers(at)postgresql(dot)org> |
| Cc: | jeevan(dot)chalke <jeevan(dot)chalke(at)enterprisedb(dot)com>, Heikki Linnakangas <hlinnaka(at)iki(dot)fi> |
| Subject: | Re:[PATCH] Speed up repeat() for larger counts |
| Date: | 2026-09-03 15:20:59 |
| Message-ID: | tencent_902E15E89BF9D99242653BC714600DA2D205@qq.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi Jan,
Thanks for pointing this out.
I updated the benchmark so that the capped strategies also use memset() for
single-byte inputs. In that path, the memset() calls are split according to
the selected block size, with CHECK_FOR_INTERRUPTS() between chunks.
I reran the benchmark three times on the same system and saw the same overall
trend in each run. The table below shows the second run.
source | repeats | output | master | doubling | nocap | cap512 | cap1k | cap2k | cap4k | cap8k | cap16k | cap32k | cap64k | cap128k | cap256k | cap512k | cap1m | cap2m | cap4m | cap8m | cap16m | fastest
--------+-----------+----------+----------+----------+--------+--------+--------+--------+--------+--------+--------+--------+--------+---------+---------+---------+--------+--------+--------+--------+--------+---------------------------
32 B | 2 | 64 B | 3 ns | 0.97x | 0.89x | 0.83x | 0.83x | 0.83x | 0.83x | 0.83x | 0.86x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | 0.89x | master (median: doubling)
32 B | 4 | 128 B | 13 ns | 1.73x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | 1.56x | doubling
32 B | 6 | 192 B | 15 ns | 1.67x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | 1.59x | doubling
32 B | 7 | 224 B | 16 ns | 1.17x | 1.12x | 1.12x | 1.12x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.19x | 1.12x | 1.14x | cap8k (median: doubling)
32 B | 8 | 256 B | 15 ns | 1.19x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | 1.15x | doubling
32 B | 9 | 288 B | 17 ns | 1.19x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | 1.16x | doubling
32 B | 12 | 384 B | 19 ns | 1.27x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | 1.25x | doubling
32 B | 16 | 512 B | 21 ns | 1.37x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | 1.35x | doubling
32 B | 64 | 2 KB | 65 ns | 2.45x | 2.19x | 2.19x | 2.28x | 2.19x | 2.19x | 2.19x | 2.19x | 2.20x | 2.19x | 2.19x | 2.19x | 2.19x | 2.19x | 2.19x | 2.19x | 2.19x | 2.19x | doubling
32 B | 4096 | 128 KB | 5.7 us | 2.90x | 2.90x | 2.52x | 2.83x | 3.06x | 3.36x | 3.69x | 3.73x | 2.94x | 2.91x | 2.91x | 2.91x | 2.91x | 2.91x | 2.91x | 2.91x | 2.91x | 2.90x | cap16k
1 B | 16777216 | 16 MB | 17.3 ms | 37.90x | 42.60x | 39.76x | 46.97x | 47.77x | 43.82x | 45.25x | 47.21x | 49.36x | 50.08x | 49.54x | 50.40x | 50.46x | 49.79x | 49.15x | 50.24x | 50.88x | 50.12x | cap8m (median: cap4m)
10 B | 1677721 | 16 MB | 1.4 ms | 2.47x | 2.64x | 3.39x | 3.27x | 2.87x | 3.36x | 3.54x | 3.64x | 3.22x | 3.29x | 3.22x | 3.30x | 2.88x | 2.25x | 2.28x | 2.55x | 2.30x | 2.67x | cap16k
100 B | 167772 | 16 MB | 557.6 us | 1.08x | 0.97x | 1.36x | 1.41x | 1.31x | 1.36x | 1.47x | 1.34x | 1.34x | 1.34x | 1.31x | 1.30x | 1.15x | 1.01x | 1.02x | 1.04x | 1.09x | 1.06x | cap8k (median: cap16k)
1 KB | 16384 | 16 MB | 374.2 us | 0.72x | 0.73x | 1.02x | 1.02x | 1.05x | 0.96x | 1.00x | 1.04x | 0.89x | 0.88x | 0.89x | 0.91x | 0.87x | 0.70x | 0.73x | 0.72x | 0.71x | 0.72x | cap2k (median: cap512)
64 KB | 256 | 16 MB | 525.7 us | 0.88x | 0.77x | 0.96x | 1.02x | 1.19x | 1.02x | 1.21x | 1.18x | 1.20x | 1.12x | 1.04x | 1.09x | 1.03x | 0.85x | 0.86x | 0.78x | 0.83x | 0.88x | cap8k (median: cap128k)
1 MB | 16 | 16 MB | 577.6 us | 1.01x | 1.01x | 1.00x | 1.03x | 1.00x | 1.03x | 1.02x | 1.01x | 1.04x | 1.05x | 1.03x | 1.05x | 1.06x | 1.03x | 1.09x | 1.07x | 1.09x | 1.09x | cap2m (median: cap8m)
1 B | 268435456 | 256 MB | 296.9 ms | 20.72x | 22.10x | 18.76x | 18.77x | 18.95x | 17.11x | 18.43x | 19.68x | 20.18x | 20.26x | 20.40x | 20.56x | 20.57x | 20.40x | 20.82x | 20.58x | 20.50x | 20.96x | nocap (median: cap16m)
10 B | 26843545 | 256 MB | 27.1 ms | 1.99x | 1.98x | 1.75x | 1.80x | 1.61x | 1.70x | 1.81x | 1.81x | 1.83x | 1.82x | 1.84x | 1.82x | 1.74x | 1.59x | 1.59x | 1.64x | 1.61x | 2.67x | cap16m
100 B | 2684354 | 256 MB | 16.7 ms | 1.14x | 1.19x | 1.04x | 1.09x | 0.92x | 1.02x | 1.07x | 1.10x | 1.12x | 1.14x | 1.13x | 1.12x | 1.01x | 0.99x | 0.98x | 0.99x | 0.90x | 1.36x | cap16m
1 KB | 262144 | 256 MB | 14.7 ms | 1.11x | 1.10x | 0.96x | 0.94x | 0.94x | 0.89x | 0.93x | 0.97x | 0.96x | 0.98x | 1.00x | 0.98x | 0.98x | 0.89x | 0.88x | 0.90x | 0.86x | 1.56x | cap16m
64 KB | 4096 | 256 MB | 15.3 ms | 1.14x | 1.16x | 0.98x | 0.98x | 0.98x | 1.00x | 0.98x | 0.99x | 1.03x | 0.98x | 1.04x | 1.02x | 1.00x | 0.92x | 0.94x | 0.94x | 0.89x | 1.66x | cap16m
1 MB | 256 | 256 MB | 16.6 ms | 1.25x | 1.22x | 0.98x | 1.00x | 1.00x | 0.98x | 0.99x | 0.99x | 1.01x | 0.99x | 0.99x | 0.99x | 1.00x | 0.99x | 0.97x | 0.97x | 0.97x | 1.95x | cap16m
1 B | 1000 | 1000 B | 1.1 us | 97.38x | 93.94x | 77.97x | 93.95x | 93.91x | 93.93x | 93.95x | 93.91x | 93.94x | 93.95x | 93.95x | 93.92x | 93.95x | 93.91x | 93.96x | 93.93x | 93.95x | 93.94x | doubling
1 B | 1000000 | 976.6 KB | 1.1 ms | 78.84x | 79.60x | 52.97x | 60.48x | 69.16x | 71.61x | 75.43x | 78.24x | 79.51x | 79.80x | 80.01x | 80.62x | 80.66x | 80.86x | 80.95x | 80.92x | 81.06x | 80.86x | cap8m (median: cap16m)
(24 rows)
The corrected single-byte results do not change the overall conclusion. For
multi-byte inputs, the preferred cap on this machine still varies noticeably
with the total output size. In particular, the 16 MB cases tend to prefer
relatively small caps, while the 256 MB cases consistently prefer cap16m.
For the single-byte cases, splitting the memset() into capped chunks also does
not seem to introduce a significant performance penalty. In some cases, the
larger capped memset variants are even slightly faster than a single large
memset().
I have attached the updated repeatbench.c. The other extension files are
unchanged.
I am also curious to see what results others get on different systems.
Thanks,
Chenhui
| Attachment | Content-Type | Size |
|---|---|---|
| repeatbench.c | application/octet-stream | 9.5 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | SATYANARAYANA NARLAPURAM | 2026-09-03 15:21:06 | Re: WAIT FOR NO_THROW option could use some documentation |
| Previous Message | Tom Lane | 2026-09-03 14:57:37 | Re: Assert failure in try_nestloop_path() |