Re: Add ASCII fast path to Unicode normalization functions

From: Andrew Dunstan <andrew(at)dunslane(dot)net>
To: Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com>
Cc: Tristan Partin <tristan(at)partin(dot)io>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Add ASCII fast path to Unicode normalization functions
Date: 2026-10-03 21:33:44
Message-ID: e588e85d-f7d7-4ba8-8f9c-e8c6bd600157@dunslane.net
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers


On 2026-10-03 Sa 3:41 PM, Ayush Tiwari wrote:
> Hi,
>
> On Sat, 3 Oct 2026 at 19:57, Andrew Dunstan <andrew(at)dunslane(dot)net> wrote:
>>
>> On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
>>>
>>> When time permits I'll send a new version, incorporating your review
>>> and Bilal's suggestion.
>>>
>>>
>> Here's v3, rebased on current master. It folds in Bilal's change to pick
>> up where the ASCII scan stopped rather than decoding the whole string, and
>> deals with Tristan's review comments:
>>
>> - The helper is now valid_ascii_prefix_len(), which works on raw bytes
>> rather than a text value, and returns the length of the valid ASCII
>> prefix it found, or the full length if the whole string is valid ASCII.
>> That gets rid of the -1 return in v2, and settles on "valid ASCII"
>> throughout, as Tristan suggested.
>>
>> - The paragraph about Unicode normalization has moved out of the helper's
>> comment and into the callers, where it actually matters.
>>
>> - As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
>> ASCII character before the non-ASCII part, since a following combining
>> mark can compose with it. ASCII characters have no decomposition and
>> are all starters, so nothing earlier can be affected, and the quick
>> check gives the same answer starting there as from the beginning of the
>> string.
>>
>> - There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
>> for normalize() (NFC and NFD) and unicode_assigned(), including a
>> combining mark right after the ASCII. It only lists lengths that give
>> the wrong answer, so the expected output is short. I checked that it
>> does fail if the last ASCII character isn't included in the part that
>> gets normalized.
>>
>> I'll add it to the next CF.
> I see CI tests failing for this patch. I haven't gone through the patch
> yet, but from the logs it looks related to missing trailing spaces in
> the newly added expected-output headings. The pg_upgrade failure seems
> to be in its regression setup, not the upgrade itself.

Sorry about that. Here's a v4 that should have the whitespace issue fixed.

cheers

andrew

--
Andrew Dunstan
EDB: https://www.enterprisedb.com

Attachment Content-Type Size
v4-0001-Add-ASCII-fast-path-to-Unicode-normalization-func.patch text/x-patch 14.5 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Zsolt Parragi 2026-10-03 22:20:18 Re: [PATCH] Refactor pgbench to make future improvements easier
Previous Message Álvaro Herrera 2026-10-03 21:32:39 Re: Coverage with make coverage-html is broken on latest Debian using lcov v2