Re: Add ASCII fast path to Unicode normalization functions

From: Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com>
To: Andrew Dunstan <andrew(at)dunslane(dot)net>
Cc: Tristan Partin <tristan(at)partin(dot)io>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Add ASCII fast path to Unicode normalization functions
Date: 2026-10-03 19:41:35
Message-ID: CAJTYsWUf4EpFX1uW_nW-T4F2NVNyb5U97OhqqCHZa8J0dzA04Q@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi,

On Sat, 3 Oct 2026 at 19:57, Andrew Dunstan <andrew(at)dunslane(dot)net> wrote:
>
>
> On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
> >
> >
> > When time permits I'll send a new version, incorporating your review
> > and Bilal's suggestion.
> >
> >
>
> Here's v3, rebased on current master. It folds in Bilal's change to pick
> up where the ASCII scan stopped rather than decoding the whole string, and
> deals with Tristan's review comments:
>
> - The helper is now valid_ascii_prefix_len(), which works on raw bytes
> rather than a text value, and returns the length of the valid ASCII
> prefix it found, or the full length if the whole string is valid ASCII.
> That gets rid of the -1 return in v2, and settles on "valid ASCII"
> throughout, as Tristan suggested.
>
> - The paragraph about Unicode normalization has moved out of the helper's
> comment and into the callers, where it actually matters.
>
> - As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
> ASCII character before the non-ASCII part, since a following combining
> mark can compose with it. ASCII characters have no decomposition and
> are all starters, so nothing earlier can be affected, and the quick
> check gives the same answer starting there as from the beginning of the
> string.
>
> - There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
> for normalize() (NFC and NFD) and unicode_assigned(), including a
> combining mark right after the ASCII. It only lists lengths that give
> the wrong answer, so the expected output is short. I checked that it
> does fail if the last ASCII character isn't included in the part that
> gets normalized.
>
> I'll add it to the next CF.

I see CI tests failing for this patch. I haven't gone through the patch
yet, but from the logs it looks related to missing trailing spaces in
the newly added expected-output headings. The pg_upgrade failure seems
to be in its regression setup, not the upgrade itself.

Regards,
Ayush

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Daniel Gustafsson 2026-10-03 19:49:25 Re: Serverside SNI support in libpq
Previous Message Greg Burd 2026-10-03 19:02:45 Re: UNDO with constant time recovery (CTR)