| From: | Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com> |
|---|---|
| To: | Andrew Dunstan <andrew(at)dunslane(dot)net> |
| Cc: | Tristan Partin <tristan(at)partin(dot)io>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Add ASCII fast path to Unicode normalization functions |
| Date: | 2026-10-03 19:41:35 |
| Message-ID: | CAJTYsWUf4EpFX1uW_nW-T4F2NVNyb5U97OhqqCHZa8J0dzA04Q@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
On Sat, 3 Oct 2026 at 19:57, Andrew Dunstan <andrew(at)dunslane(dot)net> wrote:
>
>
> On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
> >
> >
> > When time permits I'll send a new version, incorporating your review
> > and Bilal's suggestion.
> >
> >
>
> Here's v3, rebased on current master. It folds in Bilal's change to pick
> up where the ASCII scan stopped rather than decoding the whole string, and
> deals with Tristan's review comments:
>
> - The helper is now valid_ascii_prefix_len(), which works on raw bytes
> rather than a text value, and returns the length of the valid ASCII
> prefix it found, or the full length if the whole string is valid ASCII.
> That gets rid of the -1 return in v2, and settles on "valid ASCII"
> throughout, as Tristan suggested.
>
> - The paragraph about Unicode normalization has moved out of the helper's
> comment and into the callers, where it actually matters.
>
> - As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
> ASCII character before the non-ASCII part, since a following combining
> mark can compose with it. ASCII characters have no decomposition and
> are all starters, so nothing earlier can be affected, and the quick
> check gives the same answer starting there as from the beginning of the
> string.
>
> - There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
> for normalize() (NFC and NFD) and unicode_assigned(), including a
> combining mark right after the ASCII. It only lists lengths that give
> the wrong answer, so the expected output is short. I checked that it
> does fail if the last ASCII character isn't included in the part that
> gets normalized.
>
> I'll add it to the next CF.
I see CI tests failing for this patch. I haven't gone through the patch
yet, but from the logs it looks related to missing trailing spaces in
the newly added expected-output headings. The pg_upgrade failure seems
to be in its regression setup, not the upgrade itself.
Regards,
Ayush
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Daniel Gustafsson | 2026-10-03 19:49:25 | Re: Serverside SNI support in libpq |
| Previous Message | Greg Burd | 2026-10-03 19:02:45 | Re: UNDO with constant time recovery (CTR) |