Re: Add ASCII fast path to Unicode normalization functions

From: Andrew Dunstan <andrew(at)dunslane(dot)net>
To: Tristan Partin <tristan(at)partin(dot)io>
Cc: PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Add ASCII fast path to Unicode normalization functions
Date: 2026-10-03 14:27:14
Message-ID: 188d01d1-324d-4c85-955a-34b3b59be87b@dunslane.net
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers


On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
>
>
> When time permits I'll send a new version, incorporating your review
> and Bilal's suggestion.
>
>

Here's v3, rebased on current master. It folds in Bilal's change to pick
up where the ASCII scan stopped rather than decoding the whole string, and
deals with Tristan's review comments:

- The helper is now valid_ascii_prefix_len(), which works on raw bytes
  rather than a text value, and returns the length of the valid ASCII
  prefix it found, or the full length if the whole string is valid ASCII.
  That gets rid of the -1 return in v2, and settles on "valid ASCII"
  throughout, as Tristan suggested.

- The paragraph about Unicode normalization has moved out of the helper's
  comment and into the callers, where it actually matters.

- As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
  ASCII character before the non-ASCII part, since a following combining
  mark can compose with it. ASCII characters have no decomposition and
  are all starters, so nothing earlier can be affected, and the quick
  check gives the same answer starting there as from the beginning of the
  string.

- There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
  for normalize() (NFC and NFD) and unicode_assigned(), including a
  combining mark right after the ASCII. It only lists lengths that give
  the wrong answer, so the expected output is short. I checked that it
  does fail if the last ASCII character isn't included in the part that
  gets normalized.

I'll add it to the next CF.

cheers

andrew

--
Andrew Dunstan
EDB: https://www.enterprisedb.com

Attachment Content-Type Size
v3-0001-Add-ASCII-fast-path-to-Unicode-normalization-func.patch text/x-patch 14.4 KB

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Manu 2026-10-03 14:39:17 Re: BUG #19686: Rolling back SET TABLESPACE
Previous Message Andrew Dunstan 2026-10-03 14:07:45 Re: gist_trgm_ops '=' operator: planner picks it over btree, ~300x slower