| From: | Andrew Dunstan <andrew(at)dunslane(dot)net> |
|---|---|
| To: | Tristan Partin <tristan(at)partin(dot)io> |
| Cc: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Add ASCII fast path to Unicode normalization functions |
| Date: | 2026-10-03 14:27:14 |
| Message-ID: | 188d01d1-324d-4c85-955a-34b3b59be87b@dunslane.net |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
>
>
> When time permits I'll send a new version, incorporating your review
> and Bilal's suggestion.
>
>
Here's v3, rebased on current master. It folds in Bilal's change to pick
up where the ASCII scan stopped rather than decoding the whole string, and
deals with Tristan's review comments:
- The helper is now valid_ascii_prefix_len(), which works on raw bytes
rather than a text value, and returns the length of the valid ASCII
prefix it found, or the full length if the whole string is valid ASCII.
That gets rid of the -1 return in v2, and settles on "valid ASCII"
throughout, as Tristan suggested.
- The paragraph about Unicode normalization has moved out of the helper's
comment and into the callers, where it actually matters.
- As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
ASCII character before the non-ASCII part, since a following combining
mark can compose with it. ASCII characters have no decomposition and
are all starters, so nothing earlier can be affected, and the quick
check gives the same answer starting there as from the beginning of the
string.
- There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
for normalize() (NFC and NFD) and unicode_assigned(), including a
combining mark right after the ASCII. It only lists lengths that give
the wrong answer, so the expected output is short. I checked that it
does fail if the last ASCII character isn't included in the part that
gets normalized.
I'll add it to the next CF.
cheers
andrew
--
Andrew Dunstan
EDB: https://www.enterprisedb.com
| Attachment | Content-Type | Size |
|---|---|---|
| v3-0001-Add-ASCII-fast-path-to-Unicode-normalization-func.patch | text/x-patch | 14.4 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Manu | 2026-10-03 14:39:17 | Re: BUG #19686: Rolling back SET TABLESPACE |
| Previous Message | Andrew Dunstan | 2026-10-03 14:07:45 | Re: gist_trgm_ops '=' operator: planner picks it over btree, ~300x slower |