| From: | Andrew Dunstan <andrew(at)dunslane(dot)net> |
|---|---|
| To: | Ayush Tiwari <ayushtiwari(dot)slg01(at)gmail(dot)com> |
| Cc: | Tristan Partin <tristan(at)partin(dot)io>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Add ASCII fast path to Unicode normalization functions |
| Date: | 2026-10-03 21:33:44 |
| Message-ID: | e588e85d-f7d7-4ba8-8f9c-e8c6bd600157@dunslane.net |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On 2026-10-03 Sa 3:41 PM, Ayush Tiwari wrote:
> Hi,
>
> On Sat, 3 Oct 2026 at 19:57, Andrew Dunstan <andrew(at)dunslane(dot)net> wrote:
>>
>> On 2026-09-18 Fr 10:11 AM, Andrew Dunstan wrote:
>>>
>>> When time permits I'll send a new version, incorporating your review
>>> and Bilal's suggestion.
>>>
>>>
>> Here's v3, rebased on current master. It folds in Bilal's change to pick
>> up where the ASCII scan stopped rather than decoding the whole string, and
>> deals with Tristan's review comments:
>>
>> - The helper is now valid_ascii_prefix_len(), which works on raw bytes
>> rather than a text value, and returns the length of the valid ASCII
>> prefix it found, or the full length if the whole string is valid ASCII.
>> That gets rid of the -1 return in v2, and settles on "valid ASCII"
>> throughout, as Tristan suggested.
>>
>> - The paragraph about Unicode normalization has moved out of the helper's
>> comment and into the callers, where it actually matters.
>>
>> - As in Bilal's patch, normalize() and IS NORMALIZED restart with the last
>> ASCII character before the non-ASCII part, since a following combining
>> mark can compose with it. ASCII characters have no decomposition and
>> are all starters, so nothing earlier can be affected, and the quick
>> check gives the same answer starting there as from the beginning of the
>> string.
>>
>> - There's a new regression test sweeping ASCII prefixes of 0 to 40 bytes
>> for normalize() (NFC and NFD) and unicode_assigned(), including a
>> combining mark right after the ASCII. It only lists lengths that give
>> the wrong answer, so the expected output is short. I checked that it
>> does fail if the last ASCII character isn't included in the part that
>> gets normalized.
>>
>> I'll add it to the next CF.
> I see CI tests failing for this patch. I haven't gone through the patch
> yet, but from the logs it looks related to missing trailing spaces in
> the newly added expected-output headings. The pg_upgrade failure seems
> to be in its regression setup, not the upgrade itself.
Sorry about that. Here's a v4 that should have the whitespace issue fixed.
cheers
andrew
--
Andrew Dunstan
EDB: https://www.enterprisedb.com
| Attachment | Content-Type | Size |
|---|---|---|
| v4-0001-Add-ASCII-fast-path-to-Unicode-normalization-func.patch | text/x-patch | 14.5 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Zsolt Parragi | 2026-10-03 22:20:18 | Re: [PATCH] Refactor pgbench to make future improvements easier |
| Previous Message | Álvaro Herrera | 2026-10-03 21:32:39 | Re: Coverage with make coverage-html is broken on latest Debian using lcov v2 |