| From: | Laurenz Albe <laurenz(dot)albe(at)cybertec(dot)at> |
|---|---|
| To: | Thomas de Zeeuw <thomasdezeeuw(at)gmail(dot)com>, pgsql-general(at)lists(dot)postgresql(dot)org |
| Subject: | Re: Parsing of hex encoding strings |
| Date: | 2026-10-05 10:35:17 |
| Message-ID: | 088bd47f443ee4b8f657e3dbfadf611a560aa95d.camel@cybertec.at |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-general |
On Mon, 2026-10-05 at 11:50 +0200, Thomas de Zeeuw wrote:
> I’m working on a system where I need to search for hex encoded strings, for
> which I’m using full text search with the “simple” directory. This is working
> fine for the most part, but I’ve run into an issue which may or may not be
> considered a bug.
>
> When creating the ts vector and queries I want to achieve that the entire hex
> encoded string is seen as a single word. For a string such as “fe82fbb28”
> this works using the “simple” directory. However a string such as “7e82fbb28”
> is parsed as two words: “7e82” as scientific notation and “fbb28" as numword.
> I’ve put the output of ts_debug at the bottom of this email for convince.
>
> I would have expected that the entire string would be considered as one token
> considering it doesn’t have any whitespace. And since that token isn’t a valid
> scientific number as whole (though it can be, just not the examples I’ve
> shared above) I would expected it be considered a single word (numword).
> Reading the URL example from chapter 12.5
> (https://www.postgresql.org/docs/current/textsearch-parsers.html) maybe it
> could have to two tokens, one for the scientific number (“7e82”) and another
> numword token for the entire string (7e82fbb28”), this would also work for my
> use case.
>
> Finally, my question: is this considered behaviour a bug or working as intended?
That is working as intended. The problem is that you are trying to use full-text
search for something that isn't a text. Storing strings in hex encoding was a
design error. Not only does it keep you from using full-text search, but you
are also wasting storage space (unless TOAST compression takes care of that).
You could use substring search with pg_trgm, but it is not the same as
full-text search.
Yours,
Laurenz Albe
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Thomas de Zeeuw | 2026-10-05 10:58:01 | Re: Parsing of hex encoding strings |
| Previous Message | Thomas de Zeeuw | 2026-10-05 09:50:42 | Parsing of hex encoding strings |