| From: | Vaibhav Dalvi <vaibhav(dot)dalvi(at)enterprisedb(dot)com> |
|---|---|
| To: | pgsql-hackers(at)lists(dot)postgresql(dot)org, "aekorotkov(at)gmail(dot)com" <aekorotkov(at)gmail(dot)com> |
| Cc: | Andrew Dunstan <andrew(dot)dunstan(at)enterprisedb(dot)com>, "dgrowleyml(at)gmail(dot)com" <dgrowleyml(at)gmail(dot)com>, Vaibhav Dalvi <vaibhav(dot)dalvi(at)enterprisedb(dot)com> |
| Subject: | gist_trgm_ops '=' operator: planner picks it over btree, ~300x slower |
| Date: | 2026-09-02 11:11:31 |
| Message-ID: | CA+vB=AHon6EB=90pbPDik_u8+++Ty8TA6Z4-RXo_X5QCr71kcQ@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi all,
I would like to raise an old issue again. David Rowley reported
this same problem back on 2024-09-17, in this[1] thread, but it did not
get any replies.
I am hitting the same problem, so I am posting it again with three
different patches, since it can badly hurt anyone who has both a
btree and a gist_trgm_ops index on the same column.
In short: once a table has both a btree index and a gist_trgm_ops
index on the same text column, the planner sometimes picks the GiST
index for a plain equality (=) query, and that GiST plan can be
hundreds of times slower than the btree plan for the exact same
query. Below is a small, self-contained test case:
create extension if not exists pg_trgm;
create table t1 (a varchar(250), b varchar(250), c varchar(250));
create index t1_a_btree on t1 (a);
create index t1_a_gist on t1 using gist (a gist_trgm_ops);
insert into t1 select md5(a::text),md5(a::text),md5(a::text) from
generate_series(1,100000)a;
vacuum freeze analyze t1;
explain (analyze, buffers) select * from t1 where a = '1234';
QUERY PLAN
--------------------------------------------------------------------------------------------------------------------
Index Scan using t1_a_gist on t1 (cost=0.28..8.30 rows=1 width=99)
(actual time=15.186..15.187 rows=0.00 loops=1)
Index Cond: ((a)::text = '1234'::text)
Buffers: shared hit=1583
Execution Time: 15.242 ms
(4 rows)
-- disabling the GiST index makes the planner fall back to btree,
-- and the same query becomes ~355x faster:
update pg_index set indisvalid = false where
indexrelid='t1_a_gist'::regclass;
explain (analyze, buffers) select * from t1 where a = '1234';
QUERY PLAN
-------------------------------------------------------------------------------------------------------------------
Index Scan using t1_a_btree on t1 (cost=0.42..8.44 rows=1 width=99)
(actual time=0.022..0.022 rows=0.00 loops=1)
Index Cond: ((a)::text = '1234'::text)
Buffers: shared hit=3
Execution Time: 0.043 ms
(4 rows)
The estimated cost of both plans is almost the same (8.30 vs 8.44),
but the real cost is not even close. The root cause: GiST's cost
estimate for '=' on gist_trgm_ops does not reflect the real cost of
the scan, so the planner can pick GiST over a much cheaper btree
index, even for people who kept both indexes only for other reasons
(LIKE queries, for example).
I looked at three different ways to fix this:
1) v1-0001-Drop-equality-operator-from-gist_trgm_ops.patch
Stops gist_trgm_ops from offering '=' at all (gin_trgm_ops is
untouched, its cost estimator was already fixed for the same issue
in commit cd9479af2af). Small, but it changes the existing
opclass's behavior for anyone already relying on '=' through it.
2) v1-0001-gist-trgm-real-signature-stats.patch
Keeps '=' on gist_trgm_ops, and adds a new optional GiST support
function so the opclass can correct the planner's estimate using a
real measurement sampled from the index's own pages, instead of a
guess. It works, but it only helps "column = constant" conditions
(not joins), a small index still falls back to a pessimistic
guess, and it adds uncached I/O to every planning call. More
machinery than I'm comfortable with for this.
3) v1-0001-Add-gist_trgm_ops_noeq.patch
Adds a second, independently-named opclass, gist_trgm_ops_noeq,
identical to gist_trgm_ops except that it does not register '='
at all. gist_trgm_ops itself is completely untouched; anyone who
wants the safety creates new indexes with the new opclass, or
rebuilds an existing index onto it, and '=' simply cannot reach it
afterward, under any planner settings. No core or planner changes
at all, just a SQL/DDL addition.
Of the three, (3) is the one I would prefer to take forward. It is
the smallest, safest change: nothing existing changes behavior, there
is no heuristic or cost-model logic to get wrong, and it is easy to
verify its correctness just by looking at the catalog entries. (1) fixes the
regression but forces the choice on everyone using gist_trgm_ops for
'=' today, and (2) is real but has enough rough edges (documented in
that patch) that I would not want to see it committed as-is.
Would appreciate the community's view on this, or any other approach
I may have missed.
Thank you, @Andrew Dunstan <andrew(dot)dunstan(at)enterprisedb(dot)com> for the
offline inputs.
Thanks and regards,
Vaibhav Dalvi
EnterpriseDB
| Attachment | Content-Type | Size |
|---|---|---|
| v1-0001-Add-gist_trgm_ops_noeq.patch | application/octet-stream | 8.4 KB |
| v1-0001-Drop-equality-operator-from-gist_trgm_ops.patch | application/octet-stream | 5.4 KB |
| v1-0001-gist-trgm-real-signature-stats.patch | application/octet-stream | 21.2 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Nisha Moond | 2026-09-02 11:19:50 | Re: Support EXCEPT for TABLES IN SCHEMA publications |
| Previous Message | Jan Nidzwetzki | 2026-09-02 10:56:37 | Re: Many of psql's describe functions bloat cache / waste mem |