Re: pg_dump/restore failure (dependency?) on BF serinus

From: Kirill Reshke <reshkekirill(at)gmail(dot)com>
To: Tom Lane <tgl(at)sss(dot)pgh(dot)pa(dot)us>
Cc: Andres Freund <andres(at)anarazel(dot)de>, pgsql-hackers(at)postgresql(dot)org, Ashutosh Bapat <ashutosh(dot)bapat(dot)oss(at)gmail(dot)com>, Álvaro Herrera <alvherre(at)alvh(dot)no-ip(dot)org>
Subject: Re: pg_dump/restore failure (dependency?) on BF serinus
Date: 2026-10-05 11:20:04
Message-ID: CALdSSPjZviSRnOaYrfdnWNeh5zLkoT7bD3PwH98E5fhE-Cyd+g@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Tue, 8 Apr 2025 at 09:48, Tom Lane <tgl(at)sss(dot)pgh(dot)pa(dot)us> wrote:
>
> Andres Freund <andres(at)anarazel(dot)de> writes:
> > On 2025-04-08 00:11:55 -0400, Tom Lane wrote:
> >> This feels quite adjacent to my complaint here:
> >> https://www.postgresql.org/message-id/2045026.1743801143%40sss.pgh.pa.us
> >> though perhaps it's not exactly the same.
>
> > That does sound rather plausible. What an odd coincidence that it failed like
> > that so close to your email. While this specific failure probably couldn't
> > have happened much earlier, it seems that it could have as part of pg_upgrade
> > for longer.
>
> I think pg_upgrade is not vulnerable to the problem, or at least not
> the identical problem, because it doesn't expect pg_restore to load
> table data. So I think we didn't previously have any test cases
> that would expose this :-(. What I find surprising is that we
> didn't get field reports much sooner.
>
> regards, tom lane
>
>

I was able to reproduce this failure reliably on HEAD e5d25959cf87
Attached sh script (AI generated) fails 6-7 times per 100 iteration run for me.

The root cause is catalog check race and it is not pg_restore-only.

The error is real and server-side. The referenced partitioned PK index
is left permanently indisvalid=false by a race between the last two
concurrent ATTACH statements, and nothing ever retries the validation
afterwards.

So, I managed to create an isolation test with 2-level partitioning
which hits this issue in the (i believe) same way as shell reproducer
does.
The issue does not reproduce with one-level partitioning. The key
component here is that first level LEAF partitions are created
invalid. So, when two last ATTACH statements validate top-level index
validity inside ATExecAttachPartitionIdx -> validatePartitionedIndex
they count the number of other valid indexes.
With some execution ordering we can get into a schedule where 2 & 3
partition indexes are about-to-be valid (but this state is not yet
committed), and they concurrently recount that there are 2 out of 3
valid indexes, then each of this statement commits.

So, each statement runs before the other's transaction commits, so no
snapshot taken by either can ever observe the complete set.

I didn't came up with fix idea yet.

--
Best regards,
Kirill Reshke

Attachment Content-Type Size
fk_no_unique_repro.sh text/x-sh 5.7 KB
v1-0001-Add-test-for-concurrent-partitioned-index-attach-.patch application/octet-stream 14.1 KB

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Peter Eisentraut 2026-10-05 11:25:18 Re: pgindent to ignore build directories
Previous Message Matthias van de Meent 2026-10-05 10:59:32 Re: UNDO with constant time recovery (CTR)