| From: | Kirill Reshke <reshkekirill(at)gmail(dot)com> |
|---|---|
| To: | Tom Lane <tgl(at)sss(dot)pgh(dot)pa(dot)us> |
| Cc: | Andres Freund <andres(at)anarazel(dot)de>, pgsql-hackers(at)postgresql(dot)org, Ashutosh Bapat <ashutosh(dot)bapat(dot)oss(at)gmail(dot)com>, Álvaro Herrera <alvherre(at)alvh(dot)no-ip(dot)org> |
| Subject: | Re: pg_dump/restore failure (dependency?) on BF serinus |
| Date: | 2026-10-05 11:48:16 |
| Message-ID: | CALdSSPj14NGS8TWHg2UokfiPgNUQo5OMsxvkaruzqgBC1xM34g@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Mon, 5 Oct 2026 at 16:20, Kirill Reshke <reshkekirill(at)gmail(dot)com> wrote:
>
> On Tue, 8 Apr 2025 at 09:48, Tom Lane <tgl(at)sss(dot)pgh(dot)pa(dot)us> wrote:
> >
> > Andres Freund <andres(at)anarazel(dot)de> writes:
> > > On 2025-04-08 00:11:55 -0400, Tom Lane wrote:
> > >> This feels quite adjacent to my complaint here:
> > >> https://www.postgresql.org/message-id/2045026.1743801143%40sss.pgh.pa.us
> > >> though perhaps it's not exactly the same.
> >
> > > That does sound rather plausible. What an odd coincidence that it failed like
> > > that so close to your email. While this specific failure probably couldn't
> > > have happened much earlier, it seems that it could have as part of pg_upgrade
> > > for longer.
> >
> > I think pg_upgrade is not vulnerable to the problem, or at least not
> > the identical problem, because it doesn't expect pg_restore to load
> > table data. So I think we didn't previously have any test cases
> > that would expose this :-(. What I find surprising is that we
> > didn't get field reports much sooner.
> >
> > regards, tom lane
> >
> >
>
> I was able to reproduce this failure reliably on HEAD e5d25959cf87
> Attached sh script (AI generated) fails 6-7 times per 100 iteration run for me.
>
> The root cause is catalog check race and it is not pg_restore-only.
>
> The error is real and server-side. The referenced partitioned PK index
> is left permanently indisvalid=false by a race between the last two
> concurrent ATTACH statements, and nothing ever retries the validation
> afterwards.
>
> So, I managed to create an isolation test with 2-level partitioning
> which hits this issue in the (i believe) same way as shell reproducer
> does.
> The issue does not reproduce with one-level partitioning. The key
> component here is that first level LEAF partitions are created
> invalid. So, when two last ATTACH statements validate top-level index
> validity inside ATExecAttachPartitionIdx -> validatePartitionedIndex
> they count the number of other valid indexes.
> With some execution ordering we can get into a schedule where 2 & 3
> partition indexes are about-to-be valid (but this state is not yet
> committed), and they concurrently recount that there are 2 out of 3
> valid indexes, then each of this statement commits.
>
> So, each statement runs before the other's transaction commits, so no
> snapshot taken by either can ever observe the complete set.
>
> I didn't came up with fix idea yet.
>
> --
> Best regards,
> Kirill Reshke
x4m (Andrey) suggest this
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -22618,7 +22618,12 @@ validatePartitionedIndex(Relation partedIdx,
Relation partedTbl)
validatePartitionedIndex(parentIdx, parentTbl);
- relation_close(parentIdx, AccessExclusiveLock);
- relation_close(parentTbl, AccessExclusiveLock);
+ /*
+ * Keep these locks until commit, so another validator cannot miss
+ * our uncommitted changes to a descendant's indisvalid flag and leave
+ * the parent index invalid.
+ */
+ relation_close(parentIdx, NoLock);
+ relation_close(parentTbl, NoLock);
}
}
--
Best regards,
Kirill Reshke
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Peter Eisentraut | 2026-10-05 11:51:41 | Re: Fix out-of-bounds array indexing in JsonValueList |
| Previous Message | David Geier | 2026-10-05 11:48:10 | Re: Reducing relcache memory usage 2: shrink sizeof(RelationData) |