| From: | Masahiko Sawada <sawada(dot)mshk(at)gmail(dot)com> |
|---|---|
| To: | Fujii Masao <masao(dot)fujii(at)gmail(dot)com> |
| Cc: | shveta malik <shveta(dot)malik(at)gmail(dot)com>, Nisha Moond <nisha(dot)moond412(at)gmail(dot)com>, Nikolay Samokhvalov <nik(at)postgres(dot)ai>, Srinath Reddy Sadipiralla <srinath2133(at)gmail(dot)com>, pgsql-hackers <pgsql-hackers(at)postgresql(dot)org> |
| Subject: | Re: Fix "unexpected logical decoding status change" error; from concurrent logical decoding activation |
| Date: | 2026-10-08 01:09:09 |
| Message-ID: | CAD21AoAXaZSX70jFnvAtY-rhkszS5b7j0PWejw8FataTMv=DoA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Wed, Oct 7, 2026 at 5:33 PM Fujii Masao <masao(dot)fujii(at)gmail(dot)com> wrote:
>
> On Thu, Oct 8, 2026 at 4:08 AM Masahiko Sawada <sawada(dot)mshk(at)gmail(dot)com> wrote:
> > Pushed.
>
> The buildfarm member skink reported a failure in the test added by commit
> 1f2245371e3 [1].
>
> [23:05:02.489](9.575s) not ok 45 - no synced slot is left behind on
> standby5 after the re-creation test
> [23:05:02.489](0.000s) # Failed test 'no synced slot is left behind
> on standby5 after the re-creation test'
>
> The test expects the standby's synced slot to disappear after the
> corresponding slot is dropped on the primary. But, in the failed test
> case, the test found one slot instead of zero.
>
> The slot can become persistent before the primary drops it. In that case,
> WAL replay invalidates the standby slot but does not remove it, leaving one
> slot behind. This seems to explain the slot-count failure.
Right. The sync_slot was created on the standby from the re-created
remote slot as we expect, but it then became sync-ready in the same
pg_sync_replication_slots() call. I expected that after detecting the
disable/re-enable, the retry would re-create the slot but not be able
to persist it because the standby's xmin
is ahead of the remote slot's. But that's not necessarily true.
> To address this, how about running pg_sync_replication_slots() once more
> as follows? This extra sync cycle removes the obsolete slot.
>
> $primary->wait_for_replay_catchup($standby5);
> $psql_sync_slot->quit;
> + # The slot may have become sync-ready before it was dropped on the
> + # primary. In that case, replay invalidates it, and a new sync cycle
> + # must remove it from the standby.
> + $standby5->safe_psql('postgres',
> + qq[select pg_sync_replication_slots()]);
> $primary->safe_psql('postgres',
> qq[select pg_drop_replication_slot('test_slot4')]);
> wait_for_logical_decoding_disabled($primary);
I agree that it's a reasonable fix. This can handle both cases where
the synced slot is persistent and not persistent: if the synced slot
was persisted, drop_local_obsolete_slots() removes it now that the
remote slot is gone, and if it stayed temporary it has already been
removed by the time the background session finished, so the extra call
is just a no-op.
Another option would be to make it deterministic the other way around:
advance the remote slot on the primary until the synced slot becomes
sync-ready, then clean up. That's what
040_standby_failover_slots_sync.pl does for its retry case, but there
the subscription keeps consuming the remote slot for us. Nothing
consumes sync_slot here, so we would need a loop calling
pg_log_standby_snapshot() and pg_logical_slot_get_changes() until the
slot is persisted. I prefer Fujii-san's approach.
I've attached the patch.
Regards,
--
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com
| Attachment | Content-Type | Size |
|---|---|---|
| v1-0001-Stabilize-recently-added-tests-in-051_effective_w.patch | text/x-patch | 2.9 KB |
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Michael Paquier | 2026-10-08 02:04:24 | Re: Compression of bigger WAL records |
| Previous Message | Richard Guo | 2026-10-08 01:05:54 | Wrong results from a parameterized Append |