Re: 64-bit MultiXactOffset vs. 9.3->9.4 upgrade, pg_resetwal, "wraparound" msg

From: Heikki Linnakangas <hlinnaka(at)iki(dot)fi>
To: Naga Appani <nagnrik(at)gmail(dot)com>, Michael Paquier <michael(at)paquier(dot)xyz>
Cc: pgsql-hackers(at)postgresql(dot)org, Noah Misch <noah(at)leadboat(dot)com>, orlovmg(at)gmail(dot)com
Subject: Re: 64-bit MultiXactOffset vs. 9.3->9.4 upgrade, pg_resetwal, "wraparound" msg
Date: 2026-09-16 13:36:51
Message-ID: 05d064e8-219d-40bd-9fdf-652b648ba317@iki.fi
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

(adding Michael and Naga)

I'd actually attribute the below issue to commit 97b101776c "Add
pg_get_multixact_stats()". In a nutshell, MultiXactState->oldestOffset
is not updated during recovery, and that was true before widening
multixacts to 64 bits already. That's intentional and wasn't visible to
users until the pg_get_multixact_stats() function was added. If you run
pg_get_multixact_stats() in a standby, the 'num_members' and
'members_size' are calculated incorrectly, because oldestOffset is always 0.

To fix, we could update MultiXactState->oldestOffset during recovery
too, from the control file at start of recovery, and whenever we see a
XLOG_MULTIXACT_TRUNCATE_ID record. We should probably still calculate a
fresh value at end of recovery, like we currently do.

Full LLM report extracted from Noah's markdown file below:

On 28/08/2026 02:17, Noah Misch wrote:
> ### 3.3 `standby-oldestoffset-zero-stats` — `pg_get_multixact_stats()` on a hot standby reports lifetime members, not retained members
>
> **Location:** `src/backend/access/transam/multixact.c:2545-2559` and its SQL consumer
> `src/backend/utils/adt/multixactfuncs.c:127-130`.
>
> ```
> 2545 void
> 2546 GetMultiXactInfo(uint32 *multixacts, MultiXactOffset *nextOffset,
> 2547 MultiXactId *oldestMultiXactId, MultiXactOffset *oldestOffset)
> 2548 {
> 2549 MultiXactId nextMultiXactId;
> 2550
> 2551 LWLockAcquire(MultiXactGenLock, LW_SHARED);
> 2552 *nextOffset = MultiXactState->nextOffset;
> 2553 *oldestMultiXactId = MultiXactState->oldestMultiXactId;
> 2554 nextMultiXactId = MultiXactState->nextMXact;
> 2555 *oldestOffset = MultiXactState->oldestOffset;
> 2556 LWLockRelease(MultiXactGenLock);
> ```
>
> ```
> 127 GetMultiXactInfo(&multixacts, &nextOffset, &oldestMultiXactId, &oldestOffset);
> 128 members = nextOffset - oldestOffset;
> 129 membersBytes = MultiXactOffsetStorageSize(nextOffset, oldestOffset);
> ```
>
> **Mechanism.** `MultiXactState->oldestOffset` is assigned in exactly two places:
> `SetOldestOffset()` at `multixact.c:2488`, and `TruncateMultiXact()` at `multixact.c:2788`. Neither
> runs in recovery:
>
> * `TruncateMultiXact()` opens with `Assert(!RecoveryInProgress());` (`multixact.c:2685`).
> * `SetOldestOffset()` is called only from `SetMultiXactIdLimit()` at `multixact.c:2173`, which is
> *after* the early return at `multixact.c:2161-2162`:
> `if (!MultiXactState->finishedStartup) return;`. `finishedStartup` is set only by
> `TrimMultiXact()` (`multixact.c:2006`), whose sole caller is `xlog.c:6527`, at end of recovery.
> * The redo path (`multixact.c:2985-2990`) calls `SetMultiXactIdLimit(xlrec.oldestMulti, …)` — which
> correctly stores `oldestMultiXactId` at `multixact.c:2140` before the early return — and then
> `PerformMembersTruncation(xlrec.oldestOffset)` / `PerformOffsetsTruncation()`, using the WAL
> record's value directly. It never stores `xlrec.oldestOffset` into shared memory.
>
> So on a never-promoted standby, `MultiXactState->oldestOffset` keeps its zeroed shmem value for the
> life of the server while `nextOffset` is advanced continuously by redo. `num_members` therefore comes
> out as `nextOffset - 0` and `members_size` as `5 * nextOffset` bytes
> (`MultiXactOffsetStorageSize()`, `multixact_internal.h:125-134`: `MULTIXACT_MEMBERGROUP_SIZE /
> MULTIXACT_MEMBERS_PER_MEMBERGROUP` = 20/4 = 5). `num_mxids` and `oldest_multixact` are correct.
>
> The overreport is exactly the primary's current `oldestOffset`, and since bd8d9c9 made offsets
> 64-bit and monotonic, that quantity climbs for the life of the installation. Meanwhile redo *does*
> physically delete the members segments, so the standby reports gigabytes of members it does not have.
>
> **Trigger (reproduced).** Primary with `wal_level = replica`; create a table, take `FOR SHARE` from
> concurrent sessions to build multixacts; `pg_basebackup -R` a standby and start it; on the primary
> `ALTER DATABASE template0 ALLOW_CONNECTIONS true`, `VACUUM FREEZE` in every database, `CHECKPOINT`,
> so `TruncateMultiXact()` advances `oldestOffset`. Observed: primary `num_members` 0, `members_size` 0;
> standby `num_members` 3, `members_size` 15, `pg_is_in_recovery()` true. A larger run (1000 multixacts
> of 2 members, `nextOffset` 2001) showed the standby reporting 2001/10005 against a primary reporting
> 0/0. Promoting the standby corrects it immediately, confirming the `finishedStartup` mechanism.
>
> **User-visible consequence.** `SELECT * FROM pg_get_multixact_stats()` on a hot standby contradicts
> the same query on its primary for byte-identical data, and contradicts
> `doc/src/sgml/func/func-info.sgml:3010-3013`, which defines `num_members` as "the total number of
> multixact member entries **currently present in the system**" and `members_size` as "the storage
> occupied by `num_members` in the `pg_multixact/members` directory". Monitoring built on this function
> sees a standby that appears to be holding hundreds of GB of members it does not hold.
>
> **Attribution — stated honestly.** This is the weakest attribution in the set, and it survived
> 2/3 refutation votes rather than 3/3. The dissenting reading is correct on the facts and should be
> recorded: the `finishedStartup` gate and the redo path are **identical in `bd8d9c9^`**, so the
> standby's stuck `oldestOffset` predates the commit; pre-commit it was harmless because the only
> consumer, `MultiXactMemberFreezeThreshold()`, never runs in recovery. What bd8d9c9 contributed is
> deleting `MultiXactStateData.oldestOffsetKnown` and changing `GetMultiXactInfo()` from returning
> `bool` to `void` — it used to short-circuit with `*members = 0; … return false;` so callers could
> distinguish "unknown" from "zero". `97b10177` (2025-12-30) then wired the now-unqualified value into
> SQL with no `RecoveryInProgress()` guard and no NULL path. A committer may reasonably file this
> against `97b10177`; the missing "unknown" signal is bd8d9c9's.
>
> **Why it is not covered by a test.** Testing it needs a full streaming primary/standby pair plus the
> `template0`-connectable `VACUUM FREEZE` dance to move the primary's `oldestOffset` off zero — that is
> a `src/test/recovery`-style test, not something that fits in `test_slru`'s single-node harness or in
> `007_multixact_conversion.pl`. It also cannot be asserted until the intended semantics are decided:
> should the standby return NULL for these two columns, or should redo store `xlrec.oldestOffset`?
> Writing a test first would bake in whichever answer the test author guessed.
>
> **Test someone should write** (in `src/test/recovery/t/`, once the semantics are settled):
>
> ```
> init primary (wal_level=replica), create t, produce N multixacts via concurrent FOR SHARE
> pg_basebackup -R -> standby, start, wait for catchup
> on primary: ALTER DATABASE template0 ALLOW_CONNECTIONS true; VACUUM FREEZE in all dbs; CHECKPOINT
> wait_for_catchup
> is(standby: SELECT num_members FROM pg_get_multixact_stats(),
> primary: same) # or: expect NULL on the standby
> promote standby; re-run; expect it to match
> ```
>
> The middle assertion is the whole test; it fails today with standby = `nextOffset`,
> primary = `nextOffset - oldestOffset`.
>
> ---

- Heikki

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Diego 2026-09-16 13:39:39 Re: [PATCH] libpq: Add PQpassfileLookup()
Previous Message Aleksander Alekseev 2026-09-16 13:23:02 [PATCH] Remove unused PageIsPredicateLocked()