Re: Limiting WAL retained for archiving, like max_slot_wal_keep_size

From: Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>
To: Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com>
Cc: pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Limiting WAL retained for archiving, like max_slot_wal_keep_size
Date: 2026-10-05 18:45:36
Message-ID: CAEze2WgoPrWHuXyU4J6NxBT2A5qHqY8oLWohzqBQf9CY7AC6gA@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

On Mon, 5 Oct 2026 at 17:14, Joao Detomini
<joao(dot)detomini(at)enterprisedb(dot)com> wrote:
>
> Hi,
>
> While looking at barman-wal-archive we ran into something that seems
> to be missing in the server. If archiving breaks or can't keep up,
> the unarchived WAL just piles up in pg_wal until the disk fills, and
> max_slot_wal_keep_size doesn't help because it only covers
> replication slots. We considered making the archive command drop WAL
> once its queue gets too big, but then nothing tells the server, or
> anyone else, that a segment was skipped on purpose, so PITR breaks
> without a trace. Returning true from an archive_library for a segment
> it didn't archive has the same problem, since the callback only knows
> how to say "done" or "try again".
>
> What do people think about a setting like max_slot_wal_keep_size, but
> for archiving? Off by default, and once it's exceeded the server
> would let the oldest unarchived segments go and record that it did,
> maybe in pg_stat_archiver, so tools can notice the gap and take a new
> base backup. We haven't written anything yet and wanted to ask first
> whether this is something you'd want in core.

WAL is generally archived to allow PITR through basebackup+WAL archive
recovery. If we don't archive the WAL before dropping/reusing the
segments, then we won't be able to replay that WAL for the PITR.

I'm not sure such a feature (allowing removing WAL that is critical
for PITR to work, if the backlog is large enough) is an acceptable
solution: Doing so turns a guarantee (the RPO is defined by
basebackup and WAL disk size, unarchived WAL can be recovered from the
primary) into a best-effort approach (RPO is defined by basebackups:
WAL can be lost permanently).

All in all, I don't think this is a good idea. A PITR system with WAL
archiving that is allowed to lose WAL should just ignore archiving the
WAL by itself if it's too far behind; PG can't make that decision for
it.

Kind regards,

Matthias van de Meent
Databricks (https://www.databricks.com)

In response to

Responses

Browse pgsql-hackers by date

  From Date Subject
Next Message Greg Burd 2026-10-05 19:02:01 Re: Per-thread leak in ECPG's memory.c
Previous Message shihao zhong 2026-10-05 18:43:23 Re: aio: worker: Free SMGR objects when idle