Re: Limiting WAL retained for archiving, like max_slot_wal_keep_size

From: Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com>
To: Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com>
Cc: pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>
Subject: Re: Limiting WAL retained for archiving, like max_slot_wal_keep_size
Date: 2026-10-05 19:21:10
Message-ID: CABH8dKxJD_f+rJY2ctZ=k=yULRebb6bRp4_-CTtPW8arf0TEuQ@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgsql-hackers

Hi Matthias,

Thanks, that makes sense. Looking at pgarch.c again, a command that
returns success for a segment it chose to drop does already free the
space, so the tool can do this on its own. We'll handle it on the
Barman side.

Thanks,
João Marcelo

Em seg., 5 de out. de 2026 às 15:45, Matthias van de Meent <
boekewurm+postgres(at)gmail(dot)com> escreveu:

> On Mon, 5 Oct 2026 at 17:14, Joao Detomini
> <joao(dot)detomini(at)enterprisedb(dot)com> wrote:
> >
> > Hi,
> >
> > While looking at barman-wal-archive we ran into something that seems
> > to be missing in the server. If archiving breaks or can't keep up,
> > the unarchived WAL just piles up in pg_wal until the disk fills, and
> > max_slot_wal_keep_size doesn't help because it only covers
> > replication slots. We considered making the archive command drop WAL
> > once its queue gets too big, but then nothing tells the server, or
> > anyone else, that a segment was skipped on purpose, so PITR breaks
> > without a trace. Returning true from an archive_library for a segment
> > it didn't archive has the same problem, since the callback only knows
> > how to say "done" or "try again".
> >
> > What do people think about a setting like max_slot_wal_keep_size, but
> > for archiving? Off by default, and once it's exceeded the server
> > would let the oldest unarchived segments go and record that it did,
> > maybe in pg_stat_archiver, so tools can notice the gap and take a new
> > base backup. We haven't written anything yet and wanted to ask first
> > whether this is something you'd want in core.
>
> WAL is generally archived to allow PITR through basebackup+WAL archive
> recovery. If we don't archive the WAL before dropping/reusing the
> segments, then we won't be able to replay that WAL for the PITR.
>
> I'm not sure such a feature (allowing removing WAL that is critical
> for PITR to work, if the backlog is large enough) is an acceptable
> solution: Doing so turns a guarantee (the RPO is defined by
> basebackup and WAL disk size, unarchived WAL can be recovered from the
> primary) into a best-effort approach (RPO is defined by basebackups:
> WAL can be lost permanently).
>
> All in all, I don't think this is a good idea. A PITR system with WAL
> archiving that is allowed to lose WAL should just ignore archiving the
> WAL by itself if it's too far behind; PG can't make that decision for
> it.
>
>
> Kind regards,
>
> Matthias van de Meent
> Databricks (https://www.databricks.com)
>

In response to

Browse pgsql-hackers by date

  From Date Subject
Next Message Manu 2026-10-05 19:21:53 Re: Planning time quadratic in the IN-list length for "c = X AND (a, b) IN (...)" with BitmapOr
Previous Message shihao zhong 2026-10-05 19:05:31 [PG19] COPY (query) TO ... (FORMAT json) uses the table's column names