| From: | Matthias van de Meent <boekewurm+postgres(at)gmail(dot)com> |
|---|---|
| To: | Joao Detomini <joao(dot)detomini(at)enterprisedb(dot)com> |
| Cc: | pgsql-hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Limiting WAL retained for archiving, like max_slot_wal_keep_size |
| Date: | 2026-10-05 18:45:36 |
| Message-ID: | CAEze2WgoPrWHuXyU4J6NxBT2A5qHqY8oLWohzqBQf9CY7AC6gA@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Mon, 5 Oct 2026 at 17:14, Joao Detomini
<joao(dot)detomini(at)enterprisedb(dot)com> wrote:
>
> Hi,
>
> While looking at barman-wal-archive we ran into something that seems
> to be missing in the server. If archiving breaks or can't keep up,
> the unarchived WAL just piles up in pg_wal until the disk fills, and
> max_slot_wal_keep_size doesn't help because it only covers
> replication slots. We considered making the archive command drop WAL
> once its queue gets too big, but then nothing tells the server, or
> anyone else, that a segment was skipped on purpose, so PITR breaks
> without a trace. Returning true from an archive_library for a segment
> it didn't archive has the same problem, since the callback only knows
> how to say "done" or "try again".
>
> What do people think about a setting like max_slot_wal_keep_size, but
> for archiving? Off by default, and once it's exceeded the server
> would let the oldest unarchived segments go and record that it did,
> maybe in pg_stat_archiver, so tools can notice the gap and take a new
> base backup. We haven't written anything yet and wanted to ask first
> whether this is something you'd want in core.
WAL is generally archived to allow PITR through basebackup+WAL archive
recovery. If we don't archive the WAL before dropping/reusing the
segments, then we won't be able to replay that WAL for the PITR.
I'm not sure such a feature (allowing removing WAL that is critical
for PITR to work, if the backlog is large enough) is an acceptable
solution: Doing so turns a guarantee (the RPO is defined by
basebackup and WAL disk size, unarchived WAL can be recovered from the
primary) into a best-effort approach (RPO is defined by basebackups:
WAL can be lost permanently).
All in all, I don't think this is a good idea. A PITR system with WAL
archiving that is allowed to lose WAL should just ignore archiving the
WAL by itself if it's too far behind; PG can't make that decision for
it.
Kind regards,
Matthias van de Meent
Databricks (https://www.databricks.com)
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Greg Burd | 2026-10-05 19:02:01 | Re: Per-thread leak in ECPG's memory.c |
| Previous Message | shihao zhong | 2026-10-05 18:43:23 | Re: aio: worker: Free SMGR objects when idle |