| From: | Fabrízio Mello <fabrizio(at)planetscale(dot)com> |
|---|---|
| To: | Daniel Gustafsson <daniel(at)yesql(dot)se> |
| Cc: | PostgreSQL Hackers <pgsql-hackers(at)postgresql(dot)org>, "mail(at)joeconway(dot)com" <mail(at)joeconway(dot)com>, "nik(at)postgres(dot)ai" <nik(at)postgres(dot)ai>, "wolakk(at)gmail(dot)com" <wolakk(at)gmail(dot)com> |
| Subject: | Add pg_stat_log_messages: cumulative statistics about server log messages (was: Add contrib module pg_stat_log: cumulative statistics about server log messages) |
| Date: | 2026-09-28 23:20:56 |
| Message-ID: | CABo-N94G1PgQs=icPK+OvCchKqvVLsHLCv-sWDO2veYeRpo+vQ@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
On Thu, Sep 3, 2026 at 4:58 AM Daniel Gustafsson <daniel(at)yesql(dot)se> wrote:
>
> > On 18 Aug 2026, at 13:23, Fabrizio Mello <fabrizio(at)planetscale(dot)com> wrote:
>
> > * Is contrib the right home, or would people rather see this as a
> > builtin stats kind with a core pg_stat_* view? Contrib felt right for
> > a first step since it exercises the custom stats API and is entirely
> > optional, but I'm happy to go either way.
>
> IMHO, contrib is not the right home for anything new. Either we think it's
> important for production and place it in core or it should live as an external
> extension.
>
> Your email didn't motivate why this extension should be in core instead of
> maintained as an external extension?
>
Thanks Daniel, that's a fair point, and after Nikolay's strong case for
core [1] I'm convinced this belongs there: error counters by SQLSTATE
are fundamental monitoring, counters are far easier to expose to
monitoring systems than log access (which mixes diagnostics with
statements and potentially sensitive data), and external extensions
take a long time to propagate to managed platforms, if ever.
So here is v2, reworked from a contrib module into a builtin cumulative
statistics kind with a core view. Summary of what changed:
* New builtin stats kind PGSTAT_KIND_LOGMSG. Messages are now counted
directly in EmitErrorReport() when they are written to the server
log, instead of via emit_log_hook, so the previous hook-related
caveats (load order, preload requirement) are gone. As before,
log_min_messages acts as a floor on what can be tracked.
* The view is now pg_stat_log_messages, which hopefully also addresses
Kirk's point that "pg_stat_log" was too generic. Columns are
unchanged from v1: (backend_type, database, user, elevel, sqlerrcode,
sqlerrcode_name, count) plus stats_reset. Access requires
pg_read_all_stats.
* Counting is opt-in. The two previous GUCs were folded into a single
enum GUC, track_log_messages, which sets the minimum severity level
to count. The default is none, i.e. disabled.
* Reset is integrated into the usual machinery:
pg_stat_reset_shared('log_messages'). Counters persist across clean
restarts and are discarded after crash recovery, like other
cumulative statistics. Unlike v1, the dropped-entries counter and
stats_reset timestamp now persist too.
* Capacity: builtin fixed-amount stats kinds are sized at compile time
(their structs are embedded in PgStat_ShmemControl and
PgStat_Snapshot), so the previous max_entries GUC is gone. Instead
the table is sized proportionally to the number of SQLSTATE error
codes known to the server: 64 combinations per named errcode, about
16,800 entries in total. That is roughly 0.6 MB of shared memory
(allocated at startup regardless of the GUC) and the same increase in
the stats file at shutdown; a backend's snapshot only materializes
those pages if it actually reads the view. Once the table is full,
already-tracked combinations keep counting and new ones are counted
in pg_stat_get_log_messages_dropped(). I'm happy to adjust the
multiplier if people feel 64 is too generous or too tight.
* The SQLSTATE condition names (e.g. "deadlock_detected") come from a
small lookup table generated from errcodes.txt by a new
generate-errcodes-names.pl script, which also emits the errcode count
used to size the entry table.
* Documentation in monitoring.sgml and config.sgml, regression tests in
stats.sql, and persistence/crash coverage added to the existing
src/test/recovery/t/029_stats_restart.pl.
The entries are still kept in an index-based separate-chaining hash
table living entirely inside the stats block (chain links are array
indices, not pointers), since fixed-amount statistics are snapshotted
with a raw memcpy and persisted verbatim; the rationale, including why
dynahash/simplehash/dshash don't fit these constraints, is in a comment
in pgstat_logmsg.c.
Nikolay: since your alerting use case (XX000/XX001/XX002) motivated a
good part of this, I'd appreciate a look at whether the view shape and
the opt-in default work for you.
--
Fabrízio de Royes Mello
PlanetScale Postgres Core Team
| Attachment | Content-Type | Size |
|---|---|---|
| v2-0001-Add-cumulative-statistics-about-server-log-messag.patch | application/octet-stream | 54.7 KB |
| From | Date | Subject | |
|---|---|---|---|
| Previous Message | Andreas Karlsson | 2026-09-28 23:03:43 | Re: [PATCH] Add ALTER SYSTEM RELOAD |