Autofailback question

From: Adam Blomeke <adam(dot)blomeke(at)gmail(dot)com>
To: pgpool-general(at)lists(dot)postgresql(dot)org
Subject: Autofailback question
Date: 2026-07-28 17:27:34
Message-ID: CAG9Amsh4cbHtf-Afhiq0g-WVWdJyCzsku-=J9o=wqVKHsNpKvw@mail.gmail.com
Views: Whole Thread | Raw Message | Download mbox | Resend email
Thread:
Lists: pgpool-general

*Subject:* auto_failback never fires — does sr_check worker poll a node in
DOWN status?

Hello,

I'm trying to understand why auto_failback did not reattach a standby that
was healthy and streaming, and specifically whether the sr_check worker
continues to poll a backend once pgpool has marked it DOWN. I'd appreciate
a check on my understanding of the mechanism.

*Environment*

- Pgpool-II 4.6.6, 3 pgpool nodes with watchdog
- 2 PostgreSQL backends (PostgreSQL 18.3), streaming replication mode
- node 0: primary
- node 1: standby
- Persistent physical replication slots (not temporary)
- Standby is the load-balancing target for reads
(database_redirect_preference_list
= 'mydb:standby(1.0)')

*Relevant configuration*
backend_clustering_mode = streaming_replication
load_balance_mode = on
sr_check_period = 10
delay_threshold = 10485760
delay_threshold_by_time = 0
health_check_period = 1
health_check_max_retries = 0
health_check_retry_delay = 1
failover_on_backend_error = off
auto_failback = on
auto_failback_interval = 60
backend_flag0 = ALLOW_TO_FAILOVER
backend_flag1 = ALLOW_TO_FAILOVER
backend_application_name0 = 10.6.1.199
backend_application_name1 = 10.6.1.200

On the standby, max_standby_streaming_delay = 25min (deliberately, to allow
long analytical reads to complete). backend_application_name1 matches the
standby's application_name in pg_stat_replication on the primary.

*What happened*

1. A single health check failed on node 1 (health check failed on node 1
(timeout:0)). With health_check_max_retries = 0, node 1 was immediately
degenerated.
2. The replication slot was not dropped. PostgreSQL on node 1 resumed
streaming on its own and remained caught up (replay lag near zero once it
was no longer serving reads).
3. auto_failback did not reattach node 1. It remained DOWN in pgpool for
roughly 70 hours until I ran pcp_attach_node manually, which succeeded
immediately with no other change.

*The observation I can't explain*
The sr_check_worker logged node 1's replication lag every cycle right up to
the moment of degeneration:

15:55:43 sr_check_worker: Replication of node: 1 is behind ... bytes
from the primary server (node: 0)
16:00:05 sr_check_worker: Replication of node: 1 is behind ... bytes ...
16:11:41 sr_check_worker: worker process received restart request
16:11:42 sr_check_worker: process started

After the restart at degeneration, there are no further sr_check_worker
entries referring to node 1 for the entire ~70 hours it was DOWN — no lag
checks, nothing — until the manual reattach.

*My question*
My understanding is that auto_failback reattaches a node when the sr_check
worker observes it as streaming in pg_stat_replication while its backend
status is down. But if the worker stops polling node 1 once it is marked
DOWN — which the logs appear to show — how is that down + streaming
condition ever evaluated? Is the worker expected to keep querying a DOWN
backend, and if so, is there a configuration reason it might not here? Or
is auto_failback's evaluation driven by something other than the sr_check
worker's per-node polling that I'm misreading?

Thanks for any insight.
Cheers,
Adam

Responses

Browse pgpool-general by date

  From Date Subject
Next Message Tatsuo Ishii 2026-07-29 00:56:13 Re: low level protocol, implicit transactions , "idle in transaction" issue
Previous Message Achilleas Mantzios 2026-07-28 09:14:09 Re: low level protocol, implicit transactions , "idle in transaction" issue