| From: | Masahiko Sawada <sawada(dot)mshk(at)gmail(dot)com> |
|---|---|
| To: | Bharath Rupireddy <bharath(dot)rupireddyforpostgres(at)gmail(dot)com> |
| Cc: | PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org> |
| Subject: | Re: Parallel autovacuum: DROP DATABASE WITH (FORCE) fails on the parallel workers |
| Date: | 2026-09-28 18:38:32 |
| Message-ID: | CAD21AoDUTngHM7PRzNEr=EhGU47rWDQOcgFnxPWiBYW2VEQ75g@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi,
On Sun, Sep 27, 2026 at 5:00 PM Bharath Rupireddy
<bharath(dot)rupireddyforpostgres(at)gmail(dot)com> wrote:
>
> Hi,
>
> AI review found a bug in parallel autovacuum (1ff3180ca01). I checked
> it myself and it reproduces on HEAD and PG19. Patch with a test
> attached.
>
> A role with pg_signal_backend that is not a superuser gets "permission
> denied to terminate process" [1] from DROP DATABASE WITH (FORCE) while
> a parallel autovacuum on that database is in its index phase, which on
> a large table is where the vacuum spends its time. The same role can
> terminate the autovacuum worker itself, and DROP DATABASE without
> FORCE succeeds with the same vacuum running.
>
> Both the autovacuum worker and its parallel workers run as the
> bootstrap superuser, so that is not what separates them.
> TerminateOtherDBBackends() looks at the role a process published in
> its PGPROC, and only the workers published one. The autovacuum worker
> gets its user from InitializeSessionUserIdStandalone(), which sets
> AuthenticatedUserId directly and never calls SetAuthenticatedUserId(),
> so MyProc->roleId stays InvalidOid and superuser_arg() on it is false.
> Its parallel workers take the same user through ParallelWorkerMain(),
> which does call SetAuthenticatedUserId(), so they publish the
> bootstrap superuser and the check refuses them.
>
> A parallel worker of a VACUUM command is not affected, since its
> leader is a user session and the worker publishes the same role as its
> leader. Autovacuum is the only leader that publishes no role while its
> workers publish one.
>
> The fix treats a process whose lock group leader is an autovacuum
> worker the way the autovacuum worker itself is treated. The patch adds
> a test to the test_autovacuum module that fails with this error
> without the fix.
Thank you for the report and the patch!
IIUC the issue stems from the fact that the leader and its workers
advertise different roleIds (InvalidOid and BOOTSTRAP_SUPERUSERID). I
think the same issue can be reproduced in other cases. For instance,
suppose that a bgworker connecting to the database via
BackgroundWorkerInitializeConnection(dbname, NULL, 0) runs a parallel
query, the leader's roleId is InvalidOid whereas the parallel query
workers have BOOTSTRAP_SUPERUSERID. I think we should fix it as well
and backpatch the fix to 14.
I have some review comments on the proposed patch:
+ if (leader != NULL && leader != proc &&
+ leader->backendType == B_AUTOVAC_WORKER)
+ roleId = InvalidOid;
I think we should check a lock group member with its leader's roleId
instead of unconditionally using InvalidOid. That would
straightforwardly fix the inconsistency between the leader and the
workers.
if (leader != NULL && leader != proc &&
leader->databaseId == databaseId)
roleId = leader->roleId;
To fix this issue not only in autovacuum cases, the backendType check
should be removed.
Regards,
--
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com
| From | Date | Subject | |
|---|---|---|---|
| Next Message | Matheus Alcantara | 2026-09-28 18:45:26 | Re: [PATCH v1] Fix for Bug#19724 - ALTER TYPE ... ALTER ATTRIBUTE triggers internal error for base type of domain with check |
| Previous Message | Vik Fearing | 2026-09-28 17:53:18 | Re: ON EMPTY clause for aggregate and window functions |