| From: | Lucas Jeffrey <luquijeffrey(at)gmail(dot)com> |
|---|---|
| To: | pgsql-hackers(at)lists(dot)postgresql(dot)org |
| Subject: | [PROPOSAL] Isolated copy-on-write cluster forks |
| Date: | 2026-10-04 13:07:06 |
| Message-ID: | CAObkL7NYWxKWPSpmLZOkHdGmfP_6iZJom_4=ydXJFsnY2rXdKg@mail.gmail.com |
| Views: | Whole Thread | Raw Message | Download mbox | Resend email |
| Thread: | |
| Lists: | pgsql-hackers |
Hi hackers,
I'd like to propose a new PostgreSQL feature tentatively called FORK.
The main use case is to give an application, especially an automated
programming agent, an isolated PostgreSQL environment where it can
freely experiment with a complete cluster without affecting the
environment from which the fork was created.
The agent should be able to operate with unrestricted permissions
inside the fork, including running migrations, changing schemas and
data, dropping databases, and performing other destructive operations.
None of those changes should be able to affect the parent environment.
The proposed interface is intentionally small:
FORK;
ATTACH FORK <fork_id>;
DROP FORK <fork_id>;
A fork represents the complete state of a PostgreSQL cluster at the
time it is created. After that point, the parent and the fork operate
independently.
For example:
postgres=# SELECT * FROM fork_isolation_test;
id | value
----+--------------------
1 | parent-before-fork
2 | parent-before-fork
3 | only parent
(3 rows)
postgres=# FORK;
FORK
The new fork initially sees the same state:
postgres=# SELECT * FROM fork_isolation_test;
id | value
----+--------------------
1 | parent-before-fork
2 | parent-before-fork
3 | only parent
(3 rows)
Changes made afterwards belong only to the fork:
postgres=# DELETE FROM fork_isolation_test;
DELETE 3
A new connection to the parent still sees the original data.
An existing fork can be attached with:
postgres=# ATTACH FORK 102;
ATTACH FORK
postgres=# SELECT * FROM fork_isolation_test;
id | value
----+-------
(0 rows)
FORK does not require the client to establish a new connection. The
session that executes FORK continues as the same logical PostgreSQL
session. Internally, the fork has its own postmaster and backend, and
the frontend/backend protocol is redirected at the socket level.
Connections that were already connected to the parent continue
operating on the parent.
A fork can create additional forks:
main
|
+-- fork 102
|
+-- fork 103
|
+-- fork 104
A fork can modify itself and create descendants, but it can never
modify its parent or any other ancestor.
The fork is also an isolation boundary for privileges. A superuser
operating inside a fork may perform destructive operations there, but
those operations cannot affect the parent.
FORK is a cluster-level operation and is not allowed inside a transaction.
Each fork has its own WAL. There is no WAL replication between a
parent and its child after the fork is created.
The main cluster maintains the fork catalog and allocates fork IDs. A
fork only needs to see information about itself and its descendants.
A fork can be removed with:
DROP FORK <fork_id>;
Dropping a fork also removes its descendants.
Creating a fork should be atomic from the user's perspective. If the
operation fails, the fork should not remain partially created. The
main cluster should clean up any resources created before the failure.
Copy-on-write
The fork is intended to use copy-on-write functionality provided by
the underlying filesystem.
I do not propose implementing a separate CoW mechanism inside
PostgreSQL's storage layer. PostgreSQL would require the filesystem to
provide the necessary cloning or snapshot functionality. If that
capability is not available, the fork feature would not be enabled.
The goal is to avoid physically copying the entire cluster when
creating a fork. Unmodified data can remain shared between the parent
and child, while later writes create independent storage.
This differs from mechanisms such as CREATE DATABASE ... TEMPLATE ...,
since the unit being forked is the PostgreSQL cluster and its
execution environment, not a single database.
Prototype
I have implemented a prototype in a public branch:
https://github.com/LuquitasJeffrey/postgres/tree/pg_forks
The prototype is not intended to be merged in its current form.
It was built to validate the basic semantics and to see whether the
proposed model could work end to end. The implementation is highly
experimental, contains significant hacks and environment-specific
assumptions, and is not intended to represent the implementation
proposed for PostgreSQL.
I would prefer to discuss the semantics and architecture before trying
to evolve the prototype into a production-quality patch.
Questions
I'd be interested in feedback on whether this is a useful abstraction
for PostgreSQL, especially on the following points:
Is a cluster-level fork a useful abstraction distinct from existing
database cloning mechanisms?
Is using filesystem-provided CoW cloning the right boundary between
PostgreSQL and the storage layer?
What should be included in the cluster state captured by a fork?
Is maintaining fork metadata and fork IDs in the parent cluster an
appropriate model?
Are there PostgreSQL subsystems that would make an independently
running child cluster problematic?
Is the parent/child permission model reasonable, with changes limited
to the current fork and its descendants?
Are there existing PostgreSQL features or previous discussions that
overlap with this proposal?
The prototype is intended to demonstrate the idea, not to serve as a
patch for inclusion. At this stage, I am mainly interested in feedback
on the design and its interaction with PostgreSQL's existing
architecture.
Thanks,
Lucas
| From | Date | Subject | |
|---|---|---|---|
| Next Message | CharSyam | 2026-10-04 13:49:50 | [PATCH] Fix default partition validation for reordered child columns |
| Previous Message | Atsushi Ogawa | 2026-10-04 12:52:53 | Re: [PATCH] Use Boyer-Moore-Horspool for simple LIKE contains patterns |