From 0513c3535516525627219eb8d6c7e76837ff6afc Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?=C3=81lvaro=20Herrera?= <alvherre@kurilemu.de>
Date: Tue, 6 Oct 2026 17:21:58 +0200
Subject: [PATCH v4] Revamp REPACK doc refentry page

Limit the commentary under "Description" and "Parameters" to a minimum;
move the existing text to appear in the "Notes" section.  Both the
"Notes on Clustering" and "Notes on Resources" subsections, which were
under "Description", are moved to be under "Notes" instead.  Add a new
"Notes on Concurrent Operation" subsection there, which now carries some
text that was in "Parameters", and gets some additional text to
(hopefully) explain resource consumption more clearly.

Add some appropriate cross-links.

Discussion: https://postgr.es/m/CAJgoLk+dodrwuwCuXERGYwgjQzSwrKd+xgitYrL32oWDt39zvA@mail.gmail.com
---
 doc/src/sgml/ref/repack.sgml | 344 +++++++++++++++++++----------------
 1 file changed, 191 insertions(+), 153 deletions(-)

diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 346cba89c90..bea95e4c289 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -61,116 +61,17 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
 
   <para>
    If a <literal>USING INDEX</literal> clause is specified, the rows are
-   physically reordered based on information from an index.  Please see the
-   notes on clustering below.
+   physically reordered based on information from an index.  Please see
+   <xref linkend="sql-repack-notes-on-clustering"/> below.
   </para>
 
   <para>
    When a table is being repacked, an <literal>ACCESS EXCLUSIVE</literal> lock
-   is acquired on it. This prevents any other database operations (both reads
-   and writes) from operating on the table until the <command>REPACK</command>
-   is finished. If you want to keep the table accessible during the repacking,
-   consider using the <literal>CONCURRENTLY</literal> option.
+   is acquired on it, unless the <literal>CONCURRENTLY</literal> option is given.
+   Please see <xref linkend="sql-repack-notes-on-concurrently"/> for a discussion
+   on the effects of this option.
   </para>
 
-  <refsect2 id="sql-repack-notes-on-clustering" xreflabel="Notes on Clustering">
-   <title>Notes on Clustering</title>
-
-   <para>
-    If the <literal>USING INDEX</literal> clause is specified, the rows in
-    the table are physically rearranged according to the ordering implied by
-    the specified index; this is known as <firstterm>clustering</firstterm>.
-    This can have performance implications:
-    in cases where you are accessing single rows randomly within a table, the
-    actual order of the data in the table is unimportant. However, if you tend
-    to access some data more than others, and there is an index that groups
-    them together, you will benefit from using clustering.  If
-    you are requesting a range of indexed values from a table, or a single
-    indexed value that has multiple matching rows,
-    clustering will help because once the index identifies the
-    table page for the first row that matches, all other rows that match are
-    probably already on the same table page, and so you save disk accesses and
-    speed up the query.
-   </para>
-
-   <para>
-    For B-Tree indexes, the ordering used by clustering is the index's linear
-    sort order.  Other clusterable index access methods may use a different
-    ordering strategy, one which may not necessarily correspond to any SQL
-    sort order.
-    If an index name is specified in the command, that index is used and
-    is recorded as the table's clustering index.
-    (This also applies to an index given to the <command>CLUSTER</command>
-    command.)
-    If no index name is specified, then the index that has been configured as
-    the clustering one is used; if none has been configured, an error is thrown.
-    An index can be set manually using <command>ALTER TABLE ... CLUSTER ON</command>,
-    and reset with <command>ALTER TABLE ... SET WITHOUT CLUSTER</command>.
-   </para>
-
-   <para>
-    Clustering is a one-time operation: when the table is
-    subsequently updated, the changes are not clustered.  That is, no attempt
-    is made to store new or updated rows according to the clustering order.
-    (If one wishes, one can periodically recluster by issuing the command again.
-    Also, setting the table's <literal>fillfactor</literal> storage parameter
-    to less than 100% can aid in preserving cluster ordering during updates,
-    since updated rows are kept on the same page if enough space is available
-    there.)
-   </para>
-
-   <para>
-    When clustering on a B-Tree index, <command>REPACK</command> can rewrite
-    the table using either an index scan on the specified index, or a
-    sequential scan followed by sorting.  It will attempt to choose the method
-    that will be faster, based on planner cost parameters and available
-    statistical information.  When clustering on an index of an access method
-    other than B-Tree, <command>REPACK</command> always uses an index scan.
-   </para>
-
-   <para>
-    Because the planner records statistics about the ordering of tables, it is
-    advisable to specify the <literal>ANALYZE</literal> option, or to
-    run <link linkend="sql-analyze"><command>ANALYZE</command></link> on the
-    newly repacked table.  Otherwise, the planner might make poor choices of
-    query plans.
-   </para>
-
-   <para>
-    If no table name is specified in <command>REPACK USING INDEX</command>,
-    all tables which have a clustering index defined and which the calling
-    user has privileges for are processed.
-   </para>
-
-  </refsect2>
-
-  <refsect2 id="sql-repack-notes-on-resources" xreflabel="Notes on Resources">
-   <title>Notes on Resources</title>
-
-   <para>
-    When an index scan or a sequential scan without sort is used, a temporary
-    copy of the table is created that contains the table data in the index
-    order.  Temporary copies of each index on the table are created as well.
-    Therefore, you need free space on disk at least equal to the sum of the
-    table size and the index sizes.
-   </para>
-
-   <para>
-    When a sequential scan and sort is used, a temporary sort file is also
-    created, so that the peak temporary space requirement is as much as double
-    the table size, plus the index sizes.  This method is often faster than
-    the index scan method, but if the disk space requirement is intolerable,
-    you can disable this choice by temporarily setting
-    <xref linkend="guc-enable-sort"/> to <literal>off</literal>.
-   </para>
-
-   <para>
-    It is advisable to set <xref linkend="guc-maintenance-work-mem"/> to a
-    reasonably large value (but not more than the amount of RAM you can
-    dedicate to the <command>REPACK</command> operation) before repacking.
-   </para>
-  </refsect2>
-
  </refsect1>
 
  <refsect1>
@@ -211,56 +112,9 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
     <listitem>
      <para>
       Allow other transactions to use the table while it is being repacked.
+      See <xref linkend="sql-repack-notes-on-concurrently" /> for more details.
      </para>
 
-     <para>
-      Internally, <command>REPACK</command> copies the contents of the table
-      (ignoring dead tuples) into a new file, sorted by the specified index,
-      and also creates a new file for each index. Then it swaps the old and
-      new files for the table and all the indexes, and deletes the old
-      files. The <literal>ACCESS EXCLUSIVE</literal> lock is needed to make
-      sure that the old files do not change during the processing because the
-      changes would get lost due to the swap.
-     </para>
-
-     <para>
-      With the <literal>CONCURRENTLY</literal> option, the <literal>ACCESS
-      EXCLUSIVE</literal> lock is only acquired to swap the table and index
-      files. The data changes that took place during the creation of the new
-      table and index files are captured using logical decoding
-      (<xref linkend="logicaldecoding"/>) and applied before
-      the <literal>ACCESS EXCLUSIVE</literal> lock is requested. Thus the lock
-      is typically held only for the time needed to swap the files, which
-      should be pretty short. However, the time might still be noticeable if
-      too many data changes have been done to the table while
-      <command>REPACK</command> was waiting for the lock: those changes must
-      be processed just before the files are swapped, while the
-      <literal>ACCESS EXCLUSIVE</literal> lock is being held.
-     </para>
-
-     <para>
-      Note that <command>REPACK</command> with the
-      <literal>CONCURRENTLY</literal> option does not try to order the rows
-      inserted into the table after the repacking started. Also
-      note <command>REPACK</command> might fail to complete due to DDL
-      commands executed on the table by other transactions during the
-      repacking.
-     </para>
-
-     <note>
-      <para>
-       In addition to the temporary space requirements explained in
-       <xref linkend="sql-repack-notes-on-resources"/>,
-       the <literal>CONCURRENTLY</literal> option can add to the usage of
-       temporary space a bit more. The reason is that other transactions can
-       perform DML operations which cannot be applied to the new file until
-       <command>REPACK</command> has copied all the existing tuples from the
-       old file. Thus the tuples inserted into the old file during the copying
-       are also stored separately in a temporary file, until they can be
-       processed.
-      </para>
-     </note>
-
      <para>
       The <literal>CONCURRENTLY</literal> option cannot be used in the
       following cases:
@@ -376,7 +230,7 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
   </variablelist>
  </refsect1>
 
- <refsect1>
+ <refsect1 id="sql-repack-notes" xreflabel="Notes">
   <title>Notes</title>
 
    <para>
@@ -384,6 +238,23 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
     on the table.
    </para>
 
+   <para>
+    <command>REPACK</command> is primarily meant to remove bloat and, with
+    <literal>USING INDEX</literal>, to cluster the table, in order to improve
+    performance.  Although it also advances the table's
+    <structfield>relfrozenxid</structfield> and <structfield>relminmxid</structfield>,
+    it is not a good way to prevent transaction ID or multixact ID wraparound:
+    it rewrites the whole table and all of its indexes, so it takes much longer
+    than <command>VACUUM</command>, and it can fail after most of the work
+    is done, especially with <literal>CONCURRENTLY</literal> on a busy table.
+    <command>VACUUM</command> is recommended for that instead (see
+    <xref linkend="vacuum-for-wraparound"/>).
+    When a table gets close to wraparound, <command>VACUUM</command> skips index
+    vacuuming on its own (see <xref linkend="guc-vacuum-failsafe-age"/>), and its
+    <link linkend="sql-vacuum"><literal>INDEX_CLEANUP OFF</literal></link>
+    option can do the same earlier.
+   </para>
+
    <para>
     <command>REPACK</command> refuses to process a table on which an invalid
     index exists.  Such indexes must be dropped or reindexed by the user ahead
@@ -409,6 +280,173 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
     inside a transaction block.
    </para>
 
+  <refsect2 id="sql-repack-notes-on-clustering" xreflabel="Notes on Clustering">
+   <title>Notes on Clustering</title>
+
+   <para>
+    If the <literal>USING INDEX</literal> clause is specified, the rows in
+    the table are physically rearranged according to the ordering implied by
+    the specified index; this is known as <firstterm>clustering</firstterm>.
+    This can have performance implications:
+    in cases where you are accessing single rows randomly within a table, the
+    actual order of the data in the table is unimportant. However, if you tend
+    to access some data more than others, and there is an index that groups
+    them together, you will benefit from using clustering.  If
+    you are requesting a range of indexed values from a table, or a single
+    indexed value that has multiple matching rows,
+    clustering will help because once the index identifies the
+    table page for the first row that matches, all other rows that match are
+    probably already on the same table page, and so you save disk accesses and
+    speed up the query.
+   </para>
+
+   <para>
+    For B-Tree indexes, the ordering used by clustering is the index's linear
+    sort order.  Other clusterable index access methods may use a different
+    ordering strategy, one which may not necessarily correspond to any SQL
+    sort order.
+    If an index name is specified in the command, that index is used and
+    is recorded as the table's clustering index.
+    (This also applies to an index given to the <command>CLUSTER</command>
+    command.)
+    If no index name is specified, then the index that has been configured as
+    the clustering one is used; if none has been configured, an error is thrown.
+    An index can be set manually using <command>ALTER TABLE ... CLUSTER ON</command>,
+    and reset with <command>ALTER TABLE ... SET WITHOUT CLUSTER</command>.
+   </para>
+
+   <para>
+    Clustering is a one-time operation: when the table is
+    subsequently updated, the changes are not clustered.  That is, no attempt
+    is made to store new or updated rows according to the clustering order.
+    (If one wishes, one can periodically recluster by issuing the command again.
+    Also, setting the table's <literal>fillfactor</literal> storage parameter
+    to less than 100% can aid in preserving cluster ordering during updates,
+    since updated rows are kept on the same page if enough space is available
+    there.)
+   </para>
+
+   <para>
+    When clustering on a B-Tree index, <command>REPACK</command> can rewrite
+    the table using either an index scan on the specified index, or a
+    sequential scan followed by sorting.  It will attempt to choose the method
+    that will be faster, based on planner cost parameters and available
+    statistical information.  When clustering on an index of an access method
+    other than B-Tree, <command>REPACK</command> always uses an index scan.
+   </para>
+
+   <para>
+    Because the planner records statistics about the ordering of tables, it is
+    advisable to specify the <literal>ANALYZE</literal> option, or to
+    run <link linkend="sql-analyze"><command>ANALYZE</command></link> on the
+    newly repacked table.  Otherwise, the planner might make poor choices of
+    query plans.
+   </para>
+
+   <para>
+    If no table name is specified in <command>REPACK USING INDEX</command>,
+    all tables which have a clustering index defined and which the calling
+    user has privileges for are processed.
+   </para>
+
+  </refsect2>
+
+  <refsect2 id="sql-repack-notes-on-resources" xreflabel="Notes on Resources">
+   <title>Notes on Resources</title>
+
+   <para>
+    When the <literal>USING INDEX</literal> clause is omitted, or when that
+    clause is given but an index scan is chosen, a temporary
+    copy of the table is created that contains the table data in the index
+    order.  Temporary copies of each index on the table are created as well.
+    Therefore, you need free space on disk at least equal to the sum of the
+    table size and the index sizes.
+   </para>
+
+   <para>
+    When <literal>USING INDEX</literal> is given and a sequential scan and sort
+    is used, a temporary sort file is also
+    created, so that the peak temporary space requirement is as much as double
+    the table size, plus the index sizes.  This method is often faster than
+    the index scan method, but if the disk space requirement is intolerable,
+    you can disable this choice by temporarily setting
+    <xref linkend="guc-enable-sort"/> to <literal>off</literal> or omitting
+    <literal>USING INDEX</literal>.
+   </para>
+
+   <para>
+    It is advisable to set <xref linkend="guc-maintenance-work-mem"/> to a
+    reasonably large value (but not more than the amount of RAM you can
+    dedicate to the <command>REPACK</command> operation) before repacking.
+   </para>
+
+  </refsect2>
+
+  <refsect2 id="sql-repack-notes-on-concurrently" xreflabel="Notes on Concurrent Operation">
+   <title>Notes on Concurrent Operation</title>
+
+   <para>
+    <command>REPACK</command> copies the contents of the table
+    (ignoring dead tuples) into a new file, sorted by the specified index,
+    and also creates a new file for each index. It then swaps the old and
+    new files for the table and all the indexes, and deletes the old files.
+    Without the <literal>CONCURRENTLY</literal> option, an
+    <literal>ACCESS EXCLUSIVE</literal> lock is acquired at the beginning
+    and remains held throughout the operation to make sure that the old
+    files do not change during the processing; otherwise, any concurrent
+    changes would get lost due to the swap.
+   </para>
+
+   <para>
+    By contrast, when the <literal>CONCURRENTLY</literal> option is
+    specified, the bulk of the operation is run with
+    <literal>SHARE UPDATE EXCLUSIVE</literal> lock, and the
+    <literal>ACCESS EXCLUSIVE</literal> lock is only acquired
+    during the final phase to swap the table and index files.
+    The data changes that took place during the creation of the new
+    table and index files are captured using logical decoding
+    (see <xref linkend="logicaldecoding"/>) and applied before
+    the <literal>ACCESS EXCLUSIVE</literal> lock is requested.
+    Once that lock is obtained, a final pass over any remaining captured
+    concurrent data changes is done and the files are swapped.
+    Thus the lock is typically held for a short time, depending
+    on the amount of changes accumulated while the lock was being
+    waited for.
+   </para>
+
+   <para>
+    Processing of these concurrent changes requires a fixed small
+    amount of memory for each tuple concurrently updated or deleted,
+    not limited by <xref linkend="guc-maintenance-work-mem"/>;
+    if more than about 104 million rows are concurrently updated or
+    deleted during the execution of <command>REPACK</command>, the
+    command fails.
+   </para>
+
+   <para>
+    For the purposes of transaction ID wraparound
+    (see <xref linkend="vacuum-for-wraparound"/>),
+    <command>REPACK</command> is considered a single long-running
+    transaction, which prevents <command>VACUUM</command> from cleaning
+    up dead rows from other tables.
+    It is advisable to monitor <structname>pg_stat_activity.backend_xid</structname>
+    for the process running <command>REPACK</command> when repacking
+    very large tables, to avoid causing excessive bloat in other tables.
+   </para>
+
+   <para>
+    <command>REPACK (CONCURRENTLY)</command> might fail to complete if DDL
+    commands are executed on the table by other transactions during the
+    repacking.
+   </para>
+
+   <para>
+    <command>REPACK (CONCURRENTLY) USING INDEX</command>
+    does not try to order the rows inserted into the table after the
+    repacking started.
+   </para>
+
+  </refsect2>
  </refsect1>
 
  <refsect1>
-- 
2.47.3

