From: | Robert Haas <robertmhaas(at)gmail(dot)com> |
---|---|
To: | Melanie Plageman <melanieplageman(at)gmail(dot)com> |
Cc: | Andres Freund <andres(at)anarazel(dot)de>, Kirill Reshke <reshkekirill(at)gmail(dot)com>, Andrey Borodin <x4mmm(at)yandex-team(dot)ru>, PostgreSQL Hackers <pgsql-hackers(at)lists(dot)postgresql(dot)org>, Heikki Linnakangas <hlinnaka(at)iki(dot)fi> |
Subject: | Re: eliminate xl_heap_visible to reduce WAL (and eventually set VM on-access) |
Date: | 2025-09-08 20:14:47 |
Message-ID: | CA+TgmoasgmY7mzZutGisD2=3y7BwwPUS=oNsQoORKRg1r69fEA@mail.gmail.com |
Views: | Whole Thread | Raw Message | Download mbox | Resend email |
Thread: | |
Lists: | pgsql-hackers |
Reviewing 0003:
+ /*
+ * If we're only adding already frozen rows to a
previously empty
+ * page, mark it as all-frozen and update the
visibility map. We're
+ * already holding a pin on the vmbuffer.
+ */
else if (all_frozen_set)
+ {
PageSetAllVisible(page);
+ LockBuffer(vmbuffer, BUFFER_LOCK_EXCLUSIVE);
+ visibilitymap_set_vmbits(relation,
+
BufferGetBlockNumber(buffer),
+
vmbuffer,
+
VISIBILITYMAP_ALL_VISIBLE |
+
VISIBILITYMAP_ALL_FROZEN);
Locking a buffer in a critical section violates the order of
operations proposed in the 'Write-Ahead Log Coding' section of
src/backend/access/transam/README.
+ * Now read and update the VM block. Even if we skipped
updating the heap
+ * page due to the file being dropped or truncated later in
recovery, it's
+ * still safe to update the visibility map. Any WAL record that clears
+ * the visibility map bit does so before checking the page LSN, so any
+ * bits that need to be cleared will still be cleared.
+ *
+ * It is only okay to set the VM bits without holding the heap page lock
+ * because we can expect no other writers of this page.
The first paragraph of this paraphrases a similar content in
xlog_heap_visible(), but I don't see the variation in phrasing as an
improvement.
The second paragraph does not convince me at all. I see no reason to
believe that this is safe, or that it is a good idea. The code in
xlog_heap_visible() thinks its OK to unlock and relock the page to
make visibilitymap_set() happy, which is cringy but probably safe for
lack of concurrent writers, but skipping locking altogether seems
deeply unwise.
- * visibilitymap_set - set a bit in a previously pinned page
+ * visibilitymap_set - set bit(s) in a previously
pinned page and log
+ * visibilitymap_set_vmbits - set bit(s) in a pinned page
I suspect the indentation was done with a different mix of spaces and
tabs here, because this doesn't align for me.
In general, this idea makes some sense to me -- there doesn't seem to
be any particularly good reason why the visibility-map update should
be handled by a different WAL record than the all-visible flag on the
page itself. It's a little hard for me to make that statement too
conclusively without studying more of the patches than I've had time
to do today, but off the top of my head it seems to make sense.
However, I'm not sure you've taken enough care with the details here.
--
Robert Haas
EDB: http://www.enterprisedb.com
From | Date | Subject | |
---|---|---|---|
Next Message | Sami Imseih | 2025-09-08 20:34:17 | Re: GetNamedLWLockTranche crashes on Windows in normal backend |
Previous Message | Nathan Bossart | 2025-09-08 20:14:01 | Re: Should io_method=worker remain the default? |