Hi,

We chased a sporadic crash in our CI for a while and it turned out to be a live PostgreSQL bug, so here it is with a patch.

A checkpoint that invalidates an obsolete replication slot wipes the shared status flags of an unrelated backend. On an assert-enabled server that backend aborts at its next commit and the postmaster takes the whole cluster down with it. Without assertions nobody notices, and in the worst case vacuum then removes rows a standby still needs. It affects every branch from v14 up, and it is not primary-only: a standby reaches the same code during a restartpoint, and from v16 during ordinary replay.

The mechanism is a one-line indexing mistake. ReplicationSlotRelease() ends with:

        /* might not have been set when we've been a plain slot */
        LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
        MyProc->statusFlags &= ~PROC_IN_LOGICAL_DECODING;
        ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
        LWLockRelease(ProcArrayLock);

That is correct only for a process that is in the proc array. An auxiliary process never enters it, so its pgxactoff is still the zero InitProcGlobal() left there, and the store lands on the ProcGlobal->statusFlags[] entry of whichever backend owns offset 0.

The two statements are not the same operation. The bit clear applies to the caller's own private copy, which in an auxiliary process is zero anyway; the store is an assignment onto a foreign entry, so the victim loses every flag it holds, not just PROC_IN_LOGICAL_DECODING.

Auxiliary processes do get here. InvalidatePossiblyObsoleteSlot() takes the doomed slot by hand, setting MyReplicationSlot and active_pid under the slot's spinlock, and then lets go of it through ReplicationSlotRelease():

        CreateCheckPoint()                      checkpointer
        CreateRestartPoint()                    startup process
        xlog_redo()                             startup process, v16+
        ResolveRecoveryConflictWithSnapshot()   startup process, v16+
          -> InvalidateObsoleteReplicationSlots()
             -> InvalidatePossiblyObsoleteSlot()
                -> ReplicationSlotRelease()

The last two came in with 26669757b6a, which is why a standby can reach this during replay from v16 on. From v18 idle_replication_slot_timeout can trigger the invalidation as well, so max_slot_wal_keep_size is no longer needed to get there at all.

The damage surfaces later, and only sometimes. The victim has to reach the end of its transaction while running VACUUM or CREATE INDEX CONCURRENTLY. ProcArrayEndTransaction() then compares its private flags against ProcGlobal->statusFlags[]. In an --enable-cassert build the assertion there fires:

TRAP: failed Assert("proc->statusFlags == ProcGlobal->statusFlags[proc->pgxactoff]")
        LOG:  client backend was terminated by signal 6: Aborted

That is how we found it: the primary died while the checkpointer was invalidating a slot that had fallen behind max_slot_wal_keep_size, and autovacuum happened to be vacuuming pg_depend at that moment.

The victim is the proc array member with the lowest PGPROC slot number, normally the longest-connected regular backend, or an autovacuum worker when no client is connected. Whether the damage is noticed depends on what that process happens to be doing, which is why the crash looks random.

Without assertions nothing complains, and what the corruption costs depends on which flag the victim lost. Losing PROC_IN_VACUUM or PROC_IN_SAFE_IC only makes horizons more conservative. The next commit rewrites the entry anyway.

PROC_AFFECTS_ALL_HORIZONS is the bad one. InitWalSender() sets it on a walsender that connected without a database, and nothing sets it again for the life of that connection. If such a walsender is the victim, ComputeXidHorizons() stops applying its hot standby feedback xmin to data_oldest_nonremovable. Vacuum can then remove rows the standby still needs.

The attached patch adds a reproducer, src/test/recovery/t/057_slot_invalidation_statusflags.pl. It creates a physical slot nobody streams from and pushes it past max_slot_wal_keep_size. A conflicting ShareUpdateExclusiveLock parks a backend inside VACUUM. vacuum_rel() sets PROC_IN_VACUUM before it opens the relation, so that backend waits with the flag set. Then CHECKPOINT invalidates the slot.

I ran the test on REL_13_22, REL_14_24, REL_16_15, REL_17_11, REL_18_6 and master. v13 passes, everything from v14 up fails the same way. v15 and v19 carry the statement verbatim, so the range is v14 through master.

The patch skips the update unless PROC_IN_LOGICAL_DECODING is actually set. Only StartupDecodingContext() sets that flag and no auxiliary process reaches it, so nothing changes for a process in the proc array, while the checkpointer no longer executes the store at all.

With the patch applied on master I ran the core regression tests, plus the TAP suites in src/test/recovery and src/test/subscription. All of them pass.

Related, neither of them this bug:

https://postgr.es/m/tencent_CA7420C82971BC4B64F0748A7D2898A4720A%40qq.com
reports the same assertion from a different route, an exiting walsender whose pgxactoff was stale because ProcArrayRemove() had already compacted the ProcGlobal arrays. That one is v14-only, since 2f6501fa3c5 moved the slot release to before_shmem_exit() in v15, and nothing from the thread was committed.

1f2e51e3c7c, and its back-branch commits, fixed ReplicationSlotRelease() for a caller in single user mode. Same shape as this one: the function assumes its caller is an ordinary backend.


--
Best regards,
Vlad
From b600a25761f51fad938efa03f286165d215c366a Mon Sep 17 00:00:00 2001
From: Vlad Lesin <[email protected]>
Date: Wed, 23 Sep 2026 10:33:49 -0700
Subject: [PATCH v1] Fix clobbering of shared statusFlags when a slot is
 invalidated

ReplicationSlotRelease() ended with:

    MyProc->statusFlags &= ~PROC_IN_LOGICAL_DECODING;
    ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;

That is only correct for a process in the proc array.  Auxiliary
processes never join it, so their pgxactoff is still the zero it was
initialized to, and the store overwrites the entry of whichever backend
owns offset 0.

The checkpointer and the startup process both get there.
InvalidatePossiblyObsoleteSlot() acquires a slot that has fallen behind
max_slot_wal_keep_size and releases it through ReplicationSlotRelease();
the startup process does so during a restartpoint, and when replay
invalidates a slot that conflicts with recovery.

The two statements act on different processes.  The bit clear applies to
the caller's own private copy, which in an auxiliary process is zero
already; the store is an assignment rather than a bit clear, and it
lands on another backend's entry.  The backend at offset 0 therefore
loses not just PROC_IN_LOGICAL_DECODING but every flag it holds.

Skip the update unless PROC_IN_LOGICAL_DECODING is set.  Nothing but
StartupDecodingContext() sets that flag, and no auxiliary process
reaches it, so nothing changes for a process in the proc array.  Assert
as much at both places, adding AmAuxiliaryProcess(); the auxiliary
backend types are already contiguous in BackendType.

Back-patch to v14, where 5788e258bb2 introduced the array indexed by
pgxactoff.

Backpatch-through: 14
---
 src/backend/replication/logical/logical.c     |   3 +
 src/backend/replication/slot.c                |  23 +++-
 src/include/miscadmin.h                       |   3 +
 src/test/recovery/meson.build                 |   1 +
 .../t/057_slot_invalidation_statusflags.pl    | 129 ++++++++++++++++++
 5 files changed, 154 insertions(+), 5 deletions(-)
 create mode 100644 src/test/recovery/t/057_slot_invalidation_statusflags.pl

diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 98e5f1dd8f9..4d8335b4a7d 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -272,6 +272,9 @@ StartupDecodingContext(List *output_plugin_options,
 	 */
 	if (!IsTransactionOrTransactionBlock())
 	{
+		/* Only a proc array member has a pgxactoff of its own. */
+		Assert(!AmAuxiliaryProcess());
+
 		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
 		MyProc->statusFlags |= PROC_IN_LOGICAL_DECODING;
 		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
diff --git a/src/backend/replication/slot.c b/src/backend/replication/slot.c
index 63ce6d27885..663e44b2123 100644
--- a/src/backend/replication/slot.c
+++ b/src/backend/replication/slot.c
@@ -831,11 +831,24 @@ ReplicationSlotRelease(void)
 		MyReplicationSlot = NULL;
 	}
 
-	/* might not have been set when we've been a plain slot */
-	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
-	MyProc->statusFlags &= ~PROC_IN_LOGICAL_DECODING;
-	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
-	LWLockRelease(ProcArrayLock);
+	/*
+	 * This function is also reached from an auxiliary process:
+	 * InvalidatePossiblyObsoleteSlot() acquires the doomed slot on behalf of
+	 * the checkpointer, or of the startup process during a restartpoint, and
+	 * releases it here.  An auxiliary process is never entered into the proc
+	 * array, so its pgxactoff is still the zero it was initialized to, and
+	 * the store below would land on the ProcGlobal->statusFlags[] array entry
+	 * of whichever backend owns offset 0.
+	 */
+	if (MyProc->statusFlags & PROC_IN_LOGICAL_DECODING)
+	{
+		Assert(!AmAuxiliaryProcess());
+
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags &= ~PROC_IN_LOGICAL_DECODING;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
+	}
 
 	if (am_walsender)
 	{
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index 0fc59af02b9..c35c58a604f 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -410,6 +410,9 @@ extern PGDLLIMPORT BackendType MyBackendType;
 	(AmAutoVacuumLauncherProcess() || \
 	 AmLogicalSlotSyncWorkerProcess())
 
+#define AmAuxiliaryProcess() \
+	(MyBackendType >= B_ARCHIVER && MyBackendType <= B_WAL_WRITER)
+
 /*
  * Backend types that are spawned by the postmaster to serve a client or
  * replication connection. These backend types have in common that they are
diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index 72113c5ac6e..e8387ca7156 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -65,6 +65,7 @@ tests += {
       't/054_unlogged_sequence_promotion.pl',
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
+      't/057_slot_invalidation_statusflags.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/057_slot_invalidation_statusflags.pl b/src/test/recovery/t/057_slot_invalidation_statusflags.pl
new file mode 100644
index 00000000000..0fcb58761ae
--- /dev/null
+++ b/src/test/recovery/t/057_slot_invalidation_statusflags.pl
@@ -0,0 +1,129 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+#
+# A checkpoint that invalidates an obsolete replication slot must not corrupt
+# the shared ProcGlobal->statusFlags array.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+use Time::HiRes qw(usleep);
+
+my $node = PostgreSQL::Test::Cluster->new('primary');
+$node->init(allows_streaming => 1, extra => ['--wal-segsize=1']);
+
+# No checkpoint may happen on its own: the single CHECKPOINT this test issues
+# has to be the one that invalidates the slot, and it has to run while the
+# victim backend is parked inside VACUUM.
+$node->append_conf(
+	'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+min_wal_size = 2MB
+max_wal_size = 1GB
+max_slot_wal_keep_size = 1MB
+log_checkpoints = on
+
+# Cluster::init() turns this off, but a server that stays down would make
+# teardown bail out before the failure is reported.  Let it come back instead,
+# and report the crash from the log.
+restart_after_crash = on
+));
+$node->start;
+
+$node->safe_psql('postgres',
+	'CREATE TABLE vactbl AS SELECT generate_series(1, 1000) AS i');
+$node->safe_psql('postgres',
+	"SELECT pg_create_physical_replication_slot('doomed', true)");
+
+# Leave the slot -- which nobody is streaming from, so the checkpointer gets to
+# acquire it itself -- far enough behind that the next checkpoint has to
+# invalidate it.
+advance_wal($node, 2);
+
+# The control session runs over a walsender connection on purpose: its PGPROC
+# comes from the walsender free list, so it cannot take the dense-array slot
+# that the vacuuming backend below has to own.
+my $ctl = $node->background_psql('postgres', replication => 'database');
+
+# vacuum_rel() sets PROC_IN_VACUUM before it opens the relation, so a
+# conflicting ShareUpdateExclusiveLock parks the victim in exactly the state
+# the assertion at commit checks.
+$ctl->query_safe('BEGIN');
+$ctl->query_safe('LOCK TABLE vactbl IN SHARE UPDATE EXCLUSIVE MODE');
+
+my $vac = $node->background_psql('postgres', on_error_stop => 0);
+$vac->query_until(qr//, "VACUUM vactbl;\n");
+
+# Wait for it through the control session.  An ordinary poll_query_until()
+# would connect a second regular backend, and that one -- not the vacuum --
+# could end up owning dense slot 0.
+my $blocked = '';
+foreach my $i (1 .. 300)
+{
+	$blocked = $ctl->query_safe(
+		q{SELECT count(*) FROM pg_locks
+		   WHERE relation = 'vactbl'::regclass AND NOT granted});
+	last if $blocked eq '1';
+	usleep(100_000);
+}
+is($blocked, '1', 'VACUUM is parked waiting for the table lock');
+is( $ctl->query_safe(
+		q{SELECT count(*) FROM pg_stat_activity
+		   WHERE backend_type = 'client backend'}),
+	'1',
+	'the vacuuming backend is the only regular backend connected');
+
+isnt(
+	$ctl->query_safe(
+		q{SELECT wal_status FROM pg_replication_slots
+		   WHERE slot_name = 'doomed'}),
+	'lost',
+	'slot has not been invalidated yet');
+
+my $log_offset = -s $node->logfile;
+
+$ctl->query_safe('CHECKPOINT');
+
+is( $ctl->query_safe(
+		q{SELECT wal_status FROM pg_replication_slots
+		   WHERE slot_name = 'doomed'}),
+	'lost',
+	'checkpoint invalidated the obsolete slot');
+
+# Release the lock.  The VACUUM now runs to completion and commits, which is
+# where the clobbered flags are noticed.
+$ctl->query_safe('COMMIT');
+
+# A server that has just aborted takes this session down with it, so the query
+# has to be allowed to fail.
+my $vacuumed = eval { $vac->query_safe('SELECT 1') };
+
+# Older branches spell this "TRAP: FailedAssertion(", newer ones
+# "TRAP: failed Assert(".
+ok(!$node->log_contains(qr/TRAP: fail/i, $log_offset),
+	'no assertion failure while invalidating the slot');
+ok(!$node->log_contains(qr/was terminated by signal/, $log_offset),
+	'no backend was killed while invalidating the slot');
+is($vacuumed, '1', 'the parked VACUUM committed and its session survived');
+
+# Both sessions are gone if the server went down under them, so neither may be
+# allowed to take the test with it.
+eval { $vac->quit; };
+eval { $ctl->quit; };
+
+done_testing();
+
+# Burn through $n WAL segments.
+sub advance_wal
+{
+	my ($node, $n) = @_;
+
+	for (my $i = 0; $i < $n; $i++)
+	{
+		$node->safe_psql('postgres',
+			'CREATE TABLE t (); DROP TABLE t; SELECT pg_switch_wal();');
+	}
+	return;
+}

base-commit: 374522aa63af1187c15dd21db0b2c988b55c9a3a
-- 
2.52.0

Reply via email to