Sadanand Shenoy created HDDS-15962:
--------------------------------------
Summary: Remove leader readiness check on the bootstrap flow
Key: HDDS-15962
URL: https://issues.apache.org/jira/browse/HDDS-15962
Project: Apache Ozone
Issue Type: Bug
Reporter: Sadanand Shenoy
Assignee: Sadanand Shenoy
Current logic rejects checkpoint requests with 503 unless {{isLeaderReady()}}
is true.
In a degraded cluster (e.g., one OM down, one follower far behind), this
creates a loop:
# Lagging follower needs checkpoint to catch up.
# Follower requests {{/v2/dbCheckpoint}} from leader.
# Leader is {{LEADER_AND_NOT_READY}} and returns 503 due to
{{isLeaderReady()}} check.
# Follower cannot catch up, so leader readiness does not progress.
# System remains stuck until another OM is brought back.
Checkpoint creation already has consistency barriers:
- awaitDoubleBufferFlush()
- bootstrap write lock (BOOTSTRAP_LOCK)
So requiring LEADER_AND_READY here is stricter than necessary and can block
recovery. instead only reject if NOT_LEADER
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]