[ 
https://issues.apache.org/jira/browse/RATIS-2605?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Tsz-wo Sze reassigned RATIS-2605:
---------------------------------

    Component/s: server
       Assignee: Ethan Feng

>   Heartbeat AppendEntries success may incorrectly advance leader-side 
> matchIndex
> --------------------------------------------------------------------------------
>
>                 Key: RATIS-2605
>                 URL: https://issues.apache.org/jira/browse/RATIS-2605
>             Project: Ratis
>          Issue Type: Bug
>          Components: server
>            Reporter: Ethan Feng
>            Assignee: Ethan Feng
>            Priority: Major
>          Time Spent: 20m
>  Remaining Estimate: 0h
>
>  In the current AppendEntries success handling, the leader uses 
> reply.nextIndex
>   to advance the follower's matchIndex and nextIndex for all successful
>   AppendEntries replies.
>   This is not always safe for heartbeat AppendEntries.
>   On the follower side, a successful heartbeat reply returns the follower's 
> local
>   state.getNextIndex().  This value describes the follower's local log tail, 
> but
>   it does not prove that all entries up to nextIndex - 1 match the leader's 
> log.
>   A heartbeat success only proves that the request's previousLog, if present, 
> is
>   accepted by the follower.
>   This can corrupt the leader-side matchIndex when the follower has 
> uncommitted
>   entries from an old leader.
>   One possible sequence is:
>   1. A server was previously leader and appended some uncommitted entries.
>   2. A new leader is elected in a higher term and appends different entries at
>      the same indexes.
>   3. The old leader becomes a follower, but still has the old uncommitted tail
>      locally.
>   4. The new leader sends heartbeat AppendEntries to that follower.
>   5. The follower replies SUCCESS and returns its local nextIndex.
>   6. The leader treats reply.nextIndex - 1 as a matched index and advances
>      matchIndex.
>   7. Later, when real AppendEntries requests hit the conflicting tail, the
>      follower correctly replies INCONSISTENCY because previousLog does not 
> match.
>   8. The leader-side inconsistency backoff uses matchIndex + 1 as a lower 
> bound,
>      so the polluted matchIndex can prevent nextIndex from backing off to the
>      real common prefix.
>   9. The follower may remain stuck replying INCONSISTENCY instead of reaching 
> the
>      point where the conflicting tail can be truncated.
>   The root cause is that heartbeat SUCCESS and append SUCCESS have different
>   proof strength:
>   - AppendEntries with log entries proves the appended entries were accepted, 
> so
>     reply.nextIndex can be used to advance matchIndex.
>   - Heartbeat AppendEntries proves only the request previousLog, not the
>     follower's whole local tail.
>   The leader should therefore avoid advancing matchIndex from the follower's
>   local nextIndex in heartbeat SUCCESS replies.  For heartbeat SUCCESS, the
>   leader should advance at most to request.previousLog.index + 1.  Normal 
> append
>   SUCCESS can continue to use reply.nextIndex.
>   This issue can cause prolonged log replication stalls with repeated
>   INCONSISTENCY replies after leader changes, especially when an old leader 
> has
>   a divergent uncommitted log tail.
>   Proposed fix:
>   Preserve the AppendEntries request metadata together with the reply.  When
>   handling SUCCESS, use request.previousLog.index + 1 as the upper bound for
>   heartbeat requests, and keep using reply.nextIndex for requests that contain
>   log entries.
>   Test plan:
>   Add unit coverage for:
>   1. Heartbeat SUCCESS with previousLog advances only to previousLog.index + 
> 1.
>   2. Heartbeat SUCCESS without previousLog does not advance matchIndex.
>   3. Append SUCCESS still advances using reply.nextIndex.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to