[ 
https://issues.apache.org/jira/browse/KAFKA-20961?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

hyungmin OH reassigned KAFKA-20961:
-----------------------------------

    Assignee: hyungmin OH

> Loop of OffsetOutOfRangeException
> ---------------------------------
>
>                 Key: KAFKA-20961
>                 URL: https://issues.apache.org/jira/browse/KAFKA-20961
>             Project: Kafka
>          Issue Type: Bug
>          Components: replication
>    Affects Versions: 4.1.1
>            Reporter: David Dufour
>            Assignee: hyungmin OH
>            Priority: Major
>
> Hello,
> we have identified scenarios where a leader broker returns 
> OffsetOutOfRangeException error to the leader in loop.
> This is probably what was described in KAFKA-15608 but this issue was logged 
> on an old version and is still in status open.
> We are experiencing this issue very regularly as we are using JBOD and we 
> move partitions from one disk to another inside a same broker. 
> Here is an example of logs:
> {+}Leader:
> {+}INFO Attempting to replace current log Log(...) with Log(...) for xxxxxx 
> (kafka.log.LogManager) [ReplicaAlterLogDirsThread-3]
> ...\{+}
> {+}INFO The current replica is successfully replaced with the future replica 
> for xxxxxx (kafka.log.LogManager) [ReplicaAlterLogDirsThread-3]
> +Follower, few ms later:+
> INFO [ReplicaFetcher replicaId=1, leaderId=3, fetcherId=0] Reset fetch offset 
> for partition xxxxxx from 0 to the current local replica's end offset 0 
> (kafka.server.ReplicaFetcherThread) [ReplicaFetcherThread-0-3]
> INFO [ReplicaFetcher replicaId=1, leaderId=3, fetcherId=0] Current offset 0 
> for partition xxxxxx is out of range, which typically implies a leader 
> change. Reset fetch offset to 0 (kafka.server.ReplicaFetcherThread) 
> [ReplicaFetcherThread-0-3]
> ....repeat every ~12ms
> Often, the loop stops by itself after some time (traffic injected, broker 
> restart,...).
> We use Cruise Control to perform cluster rebalance. Recently, a rebalance 
> triggered a huge number of such loops that never stopped. The logs were 
> flooded and in the end, the cluster became unresponsive to the clients, 
> causing a severe outage. A rolling restart recovered the situation.
> +How to reproduce:+
> 1- create new empty topic
> 2- build some leader epoch (for ex: switch leadership from one broker to 
> another)
> 3- move partition on leader from dir1 to dir2
>  



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to