Hi all,

Following up on this: I filed
https://issues.apache.org/jira/browse/KAFKA-21190
and opened a PR with the fix and a regression test that uses the real
TierStateMachine: https://github.com/apache/kafka/pull/23650

The fix handles the equal case before the predecessor metadata lookup and
keeps the existing error when the remote range is not empty. Reviews are
welcome.

Thanks,
Jax

On Thu, Sep 24, 2026 at 10:29 AM Jax Hodgkinson <[email protected]>
wrote:

> Hi all,
>
> Before filing a JIRA, I'd like to confirm the intended KIP-1023 behavior
> for one edge case.
>
> On Kafka 4.3.0 with follower.fetch.last.tiered.offset.enable=true, we
> rebuilt an empty follower on a replacement disk. The selected fetch offset
> was 18203. That equals the leader's global log start offset, so the remote
> reconstruction range [18203, 18203) was empty.
>
> TierStateMachine still looked up remote segment metadata for 18202 (the
> selected offset minus one). The lookup returned empty, and follower
> initialization kept retrying, which left the replica out of the ISR. The
> relevant log was:
>
>
> {"source_class":"kafka.server.ReplicaFetcherThread","level":"ERROR","time":"2026-07-28T19:52:09,927","msg":"[ReplicaFetcher
> replicaId=0, leaderId=7, fetcherId=0] Error getting offset for partition
> <TOPIC-PARTITION> 
> org.apache.kafka.server.log.remote.storage.RemoteStorageException:
> Couldn't build the state from remote store for partition:
> <TOPIC-PARTITION>, currentLeaderEpoch: 406, leaderLocalLogStartOffset:
> 18203, leaderLogStartOffset: 18203, epoch: 405 as the previous remote log
> segment metadata was not found
>           at
> kafka.server.TierStateMachine.buildRemoteStorageException(TierStateMachine.java:278)
>           at
> kafka.server.TierStateMachine.lambda$buildRemoteLogAuxState$0(TierStateMachine.java:237)
>           at java.base/java.util.Optional.orElseThrow(Optional.java:403)
>           at
> kafka.server.TierStateMachine.buildRemoteLogAuxState(TierStateMachine.java:237)
>           at kafka.server.TierStateMachine.start(TierStateMachine.java:113)
>           at
> kafka.server.AbstractFetcherThread.fetchOffsetAndTruncate(AbstractFetcherThread.scala:679)
>           at
> kafka.server.AbstractFetcherThread.handleOutOfRangeError(AbstractFetcherThread.scala:749)"}
>
> Relevant code:
>
> https://github.com/apache/kafka/blob/4.3.0/core/src/main/java/kafka/server/TierStateMachine.java#L201-L238
>
> KIP-1023 says that when there are no valid remote segments, the follower
> should truncate and replicate locally. That seems to fit this case.
>
> Two existing tests cover equal positive start offsets
> (`testLastTieredOffsetWithNonZeroLSOOnNewLeader` and
> `testLastTieredOffsetWithSlowUploadNoLocalDeletion`), but both use
> `MockTierStateMachine`. The mock truncates directly to the selected offset
> and bypasses the predecessor metadata lookup, so these tests do not
> exercise the failing path.
>
> My questions:
>
> 1. When the selected offset equals the global log start offset, shouldn't
> the follower skip remote reconstruction and start a normal local fetch from
> log start?
>
> 2. If so, is an equality guard before the predecessor metadata lookup the
> right fix boundary? The existing error would stay in place for a genuinely
> non-empty remote range.
>
> Dynamically disabling follower.fetch.last.tiered.offset.enable bypassed
> this path and let the replica recover.
>
> If this matches the intended behavior, I'm happy to file an Apache JIRA
> and contribute a patch with a regression test that uses the real
> TierStateMachine.
>
> Thanks,
> Jax
>

Reply via email to