Thanks Federico! The solutions for both RS25 and RS27 sound good. Regards,
Rajini On Mon, Aug 31, 2026 at 5:12 PM Federico Valeri <[email protected]> wrote: > Hi Rajini, > > >Not sure about RS25, but you mentioned you are looking into it again, so > will wait for the response after that. > > The clamping first attempts to read LSO and LEO from the local > ReplicaManager, which is an in-memory lookup against the local > UnifiedLog instance. This works for any partition the coordinator > broker happens to host a replica for. For partitions without a local > replica, the coordinator falls back to the Admin listOffsets API, > which internally routes the request to the partition leader and > returns the offsets. We resolve all mirror partition offsets once at > the beginning of each sync cycle, then reuse them for every group. Not > one call per partition or per group. > > KIP Updated. > > > For RS27, not supporting unclean leader election for older clients may > be ok unless the fix is straightforward. > > We think it is possible to support old versions with the following > strategy: > > We detect the source cluster's fetch version dynamically. On the first > fetch request, the fetcher thread always includes lastFetchedEpoch in > the request builder. The NetworkClient then queries ApiVersions from > the source cluster and selects the latestUsableVersion. Two outcomes > are possible: > > 1. Source supports fetch v12+: the request succeeds with > lastFetchedEpoch, and truncation proceeds via the truncation-on-fetch > path. > 2. Source does not support fetch v12: NetworkClient throws > UnsupportedVersionException. We catch it, disable truncation-on-fetch, > and fall back to the OffsetsForLeaderEpochRequest path. > > This handles both old and new source clusters, including live upgrades > and downgrades. > > Unclean leader election handling: > > Previously we only handled unclean leader election (ULE) in the > truncation-on-fetch case. To support older clusters via > OffsetsForLeaderEpochRequest, we apply the same principle: log > truncation should never happen on the destination cluster during > mirroring because we only fetch committed data. If an > OffsetsForLeaderEpochResponse does trigger truncation, that truncation > must originate from a ULE on the source cluster. In that case, we > transition to the ULE_RECOVERY state and wait for all replicas to > catch up (when mirror.unclean.leader.election.enable=true). > > To guarantee that all destination truncations come from source ULEs, > we store the "last mirror offset and last mirror epoch" in the LME > record. Originally, the LME phase one truncated records with epoch > greater than the LME epoch, then phase two relied on normal > replication to handle log divergence. The problem: some truncations > are side effects of mirroring setup (e.g. reverse mirroring), not > ULEs. > > Example (cluster A mirrors from cluster B): > > Cluster A (dest) <---- Cluster B (source) > Offset 0: A (LE: 0) Offset 0: A (LE: 0) > Offset 1: A (LE: 0) Offset 1: A (LE: 0) > Offset 2: A (LE: 0) Offset 2: PID_RESET (LE: 10) > > When cluster A receives the first OffsetsForLeaderEpochResponse or > fetch response from cluster B, offset 2 on cluster A gets truncated. > This truncation is not caused by a ULE. Without the last mirror offset > and epoch in the LME record, we cannot distinguish it from a real ULE. > By storing these values in the LME record, the LOG_ALIGNMENT state > truncates only records where: record offset <= last mirror offset, AND > record epoch <= last mirror epoch. This correctly truncates offset 2 > in the example above during LOG_ALIGNMENT. After alignment completes, > any subsequent truncation on the destination is guaranteed to > originate from a source cluster ULE. > > Let us know if you are good with this solution. KIP update will follow. > > Thanks > Fede >
