SEPURI-SAI-KRISHNA commented on issue #9940: URL: https://github.com/apache/seatunnel/issues/9940#issuecomment-5690437129
I think there may be a third path to this symptom that is not covered by either of the two conclusions recorded above, and I have opened #12332 for it rather than reopening the discussion here. To be clear about what it is not, since both points above look right to me: - It is not the checkpoint restore path that #10778 fixed. It needs no restore and no failover, and reproduces on an ordinary start. - It is not the cold start with genuinely no broker-side committed offset, which as you said follows RocketMQ group-offset behavior and is not a SeaTunnel bug. The case I am describing is one where the committed offsets do exist on the broker and are simply unreadable for a moment. `RocketMqAdminUtil.currentOffsets` catches `MQClientException` and returns an empty map whenever the response code is `TOPIC_NOT_EXIST`. In RocketMQ 4.9.4, which this connector pins, `ResponseCode.TOPIC_NOT_EXIST` is 17, and 17 is also the code carried by `No topic route info in name server for the topic: <topic>`, which the name server raises for a topic that exists but whose route is momentarily unavailable. The enumerator then reads that empty map as "this group has committed nothing" and starts from the first offset. While working on an unrelated e2e problem I measured one of these route gaps directly in CI: the name server returned no route for an existing, actively used topic 10 consecutive times at about 30 second intervals, a continuous window of 4 minutes 24 seconds. With `partition.discovery.interval.millis = "1000"` as in the config above, a job re-enters that lookup often enough that it only has to be unlucky once. I want to be honest about the limits of this: I have not captured a trace of a production job rewinding through this specific path, so it is established from the code and the response code rather than from a reproduction, and your report is consistent with it but does not prove it. If you still see consumption restarting from 0 on a plain start after #10778, the useful evidence would be whether the broker still shows committed offsets for `track_report_group` at that moment, and whether any `CODE: 17` route warnings appear in the job log around startup. That would separate this from the two causes already ruled out here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
