SEPURI-SAI-KRISHNA commented on issue #9940:
URL: https://github.com/apache/seatunnel/issues/9940#issuecomment-5690437129

   I think there may be a third path to this symptom that is not covered by 
either of the two conclusions recorded above, and I have opened #12332 for it 
rather than reopening the discussion here.
   
   To be clear about what it is not, since both points above look right to me:
   
   - It is not the checkpoint restore path that #10778 fixed. It needs no 
restore and no failover, and reproduces on an ordinary start.
   - It is not the cold start with genuinely no broker-side committed offset, 
which as you said follows RocketMQ group-offset behavior and is not a SeaTunnel 
bug.
   
   The case I am describing is one where the committed offsets do exist on the 
broker and are simply unreadable for a moment. 
`RocketMqAdminUtil.currentOffsets` catches `MQClientException` and returns an 
empty map whenever the response code is `TOPIC_NOT_EXIST`. In RocketMQ 4.9.4, 
which this connector pins, `ResponseCode.TOPIC_NOT_EXIST` is 17, and 17 is also 
the code carried by `No topic route info in name server for the topic: 
<topic>`, which the name server raises for a topic that exists but whose route 
is momentarily unavailable. The enumerator then reads that empty map as "this 
group has committed nothing" and starts from the first offset.
   
   While working on an unrelated e2e problem I measured one of these route gaps 
directly in CI: the name server returned no route for an existing, actively 
used topic 10 consecutive times at about 30 second intervals, a continuous 
window of 4 minutes 24 seconds. With `partition.discovery.interval.millis = 
"1000"` as in the config above, a job re-enters that lookup often enough that 
it only has to be unlucky once.
   
   I want to be honest about the limits of this: I have not captured a trace of 
a production job rewinding through this specific path, so it is established 
from the code and the response code rather than from a reproduction, and your 
report is consistent with it but does not prove it. If you still see 
consumption restarting from 0 on a plain start after #10778, the useful 
evidence would be whether the broker still shows committed offsets for 
`track_report_group` at that moment, and whether any `CODE: 17` route warnings 
appear in the job log around startup. That would separate this from the two 
causes already ruled out here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to