m1a2st commented on code in PR #23348:
URL: https://github.com/apache/kafka/pull/23348#discussion_r3974438349


##########
clients/src/main/java/org/apache/kafka/clients/consumer/internals/AbstractHeartbeatRequestManager.java:
##########
@@ -259,13 +259,16 @@ public long maximumTimeToWait(long currentTimeMs) {
         if (pollTimer.isExpired()) {
             return 0L;
         }
-        // KAFKA-20253: mirror the guard in poll(). A heartbeat is only sent 
when the coordinator is known
-        // and the member is in a state that heartbeats. When the coordinator 
is unavailable (e.g. after a
-        // re-authentication failure) or the member should skip heartbeats 
(FATAL/FENCED/STALE/UNSUBSCRIBED),
-        // poll() returns EMPTY, so falling through to the timer-based 
branches below would return 0 (the
-        // heartbeat timer is left permanently expired) and busy-spin both the 
application and network threads.
+        // Mirror the guard in poll(). A heartbeat is only sent when the 
coordinator is known and the
+        // member is in a state that heartbeats. When the coordinator is 
unavailable (e.g. after a
+        // re-authentication failure, or while bootstrap DNS resolution is 
still in progress) or the
+        // member should skip heartbeats (FATAL/FENCED/STALE/UNSUBSCRIBED), 
poll() returns EMPTY, so
+        // falling through to the timer-based branches below would return 0 
(the heartbeat timer is left
+        // permanently expired) and busy-spin both the application and network 
threads. Wait a retry
+        // backoff rather than the heartbeat interval, because the interval is 
zero until the first
+        // heartbeat response is received, which would also busy-spin.
         if (coordinatorRequestManager.coordinator().isEmpty() || 
membershipManager().shouldSkipHeartbeat()) {
-            return heartbeatRequestState.heartbeatIntervalMs();
+            return heartbeatRequestState.retryBackoffMs();

Review Comment:
   Sorry for the late reply, I traced through this last night.
   I don't have a better improvement in mind right now. For now, I'll open a 
Jira ticket to track this issue so we can discuss and investigate it separately.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to