jkrauss82 opened a new pull request, #9153:
URL: https://github.com/apache/storm/pull/9153

   Root cause: A transient ZooKeeper connection loss during the supervisor 
heartbeat cycle caused the process to die in 629ms via 
DefaultUncaughtExceptionHandler. The Curator RetryLoop blindly slept between 
retries, racing against the ZK client's SendThread reconnection instead of 
yielding to it.
   
   Fix (two layers):
   
   1. ConnectionAwareRetryPolicy (new): A RetryPolicy wrapper that tracks the 
ZK ConnectionState via a ConnectionStateListener. When the connection is 
SUSPENDED or LOST, it calls 
CuratorFramework.blockUntilConnected(sessionTimeout) to yield to the 
SendThread's failover, instead of blind sleep+retry. On reconnection, the 
operation retries immediately on the new connection. Uses 
storm.zookeeper.session.timeout as the wait bound (no magic numbers).
   
   2. SupervisorHeartbeat safety net: Wrap the heartbeat body in try/catch so 
that a total ZK outage (beyond session timeout) skips the cycle instead of 
killing the process. A missed heartbeat is harmless: 
nimbus.supervisor.timeout.secs (30s) allows 6 missed beats.
   
   Wired up in CuratorUtils.newCurator() via AtomicReference late binding (the 
CuratorFramework doesn't exist until builder.build() returns).
   
   ## What is the purpose of the change:
   
   This fix prevents supervisor processes from dying due to transient ZooKeeper 
connection losses. During the incident, a single ZK ensemble member closing its 
socket caused the supervisor to exit in 629ms via 
`DefaultUncaughtExceptionHandler`, killing all topology workers on that node. 
The fix absorbs these transient blips by making the Curator retry policy aware 
of the ZK connection state — when `SUSPENDED` or `LOST`, it yields to the ZK 
client's SendThread reconnection via `blockUntilConnected()` instead of blind 
sleep+retry. This ensures supervisors survive network/ZK blips that are already 
being handled by the ZK client's failover mechanism.
   
   ## How was the change tested:
   
   The fix was deployed to a production cluster running 3 supervisors with 220+ 
active workers. During a subsequent network outage that caused multiple 
transient ZK connection losses, the supervisors survived without dying 
(previously they would have exited via `DefaultUncaughtExceptionHandler`). Log 
verification confirmed the `ConnectionAwareRetryPolicy` correctly detected 
`SUSPENDED` states, waited for reconnection via `blockUntilConnected()`, and 
retried operations on the new connection. No supervisors died, no topology 
workers were killed due to supervisor crashes.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to