jkrauss82 opened a new pull request, #9153: URL: https://github.com/apache/storm/pull/9153
Root cause: A transient ZooKeeper connection loss during the supervisor heartbeat cycle caused the process to die in 629ms via DefaultUncaughtExceptionHandler. The Curator RetryLoop blindly slept between retries, racing against the ZK client's SendThread reconnection instead of yielding to it. Fix (two layers): 1. ConnectionAwareRetryPolicy (new): A RetryPolicy wrapper that tracks the ZK ConnectionState via a ConnectionStateListener. When the connection is SUSPENDED or LOST, it calls CuratorFramework.blockUntilConnected(sessionTimeout) to yield to the SendThread's failover, instead of blind sleep+retry. On reconnection, the operation retries immediately on the new connection. Uses storm.zookeeper.session.timeout as the wait bound (no magic numbers). 2. SupervisorHeartbeat safety net: Wrap the heartbeat body in try/catch so that a total ZK outage (beyond session timeout) skips the cycle instead of killing the process. A missed heartbeat is harmless: nimbus.supervisor.timeout.secs (30s) allows 6 missed beats. Wired up in CuratorUtils.newCurator() via AtomicReference late binding (the CuratorFramework doesn't exist until builder.build() returns). ## What is the purpose of the change: This fix prevents supervisor processes from dying due to transient ZooKeeper connection losses. During the incident, a single ZK ensemble member closing its socket caused the supervisor to exit in 629ms via `DefaultUncaughtExceptionHandler`, killing all topology workers on that node. The fix absorbs these transient blips by making the Curator retry policy aware of the ZK connection state — when `SUSPENDED` or `LOST`, it yields to the ZK client's SendThread reconnection via `blockUntilConnected()` instead of blind sleep+retry. This ensures supervisors survive network/ZK blips that are already being handled by the ZK client's failover mechanism. ## How was the change tested: The fix was deployed to a production cluster running 3 supervisors with 220+ active workers. During a subsequent network outage that caused multiple transient ZK connection losses, the supervisors survived without dying (previously they would have exited via `DefaultUncaughtExceptionHandler`). Log verification confirmed the `ConnectionAwareRetryPolicy` correctly detected `SUSPENDED` states, waited for reconnection via `blockUntilConnected()`, and retried operations on the new connection. No supervisors died, no topology workers were killed due to supervisor crashes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
