This is an automated email from the ASF dual-hosted git repository.

rzo1 pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/storm.git


The following commit(s) were added to refs/heads/master by this push:
     new 29f9ffb38 Clarify worker/supervisor heartbeat docs and deprecate 
unused nimbus.supervisor.timeout.secs (#8984)
29f9ffb38 is described below

commit 29f9ffb38c7cea0734a50920597ed95a490fdf9c
Author: reiabreu <[email protected]>
AuthorDate: Thu Aug 20 11:06:20 2026 +0100

    Clarify worker/supervisor heartbeat docs and deprecate unused 
nimbus.supervisor.timeout.secs (#8984)
    
    * Clarify worker/supervisor heartbeat docs and deprecate unused 
nimbus.supervisor.timeout.secs
    
    Documentation still described the pre-2.0 model in which workers/tasks
    heartbeat directly into ZooKeeper. Since 2.0 (STORM-2693), worker liveness
    heartbeats are written to local disk and relayed to Nimbus over Thrift, held
    in an in-memory heartbeat cache; supervisor liveness is an ephemeral 
ZooKeeper
    node detected via session expiry.
    
    Docs:
    - Daemon-Fault-Tolerance.md: describe the actual worker heartbeat relay path
      and the ephemeral-znode supervisor liveness mechanism.
    - Lifecycle-of-a-topology.md: note (inline in the 0.7.1 walkthrough) that 
the
      ZK-directory heartbeat model was replaced in 2.0.
    - Cluster-State-Serialization.md: clarify worker heartbeats are not 
persisted
      in ZooKeeper by default (only under Pacemaker/legacy configuration).
    
    Config:
    - Because supervisor crash detection relies on ephemeral znodes, no 
Nimbus-side
      supervisor timeout check exists, so nimbus.supervisor.timeout.secs is 
never
      read. Mark the DaemonConfig constant @Deprecated, remove the misleading
      defaults.yaml entry, and drop two inert references in 
NimbusClojurePortTest.
    
    Co-Authored-By: Claude Opus 4.8 <[email protected]>
    
    * Use forRemoval/since on deprecated NIMBUS_SUPERVISOR_TIMEOUT_SECS
    
    Match Storm's existing deprecation convention (e.g. Config.java) by marking
    the unused constant @Deprecated(forRemoval = true, since = "3.0.1") and 
stating
    in the javadoc that it is scheduled for removal. 3.0.1 is the current
    development version (root pom is 3.0.1-SNAPSHOT).
    
    Co-Authored-By: Claude Opus 4.8 <[email protected]>
    
    * Target 3.1.0 milestone for NIMBUS_SUPERVISOR_TIMEOUT_SECS removal
    
    Set since = "3.1.0" (the milestone this deprecation targets) on the
    @Deprecated(forRemoval = true) annotation.
    
    Co-Authored-By: Claude Opus 4.8 <[email protected]>
    
    * Keep deprecated nimbus.supervisor.timeout.secs default with a comment
    
    An earlier commit removed the defaults.yaml entry. Since the deprecation
    targets removal in a future release (not now), keep the shipped default in
    place during the deprecation window and add a comment explaining it is 
unused
    (supervisor liveness is tracked via ephemeral ZooKeeper nodes). The entry
    itself is unchanged from master; only the explanatory comment is added.
    
    Co-Authored-By: Claude Opus 4.8 <[email protected]>
    
    ---------
    
    Co-authored-by: Claude Opus 4.8 <[email protected]>
---
 conf/defaults.yaml                                               | 1 +
 docs/Cluster-State-Serialization.md                              | 9 +++++++++
 docs/Daemon-Fault-Tolerance.md                                   | 4 +++-
 docs/Lifecycle-of-a-topology.md                                  | 1 +
 storm-server/src/main/java/org/apache/storm/DaemonConfig.java    | 7 +++++++
 .../org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java    | 2 --
 6 files changed, 21 insertions(+), 3 deletions(-)

diff --git a/conf/defaults.yaml b/conf/defaults.yaml
index 9682cf8bc..a159f6bc1 100644
--- a/conf/defaults.yaml
+++ b/conf/defaults.yaml
@@ -81,6 +81,7 @@ nimbus.thrift.tls.client.auth.required: true
 topology.worker.nimbus.thrift.client.use.tls: false
 nimbus.childopts: "-Xmx1024m"
 nimbus.task.timeout.secs: 30
+# Deprecated since 3.1.0 and unused; supervisor liveness is tracked via 
ephemeral ZooKeeper nodes, so Nimbus does not time supervisors out. Scheduled 
for removal in a future release.
 nimbus.supervisor.timeout.secs: 60
 nimbus.monitor.freq.secs: 10
 nimbus.cleanup.inbox.freq.secs: 600
diff --git a/docs/Cluster-State-Serialization.md 
b/docs/Cluster-State-Serialization.md
index b59a0b2f4..8cc404ed0 100644
--- a/docs/Cluster-State-Serialization.md
+++ b/docs/Cluster-State-Serialization.md
@@ -9,6 +9,15 @@ ZooKeeper (and other configured state stores) such as topology 
assignments, Nimb
 summaries, `StormBase` records, log configs, credentials, worker heartbeats,
 profile requests, errors, etc.
 
+> **Note on worker heartbeats.** Since 2.0 
([STORM-2693](https://issues.apache.org/jira/browse/STORM-2693)),
+> worker liveness heartbeats are, by default, *not* persisted in ZooKeeper: 
workers
+> write them to local disk, supervisors relay them to Nimbus over Thrift, and 
Nimbus
+> keeps them in an in-memory heartbeat cache. Worker heartbeats are only 
written to a
+> state store (the `WORKERBEATS_SUBTREE` path) when a heartbeat store such as 
Pacemaker
+> is configured. The serialization described below still applies to those 
stored
+> heartbeats, and to supervisor liveness (`SupervisorInfo`), which is always 
kept as an
+> ephemeral ZooKeeper node.
+
 It is distinct from
 [tuple serialization](Serialization.html), which covers payloads exchanged
 between spouts and bolts at runtime via Kryo.
diff --git a/docs/Daemon-Fault-Tolerance.md b/docs/Daemon-Fault-Tolerance.md
index 8dce601a8..b419e16cc 100644
--- a/docs/Daemon-Fault-Tolerance.md
+++ b/docs/Daemon-Fault-Tolerance.md
@@ -7,7 +7,7 @@ Storm has several different daemon processes.  Nimbus that 
schedules workers, su
 
 ## What happens when a worker dies?
 
-When a worker dies, the supervisor will restart it. If it continuously fails 
on startup and is unable to heartbeat to Nimbus, Nimbus will reschedule the 
worker.
+When a worker dies, the supervisor will restart it. Worker liveness reaches 
Nimbus indirectly: each worker writes heartbeats to local disk, and its 
supervisor relays them to Nimbus over Thrift (this replaced the pre-2.0 model 
in which workers heartbeat directly into ZooKeeper; see 
[STORM-2693](https://issues.apache.org/jira/browse/STORM-2693)). If a worker 
stops heartbeating for longer than `nimbus.task.timeout.secs`, Nimbus 
reschedules it. A freshly launched worker is given a longer gra [...]
 
 ## What happens when a node dies?
 
@@ -19,6 +19,8 @@ The Nimbus and Supervisor daemons are designed to be 
fail-fast (process self-des
 
 Most notably, no worker processes are affected by the death of Nimbus or the 
Supervisors. This is in contrast to Hadoop, where if the JobTracker dies, all 
the running jobs are lost. 
 
+Supervisor liveness is tracked differently from worker liveness. Each 
supervisor registers itself as an ephemeral ZooKeeper node (its 
`SupervisorInfo`, which also carries scheduling metadata such as ports and 
resources). When a supervisor dies, its ZooKeeper session expires and the 
ephemeral node disappears, so Nimbus detects the loss directly from ZooKeeper 
rather than by timing out heartbeats. (This is why there is no active 
Nimbus-side supervisor heartbeat-timeout setting.)
+
 ## Is Nimbus a single point of failure?
 
 If you lose the Nimbus node, the workers will still continue to function. 
Additionally, supervisors will continue to restart workers if they die. 
However, without Nimbus, workers won't be reassigned to other machines when 
necessary (like if you lose a worker machine). 
diff --git a/docs/Lifecycle-of-a-topology.md b/docs/Lifecycle-of-a-topology.md
index fe785f1e4..83e8991c0 100644
--- a/docs/Lifecycle-of-a-topology.md
+++ b/docs/Lifecycle-of-a-topology.md
@@ -34,6 +34,7 @@ First a couple of important notes about topologies:
     - Jars and configs are kept on local filesystem because they're too big 
for Zookeeper. The jar and configs are copied into the path {nimbus local 
dir}/stormdist/{topology id}
     - `setup-storm-static` writes task -> component mapping into ZK
     - `setup-heartbeats` creates a ZK "directory" in which tasks can heartbeat
+        - (**Since 2.0, STORM-2693**: workers no longer heartbeat directly 
into ZooKeeper. A worker now writes liveness heartbeats to local disk, and its 
supervisor relays them to Nimbus over Thrift. See [Daemon Fault 
Tolerance](Daemon-Fault-Tolerance.html) for the current mechanism.)
 - Nimbus calls `mk-assignment` to assign tasks to machines 
[code](https://github.com/apache/storm/blob/0.7.1/src/clj/org/apache/storm/daemon/nimbus.clj#L458)
     - Assignment record definition is here: 
[code](https://github.com/apache/storm/blob/0.7.1/src/clj/org/apache/storm/daemon/common.clj#L25)
     - Assignment contains:
diff --git a/storm-server/src/main/java/org/apache/storm/DaemonConfig.java 
b/storm-server/src/main/java/org/apache/storm/DaemonConfig.java
index 6dcc2eb37..71d3c406b 100644
--- a/storm-server/src/main/java/org/apache/storm/DaemonConfig.java
+++ b/storm-server/src/main/java/org/apache/storm/DaemonConfig.java
@@ -277,7 +277,14 @@ public class DaemonConfig implements Validated {
 
     /**
      * How long before a supervisor can go without heartbeating before nimbus 
considers it dead and stops assigning new work to it.
+     *
+     * @deprecated Unused. Supervisor liveness is tracked via an ephemeral 
ZooKeeper node (see
+     *     {@code StormClusterState#supervisorHeartbeat}); when a supervisor 
dies its ZooKeeper session
+     *     expires and the node disappears, so Nimbus detects the loss 
directly rather than by timing out
+     *     heartbeats. No code reads this value. It is scheduled for removal; 
retained for now only for
+     *     backward compatibility.
      */
+    @Deprecated(forRemoval = true, since = "3.1.0")
     @IsInteger
     @IsPositiveNumber
     public static final String NIMBUS_SUPERVISOR_TIMEOUT_SECS = 
"nimbus.supervisor.timeout.secs";
diff --git 
a/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java
 
b/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java
index 3acff939a..526cb0924 100644
--- 
a/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java
+++ 
b/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java
@@ -1717,7 +1717,6 @@ public class NimbusClojurePortTest {
                     DaemonConfig.NIMBUS_TASK_LAUNCH_SECS, 60,
                     DaemonConfig.NIMBUS_TASK_TIMEOUT_SECS, 20,
                     DaemonConfig.NIMBUS_MONITOR_FREQ_SECS, 10,
-                    DaemonConfig.NIMBUS_SUPERVISOR_TIMEOUT_SECS, 100,
                     Config.TOPOLOGY_ACKER_EXECUTORS, 0,
                     Config.TOPOLOGY_EVENTLOGGER_EXECUTORS, 0))
                 .build()) {
@@ -1821,7 +1820,6 @@ public class NimbusClojurePortTest {
                     DaemonConfig.NIMBUS_TASK_LAUNCH_SECS, 60,
                     DaemonConfig.NIMBUS_TASK_TIMEOUT_SECS, 20,
                     DaemonConfig.NIMBUS_MONITOR_FREQ_SECS, 10,
-                    DaemonConfig.NIMBUS_SUPERVISOR_TIMEOUT_SECS, 100,
                     Config.TOPOLOGY_ACKER_EXECUTORS, 0,
                     Config.TOPOLOGY_EVENTLOGGER_EXECUTORS, 0))
                 .build()) {

Reply via email to