Created a PR enhancing the documentation and using @Gianluca Graziadei suggestion to set it as Deprecated. https://github.com/apache/storm/pull/8984
I can retrofit these documentation changes into the live website documentation. On Wed, 5 Aug 2026 at 21:18, Gianluca Graziadei <[email protected]> wrote: > > Hi, > I agree with Richard for option (1). > > Karthick feel free to open a PR. > In case you want proceed, instead of removing the line what about annotating > it with the following pattern? > > @Deprecated(since = "3.0.1", forRemoval = true) > > Cheers, > > Gianluca > > > On Wed, 5 Aug 2026, 19:39 Richard Zowalla, <[email protected]> wrote: >> >> I would prefer (1) right now ;-) >> >> > Am 05.08.2026 um 00:18 schrieb Rui Abreu <[email protected]>: >> > >> > Hi! Thanks for spotting it. It seems that config was never used from >> > the beginning. >> > Storm 3.0.0 dropped Clojure in favour of Java, but the port kept the >> > logic as much as possible. >> > There used to be a comment in a Clojure source fle: >> > >> > - ;; TODO: this is broken. need to maintain a map since last time >> > - ;; supervisor hearbeats like is done for tasks >> > - ;; maybe it's ok to trust ephemeral nodes here? >> > - ;;[[id info]] >> > - ;; (when (< (time-delta (:time-secs info)) >> > - ;; (conf NIMBUS-SUPERVISOR-TIMEOUT-SECS)) >> > - ;; [[id info]] >> > - ;; ) >> > >> > https://github.com/apache/storm/commit/40e092dda0e7affe89c00552bc138a251f095dc0 >> > >> > The actual mechanism still relies on the ephemeral nodes (and always >> > had, apparently) >> > >> > - ZooKeeper session timeout (storm.zookeeper.session.timeout, default >> > 20000 ms). This is negotiated with ZK and drives when ZK deletes the >> > ephemeral node after the supervisor stops responding. >> > - Nimbus scheduler tick (nimbus.monitor.freq.secs) — how often Nimbus >> > re-reads ZK and reassigns. >> > >> > supervisor.heartbeat.frequency.secs is misleading because even if it >> > would not fire, Zookeeper client uses pings to keep the session alive, >> > according to what I can gather (this property keeps the Supervisor >> > data read by Nimbus fresh) >> > >> > @Richard Zowalla @Gianluca Graziadei @Julien Nioche would like to hear >> > on thoughts on this. >> > >> > 1- We remove the dead configs/code/documentation and just properly >> > document the actual mechanism >> > 2- We implement a different Supervisor liveness mechanism based on >> > that property that has been never used >> > >> > Either way, this is a bug. @Karthick do you want to open an issue for this? >> > >> > Thank you >> > >> > On Tue, 4 Aug 2026 at 05:29, Karthick <[email protected]> wrote: >> >> >> >> Hi, >> >> Im checking on the heartbeat flow, The below configuration is not in use, >> >> as per comment it seems. Am I missing anything? Please guide me. >> >> >> >> /** >> >> * How long before a supervisor can go without heartbeating before nimbus >> >> considers it dead and stops assigning new work to it. >> >> */ >> >> @isInteger >> >> @isPositiveNumber >> >> public static final String NIMBUS_SUPERVISOR_TIMEOUT_SECS = >> >> "nimbus.supervisor.timeout.secs"; >> >> >> >> >> >> On Sat, Jul 25, 2026 at 1:45 AM Gianluca Graziadei >> >> <[email protected]> wrote: >> >>> >> >>> Hi, >> >>> >> >>> I dug through the JIRA history and your reconciliation holds, though >> >>> with two key refinements: the bottleneck in STORM-2693 was actually >> >>> per-round read-and-recompute overhead on Nimbus rather than Zookeeper >> >>> write pressure, which is why the fix prioritized caching and supervisor >> >>> reporting over a faster store. SupervisorInfo stayed in ZK not for >> >>> liveness monitoring, but because it is shared cluster-state metadata >> >>> that any elected Nimbus leader needs to access for scheduling (content >> >>> of the serialized >> >>> https://github.com/apache/storm/blob/master/storm-client/src/jvm/org/apache/storm/generated/SupervisorInfo.java) >> >>> >> >>> Best, >> >>> >> >>> Gianluca >> >>> >> >>> >> >>> Il giorno ven 24 lug 2026 alle ore 12:25 Karthick >> >>> <[email protected]> ha scritto: >> >>>> >> >>>> Thanks, that's a clear and helpful breakdown. The push-for-liveness / >> >>>> pull-for-stats split makes sense: keeping a continuous in-memory pulse >> >>>> per worker in Nimbus (HeartbeatCache) while serving heavy stats on >> >>>> demand from ZK keeps Nimbus lightweight, and the Supervisor Thrift >> >>>> relay keeps those frequent liveness writes off ZK. >> >>>> >> >>>> One thing I left out of my original question: the supervisor liveness >> >>>> heartbeat — which is also liveness, but it still goes to ZooKeeper >> >>>> (ephemeral znode at /supervisors/<id>, via SupervisorHeartbeat), >> >>>> consumed by Nimbus under nimbus.supervisor.timeout.secs. So "liveness" >> >>>> isn't uniformly on the Thrift path. >> >>>> >> >>>> My working reconciliation is that the deciding factor is volume and >> >>>> semantics, not liveness-vs-stats: >> >>>> >> >>>> Worker/executor liveness is high-volume (potentially thousands of >> >>>> heartbeats per cluster, every ~1s) → moved off ZK onto the Supervisor >> >>>> Thrift relay to avoid write pressure. >> >>>> Supervisor liveness is low-volume (one per node, every ~5s) → cheap on >> >>>> ZK, and the ephemeral znode gives automatic crash detection when the >> >>>> session dies, which the Thrift path wouldn't provide for free. >> >>>> >> >>>> Does that match the design intent — i.e. supervisor heartbeats stayed >> >>>> on ZK deliberately because per-node volume is low and ephemeral-node >> >>>> semantics are valuable, whereas per-worker heartbeats were the actual >> >>>> ZK scaling problem? >> >>>> >> >>>> Also noted on 3.0.0's ZK read/serialization improvements — thanks for >> >>>> the pointer, will look into it. >> >>>> >> >>>> >> >>>> >> >>>> On Thu, Jul 23, 2026 at 4:53 PM Karthick <[email protected]> >> >>>> wrote: >> >>>>> >> >>>>> Hi all, >> >>>>> >> >>>>> I'm studying the Storm 2.0 heartbeat/liveness paths and want to >> >>>>> confirm my understanding of a design decision. >> >>>>> >> >>>>> As I read the 2.0 code, there are two distinct worker-originated >> >>>>> heartbeats: >> >>>>> >> >>>>> Liveness — the worker writes an LSWorkerHeartbeat to local disk >> >>>>> (Worker.doHeartBeat); the supervisor reads those files and relays a >> >>>>> batch to the leader Nimbus over Thrift (ReportWorkerHeartbeats → >> >>>>> Nimbus.sendSupervisorWorkerHeartbeats → HeartbeatCache), governed by >> >>>>> nimbus.task.timeout.secs. >> >>>>> Stats — Worker.doExecutorHeartbeats writes a heartbeat object >> >>>>> (time-secs + uptime + executor stats) to ZooKeeper, which the >> >>>>> UI/metrics consume (and which HeartbeatCache.updateFromZkHeartbeat can >> >>>>> still use for liveness on the ZK strategy). >> >>>>> >> >>>>> My understanding is that liveness was moved off ZooKeeper (the 1.x >> >>>>> model, and later Pacemaker) because high-volume per-worker heartbeat >> >>>>> writes made ZK a scaling bottleneck, and since heartbeats are >> >>>>> ephemeral they don't need ZK's persistence/consistency — so the >> >>>>> supervisor-relay-over-Thrift model removes those writes from ZK >> >>>>> entirely. >> >>>>> >> >>>>> A few questions: >> >>>>> >> >>>>> Is that the correct/primary motivation for the Thrift supervisor-relay >> >>>>> path, or were there other drivers (connection count, watch load, >> >>>>> Nimbus HA, recovery on leader change)? >> >>>>> Why do executor stats still go through ZooKeeper rather than riding >> >>>>> the same Thrift path — is it purely that stats are lower-frequency and >> >>>>> UI-oriented, or is there a stronger reason? >> >>>>> Is there a JIRA / design doc that captures this transition (beyond >> >>>>> docs/Pacemaker.md) that I could read? >> >>>>> >> >>>>> Thanks for any pointers — trying to make sure I document this >> >>>>> accurately. >>
