[jira] [Commented] (STORM-166) Highly available Nimbus

Robert Joseph Evans (JIRA) Thu, 04 Sep 2014 09:33:36 -0700

    [ 
https://issues.apache.org/jira/browse/STORM-166?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14121529#comment-14121529
 ]


Robert Joseph Evans commented on STORM-166:
-------------------------------------------

for 2) Option 1 is definitely the simplest option to start out with.  We do 
this already for our manual HA.  The state is written to a filer and if there 
is a Hardware failure we bring up a new Nimbus and have it take the same IP 
address as the old one.  It is not perfect but it does work.

Additionally we have been working on adding in a distributed cache like feature 
to storm STORM-411.  The idea would be to have the backend implement an API 
that looks like a blob store.  So if bittorrent is used it would be responsible 
to getting the minimum replication before a write is fully committed.  But for 
other things like HDFS, Swift, NFS, etc. the persistence would be built in.

> Highly available Nimbus
> -----------------------
>
>                 Key: STORM-166
>                 URL: https://issues.apache.org/jira/browse/STORM-166
>             Project: Apache Storm (Incubating)
>          Issue Type: New Feature
>            Reporter: James Xu
>            Assignee: Parth Brahmbhatt
>            Priority: Minor
>
> https://github.com/nathanmarz/storm/issues/360
> The goal of this feature is to be able to run multiple Nimbus servers so that 
> if one goes down another one will transparently take over. Here's what needs 
> to happen to implement this:
> 1. Everything currently stored on local disk on Nimbus needs to be stored in 
> a distributed and reliable fashion. A DFS is perfect for this. However, as we 
> do not want to make a DFS a mandatory requirement to run Storm, the storage 
> of these artifacts should be pluggable (default to local filesystem, but the 
> interface should support DFS). You would only be able to run multiple NImbus 
> if you use the right storage, and the storage interface chosen should have a 
> flag indicating whether it's suitable for HA mode or not. If you choose local 
> storage and try to run multiple Nimbus, one of the Nimbus's should fail to 
> launch.
> 2. Nimbus's should register themselves in Zookeeper. They should use a leader 
> election protocol to decide which one is currently responsible for launching 
> and monitoring topologies.
> 3. StormSubmitter should find the Nimbus to connect to via Zookeeper. In case 
> the leader changes during submission, it should use a retry protocol to try 
> reconnecting to the new leader and attempting submission again.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

[jira] [Commented] (STORM-166) Highly available Nimbus

Reply via email to