[
https://issues.apache.org/jira/browse/HBASE-8519?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13656471#comment-13656471
]
Jerry He commented on HBASE-8519:
---------------------------------
We probably don't need the second statement.
Checking the current cluster status is just to be safe because getting the
notification is a past event. Thoughts?
What do you think we should use for the name?
I can put comments right before the statement as this:
// We want to make sure before we stop the backup master the cluster is meant
to go down.
// We'll only stop backup master if we have received cluster notification and
the current cluster status is "not up".
We are indeed more restrictive on stopping backup master? Any thoughts if we
are too restrictive?
> Backup master will never come up if primary master dies during initialization
> -----------------------------------------------------------------------------
>
> Key: HBASE-8519
> URL: https://issues.apache.org/jira/browse/HBASE-8519
> Project: HBase
> Issue Type: Bug
> Components: master
> Affects Versions: 0.94.7, 0.95.0
> Reporter: Jerry He
> Assignee: Jerry He
> Priority: Minor
> Fix For: 0.98.0
>
> Attachments: HBASE-8519-trunk.patch
>
>
> The problem happens if primary master dies after becoming master but before
> it completes initialization and calls clusterStatusTracker.setClusterUp(),
> The backup master will try to become the master, but will shutdown itself
> promptly because it sees 'the cluster is not up'.
> This is the backup master log:
> 2013-05-09 15:08:05,568 INFO
> org.apache.hadoop.hbase.master.metrics.MasterMetrics: Initialized
> 2013-05-09 15:08:05,573 DEBUG org.apache.hadoop.hbase.master.HMaster: HMaster
> started in backup mode. Stalling until master znode is written.
> 2013-05-09 15:08:05,589 INFO
> org.apache.hadoop.hbase.zookeeper.RecoverableZooKeeper: Node /hbase/master
> already exists and this is not a retry
> 2013-05-09 15:08:05,590 INFO
> org.apache.hadoop.hbase.master.ActiveMasterManager: Adding ZNode for
> /hbase/backup-masters/xxx.com,60000,1368137285373 in backup master directory
> 2013-05-09 15:08:05,595 INFO
> org.apache.hadoop.hbase.master.ActiveMasterManager: Another master is the
> active master, xxx.com,60000,1368137283107; waiting to become the next active
> master
> 2013-05-09 15:09:45,006 DEBUG
> org.apache.hadoop.hbase.master.ActiveMasterManager: No master available.
> Notifying waiting threads
> 2013-05-09 15:09:45,006 INFO org.apache.hadoop.hbase.master.HMaster: Cluster
> went down before this master became active
> 2013-05-09 15:09:45,006 DEBUG org.apache.hadoop.hbase.master.HMaster:
> Stopping service threads
> 2013-05-09 15:09:45,006 INFO org.apache.hadoop.ipc.HBaseServer: Stopping
> server on 60000
>
> In ActiveMasterManager::blockUntilBecomingActiveMaster()
> {code}
> ..
> if (!clusterStatusTracker.isClusterUp()) {
> this.master.stop(
> "Cluster went down before this master became active");
> }
> ..
> {code}
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira