On 4/12/07, 池田淳子 <[EMAIL PROTECTED]> wrote:
Hi all,
I'm newbie, and trying to understand how or when "Split-Brain" happens.
when 1 or more nodes can't communicate with each other
there is no connection between this and an operation's "start-delay"
what you *might* be seeing is an old bug that was triggered when
start-delay > timeout
As trial, I run "Dummy" resource for now.
Heartbeat version is 2.0.8,
cib.xml, ha.cf and ha-log are attached.
See below cases, please give some advice.
*** case 01 ***
My cib.xml was created using hb_gui, so start-delay was "1m".
I run Heartbeat on 2 nodes at first, and disconnected the interconnect LAN
on DC node to lead "Split-Brain". (# ifdown eth2)
After making "Split-Brain", I would up the LAN with "ifup eth2".
If eth2 is restored after "Action Dummy01_monitor_10000 (x) confirmed" on
the former stand-by node, it seems that everything works well.
But if it's done before the confirmation, there is something wrong.
In wrong case, I couldn't generate "Split-Brain" again.
I found that pengine/tengine run on both nodes, but one node kept trying to
be DC and Dummy resource wouldn't start on that node, additionally,
failcount was incremented on that strange node.
log message is here...
WARN: do_dc_join_finalize: join-3: We are still in a transition. Delaying
until the TE completes.
in this case, I couldn't shutdown Heartbeat process without KILL command...
*** case 02 ***
I changed start-delay from 1m to 0s.
The confirmation process for monitor would work immediately, so though I put
ifdown/ifup in a row, it didn't matter.
Q1; My guess, if the interconnect LAN is down/up in a raw and some
operations aren't confirmed, one node would consider this situation as
ERROR, so update its failcaount. This case might appear when "start-delay"
is long (ex, "1m"), and cause some strange "Split-Brain".
Is this relationship between "start-delay" and "Split-Brain" correct?
*** case 03 ***
In case 02, the interconnect LAN is up after "Split-Brain" is completed, it
means each node is voted as DC.
For third test, I tried to down/up the interconnect LAN before the DC
election didn't finish.
In the result, I met "Split-Brain" again but one node stayed OFFLINE when
the interconnect LAN was up.
Q2; I know, I had better to set up two or more interconnect LANs just in
case, but are there any prefer ways to avoid case 03? ex, tuning some
parameters or something like that.
Best Regards,
Junko Ikeda
NTT DATA INTELLILINK CORPORATION
Open Source Solutions Business Unit
Open Source Business Division
Toyosu Center Building Annex, 3-3-9, Toyosu,
Koto-ku, Tokyo 135-0061, Japan
TEL : +81-3-3534-4811
FAX : +81-3-3534-4814
mailto:[EMAIL PROTECTED]
http://www.intellilink.co.jp/
_______________________________________________________
Linux-HA-Dev: [EMAIL PROTECTED]
http://lists.linux-ha.org/mailman/listinfo/linux-ha-dev
Home Page: http://linux-ha.org/
_______________________________________________________
Linux-HA-Dev: [EMAIL PROTECTED]
http://lists.linux-ha.org/mailman/listinfo/linux-ha-dev
Home Page: http://linux-ha.org/