Thanks!
Several more questions:
1. My understanding is, when the crmd gets into this mess, it will be gone
and the master heartheart control process will bring up a new instance of
crmd and it will continue to work. There will be no affect to the cluster.
Is it correct?
2. I tried to create the split-brain on a 2-node cluster and then
recovered from it. I did see the crmd killed and started again. I did see
the core file generated. But this time, I only saw the back trace of:
Core was generated by `/usr/lib64/heartbeat/crmd'.
Program terminated with signal 6, Aborted.
#0 0x00000039d642e21d in ?? ()
(gdb) bt
#0 0x00000039d642e21d in ?? ()
#1 0x00000039d642fa1e in ?? ()
#2 0x0000000000000020 in ?? ()
#3 0x0000000000000000 in ?? ()
Therefore, I am not sure if it is the same error I got like previous.
Anything I can do to get more info from the core file? Also, in the cluster
we got the first core file, it is not likely to have the split-brain
problem. Any other possible scenario could cause this problem?
3. This time, I do have the ha-debug, it looks like:
heartbeat[3167]: 2008/02/08_10:54:14 ERROR: glib: Unable to send [-1] ucast
packet: No such device
heartbeat[3167]: 2008/02/08_10:54:14 ERROR: write failure on ucast eth1.: No
such device
... ... ... ...
heartbeat[3167]: 2008/02/08_10:54:42 ERROR: write failure on ucast eth1.: No
such device
heartbeat[3167]: 2008/02/08_10:54:43 ERROR: glib: Unable to send [-1] ucast
packet: No such device
heartbeat[3167]: 2008/02/08_10:54:43 ERROR: write failure on ucast eth1.: No
such device
heartbeat[3163]: 2008/02/08_10:54:44 CRIT: Cluster node
ha1.roundbox.comreturning after partition.
heartbeat[3163]: 2008/02/08_10:54:44 info: For information on cluster
partitions, See URL: http://linux-ha.org/SplitBrain
heartbeat[3163]: 2008/02/08_10:54:44 WARN: Deadtime value may be too small.
heartbeat[3163]: 2008/02/08_10:54:44 info: See FAQ for information on tuning
deadtime.
heartbeat[3163]: 2008/02/08_10:54:44 info: URL:
http://linux-ha.org/FAQ#heavy_load
heartbeat[3163]: 2008/02/08_10:54:44 info: Link ha1.roundbox.com:eth1 up.
heartbeat[3163]: 2008/02/08_10:54:44 WARN: Late heartbeat: Node
ha1.roundbox.com: interval 31060 ms
heartbeat[3163]: 2008/02/08_10:54:44 info: Status update for node
ha1.roundbox.com: status active
heartbeat[3163]: 2008/02/08_10:54:45 WARN: Exiting /usr/lib64/heartbeat/crmd
process 3317 killed by signal 6 [SIGABRT - Abort].
heartbeat[3163]: 2008/02/08_10:54:45 ERROR: Exiting
/usr/lib64/heartbeat/crmd process 3317 dumped core
heartbeat[3163]: 2008/02/08_10:54:45 ERROR: Respawning client
"/usr/lib64/heartbeat/crmd":
heartbeat[3163]: 2008/02/08_10:54:45 info: Starting child client
"/usr/lib64/heartbeat/crmd" (502,503)
heartbeat[3442]: 2008/02/08_10:54:45 info: Starting
"/usr/lib64/heartbeat/crmd" as uid 502 gid 503 (pid 3442)
heartbeat[3163]: 2008/02/08_10:54:48 WARN: 1 lost packet(s) for [
ha1.roundbox.com] [1366:1368]
heartbeat[3163]: 2008/02/08_10:54:48 info: No pkts missing from
ha1.roundbox.com!
4. For heartbeat 2.1.3, there are 2.1.3-2 and 2.1.3-3 for Centos4. Any
big difference between them? Which one is recommended?
Thanks,
Tao
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems