On Nov 22, 2007, at 7:05 AM, <[EMAIL PROTECTED]> <[EMAIL PROTECTED] > wrote:


Hi Andrew,
Pardon me for my ignorance, but I'm new to stonith. Pls read on...

-----Original Message-----
From: [EMAIL PROTECTED] [mailto:[EMAIL PROTECTED] ] On Behalf Of Andrew Beekhof
Sent: Wednesday, November 21, 2007 8:42 PM
To: General Linux-HA mailing list
Subject: Re: [Linux-HA] Re: Linux-HA Digest, Vol 48, Issue 69


On Nov 21, 2007, at 3:36 PM, <[EMAIL PROTECTED]> <[EMAIL PROTECTED] > wrote:


Can anyone please help me debug problems with my STONITH RA. I've
slightly modified the in-built ssh RA ( changed $ssh_command=rsh and
reboot_command to "/etc/init.d/heartbeat restart".

Thats a really _really_ bad idea.

I'm trying to recover from a split brain scenario in 1:1 cluster.

split brains usually result from comms failure
in such cases anything that uses the network is not going to work
and restarting heartbeat is extremely unlikely to help, rebooting the node is not much better.

the best thing you can hope for is that the fencing operation fails which will prevent the cluster from doing anything that would kill your data.

Hoping that restarting heartbeat on the other node might sync it up with the current state of the cluster and recover from split brain. Rebooting is also an option, but don't want to be so excessive for testing. Would you have any other ideas for my case ?

I ran the stonith test on the RA and it seemed to work fine.
However , everytime I start heartbeat I get these errors and stonith
resource doesn't run. The ha-log is attached..

My questions:
1. why stonith resource doesn't start though it's passed stonith test
2. crm_verify has been run then why the asserts 3. how to know from
error code (rc) what is wrong? I've actually removed all exit 1's from
the reset operation.
4. this is heartbeat log from standby. Why does standby reset itself?
It should reset active.

Presumably because this is how you wrote the RA
My RA seems to work now, but not very consistently . I modified it a bit to return only the node name which it should be used to stonith, standby for active and active for standby node.


My cib.xml file is identical to the last thread with the changes that
frank recommended..

Thanks,
Abhi.

pengine[1808]: 2007/11/21_19:14:50 ERROR: can_run_resources: No node
supplied
pengine[1808]: 2007/11/21_19:14:50 ERROR: can_run_resources: No node
supplied
pengine[1808]: 2007/11/21_19:14:50 WARN: native_color: Resource
Stonith:1 cannot run anywhere

This is probably the issue the PE is complaining about.
I still get this error on standby where the RA failed to start , but not on the active node.


pengine[1808]: 2007/11/21_19:14:50 WARN: stage6: Scheduling Node
standby for STONITH
pengine[1808]: 2007/11/21_19:14:50 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:14:51 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:14:51 WARN: process_pe_message:
Transition 0: WARNINGs found during PE processing. PEngine Input
stored in: /var/lib/heartbeat/pengine/pe-warn-9620.bz2
pengine[1808]: 2007/11/21_19:14:51 info: process_pe_message:
Configuration WARNINGs found during PE processing.  Please run
"crm_verify -L" to identify issues.
crmd[1799]: 2007/11/21_19:14:56 ERROR: get_tag_name: Error parsing
token near usage:/usr/lib/: bad character, not in [a-zA-Z_-]
crmd[1799]: 2007/11/21_19:14:56 ERROR: crm_abort: parse_xml:
Triggered non-fatal assert at xml.c:1365 : len > 0
crmd[1799]: 2007/11/21_19:14:56 ERROR: crm_abort: find_xml_node:
Triggered non-fatal assert at xml.c:77 : root != NULL
crmd[1799]: 2007/11/21_19:14:56 ERROR: cl_get_value: wrong arugment
(__name__)
crmd[1799]: 2007/11/21_19:14:56 WARN: find_xml_node: Could not find
actions in (null).

The metadata for one of your resources is broken. I'm guessing the new one. I'm now returning exit 0 for meta-data for my OCF resources.ocf- tester helped fix this.


pengine[1808]: 2007/11/21_19:14:58 notice: cluster_option: Using
default value '-1' for cluster option 'pe-error-series-max'
pengine[1808]: 2007/11/21_19:14:58 notice: cluster_option: Using
default value '-1' for cluster option 'pe-warn-series-max'
pengine[1808]: 2007/11/21_19:15:00 ERROR: can_run_resources: No node
supplied
pengine[1808]: 2007/11/21_19:15:00 ERROR: can_run_resources: No node
supplied
pengine[1808]: 2007/11/21_19:15:00 WARN: native_color: Resource
Stonith:1 cannot run anywhere
pengine[1808]: 2007/11/21_19:15:00 WARN: stage6: Scheduling Node
standby for STONITH
pengine[1808]: 2007/11/21_19:15:00 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:00 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:01 WARN: process_pe_message:
Transition 1: WARNINGs found during PE processing. PEngine Input
stored in: /var/lib/heartbeat/pengine/pe-warn-9621.bz2
pengine[1808]: 2007/11/21_19:15:01 info: process_pe_message:
Configuration WARNINGs found during PE processing.  Please run
"crm_verify -L" to identify issues.
crmd[1799]: 2007/11/21_19:15:03 ERROR: process_lrm_event: LRM
operation Stonith:0_start_0 (call=9, rc=1) Error unknown error
pengine[1808]: 2007/11/21_19:15:07 WARN: unpack_rsc_op: Processing
failed op (Stonith:0_start_0) for Stonith:0 on standby
pengine[1808]: 2007/11/21_19:15:07 WARN: unpack_rsc_op: Handling
failed start for Stonith:0 on standby
pengine[1808]: 2007/11/21_19:15:08 WARN: native_color: Resource
Stonith:0 cannot run anywhere
pengine[1808]: 2007/11/21_19:15:09 WARN: stage6: Scheduling Node
standby for STONITH
pengine[1808]: 2007/11/21_19:15:09 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:09 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:09 WARN: process_pe_message:
Transition 2: WARNINGs found during PE processing. PEngine Input
stored in: /var/lib/heartbeat/pengine/pe-warn-9622.bz2
pengine[1808]: 2007/11/21_19:15:09 info: process_pe_message:
Configuration WARNINGs found during PE processing.  Please run
"crm_verify -L" to identify issues.
lrmd[1796]: 2007/11/21_19:15:12 WARN: G_SIG_dispatch: Dispatch
function for SIGCHLD took too long to execute: 20 ms (> 10 ms)
(GSource: 0x892cd50)
crmd[1799]: 2007/11/21_19:15:12 ERROR: process_lrm_event: LRM
operation Stonith:1_start_0 (call=12, rc=1) Error unknown error
tengine[1807]: 2007/11/21_19:15:12 WARN: status_from_rc: Action
start on standby failed (target: (null) vs. rc: 1): Error

Whatever you've done to the RA, it no longer starts.
I'd suggest OCF tester but the whole idea of a "stonith" device that
only restarts heartbeat is too horrible to contemplate.
But, I thought ocf RA and stonith RA are different. Shouldn't stonith test be used instead ?

nod, my mistake




pengine[1808]: 2007/11/21_19:15:14 notice: cluster_option: Using
default value '-1' for cluster option 'pe-error-series-max'
pengine[1808]: 2007/11/21_19:15:14 notice: cluster_option: Using
default value '-1' for cluster option 'pe-warn-series-max'
pengine[1808]: 2007/11/21_19:15:14 WARN: unpack_rsc_op: Processing
failed op (Stonith:0_start_0) for Stonith:0 on standby
pengine[1808]: 2007/11/21_19:15:14 WARN: unpack_rsc_op: Handling
failed start for Stonith:0 on standby
pengine[1808]: 2007/11/21_19:15:15 WARN: unpack_rsc_op: Processing
failed op (Stonith:1_start_0) for Stonith:1 on standby
pengine[1808]: 2007/11/21_19:15:15 WARN: unpack_rsc_op: Handling
failed start for Stonith:1 on standby
pengine[1808]: 2007/11/21_19:15:16 WARN: native_color: Resource
Stonith:1 cannot run anywhere
pengine[1808]: 2007/11/21_19:15:16 WARN: native_color: Resource
Stonith:0 cannot run anywhere
pengine[1808]: 2007/11/21_19:15:16 WARN: stage6: Scheduling Node
standby for STONITH
pengine[1808]: 2007/11/21_19:15:16 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:16 WARN: stage6: Scheduling Node
active for STONITH
pengine[1808]: 2007/11/21_19:15:16 WARN: process_pe_message:
Transition 3: WARNINGs found during PE processing. PEngine Input
stored in: /var/lib/heartbeat/pengine/pe-warn-9623.bz2
pengine[1808]: 2007/11/21_19:15:16 info: process_pe_message:
Configuration WARNINGs found during PE processing.  Please run
"crm_verify -L" to identify issues.
tengine[1807]: 2007/11/21_19:15:47 ERROR: tengine_stonith_callback:
Stonith of standby failed (2)... aborting transition.

-----Original Message-----
From: [EMAIL PROTECTED] [mailto:[EMAIL PROTECTED]
] On Behalf Of Frank
Sent: Wednesday, November 21, 2007 4:23 PM
To: [email protected]
Subject: RE: [Linux-HA] Re: Linux-HA Digest, Vol 48, Issue 69


Date: Wed, 21 Nov 2007 14:21:30 +0530
From: <[EMAIL PROTECTED]>
Subject: RE: [Linux-HA] Re: Linux-HA Digest, Vol 48, Issue 69
To: <[email protected]>
Message-ID:
        <[EMAIL PROTECTED]>
Content-Type: text/plain;       charset="iso-8859-1"


Frank,
Congratulations on getting your stonith config to work.
Unfortunately,I've not been so lucky or I don't understand stonith
very well. I've done the exact same configuration as you've but for
some reason , the stonith RA doesn't start at all. I'm using
version 2.0.8 and the ssh stonith agent to test my configuration. I
too have two nodes : active and standby, where I want active to
stonith standby and vice versa.

I've attached a snapshot of my stonith configuration for you to
look at . I would really appreciate if you could point out any
mistakes I might have made...
        

        <crm_config>
     <cluster_property_set id="cib-bootstrap-options">
       <attributes>
         <nvpair id="cib-bootstrap-options-default-resource-
stickiness" name="default-resource-stickiness" value="INFINITY"/>
         <nvpair id="cib-bootstrap-options-stonith" name="stonith-
enabled" value="TRUE"/>
         <nvpair id="cib-bootstrap-options-quorum" name="no-quorum-
policy" value="IGNORE"/>
       </attributes>
     </cluster_property_set>
   </crm_config>
   <nodes>
     <node id="6f6d5305-43e3-46f9-a8d7-7966218bab0b"
uname="standby" type="normal"/>
     <node id="bf83d809-e951-4c04-b605-28b5ead6ccfe"
uname="active" type="normal"/>
   </nodes>
   <resources>
        ...
        ...
        <clone id="stonith_active">
       <instance_attributes id="stonith_active_attr">
         <attributes>
           <nvpair id="stonith_active_clone_max" name="clone_max"
value="1"/>
           <nvpair id="stonith_active_clone_node_max"
name="clone_node_max" value="1"/>
         </attributes>
       </instance_attributes>
       <primitive id="Stonith_active" class="stonith"
type="external/rsh" provider="heartbeat">
         <operations>
           <op id="4" name="start" timeout="10s" prereq="nothing"/>
           <op id="5" name="monitor" timeout="5s" interval="15s"/>
         </operations>
         <instance_attributes id="rsh_active_instance_attr">
           <attributes>
             <nvpair id="rsh_attr_active" name="hostlist"
value="active"/>
           </attributes>
         </instance_attributes>
       </primitive>
     </clone>
     <clone id="stonith_standby">
       <instance_attributes id="stonith_standby_attr">
         <attributes>
           <nvpair id="stonith_standby_clone_max" name="clone_max"
value="1"/>
           <nvpair id="stonith_standby_clone_node_max"
name="clone_node_max" value="1"/>
         </attributes>
       </instance_attributes>
       <primitive id="Stonith_standby" class="stonith"
type="external/rsh" provider="heartbeat">
         <operations>
           <op id="6" name="start" timeout="10s" prereq="nothing"/>
           <op id="7" name="monitor" timeout="5s" interval="15s"/>
         </operations>
         <instance_attributes id="rsh_standby_instance_attr">
           <attributes>
             <nvpair id="rsh_attr_standby" name="hostlist"
value="standby"/>
           </attributes>
         </instance_attributes>
       </primitive>
     </clone>
        </resources>
        <constraints>
         <rsc_location id="proxy_stonith_active_rsc_location"
rsc="stonith_active">
       <rule id="prefered_location_proxy_stonith_active_rsc"
score="INFINITY">
         <expression
id="prefered_location_proxy_stonith_active_rsc_expr"
attribute="#uname" operation="eq" value="standby"/>
       </rule>
     </rsc_location>
     <rsc_location id="proxy_stonith_standby_rsc_location"
rsc="stonith_standby">
       <rule id="prefered_location_proxy_stonith_standby_rsc"
score="INFINITY">
         <expression
id="prefered_location_proxy_stonith_standby_rsc_expr"
attribute="#uname" operation="eq" value="active"/>
       </rule>
     </rsc_location>
   </constraints>

Regards,
Abhi.
Hi Abhi,
- first I suppose that you have set pulblic keys so that each host
can make ssh to the other without password
- define only one ssh stonith clone resource with hostlist set to
both nodes, with "clone_max" to 2 and "clone_node_max" to 1
- define 2 place constraints with the clone resource, one with
"uname eq node1" on the other with "uname eq 2"
- start stonith resources

I strongly recommend you to read the excellent tutorial 
"http://linux-ha.org/_cache/HeartbeatTutorials__LCA2007-tutorial.ppt
"  to clarify concepts.

Frank

--
Aquest missatge ha estat analitzat per MailScanner a la cerca de
virus i d'altres continguts perillosos, i es considera que està net.
For all your IT requirements visit: http://www.transtec.co.uk

_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems


The information contained in this electronic message and any
attachments to this message are intended for the exclusive use of
the addressee(s) and may contain proprietary, confidential or
privileged information. If you are not the intended recipient, you
should not disseminate, distribute or copy this e-mail. Please
notify the sender immediately and destroy all copies of this message
and any attachments.

WARNING: Computer viruses can be transmitted via email. The
recipient should check this email and any attachments for the
presence of viruses. The company accepts no liability for any damage
caused by any virus transmitted by this email.

www.wipro.com
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems


The information contained in this electronic message and any attachments to this message are intended for the exclusive use of the addressee(s) and may contain proprietary, confidential or privileged information. If you are not the intended recipient, you should not disseminate, distribute or copy this e-mail. Please notify the sender immediately and destroy all copies of this message and any attachments.

WARNING: Computer viruses can be transmitted via email. The recipient should check this email and any attachments for the presence of viruses. The company accepts no liability for any damage caused by any virus transmitted by this email.

www.wipro.com
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to