Hi,

On Fri, Jun 05, 2009 at 04:21:59PM +0200, Husemann, Harald wrote:
> Hi Dejan,
> 
> thanks for your answer, after playin' a bit more with stonith I figured 
> out the following:
> 
> - the riloe script only distinguishes between "button" and other 
> methods, so, it doesn't make a difference if the method is set to "off", 
> "power", or anything else except "button"
> - All methods (including "button") seem to rely on a running acpid on 
> the stonith'ed machine to switch it off

Yes, button is sort of "quick power button press" which is
handled by acpi. "power" should be more like pulling the power
plug.

> - If I call stonith on the command line with "-T off" the machine is 
> switched off immediately when acpid is not running, but the cluster 
> seems to use another method which relies on acpid

stonith on command line should behave in exactly the same way
like stonithd.

> - A reset works regardless of acpid

Oh, you're still running 2.1.3. external/riloe was completely
rewritten in the meantime. Definitely upgrade to 2.1.4. Or at
least update the plugin.

> So, my solution is now to stop acpid and set the stonith_method to 
> "reboot" instead of "poweroff".
> With this, the fenced node is rebooted by the DC when the communication 
> is disturbed.
> 
> Hm... I'd like to have it switched off instead of rebooting to prevent 
> resources from falling back to a probably defect node, but it seems that 
> this is impossible with riloe.

It should work with the new version.


Thanks,

Dejan

> Now, I'd take a look at the link and configure the second stonith resource
> 
> Thanks, and regards,
> 
> Harald
> 
> Dejan Muhamedagic schrieb:
> > Hi,
> > 
> > On Thu, Jun 04, 2009 at 06:04:04PM +0200, Husemann, Harald wrote:
> >> Hi list,
> >>
> >> I have a Heartbeat v2 cluster build out of two HP DL380, each equipped 
> >> with an iLO. Both nodes can see each other's iLO interface, and STONITH 
> >> works almost as expected - when I cut off the communication of the nodes 
> >> (by blocking port 629 on the DC with iptables), it sends a powerdown 
> >> request to the other node's iLO.
> >> The other node starts shutting down, and now things get funny: When it 
> >> comes to shutdown heartbeat itself, the node wants to tell the DC that 
> >> it's going down - but it can't reach the DC, so, the shutdown script 
> >> goes in an endless loop.
> > 
> > Perhaps try with another method. Can't recall which one, but one
> > of the methods depends on the system OS to do an orderly
> > shutdown. You should use another ilo_powerdown_method (power is
> > the default, I guess that that should be OK). BTW, it's not
> > really good that fencing depends on the OS cooperation.
> > 
> >> When I re-enable the internal communication (i. e., drop the iptables 
> >> rule), the node deregisters itself at the DC, continues the shutdown 
> >> process and finally switches off the power.
> >>
> >> Hmmm... Any ideas how I can force the DC to *really* powering off the 
> >> other node, or prevent this one from trying to deregister itself??
> >>
> >> Another (maybe stupid) question: I learned that I can use clone 
> >> resources for stonith, and that I should do this to ease things and make 
> >> the config better readable. Okay, but how to do this?? All examples I've 
> >> found deal with ibmhc which has only one IP address for the Stonith 
> >> device, but my iLO has different adresses for each node. Is it possible 
> >> to use cloned resources here? Maybe someone can give me an example for 
> >> this?
> > 
> > No, since you have just two nodes (right?) and the configuration
> > is different for each node. You can check this for more info:
> > 
> > http://clusterlabs.org/mediawiki/images/f/f2/Crm_fencing.pdf
> > 
> > Thanks,
> > 
> > Dejan
> > 
> > 
> >> Thanks + have a nice hackin',
> >>
> >> Harald
> >>
> >> Some infos of the hardware and software in use:
> >>
> >> HP DL380, iLO-version 1.84
> >> HA Version 2.1.3, CRM Version 2.0 (CIB feature set 2.0)
> >> OS CentOS 5.3 (final)
> >>
> >> The "interesting" parts of my CIB:
> >>
> >> ===================/snip/==========================
> >>   <cluster_property_set id="cib-bootstrap-options">
> >>           <attributes>
> >>             <nvpair id="cib-bootstrap-options-dc-version" 
> >> name="dc-version" value="2.1.3-node: 
> >> 552305612591183b1628baa5bc6e903e0f1e26a3"/>
> >>             <nvpair id="cib-bootstrap-options-last-lrm-refresh" 
> >> name="last-lrm-refresh" value="1244115187"/>
> >>             <nvpair id="cib-bootstrap-options-stonith-enabled" 
> >> name="stonith-enabled" value="true"/>
> >>             <nvpair id="cib-bootstrap-options-stonith-action" 
> >> name="stonith-action" value="poweroff"/>
> >>           </attributes>
> >>         </cluster_property_set>
> >> (...)
> >>   <primitive id="rs_stonith-db1" class="stonith" type="external/riloe" 
> >> provider="heartbeat">
> >>           <meta_attributes id="rs_stonith-db1_meta_attrs">
> >>             <attributes>
> >>               <nvpair id="rs_stonith-db1_metaattr_target_role" 
> >> name="target_role" value="started"/>
> >>             </attributes>
> >>           </meta_attributes>
> >>           <instance_attributes id="rs_stonith-db1_instance_attrs">
> >>             <attributes>
> >>               <nvpair id="a83a7329-5cda-4dbd-9826-7e20bbda3835" 
> >> name="hostlist" value="mat-db-1.***"/>
> >>               <nvpair id="052b8d9c-309b-42ce-8d07-d17e612172f9" 
> >> name="ilo_hostname" value="mat-db-1.***"/>
> >>               <nvpair id="6da9867d-8d45-4344-8fff-4cfda8cce93d" 
> >> name="ilo_user" value="stonith"/>
> >>               <nvpair id="65dfb20e-0c60-41fb-bf76-ce207c8efd47" 
> >> name="ilo_password" value="******"/>
> >>               <nvpair id="831ff8dc-55ac-46ea-b160-f4a270dd9452" 
> >> name="ilo_can_reset" value="1"/>
> >>               <nvpair id="46649c31-7fbb-4213-a290-ef71e87b970b" 
> >> name="ilo_protocol" value="2.0"/>
> >>               <nvpair id="256ccbd8-1a7f-495e-af44-ebf4be4edca3" 
> >> name="ilo_powerdown_method" value="button"/>
> >>             </attributes>
> >>           </instance_attributes>
> >>         </primitive>
> >> (...)
> >> <constraints>
> >>         <rsc_location id="location_stonith-db1" rsc="rs_stonith-db1">
> >>           <rule id="prefered_location_stonith-db1" score="INFINITY" 
> >> boolean_op="and">
> >>             <expression attribute="#uname" 
> >> id="b2bb572c-4e9c-499d-8cb4-9d71cc924268" operation="eq" 
> >> value="mat-db-2.materna-com.de"/>
> >>           </rule>
> >>         </rsc_location>
> >>       </constraints>
> >> ============/snap/==================================================
> >>
> >>
> >> -- 
> >> Harald Husemann
> >> Netzwerk- und Systemadministrator
> >> Operation Management Center (OMC)
> >> MATERNA GmbH
> >> Information & Communications
> >>
> >> Westfalendamm 98
> >> 44141 Dortmund
> >>
> >> Gesch?ftsf?hrer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig
> >> Amtsgericht Dortmund HRB 5839
> >>
> >> Tel: +49 231 9505 222
> >> Fax: +49 231 9505 100
> >> www.annyway.com <http://www.annyway.com/>
> >> www.materna.com <http://www.materna.com/>
> >> _______________________________________________
> >> Linux-HA mailing list
> >> [email protected]
> >> http://lists.linux-ha.org/mailman/listinfo/linux-ha
> >> See also: http://linux-ha.org/ReportingProblems
> > _______________________________________________
> > Linux-HA mailing list
> > [email protected]
> > http://lists.linux-ha.org/mailman/listinfo/linux-ha
> > See also: http://linux-ha.org/ReportingProblems
> > 
> 
> -- 
> Harald Husemann
> Netzwerk- und Systemadministrator
> Operation Management Center (OMC)
> MATERNA GmbH
> Information & Communications
> 
> Westfalendamm 98
> 44141 Dortmund
> 
> Gesch?ftsf?hrer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig
> Amtsgericht Dortmund HRB 5839
> 
> Tel: +49 231 9505 222
> Fax: +49 231 9505 100
> www.annyway.com <http://www.annyway.com/>
> www.materna.com <http://www.materna.com/>
> _______________________________________________
> Linux-HA mailing list
> [email protected]
> http://lists.linux-ha.org/mailman/listinfo/linux-ha
> See also: http://linux-ha.org/ReportingProblems
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to