Il giorno Mar 20 Set 2011 17:54:58 CEST, Dejan Muhamedagic ha scritto:
[...]
>> And I completely agree with this, but in an environment like mine, where
>> a single resource failure might involve all the others (with fence) is
>> wronk to keep this kind of settings. Do you agree with me?
> No. If the resource cannot stop, then something's wrong either
> with the resource or with the RA. And needs to be fixed.

And this is for sure. But I cannot make all resources on a node stop and 
migrate (because of fence) just because one of them has failed.
Those resources are not connected one to each other, so it is more 
reasonable to keep the survived resources alive, or at least live 
migrating them (on-fail=standby) to the other node and THEN reboot the 
first one.

[...]
> If a resource fails to stop, then on-fail="stop" cannot possibly
> help. Furthermore, you basically make this resource less
> available (the cluster won't try to recover it). Must be that
> I'm missing something.
> At any rate, I don't think that you need to fiddle with the
> on-fail attribute, but see what's wrong with the RA or libvirt
> or the combination of the two.
> Thanks,
> Dejan

Then my question is why the attribute on-fail was created? I repeat, I 
totally agree with the fact that there are some problems with the RA but 
until I find out exactly what's wrong I had to take care of all my 
cluster's resources and so the most reasonable thing is to keep one 
single failed vm stopped.

Thanks,

-- 
RaSca
Mia Mamma Usa Linux: Niente รจ impossibile da capire, se lo spieghi bene!
[email protected]
http://www.miamammausalinux.org

_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to