On Mon, May 2, 2011 at 3:51 PM, Lars Ellenberg <[email protected]> wrote: > On Mon, May 02, 2011 at 01:20:16PM +0200, Andrew Beekhof wrote: >> On Mon, May 2, 2011 at 8:27 AM, Ulrich Windl >> <[email protected]> wrote: >> >>>> Andrew Beekhof <[email protected]> schrieb am 29.04.2011 um 09:31 in >> >>>> Nachricht >> > <[email protected]>: >> >> On Fri, Apr 29, 2011 at 9:27 AM, Dominik Klein <[email protected]> >> >> wrote: >> >> > It waits $dampen before changes are pushed to the cib. So that >> >> > eventually occuring icmp hickups do not produce an unintended failover. >> >> > >> >> > At least that's my understanding. >> >> >> >> correcto >> > >> > Hi! >> > >> > Strange: So the update is basically just delayed by that amount of >> > time? I see no advantage: If you put a bad value to the CIB >> > immediately or after some delay, the value won't get better by that. >> > "Damping" siggests some filtering to me, but you are saying your are >> > not filtering the values, but just delaying them. Right? >> >> Only the "current" value is written. >> So the cluster will tolerate "minor" outages provided they last for >> less than the dampen interval and the monitor frequency is high >> enough. > > Still, we may get a spurious failover in this case: > > reachability: > +++++++++_____++++++++++++_____________________________ > Node A monitoring intervals: > + - + + + - - - - - > Node B monitoring intervals: > + + - + + - - - - - > "dampening" interval: |---------| > > Note how the "dampening" helps to ignore the first network "glitch". > > But for the "permanent" network problem, we may get spurious failover:
Then your dampen setting is too short or interval too long :-) > > One dampening interval after node B notices loss of reachability, > it will trigger a PE run, potentially moving things from B to > A, because on A, the reachability (in the CIB) is still ok. > > Shortly thereafter, the dampening interval on A also expires, and the > CIB will be updated with "A cannot reach out there either". > > Any resource migrations triggered by "B cannot reach out there" > are now recognized as spurious. > > Question is, how could we avoid them? > > "ipfail" used to "ask the peer", wait for the peer to notice the new > situation as well, and only then trigger actions. > > We could possibly store a short history of values, and actually do some > "filtering" arithmetic with them. Not sure if this should be done > inside or outside of the CIB. Probably outside. Yes, outside :-) One of these attrd needs a rewrite :-( _______________________________________________ Linux-HA mailing list [email protected] http://lists.linux-ha.org/mailman/listinfo/linux-ha See also: http://linux-ha.org/ReportingProblems
