IMHO, since a relatively large amount of time will have passed for snmp_1 since the 1st knock and the 2nd knock, it should fail to the same state as snmp_2 and url_1. Other wise you'll be displaying a state that is most likely no longer true.
In fact, this may be a case for retrying ALL the checks during the 2nd knock, even if they were good during the first round. Although, unless you set up a recursive test to walk the failures backwards during a 2nd knock, you'd ultimately test all checks twice. On the other hand, that would also be a VERY cool thing to do during a second knock (retest the entire down chain of failed tests) by starting at the lowest level and working backwards until you get an UP. Michael D. Shook Technical Analyst Saddle Creek Corporation [EMAIL PROTECTED] 863 668 4477 (work) 863 860 4070 (cell) 863 665 1261 (fax) www.saddlecrk.com > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 3:40 PM > To: Servers Alive Discussion List > Subject: [SA-list] RE: [SA-list] RE: [SA-list] RE: [SA-list] > What if all checks fail? > > I could see a case for either "Not Available" or "OK" > > OK because the last time you checked it (first pass through) > it was OK. > On the Second Knock the parent failed, so you would normally > mark it as > Not Available. > > > -tom > > > > > > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 3:25 PM > To: Servers Alive Discussion List > Subject: [SA-list] RE: [SA-list] RE: [SA-list] What if all > checks fail? > > OK I understand. > BUT what to do for snmp_1 on host? > > > > > Dirk Bulinckx. > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 8:05 PM > To: Servers Alive Discussion List > Subject: [SA-list] RE: [SA-list] What if all checks fail? > > > Actually Dirk, Here is the process now: > > First Pass: > Ping Host > snmp_1 on host > [Host goes down here] > snmp_2 on host (Failed - queue for second knock) > url_1 on host (failed - queue for second knock) > > Second Knock: > snmp_2 on host (Failed - send alert) > url_1 on host (failed - send alert) > > > > What was being asked was: could it do this? > > First Pass: > Ping Host > snmp_1 on host > [host goes down here] > snmp_2 on host (Failed - queue for second knock) > url_1 on host (failed - queue for second knock) > > Second Knock: > PING HOST (since it is a "parent" of a second knock item) > FAILS - send alert > snmp_2 on host DON'T RECHECK - PARENT IS DOWN > url_1 on host DON'T RECHECK - PARENT IS DOWN > > OR - to put it another way: An item being queued for a > SECOND KNOCK should automatically queue its "parent" item. > This would have to be recursive right back up to the root and > ideally could be a configuration options (second knock on parent: y/n) > > > -Tom > > > > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 1:51 PM > To: Servers Alive Discussion List > Subject: [SA-list] RE: [SA-list] RE: [SA-list] What if all > checks fail? > > If the parent check would be down, then it would not check > the childs or do the recheck on second knock. > The way it works is that the parent is checked and if the > parent is up then the childs are checked too. > The only way to be able to do it the way you want it (or the > way I understand that you want it), is that we would redo the > parent EACH time a child needs to be checked. > > Example: > > ping host > snmp_1 on host > snmp_2 on host > url_1 on host > > That would become > ping host > snmp_1 on host > ping host > snmp_2 on host > ping host > url_1 on host > > > > Dirk Bulinckx. > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 7:41 PM > To: Servers Alive Discussion List > Subject: [SA-list] RE: [SA-list] What if all checks fail? > > Try to adjust the logic of the check dependencies to prevent > this from happening. Servers Alive is a great product, but > it has it's limitations, and one often has to think a little for it. > > I do a lot of what I call "sanity" checks that have to work > in order for my real checks to even be executed. The sanity > checks do not alert and their results are not output to web > pages. Most are pings. I weave them into my check cycle to > make sure that checks that require network or VPN > connectivity to work, will not fail just because the > connection is down. > > Servers Alive has a limitation where if it is in the middle > of the of a series of checks on a server, that are being > performed over say VPN, and the VPN goes down, all subsequent > checks will fail. If there are 25 checks of a server, and > the VPN drops after check 3, that's 22 alerts going out. > > Even if you use the second knock feature, SA does not have > the ability to realize "Hey the ping check that all these > other's depend on is down, so I won't bother checking it's > "child" checks." Rather, it just re-checks all the checks > that were flagged as down on the first pass, and if they are > still down (which they will be because the underlying > connection is still > down) it fires the alerts. > > I resist the urge to jump on the "SA future features > discussions" that seem to blossom on this list every 10 days, > but if I could change on thing about Servers Alive, it would > be it's ability on second knock to act based on the status of > a parent check. > > In general: do as few checks as you have to, and only alert > when you have to. > > Also, consider purchasing more than one copy of SA, and run > it locally on the server/LAN being checked. This improves > performance and reliability dramatically. > > Hope this helps. > > David > > > -----Original Message----- > From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] > Sent: Thursday, May 18, 2006 11:25 AM > To: Servers Alive Discussion List > Subject: [SA-list] What if all checks fail? > > I am using SNPP to notify support personnel when a system's > checks fail. > > > I am going to build another SA system for a new facility with > a dedicated paging server. > > The catch is that the new configuration won't have the SA > system separated from the paging server by the LAN, they will > be attached with a crossover Ethernet cable on second NIC > ports so the pages should never fail. > > So here's my question... > > What if all LAN based checks fail (ping, NT services, etc.)? (i.e. > because the SA host's primary nic / cable / switch port > fails) The redundant network connection to the paging server > will likely ensure that everyone will be paged because > everything is down. > > Can I tune something in SA to not take action or take a > custom action if say 30% or more of the checks fail? ( I > would rather send one page to myself saying ... "Hey, most of > the checks have failed. The ServersAlive host may be wigging out!" ) > > Thanks in advance! > > Robert > > > To unsubscribe send a message with UNSUBSCRIBE as subject to > [email protected] To unsubscribe send a message with > UNSUBSCRIBE as subject to [email protected] To unsubscribe > send a message with UNSUBSCRIBE as subject to > [email protected] To unsubscribe send a message with > UNSUBSCRIBE as subject to [email protected] To unsubscribe > send a message with UNSUBSCRIBE as subject to > [email protected] To unsubscribe send a message with > UNSUBSCRIBE as subject to [email protected] > > -------------------------------------- The information contained in this message is intended only for the use of the addressee. If the reader of this message is not the intended recipient or agent of the intended recipient, you are hereby notified that any dissemination, distribution, or copying of the message is strictly prohibited. To unsubscribe send a message with UNSUBSCRIBE as subject to [email protected]
