Try to adjust the logic of the check dependencies to prevent this from happening. Servers Alive is a great product, but it has it's limitations, and one often has to think a little for it.
I do a lot of what I call "sanity" checks that have to work in order for my real checks to even be executed. The sanity checks do not alert and their results are not output to web pages. Most are pings. I weave them into my check cycle to make sure that checks that require network or VPN connectivity to work, will not fail just because the connection is down. Servers Alive has a limitation where if it is in the middle of the of a series of checks on a server, that are being performed over say VPN, and the VPN goes down, all subsequent checks will fail. If there are 25 checks of a server, and the VPN drops after check 3, that's 22 alerts going out. Even if you use the second knock feature, SA does not have the ability to realize "Hey the ping check that all these other's depend on is down, so I won't bother checking it's "child" checks." Rather, it just re-checks all the checks that were flagged as down on the first pass, and if they are still down (which they will be because the underlying connection is still down) it fires the alerts. I resist the urge to jump on the "SA future features discussions" that seem to blossom on this list every 10 days, but if I could change on thing about Servers Alive, it would be it's ability on second knock to act based on the status of a parent check. In general: do as few checks as you have to, and only alert when you have to. Also, consider purchasing more than one copy of SA, and run it locally on the server/LAN being checked. This improves performance and reliability dramatically. Hope this helps. David -----Original Message----- From: Servers Alive Discussion List [mailto:[EMAIL PROTECTED] Sent: Thursday, May 18, 2006 11:25 AM To: Servers Alive Discussion List Subject: [SA-list] What if all checks fail? I am using SNPP to notify support personnel when a system's checks fail. I am going to build another SA system for a new facility with a dedicated paging server. The catch is that the new configuration won't have the SA system separated from the paging server by the LAN, they will be attached with a crossover Ethernet cable on second NIC ports so the pages should never fail. So here's my question... What if all LAN based checks fail (ping, NT services, etc.)? (i.e. because the SA host's primary nic / cable / switch port fails) The redundant network connection to the paging server will likely ensure that everyone will be paged because everything is down. Can I tune something in SA to not take action or take a custom action if say 30% or more of the checks fail? ( I would rather send one page to myself saying ... "Hey, most of the checks have failed. The ServersAlive host may be wigging out!" ) Thanks in advance! Robert To unsubscribe send a message with UNSUBSCRIBE as subject to [email protected] To unsubscribe send a message with UNSUBSCRIBE as subject to [email protected]
