Go to All Forums

Cooldown Timer on objects with dependencies from triggering down alerts

In our branch offices we have a typical network stack - Router - Switch - Server

The WAN/LAN team is a different team than the Server team

The server is set to have a dependency of the switch, and the switch is set to have a dependency of the router. When the power at the location goes out, once the battery is depleted and the equipment goes offline, the server does not alert/call as the switch and router are also down. The LAN/WAN team gets all notifications. When power is restored, the second the Router and Switch come back up, the Server alert fires "Down for 21 hours (or however long the power outage has been) IMMEDIATLEY after the switch is pingable, before the switch has started authenticating clients or the server is fully booted to the OS.

There should be a cooldown time that's configurable to prevent this.

Support suggested using the notification delay function, but that applies to all instances of the alert, not just alerts tied to a dependency so I would be delaying a server crash notification event where only the server was involved by however long I pick for my delay time. Which is a no go for us.

Like (1) Reply
Replies (1)

Hi there,

Thanks for the detailed write-up, and for including Support's suggested workaround along with why it doesn't fit, that's genuinely helpful context.

The distinction you're drawing is clear, notification delay is a blanket setting applied to every instance of an alert on that monitor, while what you need is scoped specifically to the moment a dependency clears, since network gear and servers recover on very different timelines after a shared outage. Delaying all server-down alerts to smooth over a recovery-timing gap would mean delaying a genuine, unrelated server crash too, which isn't an acceptable tradeoff.

This is a more precise ask than a general cooldown, a delay tied specifically to dependency-clearing events, not to the alert type as a whole.
I've logged this for the team working on dependency and notification logic. I don't have a timeline to share yet, but this kind of detail (the exact failure sequence, and why the existing workaround doesn't hold up) makes it much easier to scope properly.

Appreciate you taking the time to lay this out so clearly.

Regards,
Jenzo
Site24x7
Like (0) Reply

Was this post helpful?