Skip to content

Notification Fatigue Prevention

Warden prevents alert fatigue with several mechanisms that work together. A notification only fires when a problem is confirmed, sustained, not already covered by a bigger incident, not from a monitor that has been shouting all day, not flapping, and not in cooldown.

Confirming a monitor is down and telling you about it are two different moments. When a monitor is confirmed down Warden opens an outage and says nothing; the outage is announced only if it is still down after the sustained window. A blip that resolves inside that window is recorded in the history and the daily digest, and never interrupts anyone.

00:00 check fails, threshold met → outage opens, silent
00:02 still down → still silent
00:03 still down → ALERT: "Monitor is down (Status: 503) — down for 3m"
00:33 still down → reminder: "Still down after 33m"
01:33 still down → reminder: "Still down after 1h33m"
01:40 recovers → recovery sent, because the outage had been announced

The same ladder applies to degraded (high latency) outages.

A recovery is announced only if the outage itself was. Under the default policy most short outages are never announced, so telling you they recovered would be pure noise.

Default: announce after 180 seconds, first reminder after 30 minutes, then every 60 minutes. Set the sustained window to 0 to alert the moment an outage opens, or the reminder interval to 0 to turn reminders off.

A single failed check doesn’t trigger an alert — it could be a transient blip. Warden waits for N consecutive failures before confirming a monitor is down. The sustained-outage window then determines when it can notify. Same logic applies to degraded (high latency) checks.

Check 1: DOWN → count=1/3 → no alert
Check 2: DOWN → count=2/3 → no alert
Check 3: DOWN → count=3/3 → CONFIRMED → outage opens; sustained window starts
Check 4: UP → RECOVERED → recovery only if outage was announced, counter reset

Default: 3 consecutive failures. Set to 1 to confirm on the first failure; the sustained window still applies.

After a flapping or stabilized alert fires, repeats of the same event type are suppressed for a cooldown period.

Down and degraded no longer use the cooldown: how often an ongoing outage repeats is governed by the reminder interval in the sustained ladder above, which knows how long the outage has actually lasted. The cooldown only ever knew how long ago the last message was.

Default: 30 minutes. Set to 0 to disable.

If a monitor rapidly oscillates between UP and DOWN, Warden detects it as “flapping” and suppresses all notifications until the monitor stabilizes. You get a single “flapping” alert when it starts and a “stabilized” alert when it stops.

It works by measuring the percentage of state transitions in a sliding window. Uses hysteresis (start threshold: 25%, stop threshold: 20%) so the detection itself doesn’t oscillate.

Default: enabled, 25% threshold over last 21 checks.

See Adaptive Latency Baselines for the complete guide, including learning time, threshold precedence, MCP inspection, and what happens when Warden moves to a different network or region.

“Slow” is not a single number. A health check that answers in 254ms and a homepage that answers in 427ms are not degraded at the same point, and a fixed global threshold gets both wrong: too high for the fast one, too low for the slow one.

Warden learns each monitor’s own p50 and p95 from its recent successful checks and marks it degraded above max(p95 x 1.5, p95 + 100ms). The multiplier governs for normal services; the 100ms floor takes over for very fast targets, so a service whose p95 is 8ms is not called degraded at 12ms.

service normally at 254ms, p95 420ms → degraded above 630ms
service normally at 5ms, p95 8ms → degraded above 108ms

The alert message says what normal is, because a bare “>630ms” gives the reader no way to judge it:

High latency detected (>630ms, normally ~254ms)

Only successful checks feed the baseline — a failed check’s latency is how long it took to fail, which would inflate “normal” exactly when the monitor is in trouble. Baselines are recomputed hourly over a 7-day window and persisted, so a restart does not spend an hour with no idea what normal looks like.

Precedence. A per-monitor latency threshold set by hand always wins: it is a deliberate statement about that service, usually an SLA, and nothing learned may quietly overrule it. Failing that, the monitor’s own baseline. Failing that — a new monitor, or one with under 200 successful checks — the fixed global threshold.

Repointing a monitor at a different target does not reset its baseline. Warden keeps a monitor’s history across edits rather than discarding it, so the old target’s checks remain in the window and the baseline re-learns gradually as the window rolls past the change. If the new target is fast or slow enough that the old numbers matter in the meantime, set the per-monitor threshold, which outranks the baseline.

Set High Latency → Learn what is normal for each monitor to off to go back to a single fixed threshold for everything.

Monitors that fail together are usually one thing failing. When enough of a group opens an outage inside the correlation window, Warden sends one message naming the group and the affected monitors instead of one per monitor.

“Enough” is a percentage of the group with an absolute floor, so it still means something whether the group has 12 monitors or 200: max(3, 30% of the group).

19:29 eleven of twelve production monitors start failing, all with 404
19:32 ALERT: "11 of 12 monitors in Production APIs are down, for 3m: ..."
20:02 reminder: one message, not eleven

Reminders follow the incident, not the monitor: outages announced together share a correlation id and are reminded about together.

Broad failures are grouped to avoid one alert per target. A high percentage of failing monitors alone does not establish that Warden’s Internet connection is broken. Recent failures across at least three distinct hostnames in the same DNS/TCP/TLS phase provide limited connectivity evidence; queue pressure is reported separately. See HTTP diagnostics for the evidence and its limits.

A monitor that has already interrupted you three times in 24 hours is describing itself, not an event. On the third alert Warden says so explicitly and then stops sending individual alerts for it:

Checkout API has alerted 3 times in the last 24h and is down again.
Muting its individual alerts until it settles — it stays in the daily
digest and on the dashboard.

Subsequent outages are still recorded, still shown, still in the digest. They just do not interrupt. Because those outages are never marked as announced, their recoveries stay quiet too. Set the limit to 0 to turn the damping off.

Any monitor can have its alerts muted from its details panel (Alerts → Mute Alerts). A muted monitor is still checked, still records outages, and still appears in the digest and on the dashboard — it simply never interrupts anyone. This is the right tool for test and staging targets that should not wake anyone.

All settings live in Settings on the dashboard. Changes apply immediately to all running monitors.

SettingDefaultRange
Confirmation threshold31-100
Announce after (seconds)1800-86400
First reminder (minutes)300-10080
Repeat reminder (minutes)600-10080
Cooldown minutes300-1440
Correlation window (seconds)3000-86400
Correlation minimum monitors31-1000
Correlation group share (%)301-100
Probe-wide share (%)801-100
Repeat-offender limit30-1000
Repeat-offender window (minutes)14401-43200
Adaptive latency thresholdstruetrue/false
Latency baseline window (days)71-90
Latency baseline minimum samples2001-1000000
Degraded at (% of p95)150100-10000
Degraded floor above p95 (ms)1000-60000
Flap detection enabledtruetrue/false
Flap window (checks)213-100
Flap threshold (%)251-100

Confirmation threshold, cooldown and latency threshold can be overridden on individual monitors (in the monitor’s Advanced Settings). This lets you set threshold=1 on critical monitors while keeping threshold=5 on less important ones. When not set, the global default is used — except for latency, where an unset value means the monitor’s own learned baseline is used instead.

Alerts muted is also per-monitor, from the monitor’s details panel.

Flap detection, correlation and repeat-offender settings are global only.

Selecting an event under Daily Digest → Include in the digest controls what the daily summary covers. It used to divert the event, so choosing “Down” there silenced outage alerts entirely — an easy way to end up with no notifications at all without realising it.

Those are now two independent decisions:

  • What the digest covers — Daily Digest → Include in the digest.
  • What reaches you immediately — the per-event toggles under Event Types.

An event can be in both. To stop an immediate alert, turn the event off; putting it in the digest no longer does that.