Skip to content

Monitors

Monitors are the core of Warden. Choose a check type, give it a name and target, and Warden starts checking it immediately — your first result appears within seconds.

TypeTargetUp when
HTTPhttps://example.com/healthThe response status code is accepted
TCPdb.internal:5432A connection to the port succeeds
Ping192.168.1.1The host replies to an ICMP echo request
DNSexample.comThe selected record lookup returns an answer
DockerA container selected from a connected Docker hostThe container is running and healthy

All types use the same check history, incidents, thresholds, retries and notification logic. Ping may require an ICMP sysctl or NET_RAW capability on Kubernetes. Warden can reach private network targets from the host where it runs, but blocks link-local cloud metadata ranges.

Docker hosts are connected once and reused by container monitors. Warden stores the container name so the monitor survives normal Compose recreation. Access to the Docker socket is equivalent to powerful host control; use a restricted socket proxy rather than relying on a :ro mount.

See Monitor types for ICMP permissions, Docker host setup and target-specific options.

  1. From the dashboard, click Add Monitor
  2. Fill in:
    • Name — a unique display name
    • Type — HTTP, TCP, Ping, DNS or Docker
    • Target — a URL, host:port, host/IP, DNS name or container from a connected Docker host
    • Group — which group this monitor belongs to (use Default if you don’t need groups)
    • Interval — how often to check (default: every 60 seconds)
  3. Click Create

That’s it. Warden runs the first check immediately and shows the result on the dashboard.

StateMeaning
UpResponding successfully within latency threshold
DegradedResponding but slower than the latency threshold
DownFailing checks (confirmed after consecutive failures)
PausedMonitoring stopped — no checks running
MaintenanceGroup is in a maintenance window — checks run but notifications are suppressed

Pause a monitor to temporarily stop checks without deleting it. Resume to start checking again immediately. Historical data is preserved.

Most monitors work great with the defaults. These settings are available when you need more control.

Override the global defaults from Settings for individual monitors:

FieldRangeGlobal DefaultDescription
Confirmation Threshold1–1003Failed checks before confirming down
Notification Cooldown0–1440 min30 minPause between repeated flap/stabilized events
Latency Threshold1+ ms1000 msLatency above this = degraded

Customize the HTTP request Warden sends:

FieldDefaultDescription
MethodGETHTTP method (GET, HEAD, POST, PUT, DELETE)
Headers—Custom request headers (max 50)
Body—Request body for POST/PUT (max 10 KB)
Timeout5sRequest timeout (1–120 seconds)
Follow RedirectsYesWhether to follow HTTP redirects
Accepted Status Codes< 400Custom success criteria
Automatic RetryOn for new HTTP monitors without custom request configurationAt most one eligible transient retry within the original timeout
Manual Retry Count0Explicit retries on failure (0–5, 1s delay between); cannot combine with automatic retry

Timeout and retry count also apply to TCP, ping and DNS. DNS monitors additionally support A, AAAA, MX, NS and TXT records plus an optional custom resolver.

By default, any status code below 400 is considered successful. Customize with individual codes, ranges, or both:

  • 200,201,301
  • 200-299
  • 200-299,301,302

Open the Checks tab for a 24-hour failure and recovery summary and paginated checks. Automatic HTTP mode records retry evidence; existing/manual monitors need HTTP_DIAGNOSTICS_ENABLED=true for additional tracing. HTTP diagnostics explains timing phases, retry eligibility and the optional second probe.

Each monitor tracks:

  • Uptime over 24 hours, 7 days, and 30 days (percentage of successful checks)
  • Latency displayed in 4 time ranges: 1 hour, 24 hours, 7 days, and 30 days

When no explicit monitor threshold exists, Warden can learn a P50/P95 baseline and derive a degraded threshold for that service. See Adaptive latency.

Warden doesn’t alert on the first failure. It first confirms the state, then waits for the configured sustained-outage window before announcing it.

Confirmation threshold: By default, 3 consecutive failures are required before a monitor is confirmed down and an outage opens. Change this globally in Settings or per monitor.

Sustained alerting: By default, a confirmed outage must remain open for 180 seconds before the first notification. Ongoing outages follow first and repeat reminder intervals.

Notification cooldown: Repeated flapping and stabilized events are suppressed for 30 minutes by default. Down and degraded reminders use the sustained-alert ladder instead.

Flap detection: If a monitor keeps flipping between up and down, Warden recognizes the instability and sends a single “flapping” alert instead of many individual alerts. Once the monitor stabilizes, you get a “stabilized” notification.

Recovery confirmation: Warden can require consecutive successful checks before confirming recovery, preventing false recovery alerts. Default is 1 check, configurable up to 20 in Settings.

For HTTPS monitors, Warden automatically tracks SSL certificate expiry and alerts you at 30, 14, 7, and 1 day before expiry. SSL alerts are sent once per threshold during a mid-day window in your timezone.

EventTrigger
downMonitor confirmed down
upMonitor recovered
degradedLatency exceeded threshold
recoveredLatency returned to normal
flappingRapid state changes detected
stabilizedFlapping stopped
ssl_expiringSSL certificate approaching expiry

Each event type can be toggled on/off and independently included in a daily digest in Notifications.

Check history is automatically cleaned up based on the retention period in Settings (default: 365 days, range 1–3650). Monitor metadata, events, and outage records are kept indefinitely.

Deleting a monitor permanently removes it along with all check history, events, and outage records. This cannot be undone.