Know what broke without the alert storm.
Know what your cluster costs.
Two open-source tools for the people who run infrastructure. One is shipping and you can run it today. The other is honestly still being built.
Reliability monitoring that explains itself.
Warden confirms sustained failures, learns normal latency for each service, detects patterns and keeps enough evidence to explain the alert. It stays small enough to self-host without turning into a general observability platform.
Connect an MCP-compatible assistant and ask what is down, why a check failed or which certificates need attention. With a separate editor key, it can create and organize monitors without delete access. See AI control.
-
1.0
It checks the service
HTTP and HTTPS endpoints, TCP ports, ICMP hosts, DNS records and Docker containers, from 10 seconds up. HTTP checks let you set the method, headers, body, timeout, retries and healthy status codes per monitor.
- Checks
- HTTP · TCP · ping · DNS · Docker
- Interval
- from 10s
-
2.0
It learns normal before it calls something slow
Warden learns P50 and P95 for each monitor. One bad response is not an alert: it confirms failure or degradation, waits for the condition to persist, groups related failures and quiets flapping until it settles.
- Signal
- P50/P95 baseline · explicit override
- Recovery
- confirmed, not assumed
-
3.0
It keeps the evidence and makes it operable
Alerts go to Slack, a JSON webhook or email. Operators get failure evidence, incidents and patterns; a role-scoped MCP assistant can investigate without receiving delete or credential-management tools.
- Alerts
- Slack · webhook · email
- Interfaces
- dashboard · REST API · MCP · status page
Kubernetes cost intelligence.
A cluster spends money in places nobody remembers deploying. Recon is meant to map the spend to the namespace and tell you what to cut. Meant to — it is being built, there is nothing to run yet, and this section will stay this short until there is.
- Cost broken down per namespace Planned
- Idle resources surfaced Planned
- Right-sizing recommendations Planned
What isn’t built yet.
Published so you can decide with the real picture, and without dates, because a date we invented would be worth exactly as much as a feature we invented. If one of these is a dealbreaker, say so on the call and we’ll tell you straight whether to wait.
-
Checks from more than one place
PlannedEvery check runs from the single host Warden lives on. One host means one opinion about whether a site is up. Remote probes are the next big piece of work.
-
On-call rotations and escalation
PlannedAlerts fire into Slack, webhooks and email today, and every enabled channel gets every event. Rotations, escalation policies and routing a client’s alerts to their own channel are on the list.
-
Recon — Kubernetes cost intelligence
In developmentNamespace-level cost breakdown, idle resource detection and right-sizing. Being built now; nothing to install yet.
Bring your client list.
Subscribe once and we’ll prepare your managed Warden instance, including hosting, updates and backups, within one business day.
Start managed Warden