Know what broke without the alert storm.
Know what your cluster costs.

Two open-source tools for the people who run infrastructure. One is shipping and you can run it today. The other is honestly still being built.

Warden
Shipping now

Reliability monitoring that explains itself.

Warden confirms sustained failures, learns normal latency for each service, detects patterns and keeps enough evidence to explain the alert. It stays small enough to self-host without turning into a general observability platform.

Connect an MCP-compatible assistant and ask what is down, why a check failed or which certificates need attention. With a separate editor key, it can create and organize monitors without delete access. See AI control.

  1. 1.0

    It checks the service

    HTTP and HTTPS endpoints, TCP ports, ICMP hosts, DNS records and Docker containers, from 10 seconds up. HTTP checks let you set the method, headers, body, timeout, retries and healthy status codes per monitor.

    Checks
    HTTP · TCP · ping · DNS · Docker
    Interval
    from 10s
  2. 2.0

    It learns normal before it calls something slow

    Warden learns P50 and P95 for each monitor. One bad response is not an alert: it confirms failure or degradation, waits for the condition to persist, groups related failures and quiets flapping until it settles.

    Signal
    P50/P95 baseline · explicit override
    Recovery
    confirmed, not assumed
  3. 3.0

    It keeps the evidence and makes it operable

    Alerts go to Slack, a JSON webhook or email. Operators get failure evidence, incidents and patterns; a role-scoped MCP assistant can investigate without receiving delete or credential-management tools.

    Alerts
    Slack · webhook · email
    Interfaces
    dashboard · REST API · MCP · status page
10s
fastest check
4
roles
team seats
AGPL-3.0
licence
Recon
In development

Kubernetes cost intelligence.

A cluster spends money in places nobody remembers deploying. Recon is meant to map the spend to the namespace and tell you what to cut. Meant to — it is being built, there is nothing to run yet, and this section will stay this short until there is.

  • Cost broken down per namespace Planned
  • Idle resources surfaced Planned
  • Right-sizing recommendations Planned

What isn’t built yet.

Published so you can decide with the real picture, and without dates, because a date we invented would be worth exactly as much as a feature we invented. If one of these is a dealbreaker, say so on the call and we’ll tell you straight whether to wait.

  • Checks from more than one place

    Planned

    Every check runs from the single host Warden lives on. One host means one opinion about whether a site is up. Remote probes are the next big piece of work.

  • On-call rotations and escalation

    Planned

    Alerts fire into Slack, webhooks and email today, and every enabled channel gets every event. Rotations, escalation policies and routing a client’s alerts to their own channel are on the list.

  • Recon — Kubernetes cost intelligence

    In development

    Namespace-level cost breakdown, idle resource detection and right-sizing. Being built now; nothing to install yet.

Bring your client list.

Subscribe once and we’ll prepare your managed Warden instance, including hosting, updates and backups, within one business day.

Start managed Warden