Know what broke without the alert storm.
Know what your cluster costs.

Two open-source tools for the people who run infrastructure. One is shipping and you can run it today. The other is honestly still being built.

Warden
Shipping now

Reliability monitoring that explains itself.

Warden confirms sustained failures, learns normal latency for each service, detects patterns and keeps enough evidence to explain the alert. Uptime is the starting point: service metrics, logs and distributed traces are planned extensions toward a complete observability product.

Connect an MCP-compatible assistant and ask what is down, why a check failed or which certificates need attention. With a separate editor key, it can create and organize monitors without delete access. See AI control.

  1. 1.0

    It checks the service

    HTTP and HTTPS endpoints, TCP ports, ICMP hosts, DNS records and Docker containers, from 10 seconds up. HTTP checks let you set the method, headers, body, timeout, retries and healthy status codes per monitor.

    Checks
    HTTP · TCP · ping · DNS · Docker
    Interval
    from 10s
  2. 2.0

    It learns normal before it calls something slow

    Warden learns P50 and P95 for each monitor. One bad response is not an alert: it confirms failure or degradation, waits for the condition to persist, groups related failures and quiets flapping until it settles.

    Signal
    P50/P95 baseline · explicit override
    Recovery
    confirmed, not assumed
  3. 3.0

    It keeps the evidence and makes it operable

    Alerts go to Slack, a JSON webhook or email. Operators get failure evidence, incidents and patterns; a role-scoped MCP assistant can investigate without receiving delete or credential-management tools.

    Alerts
    Slack · webhook · email
    Interfaces
    dashboard · REST API · MCP · status page
10s
fastest check
4
roles
∞
team seats
AGPL-3.0
licence

Uptime first. Full observability is the direction.

Warden starts with knowing whether a service is available and responding normally. The roadmap expands into service metrics, then logs and traces, so teams can investigate service health in one product.

  1. Available today

    Uptime and response time

    Endpoint checks, adaptive latency, confirmed incidents, notifications, status pages and role-scoped MCP operations.

  2. Planned next

    Service metrics

    Expand from external checks to measurements from the services themselves.

  3. Planned

    Logs

    Add application log collection and investigation alongside service health.

  4. Planned

    Distributed traces

    Follow requests across services to investigate where time is spent.

Metrics ingestion, log storage/search and distributed tracing are not available in Warden today. There are no committed release dates or announced prices for these planned capabilities. Choose the current managed plan for its uptime features.

Explore the foundations: metrics, logs and traces.

Recon
In development

Kubernetes cost intelligence.

A cluster spends money in places nobody remembers deploying. Recon is meant to map the spend to the namespace and tell you what to cut. Meant to — it is being built, there is nothing to run yet, and this section will stay this short until there is.

  • Cost broken down per namespace Planned
  • Idle resources surfaced Planned
  • Right-sizing recommendations Planned

What isn’t built yet.

Published so you can decide with the real picture, and without dates, because a date we invented would be worth exactly as much as a feature we invented. If one of these is a dealbreaker, say so on the call and we’ll tell you straight whether to wait.

  • Checks from more than one place

    Planned

    Every check runs from the single host Warden lives on. One host means one opinion about whether a site is up. Remote probes are planned.

  • On-call rotations and escalation

    Planned

    Alerts fire into Slack, webhooks and email today, and every enabled channel gets every event. Rotations, escalation policies and routing a client’s alerts to their own channel are on the list.

  • Recon — Kubernetes cost intelligence

    In development

    Namespace-level cost breakdown, idle resource detection and right-sizing. Being built now; nothing to install yet.

Bring your client list.

Subscribe once and we’ll prepare your managed Warden instance, including hosting, updates and backups, within one business day.

Start managed Warden