Website Downtime Tracking: How to Measure and Report Outages

Track website outages with a downtime report template, explicit observation gaps and separate down/degraded records. Includes a worked availability example.

Updated · Project Helena · 5 min read ·
uptime monitoring downtime incident management

To track website downtime, record each outage’s observed start, recovery, duration, affected endpoint and observer location. Keep missing observations and slow-but-successful responses separate from confirmed downtime. A useful report explains what was measured, not just a percentage.

Use the downtime report template below for a monthly summary and incident log. The example values are illustrative, not Warden benchmark results or customer data.

What is Downtime Tracking?

Downtime tracking is the systematic recording of every period when your service is unavailable or degraded. It goes beyond uptime monitoring (which detects incidents in real-time) to provide historical data for SLA reporting, trend analysis, and capacity planning.

A complete downtime record includes:

  • Start time — When the outage began (detected by monitoring)
  • End time — When the service recovered
  • Duration — Total downtime in minutes/seconds
  • Impact — Which services/users were affected
  • Root cause — What caused the outage
  • Category — Infrastructure, deployment, dependency, etc.

How to Measure Downtime

External uptime monitoring tools sample your service from an observer’s network. They record observed failures but can miss an outage between checks or a problem affecting another region.

For equally spaced samples with complete coverage, a rough estimate is:

Estimated downtime = Number of failed checks × Check interval

At one-minute intervals, five consecutive failures suggest roughly five minutes of downtime. This does not identify the exact moment the outage started or ended. Preserve last success, first failure, last failure and first recovery timestamps, as well as confirmation settings. When samples are missing or intervals change, use observed incident boundaries and mark uncertainty instead of multiplying counts blindly.

Method 2: Log Analysis

Parse your server access logs and error logs to identify periods with elevated error rates or zero traffic. This catches issues monitoring might miss but requires log infrastructure.

Method 3: Real User Monitoring (RUM)

Collect availability data from actual user sessions. This shows real user impact but can’t detect outages during low-traffic periods (middle of the night, holidays).

Combining Methods

Best practice is external monitoring as your primary source of truth, supplemented by log data and RUM for context. The monitoring tool defines when downtime starts and ends; logs and RUM explain the impact.

Calculating Availability

The standard formula:

Availability % = ((Total minutes - Downtime minutes) / Total minutes) x 100

For a 30-day month (43,200 minutes) with 45 minutes of downtime:

((43,200 - 45) / 43,200) x 100 = 99.896%

Use the uptime calculator to convert between downtime duration and availability percentage.

Creating Downtime Reports

Downtime report template

Copy this structure into your incident record, or download the Markdown template.

Service / endpoint:
Reporting period and timezone:
Observer location and check interval:
Success rule and timeout:
Confirmation and notification settings:
Planned maintenance policy:
Observed downtime:
Observed degradation (reported separately):
Missing observations / monitor downtime:
Availability formula and denominator:
Incident timestamps, impact and evidence:
Known cause, or "not established":
Corrective action, owner and review date:

Suppose a fully observed 30-day month contains two non-overlapping outages of 12 and 33 minutes. Downtime is 45 minutes and time-based availability is about 99.896%. A separate 20-minute latency degradation does not automatically add 20 minutes of downtime: the agreed success rule determines that. If the observer was offline for two hours, report the gap; do not quietly count it as healthy time.

Monthly SLA Report

Include:

  1. Overall availability — The headline number (e.g., 99.95%)
  2. Incident summary — Each outage with duration, impact, and root cause
  3. Trend chart — Monthly availability over the last 12 months
  4. Error budget status — How much of your error budget was consumed
  5. Action items — What you’re doing to prevent recurrence

Incident Report (Per Outage)

For each significant outage, create a postmortem:

  • Timeline of events
  • Root cause analysis
  • Impact assessment (users affected, financial cost)
  • Corrective actions with owners and deadlines

Tracking Downtime Over Time

The most useful metric isn’t a single month’s availability, it’s the trend. Track:

  • Incidents per month — Is frequency increasing or decreasing?
  • Mean Time Between Failures (MTBF) — Average time between incidents. Higher is better
  • Mean Time To Detect (MTTD) — How fast you find problems. Determined by check interval
  • Mean Time To Resolve (MTTR) — How fast you fix problems. Improved by runbooks and automation
  • Availability trend — Month-over-month and quarter-over-quarter

A service with 99.9% availability and improving MTTR is in better shape than a service with 99.99% availability and worsening incident frequency.

Tools for Downtime Tracking

  1. Uptime monitoring tools — Warden, Uptime Robot, Better Uptime. These are your primary data source
  2. Status page platforms — Public-facing incident history for customers
  3. Incident management tools — PagerDuty, Opsgenie. Track incident lifecycle
  4. Spreadsheets — Surprisingly effective for small teams. Log every incident manually
  5. Custom dashboards — Grafana, Datadog. Visualize trends from monitoring data

Common Mistakes

  • Not tracking scheduled maintenance — Even planned downtime should be recorded (but flagged as planned)
  • Rounding generously — 4 minutes and 20 seconds of downtime is not “about 4 minutes.” Precision matters for SLA calculations
  • Ignoring partial outages — If 30% of users can’t access your API, that’s partial downtime worth tracking
  • No root cause categorization — Without categories, you can’t identify systemic issues (e.g., “50% of outages are deployment-related”)
  • Tracking without action — Data without follow-up actions is just noise. Every report should include improvement actions

Collect the evidence with Warden

Warden records monitor state, outage history, latency and incidents, with public or restricted status pages for sharing service health. Set up an endpoint monitor and verify a controlled failure and recovery before you depend on its history.

Use those observations to prepare the report above. Warden is not an automatic contractual SLA auditor, and the Markdown download is a blank template, not an export from your instance. Logs and traces can help establish root cause; Warden’s external checks alone cannot prove it.

For a latency commitment, use the API response-time SLA guide alongside this availability report. Self-host Warden or choose managed hosting for $49/month.


Related tools:

Monitor your services with Warden

Catch sustained outages and slow responses, keep incident evidence, and share status with your team. Warden by Project Helena is open-source uptime monitoring with adaptive P50/P95 latency and role-scoped AI operations.

Managed hosting includes one Warden instance, updates and backups. $49/month, no setup fee; access within one business day after successful payment.

Check product fit and capabilities · Uptime today; metrics, logs and traces planned