An API response-time SLA needs two numbers: a latency threshold and the percentage of requests that must meet it. “99% of requests under 200 ms” is measurable. “The API should be fast” is not.
Availability and response time are separate signals. A request can return 200 OK and still violate its latency target. To track API response-time SLAs, define the endpoint, threshold, required percentage, observation window and failure policy before choosing the alert.
Warden fits the synthetic monitoring part: it checks endpoints from the host where you run it, records availability and response time, and alerts on sustained failures or degradation. It does not measure every customer request or produce a contractual SLA compliance report automatically.
A response-time objective you can calculate
For a set of measurements:
compliance = eligible requests meeting the success and latency rule / all eligible requests × 100For example, define success as an accepted HTTP status within 200 ms, and count timeouts as violations. If 9,940 of 10,000 eligible requests meet both conditions, compliance is 99.4%. That passes a 99% objective but fails a 99.9% objective. Keep the observation window in the result: a one-hour sample does not establish monthly compliance.
The API response time SLA calculator evaluates numeric durations you provide. It has no HTTP outcome data: a fast 500 cannot be distinguished from a fast 200. Use it to explore latency distributions, and evaluate the full success rule separately for an SLA report.
Define the measurement method too: observer location, request path, HTTP method, timeout, authentication, check frequency and evaluation window. Changing any of them changes the result.
Synthetic checks or real requests?
| Measurement | Answers | Limitation |
|---|---|---|
| Scheduled HTTP checks | Can this observer reach this endpoint, and how long does it take? | Samples one network path at a fixed cadence |
| Server request metrics | What fraction of actual requests meets the target? | Needs complete instrumentation and a defined eligibility rule |
| Browser monitoring | Can a user load and interact with the application? | Measures a different workflow from an API health request |
A monitor that succeeds every minute can miss failures between checks. Equally, a failure from one observer does not establish a worldwide outage. Use application request metrics when the promise covers all production traffic; keep synthetic checks as an independent operational signal.
Why P50, P95 and P99 still matter
Compliance gives a pass/fail result. Percentiles explain the distribution:
- P50 is the typical response.
- P95 is a practical view of the slow tail.
- P99 exposes the experience of the slowest 1%.
An API can pass “99% under 500 ms” while its normal response time gradually doubles from 80 ms to 160 ms. It still passes, but the regression is real. This is why a fixed SLA threshold and a learned baseline answer different questions.
A monitoring policy that avoids noise
Use a fixed per-monitor threshold when a contract or product requirement defines one. Otherwise, learn the service’s normal latency and alert on sustained deviation. Never open an incident from one slow sample: confirm the condition across consecutive checks or a sustained window.
Warden records response time on each check, learns P50 and P95 per monitor, and computes a default degraded threshold from that baseline. An explicit monitor threshold takes precedence. Degraded means reachable but abnormally slow; down remains a separate state.
For example, a learned P95 of 240 ms produces a default degraded threshold of max(240 × 1.5, 240 + 100) = 360 ms. If your target is 200 ms, set the explicit monitor threshold to 200 ms. Letting a learned baseline drift upward must not redefine a contractual target. The adaptive latency guide documents the current defaults.
Set up response-time monitoring in Warden
- Install Warden and create an HTTP monitor for a read-only endpoint you control.
- Configure the accepted status, timeout, authentication and redirect policy. A login redirect returning
200is not proof that the authenticated endpoint is healthy. - Pick a check interval and a per-monitor latency threshold that match the operational target. Use adaptive learning when you want to detect changes from normal rather than enforce a fixed limit.
- Configure notification confirmation and the sustained-alert window. These affect when you are interrupted; a sample can violate the target before an alert is sent.
- Test down, slow and recovered conditions on a non-production endpoint. Confirm delivery in the destination channel and inspect the recorded history.
Use the endpoint setup walkthrough for a concrete configuration. If you prefer not to operate the monitor, Managed Warden is $49/month. It uses the same monitoring model; managed hosting does not turn sampled checks into all-request SLA accounting.
What to put in the SLA document
Write down:
- The endpoint and request configuration.
- The threshold and required compliance.
- The evaluation window.
- Whether planned maintenance is excluded.
- Where the observer runs.
- How timeouts and failed requests count.
- The evidence retained for disputes.
That final point matters. A monthly percentage without check history is difficult to audit. Keep the measurements, incident boundaries and configuration that produced the number.
Start with one important endpoint
Choose a health or read-only API endpoint that represents real dependencies. Monitor availability and latency every 30–60 seconds, then review the first week of data before setting a tight threshold. If you already have a contractual target, configure it immediately and compare it with the observed P95.
Common questions
Does HTTP 200 mean the API meets its response-time SLA?
No. A successful status and a latency limit are separate conditions. A 900 ms response can violate a 200 ms target while the endpoint remains reachable.
Is a latency alert the same as an SLA breach?
No. An alert evaluates an operational rule. A breach evaluates the agreed population and window, such as 99% of eligible requests within 200 ms over a calendar month. Confirmation, cooldowns and maintenance rules must not silently change that agreement.
Can Warden show which database query caused the delay?
No. Warden observes the request from outside the application. Use logs, metrics or distributed tracing to identify internal causes, and correlate them with the time Warden observed the regression.
Product behavior checked against Warden v0.8.0 adaptive latency documentation.