P50 is the median: half of measurements are at or below it. P95 is the value at or below which 95% finish. P99 does the same for 99%. Percentiles show the shape of latency better than an average, which a few very slow responses can distort.
Use each percentile for a different question
- P50: What does a typical check look like?
- P95: Is the normally slow tail getting worse?
- P99: What do the worst non-extreme requests experience?
Do not treat P99 from 20 samples as meaningful. One observation represents 5% of that dataset. Tail percentiles need enough samples and a clearly defined window.
Example
Suppose 100 API checks have a P50 of 85 ms, P95 of 180 ms and P99 of 900 ms. The typical experience is fast, but at least one request is dramatically slower. The average alone may not make that pattern obvious.
Paste your values into the P50/P95/P99 latency calculator to calculate nearest-rank percentiles locally in your browser.
Which one should trigger an alert?
For a contractual target, alert on its actual compliance rule. For operational degradation, P95 is often a stable baseline because it includes routinely slow responses without letting the most extreme point define normal.
Warden learns P50 and P95 from successful checks for each monitor. Failed checks are excluded because a timeout duration does not describe healthy service performance. The learned baseline complements—not replaces—an explicit response-time target.
Keep the comparison valid
Percentiles only compare cleanly when the endpoint, request configuration, observer location and time window remain consistent. After moving the observer to another network or region, relearn the baseline instead of assuming the service regressed.