Setting Latency Thresholds That Actually Catch Real Problems
Latency Monitoring Is More Than a Single Number
Monitoring whether an API returns HTTP 200 is necessary but not sufficient. A service can be "up" while responding so slowly that users abandon requests, checkout flows fail, and mobile apps time out. Latency thresholds turn raw response times into actionable alerts before user experience collapses.
The challenge is setting thresholds that catch genuine problems without waking your team for every network hiccup. This guide covers practical approaches to defining, measuring, and tuning latency limits for production APIs.
Start With a Baseline
Before setting thresholds, observe normal behavior. Run health checks for at least one to two weeks and record:
- Typical response times during peak and off-peak hours
- Seasonal patterns (end-of-month traffic, marketing campaigns)
- Dependency-related spikes (database maintenance, cache cold starts)
Your threshold should sit above normal variance but below user pain. If p95 latency during business hours is 180ms, a 400ms alert threshold may be reasonable; 250ms might generate constant noise.
Think in Percentiles, Not Averages
Averages hide tail latency. An API averaging 120ms might still have 5% of requests exceeding 2 seconds—enough to frustrate users on slow connections. When evaluating thresholds, look at p95 and p99 response times, not just mean values.
For external uptime monitoring, consecutive slow responses often matter more than a single spike. Require two or three slow checks in a row before alerting to filter transient blips.
Align Thresholds With SLAs and User Expectations
Different endpoints have different expectations:
- Authentication APIs: Users tolerate little delay—often under 300ms
- Search and listing endpoints: 500ms–1s may be acceptable
- Report generation: Seconds may be fine if UX sets expectations
Document SLA targets per service tier. Your monitoring thresholds should trigger before SLAs are breached, giving time to investigate and remediate.
Separate Availability From Performance Alerts
Down alerts (connection failures, 5xx errors) should fire immediately. Latency alerts can use higher consecutive failure counts or longer evaluation windows. Treating both identically leads to either missed slowdowns or excessive noise.
Tuning Over Time
Review alert history monthly. Ask:
- Did latency alerts correlate with user complaints or support tickets?
- Which alerts were false positives during CDN or DNS blips?
- Did we miss incidents because thresholds were too loose?
Adjust thresholds incrementally. Tighten limits after repeated user-impacting incidents; loosen them when alert volume exceeds your team's capacity to respond meaningfully.
Latency Monitoring With TwoPulse
Configure each monitored service with a maximum acceptable latency alongside expected HTTP status codes. Heartbeat checks record response time on every probe, surface slow services on your dashboard, and trigger alerts when thresholds are breached—giving you performance visibility without a complex APM deployment.
Well-tuned latency thresholds bridge the gap between "service is up" and "service is usable." Invest time in baselines and iteration, and your team will catch slowdowns before they become outages in the eyes of your users.