A request can fail even when the service at the other end is healthy. A connection drops, a route changes, or a resolver returns an old answer. If every failed request produced an alert, a small team would spend too much time checking problems it cannot reproduce.
Tern runs network checks from Frankfurt, London, Virginia, Oregon, Singapore and Sydney. When a check fails, we ask a second region to repeat it before opening an incident and sending alerts. Two observations give us more useful evidence than one.
What the second request tells us
Suppose a request from London times out. We make the confirmation request from another region, with the same URL and check conditions. If that request also fails, we have reason to treat the service as unavailable. The incident records both observations so the person on call can see why we paged them.
If the confirmation succeeds, we keep the failed result in the check history. It still matters. A run of isolated failures may point to a regional routing issue, even when it does not meet our threshold for an incident. The response-time chart can help you see whether the same region keeps behaving differently.
This decision has a cost. Confirmation adds time between the first failure and the alert. A 30-second interval describes how often a paid check runs; it does not promise an alert within exactly 30 seconds. Request timeouts, confirmation and message delivery all take time. We think that delay is worth stating clearly.
Where this rule needs care
Some endpoints deliberately restrict access by location. Others return different content depending on where a request starts. A successful confirmation does not prove that every visitor can use your service. It tells us that the failure was not reproduced in the second region.
Choose a check target that represents the service you want to watch. For an API, that might be a small authenticated request with a predictable response. For a website, a keyword check can catch an error page that still returns a successful HTTP code. Keep credentials scoped to the monitoring task.
Scheduled jobs need a different observation. There is no page to request when a backup misses its deadline. We use the heartbeat deadline, allow the configured grace period, then have a second regional worker confirm that no heartbeat arrived.
The aim is a page that gives someone a reason to act. We keep the underlying results visible because confirmation is a decision rule, and your team should be able to inspect the evidence behind it.



