Learn

Multi-region confirmation explained

GlossaryUpdated

Multi-region confirmation is a rule that an outage is only declared, and alerts only sent, when a minimum number of monitoring locations have each seen the target fail. Combined with a per-location fail threshold, it filters out network blips and single-region problems so that a page at three in the morning means the service is genuinely down.

What is a false positive in uptime monitoring?

A false positive is an alert for an outage that did not happen to your users. A single request can fail because of a dropped packet, a transient DNS hiccup, a routing change between the monitoring location and your server, or a rate limiter that briefly refused one more connection. None of these mean your site is down, yet a naive monitor that alerts on the first failed check will page you for every one of them.

False positives are not merely annoying: each one trains the team to ignore alerts, and an ignored alert eventually becomes a real outage that nobody acts on. The two main tools for reducing them are fail thresholds and multi-region confirmation.

Fail thresholds

A fail threshold is the number of consecutive failed checks from one location before that location considers the target down. With a threshold of two, a single failed check is treated as suspicious but not conclusive; the location checks again, usually sooner than its normal interval, and only if the second check also fails does it report a failure.

Thresholds trade speed for certainty. A threshold of one alerts fastest but passes every blip through; two removes most transient errors for one extra check's worth of delay. Higher values suit flaky targets or very short intervals, where three consecutive ten-second failures still add up to under a minute.

Confirmation quorum

A fail threshold protects against a momentary error, not against a problem specific to one monitoring location: a congested route, a regional DNS resolver returning stale answers, or your firewall blocking one set of IP addresses. Then every check from that location fails consistently while users everywhere else are fine.

Confirmation quorum solves this by asking how many locations must independently agree before an incident is opened. With three selected regions and a quorum of two, a single location losing connectivity never pages anyone, because the other two still see the service up. A quorum of one disables confirmation; a quorum equal to the number of regions is the strictest and the slowest to alert. Two of three is a sensible production default: it tolerates one bad location while still alerting within a couple of check cycles.

Recovery conditions

Declaring an outage over is as important as declaring it started. If recovery is announced as soon as any single location sees a success, a service that is flapping, or that has come back for one region but not others, will generate a stream of down and recovered messages. A stricter and more useful rule is that recovery is declared only when every selected region is healthy again. That way a recovery message means the service is reachable from everywhere you measure, and the reported downtime covers the full span of the incident.

While a monitor is down, good systems check more often than the configured interval, so recovery is noticed within a minute or so even on monitors that normally run every five minutes.

Choosing settings for your service

There is no universal answer, but a few patterns cover most cases.

  • Critical production: a 10 to 30 second interval, fail threshold 2 and confirmation 2 across three regions. Outages are confirmed in about a minute with very few false alerts.
  • Standard customer-facing sites: a 60 second interval, threshold 2 and confirmation 2 across two or three regions.
  • Internal tools and side projects: a 5 minute interval, threshold 2 and a single region. Occasional false positives are acceptable when nobody is being woken up.

How MoniterMySite handles this

MoniterMySite checks from five regions on four continents (New York, San Francisco, Frankfurt, Singapore and Sydney) and rotates checks between the regions selected for each monitor, so consecutive checks come from different places. Each monitor has a fail threshold (default 2, up to 10) and a confirmation setting that specifies how many regions must agree before an incident is opened. Recovery is declared only when every selected region is healthy again, and a downed monitor is re-checked every minute whatever its normal interval. The Starter plan checks from one region, Launch from up to three, and Growth and Summit from all regions.

Frequently asked questions

Does multi-region confirmation make alerts slower?

Slightly. Confirming from a second region typically adds one check cycle. At a 30-second interval with a threshold of two, an outage is normally confirmed within about a minute.

What if my site is only meant to be reachable from one country?

Select only monitoring regions that are supposed to reach it, or allowlist the monitoring IP addresses in your firewall. A geo-block that refuses a monitoring region will otherwise look like a permanent regional outage.

Why did one region report a failure while the others did not?

Usually a network path or DNS problem between that region and your server, or a firewall rule that affects one set of addresses. With confirmation enabled this is recorded in the check history but does not open an incident unless a second region agrees.

See it in practice.

Add your first monitor on the free plan in under a minute. Multi-region confirmation is included from the Launch plan.