Guides

How to Set Up Uptime Monitoring: Checklist

By the MoniterMySite teamPublished 9 min read

Most monitoring setups fail in one of two ways: they page for blips, or they watch a homepage while checkout is broken. This checklist covers what to monitor, how often, from where, and how to route the alerts.

The short answer

Monitor the handful of things that break most and hurt most (homepage, login, checkout or API health, DNS, SSL, scheduled jobs), check them from more than one region, require two consecutive failures and agreement between regions before anyone is paged, send critical alerts to a pager and the rest to chat, mute planned work with maintenance windows, and publish a status page so customers stop asking. Then review it monthly. The steps below use the settings MoniterMySite exposes; the principles transfer to any tool.

1. Decide what to monitor first

Start from the user's point of view, not the infrastructure diagram. The first six monitors should cover the paths a customer walks and the things that fail silently.

  • Homepage: an HTTP(S) monitor on the public URL. Leave the expected status at the default (any 2xx), set a 10-second timeout, and enable SSL validation so a broken certificate chain counts as down.
  • Login page: a Keyword monitor that requires a phrase from the real page, such as “Forgot your password”. A framework error page often returns 200; this is how you catch it.
  • Checkout or API health endpoint: an HTTP monitor on an endpoint that touches the database, expected status set to 200 only. If it needs a header or bearer token, put it on the monitor rather than exempting the endpoint from auth.
  • DNS: a DNS monitor on apex and www with the record type (A, AAAA or CNAME) and expected value, so a registrar hijack or botched change is caught while the old answer is still cached elsewhere. Add MX if you run mail.
  • SSL certificate: an SSL monitor per public hostname. Alerts go out 30, 14, 7, 3 and 1 day before expiry plus your own threshold; 14 days leaves room for a renewal script to fail twice.
  • Scheduled jobs: a Heartbeat monitor for the nightly backup, the invoice run and any queue worker. The job pings a private URL when it finishes; no ping inside the expected interval plus grace period means an alert. Ping only on success: `backup.sh && curl -fsS <ping-url>`.
  • Then, if relevant: a Port monitor for mail, database or any TCP service customers reach directly, and a Ping monitor for office or edge devices.

2. Choose intervals and regions

Interval sets the longest an outage can go unnoticed. The fastest interval depends on the plan: Starter every 5 minutes, Launch every 60 seconds, Growth every 30 seconds, Summit every 10 seconds. Once a monitor is down it is re-checked every minute regardless of interval, so recovery is seen quickly on every plan.

Match the interval to the cost of a missed minute. A marketing site or internal tool is fine at 5 minutes. A checkout, a partner-facing API or anything with an SLA wants 30 to 60 seconds. Ten-second checks are for the few endpoints where a minute of silent downtime is measured in money.

Regions decide whether you can tell a real outage from a bad network path. Starter checks from one region, Launch from up to three, Growth and Summit from all five (New York, San Francisco, Frankfurt, Singapore, Sydney). Checks rotate between the selected regions, and every check, incident and alert records where it was observed from. Pick regions where your customers are, plus one elsewhere as a tie-breaker.

3. Set fail thresholds and confirmation

Two settings filter blips. The fail threshold is the number of consecutive failed checks from one region before that region counts as down; the default of 2 removes one-off packet loss. Confirmation is how many regions must agree before an incident opens and alerts go out. With three regions and confirmation 2, one location losing connectivity never pages anyone, and recovery is declared only when every selected region is healthy again.

A working rule: critical production at a 10 to 30 second interval, threshold 2, confirmation 2 across three regions; internal tools at 5 minutes, threshold 2, one region. Resist threshold 1. It feels safer and it is how teams end up muting the channel.

4. Catch soft errors with keyword checks

The most expensive outages return 200. An empty product grid, a login form that posts to nowhere, a maintenance page left switched on: a plain HTTP check calls all of these healthy. A Keyword monitor requires a word or phrase to be present (or absent) in the response, so pick text only a working page contains: a price, a footer build number, the first product name.

Two details matter. Matching is case-insensitive but exact and runs against the raw HTML, so text rendered by JavaScript will not match; use a server-rendered string or a health endpoint. And for endpoints that should reject anonymous requests, set the expected status to 401 or 403; a format such as 200-299,301 covers redirects.

5. Add response-time thresholds

Up is not the same as usable. Set a slow-response threshold on HTTP and Keyword monitors so you hear when a page answers slowly; one alert is sent when the threshold is first crossed, not on every slow check. Start at roughly double the normal p95 and tighten later. Keep the hard timeout at 10 seconds for APIs and up to 30 for slow legacy pages; anything longer counts as failed.

6. Route alerts deliberately

Decide who is woken for what before the first incident. Email and webhook contacts are available on every plan; Slack, Discord, Microsoft Teams, Google Chat and Telegram from Launch; PagerDuty from Growth. Create one contact per destination and attach contacts per monitor, not globally.

  • Critical monitors (checkout, API health, DNS) go to PagerDuty. A down alert opens a PagerDuty incident with the monitor name, regions and error; recovery resolves it automatically.
  • Everything else goes to a chat channel, where recovery and slow-response messages are context rather than noise.
  • Turn on Down alerts only for the paging contact so recoveries and slow responses never trigger it.
  • Use a webhook for automation. Each event (down, up, slow, ssl_expiry, test) arrives as a JSON POST, retried three times if your endpoint does not answer 2xx within 10 seconds, and HMAC-SHA256 signed when you set a secret.
  • Watch the e-mail budget: Starter allows 20 alert e-mails per day, so point chatty monitors at a webhook or chat channel.
  • Press the test button on every contact and check the delivery status on the Alert contacts page. An untested channel is a guess.

7. Mute planned work with maintenance windows

A deploy that takes the site down for ninety seconds is not an incident and should not count against uptime. Under Maintenance, define windows during which checks keep running but alerts are muted and downtime is excluded from uptime figures. Windows can be one-off or repeat daily, weekly or monthly, for all monitors or a selection. Create the weekly patch window now, before someone is paged at 03:00 by their own change.

8. Publish a status page

A status page is the cheapest support engineer you will ever hire. Create one under Status pages, add components with customer-friendly names (“Website”, “API”) rather than hostnames, group them, and choose what to show: uptime bars for 7, 30 or 90 days, uptime percentage, response times and incident reasons. Visitors subscribe by email and receive every update you post.

Starter includes one public page. From Launch you can serve it at status.yourcompany.com with a single CNAME record (use a subdomain; a bare domain cannot carry a CNAME at most providers). Tick Hide from search engines for internal pages, and on Growth and above add a password for pages meant only for customers under contract.

9. Allowlist the probes

If a WAF or rate limiter blocks unknown clients, you will see a monitor reporting 403 or 429 while the site is fine. Checks identify themselves with the user agent MoniterMySite/1.0 and come from one published IP address per region, listed in the docs under Allowlisting. Allow them in the firewall and filter the user agent out of analytics.

10. Hold a monthly review

Monitoring rots quietly. Book thirty minutes a month and work through the same list:

  • Open each incident from the month. For every false positive, raise the fail threshold or add a region rather than muting the contact.
  • Delete monitors for services that no longer exist; add monitors for anything launched since.
  • Send a test to every alert contact and confirm delivery status is green.
  • Check the SSL monitor list against the certificates you serve; new subdomains are the usual gap.
  • Compare uptime percentages like with like: a 5-minute monitor shows a bigger dip than a 30-second one for the same incident.
  • Confirm maintenance windows still match the deploy schedule.
  • Tag everything (prod, api, customer-x) so the next review takes fifteen minutes.

The checklist

Copy into your runbook.

  • HTTP monitor on the homepage: 2xx expected, 10 s timeout, SSL validation on.
  • Keyword monitor on login with a phrase only the working page contains.
  • HTTP monitor on checkout or an API health endpoint that touches the database, expected status 200.
  • DNS monitors on apex, www and MX with expected values.
  • SSL monitor per public hostname, warning at 14 days.
  • Heartbeat monitor per cron job, backup and queue worker, pinged only on success.
  • Interval: 30–60 s for revenue paths, 5 min for everything else.
  • Three regions on critical monitors, fail threshold 2, confirmation 2.
  • Slow-response threshold at roughly 2× normal p95.
  • PagerDuty contact with Down alerts only on critical monitors; chat contact for the rest.
  • Signed webhook for automation; test sent to every contact.
  • Weekly maintenance window matching the deploy schedule.
  • Status page with customer-named components, email subscriptions on, custom domain from Launch.
  • Probe IPs allowlisted; user agent filtered from analytics.
  • Monthly review booked in the calendar.

Frequently asked questions

How many monitors do I actually need?

Fewer than you think. Six well-chosen monitors (homepage, login, checkout or API health, DNS, SSL, scheduled jobs) cover most of what customers notice. Add more as incident reviews reveal gaps.

Should I set the fail threshold to 1 so I hear about everything?

No. Threshold 1 turns every dropped packet into a page, and within weeks the channel is muted. Threshold 2 with confirmation from a second region catches real outages a check later and almost never cries wolf.

What is the difference between a maintenance window and a scheduled maintenance announcement?

A maintenance window mutes alerts and excludes downtime from uptime figures; it is for your team. A scheduled maintenance announcement appears on the status page and is emailed to subscribers; it is for your customers.

Can I monitor an endpoint that requires authentication?

Yes. HTTP monitors support headers, basic auth and bearer tokens; credentials are encrypted at rest and shown masked. Or set the expected status to 401 to confirm the endpoint answers without exposing it.

Start monitoring in 30 seconds.

Nothing to install. No credit card. 10 monitors free, forever.