Guides

Website down? A 15-minute incident playbook

By the MoniterMySite teamPublished 9 min read

The first fifteen minutes of an outage decide how long it lasts and how many customers notice. A minute-by-minute playbook: confirm, communicate, locate the layer, decide, recover, learn.

The short answer

Confirm the outage is real from a second region, post an Investigating update on the status page within five minutes, then work down the stack in a fixed order (DNS, SSL, CDN or WAF, origin, dependencies) until the recorded error matches a layer. If the outage lines up with a deploy and rollback is one command, roll back; otherwise fix forward. Post Identified, Monitoring and Resolved as you go, let the monitor confirm recovery from every region, and write the blameless postmortem while the timeline is fresh. The clock below assumes an alert has just arrived.

Minutes 0–2: confirm it is real

Open the alert before touching anything. A MoniterMySite down alert names the monitor, the error (for example “Unexpected HTTP status 503” or a timeout) and the regions that observed it, and the incident page shows a per-region timeline. If confirmation was set to 2, the outage has already been seen from two places and you can skip the “is it just me?” step entirely. If you run a single-region monitor, that is the moment to curl the URL from somewhere else: a laptop on mobile data, a cloud shell, or a colleague in another city.

Read the regions. All five failing with the same error means the service or something in front of it is down. One region failing while the others pass points at a network path, a geo-blocking rule or a regional CDN edge, which is a different fix and usually a less urgent one.

Minutes 2–5: say something

Post before you diagnose. Open the status page, create an incident with an impact level, set it to Investigating and write two sentences: what customers will see, and that you are on it. Subscribers are emailed immediately and the update is timestamped on the page, so support can point every ticket at one link instead of answering each one. Promise the next update in a fixed time (“next update within 15 minutes”) and keep that promise even if the news is “still investigating”.

If your monitors already feed the page, the affected components have been marked down automatically and the incident history shows the start time; your post adds the human context. Keep it plain: no root cause guesses yet, no apologies longer than the facts.

Minutes 5–12: find the layer

Work top-down in the same order every time, and stop at the first layer whose symptom matches the recorded error. The point of the order is that each layer rules out everything above it.

Symptom in the alertLikely layerFirst check
DNS resolution failed, or wrong IP returnedDNSdig the name from two resolvers; compare with the DNS monitor's expected value; check registrar and nameserver changes
Certificate expired, name mismatch, handshake errorSSL / TLSopenssl s_client to the host; look at the SSL monitor's days-to-expiry; check the renewal job's heartbeat
403, 429, or a 5xx with the CDN's brandingCDN / WAFFetch the origin directly, bypassing the CDN; look for a new firewall rule, rate limit or an edge outage in the provider's own status page
503, 502, timeouts, or connection refused from the originOrigin / applicationCheck the last deploy time against the incident start; look at app and web server logs, disk, memory and the database connection pool
Page returns 200 but the keyword is missingApplication or dependencyOpen the page; a blank grid, empty JSON or an error template usually means a database, cache or third-party API dependency

Minutes 5–12, continued: the four layers in detail

DNS first because nothing else matters if the name does not resolve. Confirm from a resolver you do not control (a public one, not the office router). A DNS monitor with expected values will have told you exactly which record changed; without it, compare today's answer with yesterday's from the registrar's change log. Remember TTLs: a bad record can look fine in one region and broken in another for an hour.

SSL second because an expired or mis-issued certificate takes the whole site down for every browser at once, yet the server itself is healthy. SSL expiry alerts at 30, 14, 7, 3 and 1 days are there to make this layer boring; if you are here anyway, the fix is a renewal, and the follow-up is a heartbeat on the renewal job.

CDN and WAF third. A sudden 403 or 429 from all regions right after a security rule change, or a 5xx page carrying the CDN's own branding, means the origin may be fine. Fetch the origin directly (by IP with a Host header, or through a bypass hostname) to prove it. If one region alone is failing, suspect an edge location before anything else.

Origin last, because it is where most outages actually live and where the fix takes longest. Line the incident start time up with the deploy log; most application outages begin within a few minutes of a release, a config change or a certificate rotation. Then read the application log for the first error after that timestamp, not the most recent one.

Minute 12: roll back or fix forward

The rule is simple and worth agreeing in calm weather. If the outage coincides with a deploy and rolling back is a single, rehearsed command, roll back now and investigate later; the cost of being wrong is one more deploy. Fix forward when rollback is unsafe (a database migration has already run, data has been written in the new format) or when the cause is external (DNS, certificate, CDN, a third-party API), because there is nothing to roll back to.

Whichever you choose, post Identified on the status page with one sentence on the cause in customer terms (“a configuration change to our CDN”) and the action under way. Avoid the word “fix” until the monitor agrees with you.

Minutes 12–15: let monitoring confirm recovery

After the change, do not declare victory from one browser. A down monitor is re-checked every minute, and recovery is declared only when every selected region is healthy again; the recovery alert carries the total downtime, and a PagerDuty incident opened by the alert resolves itself at the same moment. Move the status page incident to Monitoring when the first regions turn green and to Resolved when all of them have. Subscribers get both updates, and the incident history keeps the timeline.

If recovery is partial, with one region still failing, you are back at the CDN or network layer for that region and the incident stays open.

Afterwards: the blameless postmortem

Write it within two working days, while memory and logs are both intact, and write it without names. The question is what allowed the failure, not who typed the command. A useful template fits on one page:

  • Timeline: first failed check, alert sent, acknowledged, status page posted, cause identified, change applied, recovery confirmed. Pull the timestamps from the incident record rather than memory.
  • Impact: which components, from which regions, for how long, and how many subscribers were notified.
  • Cause: the technical trigger and the condition that let it reach production.
  • What went well: usually detection, if confirmation was set; sometimes communication.
  • Action items with owners and dates: a missing monitor, a keyword check on a page that returned 200 while broken, a heartbeat on a renewal job, a rehearsed rollback, a maintenance window for the change that caused it.
  • One sentence on whether the fail threshold, confirmation count or interval should change. Most false alarms and most slow detections are fixed here.

How monitoring data shortens each step

Each minute above gets shorter when the monitor records more than “down”:

  • Confirmation (minutes 0–2): the alert already lists the regions that saw the failure, so the second-opinion step is done before you open a terminal.
  • Communication (2–5): components tied to monitors flip on the page automatically, with the start time, so the Investigating post is two sentences, not a diagnosis.
  • Locating the layer (5–12): the recorded error (status code, timeout, DNS or TLS failure, missing keyword) maps directly onto the table above, and DNS and SSL monitors with expected values name the changed record or the expiring certificate outright.
  • Deciding (12): per-region timelines distinguish an edge problem from an origin problem, which is the difference between calling the CDN and rolling back a deploy.
  • Recovery (12–15): one-minute re-checks and all-regions-healthy recovery mean you stop guessing when it is safe to post Resolved.
  • Postmortem: the incident record, alert delivery log and daily uptime statistics supply the timeline, so the write-up is analysis rather than archaeology.

The 15-minute playbook

Pin this next to the on-call rota.

  • 0–2 min: read the alert's error and regions; confirm from a second region if the monitor only has one.
  • 2–5 min: post Investigating on the status page with impact and next-update time.
  • 5–12 min: DNS → SSL → CDN/WAF → origin → dependencies; stop at the first layer matching the recorded error.
  • 12 min: deploy-correlated and rollback rehearsed → roll back; otherwise fix forward. Post Identified.
  • 12–15 min: wait for every region to pass; post Monitoring, then Resolved after the recovery alert.
  • Within 2 days: blameless postmortem with timeline, cause, action items and one monitoring change.

Frequently asked questions

How do I know whether the outage is my site or the monitoring service?

Look at the regions. If every region reports the same error at the same time, the problem is on your side or in front of it. If one region fails while the rest pass, suspect a network path or a regional edge before your origin. Requiring confirmation from two regions before an alert removes most monitoring-side noise.

Should I post on the status page before I know the cause?

Yes. Customers want to know that you know, not what the root cause is. An Investigating update with the user-visible impact and a time for the next update, posted within five minutes, does more for trust than a perfect explanation an hour later.

When is rolling back the wrong call?

When a database migration has already run, when data has been written in a format the old version cannot read, or when the cause is outside the deploy (DNS, certificate, CDN, a third-party API). In those cases there is nothing safe to roll back to, so fix forward.

What should the postmortem change in monitoring?

Almost every incident suggests one: a keyword check on a page that returned 200 while broken, a heartbeat on the job that silently stopped, a DNS or SSL monitor with expected values, a maintenance window for the change that triggered the page, or a different fail threshold. Pick one and ship it the same week.

Start monitoring in 30 seconds.

Nothing to install. No credit card. 10 monitors free, forever.