Learn

Heartbeat monitoring for cron jobs

GlossaryUpdated

Heartbeat monitoring, sometimes called a dead man's switch, reverses the usual direction of uptime checks. Instead of a monitor calling your service, your scheduled job calls a private URL each time it finishes successfully. If a ping does not arrive within the expected interval plus a grace period, the monitor raises an alert.

Why ordinary monitoring misses cron jobs

An HTTP or ping monitor checks something that is supposed to be always on. A nightly backup, an hourly report, a queue worker or a data sync is different: it is supposed to run, finish and go quiet until next time. There is no URL to check at two in the morning, and if the job silently stops running, nothing on the server looks any different. The server is up, the cron daemon is up, and the backup directory simply stops gaining new files.

A job can stop because the crontab was lost in a rebuild, a credential expired, disk filled up, the container that ran it was never redeployed, or someone commented it out and forgot. The failure is usually discovered weeks later, when the backup is needed and the newest copy is stale.

How heartbeat monitoring works

A heartbeat monitor expects a signal rather than sending one. You tell it how often the job should run, say once a day, and how much lateness to tolerate, say thirty minutes. It gives you a unique, private URL. At the end of the job, the job makes a single HTTP request to that URL. Each request resets the timer.

If the timer runs past the expected interval and the grace period without a request arriving, the monitor goes down, an incident is opened and alerts go out exactly as they would for a website outage. The next successful ping closes the incident. Because the monitor only knows about absence, the job has to report success, not merely that it started; a job that starts and crashes halfway should produce no ping at all.

Expected interval and grace period

The expected interval is the schedule of the job: every five minutes, every hour, every day. The grace period absorbs normal variation. A backup that usually takes forty minutes but occasionally takes seventy should have a grace period that covers the slow case, otherwise you will be alerted for jobs that are merely slow. Conversely, a grace period longer than the interval itself means a missed run can go unnoticed for a whole cycle.

A useful rule is to set the grace period to the longest run time you have seen plus a margin. For a job that must never be late, such as a payment reconciliation, a small grace period is correct and an alert for lateness is a feature.

How to monitor a cron job

The simplest integration is one extra command on the crontab line. Run the job, and only if it exits successfully, call the ping URL with curl. The && operator ensures the ping is skipped when the job fails, so a failed backup is reported as a missed heartbeat.

0 2 * * * /usr/local/bin/backup.sh && curl -fsS https://monitermysite.com/api/heartbeat/<your-token> > /dev/null

The flags matter: -f makes curl return a non-zero exit code on an HTTP error, -s silences progress output so cron does not email it to you, and -S still prints errors if the ping itself fails. Redirecting to /dev/null keeps the crontab quiet on success. The same pattern works in CI pipelines, systemd timers, Kubernetes CronJobs and any language that can make an HTTP request: finish the work, verify it succeeded, then ping.

What makes a good heartbeat

Treat the ping as a statement that the job completed correctly. Put it after the verification step, not after the launch step, and never ping from a catch block that swallows errors. Keep the ping URL private, because anyone who has it can mark the job as healthy. Give each job its own monitor rather than sharing one URL, so a missed run names the job that missed.

How MoniterMySite handles this

Heartbeat is one of the seven MoniterMySite monitor types and is included on every plan, including the free Starter plan. You set an expected interval and a grace period, copy the private ping URL from the monitor page and call it at the end of the job; GET and POST both work. When no ping arrives within the interval plus the grace period the monitor goes down, an incident opens and your attached contacts are alerted; the next ping resolves it.

Frequently asked questions

What happens if the job runs but fails?

With the && pattern the ping is only sent when the job exits with status zero, so a failing job produces no ping and is reported as a missed heartbeat once the grace period passes. Make sure your script exits non-zero on failure.

Can I monitor a job that runs every minute?

Yes. Set the expected interval to one minute and a short grace period. For very frequent jobs consider whether a single missed run should alert, or whether a longer grace period that catches a sustained stoppage is more useful.

Is a heartbeat the same as a cron monitor or dead man's switch?

The terms are used interchangeably. All describe a monitor that alerts on the absence of an expected signal rather than on a failed request.

See it in practice.

Add your first monitor on the free plan in under a minute. Multi-region confirmation is included from the Launch plan.