Cron monitoring: report successful completion, not just a running server
Design heartbeat checks that catch missing, failed and stuck background jobs without treating a running host as proof of a successful task.
By the UptimeMonitor360 team · · 3 min read
A healthy server can miss every scheduled job
An HTTP monitor tells you whether its endpoint responded. It cannot establish that yesterday's backup finished, an invoice export reached its destination or a queue consumer made progress. A scheduler can stop while the web application remains healthy. Monitor the job's expected result separately from the machine that runs it.
A heartbeat reverses the direction of the observation: the job reports an event to the monitoring service. The monitor alerts when the expected signal does not arrive within its schedule and grace period. That catches absence, which a success-only application log can make surprisingly hard to notice.
Choose the event that means success
Send a success signal only after the task's meaningful work completes and its result passes the checks you need. For a backup, writing a file may be insufficient; a non-empty, readable archive and a separate restore exercise provide stronger evidence. For an export, verify the destination accepted the transfer rather than merely opening a connection.
A start signal is useful for measuring runtime, but it is not completion. Pair it with a maximum expected duration so a process stuck halfway through is detectable. A failure signal should carry a bounded diagnostic such as an exit code, never passwords, access tokens or a full customer record.
Set an explicit grace period
If a job runs every 15 minutes and sometimes waits three minutes for a lock, an immediate missing-heartbeat alert will be noisy. Choose grace from observed scheduling jitter and the delay the business can tolerate. Very long grace suppresses noise by delaying detection; it is a tradeoff, not free reliability.
Check the timezone used by the actual scheduler. Calendar schedules can behave differently during daylight-saving changes. The cron expression explainer helps inspect the expression, but the host scheduler's timezone and implementation still determine execution.
Handle delivery retries without repeating the job
The monitoring endpoint might be temporarily unavailable after the business task has completed. Retry the heartbeat with a bounded timeout; do not repeat an irreversible business transaction just because the notification failed. Use a stable per-run idempotency key where supported so repeated delivery does not become multiple successful runs.
Keep the heartbeat URL secret. Anyone who knows a bearer-style heartbeat URL may be able to submit a false signal. Store it in the job's protected environment, avoid command tracing that prints it, and rotate it when access changes. Do not place it in a public screenshot or source repository.
Test all three failure modes
- Stop the schedule in a controlled test and confirm a missing-run alert.
- Make a test execution fail and confirm a failure event.
- Simulate a run exceeding its allowed duration and confirm the stuck-run behaviour.
Restore the schedule and verify recovery. Route alerts to someone who can investigate the job, with a concise description of its purpose. The cron failure guide explains how these checks complement ordinary uptime monitoring.