Skip to content

Backup alerts technicians don't mute

· Practice · 7 min read · BackupSentinel

Every backup alerting setup starts out well meant. Every failure goes to the team channel, every warning too, and for a week everyone reads them. Then a job flaps all night, a NAS sends the same warning forty times, and someone mutes the channel. From then on, the alerts are still being sent, and nobody is reading them.

A muted channel is worse than no channel, because it looks like coverage. The fix is not more alerts or louder ones. It is fewer alerts that each mean something, sent to the person who can act on them, at a volume a person can keep up with.

This post sets out the practices that keep backup alerts read. They apply whatever tool sends them.

One alert per job, not per email

The unit of an alert is a job that needs attention, not an email that arrived. If a job fails at 01:00, retries at 02:00 and fails again, that is one problem, and it should be one alert.

So an alert opens when a job goes from fine to not fine, and stays open while it stays that way. Further failures from the same job belong to that open alert and send nothing new. A failure streak ("failed four times in a row") is useful to see wherever the alert is listed. Four separate notifications are not.

The rule that follows is simple: if nothing has changed for a person to do, do not notify them again.

Two severities, and what they mean

Backup alerts need very few levels. Two cover almost everything:

  • Critical means data is not being protected: the backup failed, or it did not run and its report never arrived. Someone should look today, often tonight.
  • Warning means the backup ran but something is off: some files were skipped, a retry was needed, it completed with errors. Someone should look during working hours.

Resist a third "info" level for alerts. Successes, size changes and routine notices belong on a dashboard or in a summary, not in a channel where they dilute the two levels that matter.

A missing report deserves the same severity as a failure. A job that stopped running has protected nothing, and it is the one that never sends a failure email of its own. When a report counts as missing is a question of its own, covered in RPO, grace windows and when a backup is really late.

Escalate in place, do not duplicate

Say a nightly job completes with warnings on Monday. A warning alert opens. On Tuesday it fails outright.

The wrong outcome is two alerts: an old warning someone has already acknowledged, and a new failure that looks like a separate problem. The right outcome is that the same alert is raised to critical, sent again because it is now worse news, and its acknowledgement is cleared, because whoever said "I will look at the warning" did not sign up for a failure.

Going the other way, from failure to warning, is not news worth sending. The alert stays open at its highest severity until the job recovers.

Tell people when it is fixed

An alert with no ending leaves people guessing. When a job sends a good report again, the alert should close by itself and a short recovery notice should go to the same places the alert went: "Recovered: Nightly File Servers is healthy again."

Two details make recovery notices trustworthy:

  • Send them only where the original alert was delivered. A recovery for an alert the channel never received is noise.
  • When someone resolves an alert by hand, say who did it, so the rest of the team stops looking.

Quiet hours for warnings only

Nobody needs a "completed with warnings" notice at 03:00. Quiet hours, in the team's own timezone, are a good way to keep night-time notifications to the ones that matter.

But quiet hours should apply to warnings only. Critical alerts, a failed backup or a missing one, go out at any hour. If a failure at night is not worth waking anyone for, that is a decision about who is on call, not about the alert.

Decide what happens to warnings held back by quiet hours. Delivering all of them in a burst at 08:00 recreates the noise problem in the morning. Dropping them from the channel and leaving them visible in the queue and the daily summary is usually enough.

Escalate critical alerts nobody has picked up

A critical alert can be sent, delivered and still missed: the person on duty is with a client, the message scrolled away, it went to a shared channel where everyone assumed someone else had it.

Acknowledgement closes that gap. Once someone acknowledges an alert, the team knows it is owned. If a critical alert is still unacknowledged after a set time, send it once more to a second place: a manager, an on-call channel, a pager. Pick the delay to match how quickly you need a response; thirty minutes to two hours suits most backup work.

Send the escalation once. Repeating it every few minutes is how escalation channels get muted too.

Cap the flapping jobs

Some jobs flap: a VPN that drops every other hour, a target that is nearly full, a job that fails and succeeds alternately. Each recovery closes the alert and each failure opens a new one, so the one-alert-per-job rule does not help.

A cap does. Limit how many notifications a single job can send in a rolling day, and when it hits the limit, stop sending and show clearly that it has been paused. The job is still broken and still listed; it has just lost the right to interrupt people. A flapping job is one problem that needs one person to fix it, not a stream of messages for everyone.

A digest instead of noise

Plenty of what is worth knowing is not worth interrupting someone for: a warning that cleared itself, a job that has been acknowledged and is being worked on, a backup that shrank by half overnight, reports that did not match any job.

Put that in one daily email, early in the morning, sent to the person who runs the service. It answers "what happened overnight?" in one read, and it means the channels can stay reserved for things that need action now.

Route by client

An MSP with forty clients rarely has one team that handles all of them. Alerts should follow the responsibility:

  • Limit each channel to the clients its team looks after, so a channel's members see only their own clients' alerts.
  • Send critical alerts for clients with out-of-hours cover to the people on call for that client.
  • Keep one channel for escalations only, so it stays quiet enough to be noticed.

A technician who receives only the alerts they can act on has no reason to mute anything.

A short checklist

  • One open alert per job; repeats update it rather than notify again.
  • Two severities: critical for failed and missing, warning for everything that ran with problems.
  • A warning that becomes a failure is raised in place and sent again.
  • Recovery notices go only where the alert went.
  • Quiet hours hold back warnings, never critical alerts.
  • Unacknowledged critical alerts escalate once, to a second place.
  • Flapping jobs hit a cap and say so.
  • Everything else goes in a daily digest.
  • Channels are scoped to the clients their team owns.

How BackupSentinel does it

BackupSentinel keeps one open alert per job. A job that keeps failing stays on that alert and does not send again; a warning that becomes a failure is escalated in the same alert, with its severity raised, its acknowledgement and snooze cleared, and the alert sent again. A missing report opens a critical alert. When a good report arrives the alert resolves itself and "Recovered: Nightly File Servers is healthy again" goes to the channels that received it; a manual resolve sends "Resolved by" with the person's name.

Alerts go to Email, Slack, Microsoft Teams and PagerDuty (critical alerts only). Quiet hours follow the workspace timezone: warnings in quiet hours are not sent and are recorded as skipped, while critical alerts and escalations always go out. A critical alert that is not acknowledged or snoozed after 30, 60, 120, 240 or 480 minutes, as you choose, is sent once more to an escalation channel as "Not acknowledged". Channels can be limited to some clients or set to escalation-only, and a client can send its critical alerts to up to five on-call contacts. Alert emails are capped at 100 per backup in a trailing 24 hours, and a banner says how many flapping backups have had their email alerts paused. Size drops are never sent to channels at all. A daily digest goes to the workspace owner at 07:00.

Everything that needs a person, alerting or not, sits in Needs attention, where alerts can be acknowledged, snoozed, assigned and resolved with a note. The channel settings are covered in Alerts and on the alerts feature page.

Find out what your backups are not telling you.

Start the free trial, point one backup's report at its client address, and see it turn Healthy, or Missing.