Skip to content

Product

Everything a technician needs to trust the backups.

BackupSentinel reads the result emails your backup software already sends and keeps one status per backup job, per client. One list in the morning shows what failed, went missing or needs a look.

The morning question, answered in one list.

Are all backups healthy? If not, which ones, and why? The Overview opens with one sentence: how many items need you, split into failed, warnings and missing. Below it sits Needs attention: one open alert per job, ordered by state, then severity, then kind, then age.

Every line says what happened in plain words: “Failed 4 times in a row”, “Last good backup 3 d ago”, “No report for 26 h”. Acknowledge an item so the team knows it is yours, snooze it for 1 h, 4 h or 24 h, or resolve it with a note.

It also lists a backup that reported success but wrote half its usual size or less (against the median of its previous seven runs), and alert channels that stopped delivering.

Acknowledge
The item stays listed, marked as yours, until the backup recovers.
Snooze
For 1 h, 4 h or 24 h. When the time is up it is open again, unless the backup has recovered.
Resolve
With a note. The next problem on the same backup shows the note as a known fix.
Recover
The next OK report closes the alert by itself and sends a recovery notice.
Keyboard
j and k move, a acknowledges, s snoozes, r resolves; press ? to see every shortcut.

Jobs with a heartbeat.

Every backup job is one row: status, source, last run, next expected and the last seven results as a strip. A flaky job shows as a broken strip before it becomes a failed one. Filter by client or state and export as CSV when a customer asks what is being watched.

  • Six states: Healthy, Warning, Failed, Missing, Waiting, Paused
  • Status, last run and next expected update with every report read
  • The jobs list and Needs attention both export as CSV

When a report does not arrive.

Each job has a cadence, daily by default. Every accepted report sets the next expected time from it, and a check runs every 15 minutes.

A job turns Missing once its report is late by more than its grace: a quarter of the cadence, between 15 minutes and 2 hours, for cadences up to 8 hours; 10% and at least 2 hours above that. Until then the job shows a “late” hint, so nobody is alerted about a backup that is still finishing.

Check
every 15 min
Daily job
2 h 24 min grace
Hourly job
15 min grace
SRV-FS01 · Exchange backupDaily · report due 01:00
Missing

Grace: 2 h 24 min for a daily job A check every 15 minutes

Every report is read by a rule.

Ten built-in rules read the common products. Anything else falls to a generic rule that looks for failed, warning and success, or to a rule you write.

A custom rule adds its own patterns and keywords, and an in-browser tester runs the real parser on a pasted email before you save. Customize copies a built-in rule; finished rules can be submitted to a reviewed community library.

Subject pattern
A regular expression that decides which emails the rule reads
Sender pattern
Optional, limits the rule to one sender
Keywords
One list each for OK, Warning and Failed
Field patterns
Job name and size, read from the subject or body

Emails that match no job, or that no rule can read, wait in the Inbox: assign them to a job, start monitoring a new job from the email, or ignore them. A report replaced by a newer one is marked superseded.

Told once, in the right place.

An alert belongs to a job, not to an email: one open item per job, carried by the channels you choose, with a digest for the rest.

  • Channels

    Email, Slack, Microsoft Teams and PagerDuty (Events API v2, auto-resolve). Failed deliveries retry every 10 minutes; each channel has a Send test button.

  • Behaviour

    One open alert per job. A warning that becomes a failure escalates the same alert. When the next report is OK the alert resolves and a recovery notice goes out. Paused clients keep storing reports and send no alerts.

  • Digest

    A daily digest email goes to the workspace owner at 07:00 workspace time with what is open. One click unsubscribes.

The whole estate at a glance.

Two views across every client at once: one for the screen on the office wall, one for the last thirty days.

TV mode

The Overview as a wallboard for a screen in the office: the status counts and the Needs attention list in type you can read from across the room, with the affected clients and the latest reports beside them. Open it with the TV mode button on the Overview, in a browser signed in to your workspace.

Refresh
Every 10 minutes, in place. A hidden tab waits and refreshes as soon as it is shown.
Stale data
After 25 minutes without a refresh it says DATA MAY BE STALE; after an hour the screen dims.
Screen
Made for 16:9 displays and scales with them. One click for full screen.
TV mode with sample data: large counts of 5 failed, 2 missing, 2 warning, 1 waiting, 4 paused and 18 healthy backups, the Needs attention list in large type, the affected clients and the latest reports.

Backup Matrix

One row per backup, one cell per day for 30 days, coloured by that day’s result. Step back a month at a time, filter by client, device or current status, and sort by the backups with the most issues first.

The records around every job.

  • Clients

    SLA tier standard 99.0, gold 99.5, platinum 99.9 or a custom target. Contacts with a role, an email address and an on-call flag; a timezone per client. Pause a client for 24 hours, 7 days, until a date or until resumed: reports are still stored, no alerts go out.

  • Devices

    Created from the first report or by hand, with addresses, model, serial, location, warranty and notes. Labelled logins are encrypted, revealed only to owners and admins, and every reveal is logged. Free and unlimited.

  • Reports

    A monthly client report with success rate, missing reports, incidents and median time to recover, as CSV. A health summary per client, and scheduled summaries by email daily, weekly, monthly or quarterly.

Who can do what.

Roles
Owner, admin, member and viewer. Owners and admins read the activity log and reveal device logins. Seats are unlimited on every plan.
Two-factor
TOTP codes from an authenticator app, per account, and enforceable workspace-wide.
Sign-in
Email and password, Microsoft or Google. Sign-in forms are protected by Cloudflare Turnstile.
Audit and export
An activity log for owners and admins, a JSON export of the workspace data, and self-serve workspace deletion. API keys are stored hashed and limited to 60 requests per minute.

See your own backups in it.

Point one client's report email at the address you get on sign-up and the first job appears with its first result.