Skip to content

RPO, grace windows and when a backup is really late

· Practice · 7 min read · BackupSentinel

A backup that fails sends an email saying so. A backup that does not run sends nothing, and the only way to notice is to know when a report was due and see that it has not come. That raises a question that sounds simple and is not: when exactly is a backup late?

Answer too early and every slow night produces a false alarm. Answer too late and a job can be missing for a day before anyone looks. This post works through the arithmetic, with real times, and the limits of the answer.

RPO, and what detection adds to it

A recovery point objective (RPO) is the most data, measured in time, that a client can afford to lose. A nightly backup gives an RPO of about a day: if the server dies just before tonight's run, the newest restore point is last night's.

That figure assumes every run happens. Once a run is missed, the real recovery point keeps getting older until someone notices and fixes the cause. So the number that matters in practice is the backup interval plus the time it takes to notice a missing run. Detection is part of the RPO, and the rest of this post is about keeping that part small without drowning in false alarms.

Interval or schedule

There are two ways to describe what a job should do.

  • A schedule says when it runs: weekdays at 22:00, the first Sunday of the month at 02:00. To check it, the monitor needs the same calendar, the same timezone and the same idea of what happens when the clocks change.
  • An interval says how often a report should arrive: at least once a day, at least once a week. To check it, the monitor needs only the time of the last report.

Schedules are more precise. Intervals are much harder to get wrong: a schedule copied into a monitoring tool has to be kept in step with the backup software every time someone moves a job, while an interval only cares that the job keeps reporting.

Count from when the report arrives

With an interval, the next report is due one interval after the last one arrived. That has a few useful properties:

  • You can work out the deadline yourself. Last report at 23:10 on Monday, daily interval: due at 23:10 on Tuesday. There is nothing hidden to reason about.
  • Moving a job moves its deadline. After one run at the new time, the deadline follows. Moving the nightly run from 22:00 to 20:00 never raises an alert, because early reports are never a problem; moving it later by more than the grace raises one, once.
  • Clock changes are absorbed. On the night the clocks go back, a job that runs at 22:00 local time reports 25 hours after the previous one. A grace window of more than an hour covers that without any timezone logic.

The price is that the deadline also moves when a job drifts later, which is one of the limits covered below.

Grace: how much slack to allow

Reports rarely arrive at the same minute: backups take longer with more data, retries push the result out, mail servers queue. A grace window after the due time absorbs that.

Grace works best in proportion to the interval. A fixed hour is far too generous for an hourly job and far too tight for a weekly one. A rule that holds up:

  • Intervals up to 8 hours: a quarter of the interval, at least 15 minutes and at most 2 hours.
  • Longer intervals: a tenth of the interval, at least 2 hours.

Worked through with real times:

  • Hourly. Last report at 10:05. Due at 11:05, grace 15 minutes, so the job is missing from 11:20.
  • Daily. Last report at 23:10 on Monday. Due at 23:10 on Tuesday, grace 2 h 24 min, so the job is missing from 01:34 on Wednesday.
  • Weekly. Last report at 03:00 on Sunday. Due at 03:00 the following Sunday, grace 16 h 48 min, so the job is missing from 19:48 that Sunday.
  • Monthly (30 days). Last report at 02:00 on 1 October. Due at 02:00 on 31 October, grace 3 days, so the job is missing from 02:00 on 3 November. A job that runs on the 1st of each month reports on 1 November, a day after the due time and well inside the grace, so 31-day months do not raise false alarms.

Add how often the check runs. If the monitor looks every 15 minutes, a daily job that stopped after Monday's run is flagged by 01:49 on Wednesday at the latest. By then the newest restore point is more than 26 and a half hours old. That is the honest RPO of a nightly backup with good monitoring, and it is why a client who needs a firm 24-hour RPO needs backups more often than once a day.

Weekday-only jobs

The interval should cover the longest normal gap between two reports, not the usual one. For a job that runs Monday to Friday nights, the longest normal gap is Friday night to Monday night: about 72 hours.

No daily-style interval covers that. Daily plus grace is 26 h 24 min; every two days plus grace is 52 h 48 min. With a daily interval and a report at 22:30 on Friday, the job is missing from 00:54 on Sunday until Monday night's report arrives. There are three ways to handle it:

  1. Run it every day. Often the simplest answer. Weekend runs of an incremental job are usually small, and the RPO over the weekend improves too.
  2. Keep the daily interval and expect the weekend alert. It clears itself on Monday night, but only works if everyone knows why it fires.
  3. Use a weekly interval. No weekend alert, but a job that stops on a Tuesday is not flagged until well into the following week. That is rarely a good trade.

Long-running jobs

The report arrives when the job ends, so a run that takes longer arrives later. Say a daily job usually reports around 01:00, but the weekly full backup on Saturday night runs until 07:00 on Sunday. The gap from Saturday's 01:00 report to Sunday's 07:00 report is 30 hours, longer than the 26 h 24 min that daily plus grace allows. The job is flagged as missing at 03:24 on Sunday and recovers at 07:00.

Ways to fix it, roughly in order of preference:

  • Start the long run earlier, so it finishes near the usual time.
  • Split the full backup into its own job where the product allows it, with its own weekly expectation.
  • Lengthen the interval to every two days, accepting that a genuine stop is noticed a day later.

Once Sunday's late report is in, Monday's arrives after only 18 hours. That is early, and early is always fine.

The honest limits of interval-based detection

Intervals are hard to get wrong, but they leave gaps:

  • No calendar. An interval cannot say "weekdays only" or "the first Sunday of the month". Those jobs need a looser interval or an accepted false alarm.
  • The deadline follows the last report. A job that runs a little later each night pushes its own deadline later each time and is never flagged, even if it ends up running in the middle of the working day.
  • Some lag is built in. Detection is at best the interval plus the grace plus the gap between checks.
  • Arrival is not quality. A report that arrives on time says the job ran and what it said about itself. It does not prove the backup can be restored, or that it still includes everything it should. A job that reports success while backing up far less than usual needs a different check, such as comparing its size with earlier runs.

None of these are reasons to skip missing-report detection. They are reasons to choose intervals deliberately and to restore-test now and then.

How BackupSentinel does it

Every backup in BackupSentinel has an Expected every interval: hourly, every 2, 3, 4, 6, 8 or 12 hours, daily, every 2 days, weekly or monthly, with daily as the default. The next report is due one interval after the last one was received, with the grace window described above. Every 15 minutes, a backup past its window turns Missing and opens a critical alert; if it already has an open warning, that alert is escalated to critical instead.

Grace window and Missing time for each interval
Expected everyGraceMissing after
Hourly15 min1 h 15 min
Every 2 hours30 min2 h 30 min
Every 3 hours45 min3 h 45 min
Every 4 hours1 h5 h
Every 6 hours1 h 30 min7 h 30 min
Every 8 hours2 h10 h
Every 12 hours2 h14 h
Daily2 h 24 min26 h 24 min
Every 2 days4 h 48 min2 d 4 h 48 min
Weekly16 h 48 min7 d 16 h 48 min
Monthly3 d33 d

A new backup counts from when it was created, so one that never reports is caught too. Changing the interval restarts the clock from the last report, or from now if there is none. Once a backup has five gaps between reports, its page suggests an interval from the median gap; check it against the longest normal gap before you accept it. Paused and archived clients are not checked, and resuming one restarts the clocks. Synology integrity-check reports never count as a backup run, so they do not reset the clock. Backups that shrink by half or more against their recent runs are flagged in Needs attention.

The detail is in Job statuses and missing reports and on the missing reports feature page. For why success reports matter in the first place, see A backup can fail by saying nothing; for what to do once a job is flagged, see Backup alerts technicians don't mute.

Find out what your backups are not telling you.

Start the free trial, point one backup's report at its client address, and see it turn Healthy, or Missing.