Skip to main content

Open Backup Failure Tracker

Running record of clients whose backups are currently failing, how long they have been failing, and which leg (local or cloud) is broken. This is an operational log, not a standard — it is expected to change daily. For the policy this measures against, see Backup & Data Protection Standards.

Why this page exists. Time-since-last-successful-backup is the only backup health signal DTC monitors routinely. If nobody writes down when the clock started, "it's been failing a while" is the best answer anyone can give — and that is not an answer you can take to a client, an insurer, or an auditor. Record the date of the last known-good backup, not the number of days. Days rot; dates don't.


Current Open Failures

As of 2026-08-31.

Client / Scope What's failing Last known-good Days down Owner Status
Bunin, Dr. Kevin — Burke Backup failing (leg not yet confirmed) 2026-08-14 17 Longest-running failure in the fleet. Needs diagnosis before anything else on this list.
MoCo Cloud leg only. Local leg is succeeding. Local: current. Cloud: TBD ~7 Root cause is site bandwidth, not backup config. See MoCo — bandwidth-bound below.
Remaining fleet — servers and pan/imaging workstations on the backup list Failing across the board ~2026-08-24 ~7 Roster not yet enumerated. A fleet-wide ~7-day break starting around the same date suggests a common cause, not 20 independent faults.

Two things to nail down before this table is trustworthy:

  1. The roster. "Everyone else in the backup list" needs to become actual device and client names, each with its own last-known-good date. A single "~7 days" row hides the outliers.
  2. The Bunin arithmetic. 2026-08-14 to 2026-08-31 is 17 days, not 14. Either the last-good date is 2026-08-17, or the 14-day figure was read off a dashboard a few days ago and has since drifted. Confirm which, because the number that matters for a client conversation is the date.

MoCo — bandwidth-bound, not broken

MoCo's local backup leg is succeeding; only the cloud leg fails. That is a network capacity problem wearing a backup problem's clothes, and it does not get fixed in the backup console. Do this first, regardless of which option is chosen. Get MoCo's line-of-business database extracted to a logical dump and shipped offsite nightly. A compressed dump — and certainly a stream of transaction log records — fits on MoCo's circuit no matter how bad it is, because the constraint has only ever been image-sized transfers. Right now MoCo has no offsite copy of the one thing that is irreplaceable, and that is fixable this week without waiting on a carrier or a standards decision. See Backup Methodology & Direction for the mechanics.

Then, for the image and the rest of the data, there are two ways out and they are not equivalent:

Option A — fix the internet. Get MoCo enough reliable upstream bandwidth to push a nightly image to cloud inside the overnight window. Keeps them on the standard hybrid image model, keeps the tier definitions clean, and stops this recurring. Costs the client money every month and depends on what the local carriers will actually sell them.

Option B — split the legs. Local image stays local (fast restore, bare-metal capability preserved). Cloud stops receiving the full image and instead receives the logical database extract plus file- and object-level data. Far smaller nightly delta, survives a thin pipe, and the offsite copy becomes the data you would actually restore rather than a container you have to mount first. Costs less, but it means MoCo no longer has an offsite bare-metal image, and that has to be a stated, accepted change to their recovery expectations — written down and agreed to, not quietly implemented.

Option B is the direction the broader model is heading anyway, and MoCo is the natural pilot for it — bandwidth-bound, healthy local leg, failing cloud leg. Deciding it for MoCo alone, ahead of a standards change, is a per-client exception; deciding it as the pattern is a standards change. Either way, the database extract above is not part of that decision — it happens now.


Escalation Thresholds

Time since last good backup Action
24–48 hours Normal triage. Work the platform-specific error runbook.
48 hours+ Follow Hybrid Cloud Backup No Success In 48 Hours. Ticket must exist and be assigned.
7 days+ Client is outside every tier's RPO. Notify the account manager. The client needs to be told.
14 days+ Treat as a service failure, not a ticket. Escalate to T3 and management; assume the fix is not going to be a console setting.

The 7-day line matters because Tier 3 (HCB-MSP) carries a 1-day RPO. Seven days of failure is not a degraded backup — it is no backup, for a week, at a client who is paying for one.


How to Maintain This Page

  • Add a row when a client passes 48 hours without a successful backup.
  • Record the date of the last known-good backup and let the day count be derived. Update the "As of" date when you touch the table.
  • Note which leg fails. Local-only and cloud-only failures have different causes and different urgency; a hybrid client with a healthy local leg is in far better shape than the row length suggests.
  • Remove a row only after a successful backup, not after a change that should have fixed it. Restarting a service is not a resolution; a green job is.
  • When a client comes off this list, note the root cause in the ticket. Repeated appearances by the same site are the signal worth acting on.