What Is MTTR (Mean Time to Recovery)?

Also called: Mean time to restore

Related problems: Outages last too long once they happen; Provider reports MTTR but we don't know what it measures; Need a metric for how fast our team or provider fixes things; Leadership wants to know how quickly we bounce back from incidents

Mean time to recovery (MTTR) is the average time it takes to restore a system or service to normal operation after a failure. It is calculated by adding up the downtime from a set of incidents and dividing by the number of incidents. MTTR describes how quickly a team, provider or system bounces back, and together with how often failures happen, it decides the uptime users actually experience.

The letters MTTR are also used for mean time to repair, respond, resolve and remediate; this entry covers recovery and explains the others below.

At a glance

  • MTTR measures the average duration of outages, from failure (or detection) to restored service.
  • The acronym has several expansions, so confirm which one a vendor, report or contract uses.
  • Where the clock starts and stops changes the number more than anything else.
  • Fewer failures and faster recovery both raise uptime; MTTR tracks the second.
  • Averages can hide long outliers, so look at the longest incidents as well as the mean.

What problem it solves

When something breaks, the cost to the business grows with every minute it stays broken. MTTR gives a team or a provider a way to measure that and improve it over time. It also helps buyers compare how well different services and teams handle failure, not just how often they fail.

For providers, MTTR turns a vague promise to “fix things quickly” into a number that can be reported and, in some contracts, committed to. For internal teams, it highlights where recovery is slow: late detection, unclear ownership, waiting on a carrier, missing spare parts or backups that take hours to restore.

How it works

The calculation. Take the total time systems were down across a period’s incidents and divide it by the number of incidents. Four outages lasting 20, 40, 60 and 120 minutes give an MTTR of 60 minutes.

Defining the clock. The start point might be the moment of failure, the moment monitoring detected it, or the moment a ticket was opened. The end point might be when service was restored (perhaps by a workaround or failover) or when the underlying fault was fully fixed. Two providers measuring different stretches can’t be compared directly.

The stages inside it. Recovery time usually includes detection, diagnosis, the fix or failover, and verification. Mean time to detect (MTTD) covers the first stage on its own, which is why the two are often reported together.

Reporting. A network operations center (NOC) or operations team tracks MTTR per service, per priority level and over time. Many teams also review the longest incidents individually, since a single long outage can matter more than a good average.

When it matters for buyers

  • When reviewing a provider’s SLA. Many service level agreements (SLAs) commit to response times rather than restoration times. Know which one you are buying.
  • After a major outage. Ask how long each stage took and what will shorten it next time.
  • When choosing between redundancy and faster repair. A failover design can restore service in seconds even if the failed part takes a day to fix.
  • When outsourcing monitoring. Our network operations center overview covers what outsourced monitoring and response typically include.

Questions to ask vendors

  • Which MTTR do you report: recovery, repair, respond or resolve?
  • When does the clock start and stop in your measurement?
  • What was your MTTR for incidents like ours over the last year, and what was the longest outage?
  • Do you commit to restoration times in the contract, or only to response times?
  • What spare parts, failover capacity or field staff do you keep near our sites?
  • How are major incidents reviewed, and do we receive the findings?

How it differs from MTTD and other MTTR metrics

Mean time to detect (MTTD) measures how long it takes to notice a problem; recovery time covers what comes after, and in many definitions includes detection as well. The other MTTR expansions measure different stretches. Mean time to repair often refers to fixing the failed component itself, which can take longer than restoring service by a workaround. Mean time to respond usually means the time until someone starts working the issue, and in security it often describes the time to contain a threat. Mean time to resolve typically runs until the root cause is fixed and the incident closed, and mean time to remediate is common in vulnerability and security reporting. When a vendor quotes MTTR, ask for the definition before comparing numbers, and for security incidents see also incident response (IR).

Frequently Asked Questions

What does MTTR stand for?
It depends who is using it. In operations it usually means mean time to recovery or restore. It is also used for mean time to repair, to respond, to resolve and, in security, to remediate. Each measures a different stretch of an incident, so ask which one a report or contract means.
How is MTTR calculated?
Add up the downtime of all incidents in a period and divide by the number of incidents. The result depends heavily on when the clock starts (failure, detection or ticket creation) and when it stops (service restored or fully resolved).
What is a good MTTR?
There is no universal number. It depends on the system, how critical it is and how recovery is designed. Track your own trend and set targets per service based on what downtime costs the business.
How is MTTR different from an SLA response time?
A response time commits to how quickly a provider acknowledges or starts work on an issue. MTTR measures how long it actually took to restore service. A provider can meet a 15-minute response target and still take hours to recover.
How do you reduce MTTR?
Common levers are faster detection and alerting, clear runbooks and escalation paths, spare parts or failover capacity, tested backups and well-rehearsed recovery. Redundant design often helps the most because service can come back before the failed part is repaired.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.