Mean time to recovery (MTTR) is the average time it takes to restore a system or service to normal operation after a failure. It is calculated by adding up the downtime from a set of incidents and dividing by the number of incidents. MTTR describes how quickly a team, provider or system bounces back, and together with how often failures happen, it decides the uptime users actually experience.
The letters MTTR are also used for mean time to repair, respond, resolve and remediate; this entry covers recovery and explains the others below.
At a glance
- MTTR measures the average duration of outages, from failure (or detection) to restored service.
- The acronym has several expansions, so confirm which one a vendor, report or contract uses.
- Where the clock starts and stops changes the number more than anything else.
- Fewer failures and faster recovery both raise uptime; MTTR tracks the second.
- Averages can hide long outliers, so look at the longest incidents as well as the mean.
What problem it solves
When something breaks, the cost to the business grows with every minute it stays broken. MTTR gives a team or a provider a way to measure that and improve it over time. It also helps buyers compare how well different services and teams handle failure, not just how often they fail.
For providers, MTTR turns a vague promise to “fix things quickly” into a number that can be reported and, in some contracts, committed to. For internal teams, it highlights where recovery is slow: late detection, unclear ownership, waiting on a carrier, missing spare parts or backups that take hours to restore.
How it works
The calculation. Take the total time systems were down across a period’s incidents and divide it by the number of incidents. Four outages lasting 20, 40, 60 and 120 minutes give an MTTR of 60 minutes.
Defining the clock. The start point might be the moment of failure, the moment monitoring detected it, or the moment a ticket was opened. The end point might be when service was restored (perhaps by a workaround or failover) or when the underlying fault was fully fixed. Two providers measuring different stretches can’t be compared directly.
The stages inside it. Recovery time usually includes detection, diagnosis, the fix or failover, and verification. Mean time to detect (MTTD) covers the first stage on its own, which is why the two are often reported together.
Reporting. A network operations center (NOC) or operations team tracks MTTR per service, per priority level and over time. Many teams also review the longest incidents individually, since a single long outage can matter more than a good average.
When it matters for buyers
- When reviewing a provider’s SLA. Many service level agreements (SLAs) commit to response times rather than restoration times. Know which one you are buying.
- After a major outage. Ask how long each stage took and what will shorten it next time.
- When choosing between redundancy and faster repair. A failover design can restore service in seconds even if the failed part takes a day to fix.
- When outsourcing monitoring. Our network operations center overview covers what outsourced monitoring and response typically include.
Questions to ask vendors
- Which MTTR do you report: recovery, repair, respond or resolve?
- When does the clock start and stop in your measurement?
- What was your MTTR for incidents like ours over the last year, and what was the longest outage?
- Do you commit to restoration times in the contract, or only to response times?
- What spare parts, failover capacity or field staff do you keep near our sites?
- How are major incidents reviewed, and do we receive the findings?
How it differs from MTTD and other MTTR metrics
Mean time to detect (MTTD) measures how long it takes to notice a problem; recovery time covers what comes after, and in many definitions includes detection as well. The other MTTR expansions measure different stretches. Mean time to repair often refers to fixing the failed component itself, which can take longer than restoring service by a workaround. Mean time to respond usually means the time until someone starts working the issue, and in security it often describes the time to contain a threat. Mean time to resolve typically runs until the root cause is fixed and the incident closed, and mean time to remediate is common in vulnerability and security reporting. When a vendor quotes MTTR, ask for the definition before comparing numbers, and for security incidents see also incident response (IR).
