Mean time between failures (MTBF) is a reliability measure: the average amount of operating time that passes between one failure and the next for a repairable system or component, such as a router, a server or a whole service. It is a statistical figure for how often things break across a population or over a period, not a promise of how long a single device will run. Paired with mean time to recovery (MTTR), which measures how long a fix takes, it describes how much downtime you can expect.
At a glance
- MTBF answers “how often does this fail?”; MTTR answers “how long until it works again?”
- It is calculated as operating time divided by the number of failures in that time.
- Vendor MTBF figures are statistical estimates under stated conditions, not a guaranteed lifetime for any one unit.
- It is most useful for comparing equipment, planning spares and deciding where redundancy pays off.
- Your own field MTBF, measured from tickets and incident logs, often tells you more than a datasheet.
What problem it solves
When equipment or a service keeps going down, the first question is whether that is normal. MTBF gives you a way to put a number on reliability so you can compare one switch model with another, one site with another, or this year with last year. Without it, decisions about refreshing hardware, stocking spares or paying for a redundant link rest on anecdotes and whichever outage people remember most.
It also helps you read vendor claims. Datasheets often print very large MTBF figures. Knowing what the number actually describes stops a buyer from treating “300,000 hours” as “this box will run for 34 years”, and points the conversation toward warranty, support and replacement terms, which are what the vendor actually commits to.
How it works
The calculation. Take the total time a system or a group of identical units was operating and divide by the number of failures in that time. Ten units running for 1,000 hours each with two failures between them gives 10,000 operating hours over two failures, an MTBF of 5,000 hours. The answer depends on what counts as a failure, whether partial degradation is included, and whether downtime is excluded from operating time, so definitions matter.
Where vendor figures come from. Manufacturers rarely have decades of field data for a new product. Published MTBF figures are often predictions from component reliability models or accelerated testing at stated temperatures and duty cycles. They are reasonable for comparing similar products tested the same way, and much less reliable as a forecast for your environment. Heat, dust, power quality, firmware and how hard the device is worked all change real results.
Fleet, not unit. MTBF describes a population during its useful life. Failure rates tend to be higher early (manufacturing defects) and late (wear-out), and a single MTBF figure usually assumes the flat middle. That is why a device with a very high MTBF can still fail in its first month, and why MTBF says little about when equipment reaches end of life.
MTBF and availability. A common approximation is that availability equals MTBF divided by MTBF plus MTTR. A system that fails once every 1,000 hours and takes 1 hour to restore is up roughly 99.9% of the time. You can improve that by failing less often, by recovering faster, or by designing so that one failure doesn’t take the service down at all.
When it matters for buyers
- When choosing hardware. Comparing MTBF across similar products, tested under similar conditions, is one input alongside warranty, support and total cost.
- When planning spares and field support. Expected failure rates across a fleet of sites tell you how many spare units to stock and how much on-site support you need.
- When deciding on redundancy. If a single point of failure fails rarely but takes days to replace, redundancy such as N+1 may be worth more than a better MTBF.
- When a provider reports reliability. A managed service or network operations center may report MTBF per site or device class; ask how failures are counted.
- When a “reliable” product keeps failing. Your own measured MTBF is evidence for an RMA, a refresh or a vendor conversation.
Questions to ask vendors
- How was this MTBF figure derived: field data, accelerated testing or a prediction model?
- What operating conditions (temperature, duty cycle, power) does it assume?
- Does it cover the whole system or a single component?
- What counts as a failure in your figure, and in your reporting to us?
- What does the warranty and advance replacement commitment actually say, and how fast are spares delivered?
- Can you report failure rates for our installed base, not just the datasheet number?
For multi-site estates, our field support overview covers how providers stock spares and dispatch technicians when hardware fails.
How it differs from MTTR
MTBF and mean time to recovery (MTTR) measure opposite halves of the same cycle. MTBF is the average time from one failure to the next: it is about how often things break. MTTR is the average time from a failure to restored service: it is about how quickly you recover. A product can have an excellent MTBF and still cause long outages if parts take days to arrive, and a fragile product can be tolerable if failover restores service in seconds. Looking at both, along with mean time to detect (MTTD), gives a far better picture of uptime than either number alone, and high availability designs aim to make the effect of each failure small whatever the MTBF.
