What is MTTR

What Is MTTR? The Four Meanings Behind One Acronym

What is MTTR? Ask three engineers and you may get three different answers, because the acronym covers four separate metrics. Two providers can quote the same MTTR and measure different things.

This guide defines the four meanings of MTTR, shows how each is calculated, and explains what the number means when the thing that failed is a physical server. It also shows why a quoted MTTR is an average, not a promise.

📖 Looking at MTTR as part of a provider’s uptime commitment?

Read Uptime SLA: What “99.9%” Actually Commits a Provider To for how recovery figures sit alongside uptime percentages in a contract.


What Is MTTR? The Short Answer

MTTR stands for four different metrics: mean time to repair, mean time to recovery, mean time to respond, and mean time to resolve. Each starts and stops the clock at a slightly different moment.

All four are averages. You add up the time spent across a set of incidents and divide by the number of incidents. What differs is what counts as time spent. According to Atlassian’s incident management guide, the four metrics overlap but each has its own meaning, so a team should confirm which one it tracks before it tracks anything.

How Do the Four Meanings Differ?

They differ in where the clock starts and stops, and in what each one is best used to judge.

The four meanings of MTTR

Meaning The clock usually runs from… to… Best used to judge
RepairStart of repair work to a tested, working systemHow efficient the repair work itself is
RecoveryThe failure to full service restoredThe whole outage, as users experience it
RespondAn alert to the first response (some guides count until fixed)How quickly the team reacts
ResolveDetection to a fix that stops the failure recurringHow well repeat incidents are prevented

Table listing the four meanings of MTTR, repair, recovery, respond, and resolve, with where each clock usually starts and stops and what each is best used to judge.

The definitions also drift within each meaning. For example, one guide counts “respond” from the first alert until the system is fully working again, while others stop the clock at the first qualified response. Even one vendor guide, Splunk’s, is titled “Mean Time to Repair” and then defines “Mean Time To Recover” in its first paragraph. As a result, the definition matters more than the number.

One incident, four MTTR numbers

Illustrative timeline. Where each clock starts and stops varies between organisations.

Recovery
failure to service restored
Resolve
detection to root cause closed
Respond
Repair
fix and test

Milestones, in order (not to scale): failure, detected, first response, service restored, root cause closed.

Timeline diagram of a single incident showing four bars: recovery running from failure to service restored, resolve from detection to root cause closed, respond as a short bar from detection to first response, and repair from the start of the fix until testing is complete.

Same incident, four numbers. In this illustration, the respond clock is short, recovery runs until service returns, and resolve keeps running until the root cause is closed.


How Do You Calculate MTTR?

MTTR is the total time across a set of incidents divided by the number of incidents. For recovery MTTR, the time is downtime.

Take three outages of 40, 95, and 25 minutes. Together they add up to 160 minutes, so the mean is 53 minutes.

A worked example (illustrative figures)

Incident Downtime
Outage 140 minutes
Outage 295 minutes
Outage 325 minutes
Mean (MTTR)160 ÷ 3 = 53 minutes
Median40 minutes

Table showing three outages of 40, 95, and 25 minutes, a mean of 53 minutes calculated as 160 divided by 3, and a median of 40 minutes.

However, the middle incident took only 40 minutes. One long outage pulls the mean upwards, which is why many teams track the median, or the slowest incidents, next to the mean.


What Is MTTR for a Dedicated Server?

On a dedicated server, MTTR usually comes down to two things: how quickly failed hardware is replaced, and how quickly you restore the software and data on it. The first depends mostly on the provider. The second depends mostly on you.

Parts availability matters more than most buyers expect. According to Splunk, MTTR also includes the lead time for parts that are not readily available, and that lead time can significantly extend the repair.

Who controls which part of the clock

Phase What decides how long it takes Typically controlled by
Detecting the faultMonitoring and alertingProvider for hardware and network, you for the application
Replacing failed hardwareSpare parts on site, technician availability, parts lead timeProvider
Restoring the systemBackups, runbooks, and a recently tested restoreYou, on an unmanaged server
Avoiding the outageRedundancy in the build, such as a RAID 1 mirrorThe configuration you choose

Table listing four phases of recovering a dedicated server, detecting the fault, replacing hardware, restoring the system, and avoiding the outage, with what decides the duration of each and who typically controls it.

Redundancy shows why the four meanings differ. When one drive in a RAID 1 mirror fails, the service keeps running, so recovery time is zero. Repair time is not zero: the array stays exposed until the replacement drive is in and rebuilt. Consequently, a single incident can score zero on one MTTR and several hours on another.

📖 What starts the clock in the first place

Read What Happens When a Server Crashes? for the causes of failure that start the MTTR clock.


Why Is a Provider’s MTTR Not a Guarantee?

A provider’s MTTR is an average of past incidents, not a promise about the next one. As Splunk’s guide explains, a vendor claiming a 24-hour MTTR means repairs usually take about a day, and individual incidents can take more or less. Individual incidents can take more or less. Splunk also notes that most SLAs mention MTTR in some form, so an MTTR in a contract is only as strong as the wording around it.

Therefore, ask four questions before trusting the figure:

  • Which MTTR is it? Repair, recovery, respond, or resolve.
  • When does the clock start? At the failure, at detection, or when a ticket opens.
  • Is it a mean or a median, and over what period? A single bad month can hide inside a yearly average.
  • Is it in the contract, and is there a written replacement window? A figure on a sales page commits the provider to nothing. A written window with a remedy for missing it does.

Our guide to uptime SLAs covers how these commitments sit alongside uptime percentages.


How Do You Reduce MTTR?

You reduce MTTR by shortening each phase separately: detect failures sooner, respond faster, and make the fix and the restore routine. Four habits do most of the work.

Monitoring and alerting. You cannot fix what you have not noticed. Splunk lists proactive monitoring and alerting first among the ways to lower MTTR.

Runbooks and tested restores. In practice, recovery time depends less on hardware and more on whether the team can find the runbook and has tested the restore. Setting a recovery target first gives the restore a number to beat, as RPO and RTO Explained describes.

Redundancy. A mirrored drive removes a whole class of outage from the count, because the service never stops.

Root cause reviews. After each incident, fix the cause, not only the symptom. This is what pulls the resolve number down over time

📖 Shortening the detection phase

Read Best Tools to Monitor Dedicated Server Performance for the monitoring stack that makes failures visible early.


How Does MTTR Relate to MTBF and Uptime?

MTBF measures how often a system fails, MTTR measures how long each failure lasts, and together they set availability. Atlassian gives the formula: availability equals MTBF divided by the sum of MTBF and MTTR. Use the recovery version of MTTR here, because it counts the whole outage.

For example, Atlassian gives about 15,000 hours as an MTBF that servers may have. That is roughly one failure every 20 months, or 0.6 failures a year. The table shows what different recovery times do to that same server.

Same failure rate, different recovery time (MTBF of 15,000 hours)

MTTR (recovery) Availability Expected downtime per year
24 hours (next-day replacement)99.840%14.0 hours
8 hours99.947%4.7 hours
4 hours99.973%2.3 hours
1 hour99.993%About 35 minutes

Table showing availability and expected yearly downtime for a server with an MTBF of 15,000 hours at four recovery times: 24 hours gives 99.840 percent, 8 hours gives 99.947 percent, 4 hours gives 99.973 percent, and 1 hour gives 99.993 percent.

The same failure rate produces very different results. A 24-hour replacement leaves this server below 99.9%, while an 8-hour replacement already clears it. The uptime table in our SLA guide shows what each of those percentages allows per year.

This model treats every failure as a full outage. In practice, redundancy such as a RAID 1 mirror removes part of those failures from the count, which is the reason to build it.

Recovery Starts With Full Control

Swify dedicated servers give you full root access to build the monitoring, runbooks, and backup routines that keep your own recovery time short, on enterprise HP hardware in an AMS-IX connected Netherlands data centre.

→ Explore Swify Dedicated Servers


Frequently Asked Questions

What is MTTR, and what does it stand for?

MTTR stands for mean time to repair, recovery, respond, or resolve. All four average the time spent across incidents, but they start and stop the clock at different moments, so the definition matters as much as the number.

Read Uptime SLA: What “99.9%” Actually Commits a Provider To for how MTTR fits alongside uptime commitments.


How do you calculate MTTR?

Divide the total time across incidents by the number of incidents. For recovery MTTR, that means total downtime divided by the number of outages: three outages of 40, 95, and 25 minutes add up to 160 minutes, or about 53 minutes per incident.


What is a good MTTR?

There is no universal good MTTR, because it depends on which MTTR is measured and how severe the incident is. Atlassian’s guide suggests that anything below an hour is great for IT and security teams, and under five hours for manufacturing systems. Replacing failed hardware is a different task from restarting a service, so the most useful comparison is your own trend, measured the same way each time.

Read Best Tools to Monitor Dedicated Server Performance for tracking it consistently.


What is the difference between MTTR and RTO?

RTO is the target and MTTR is the measurement. RTO is the longest downtime the business can accept, while MTTR is how long recoveries actually take on average. If your average recovery is longer than your RTO, the objective is not being met.

Read RPO and RTO Explained: Setting Recovery Objectives for how to set the target.


Is a provider’s MTTR the same as an SLA guarantee?

No. A quoted MTTR is a typical figure from past incidents, while an SLA guarantee is a contractual commitment with defined remedies. Splunk points out that a vendor claiming a 24-hour MTTR means repairs usually take about that long, not that every repair will.

Read Uptime SLA: What “99.9%” Actually Commits a Provider To for what a written commitment should contain.


How can I reduce MTTR on a dedicated server?

Shorten each phase: monitor so failures are detected quickly, keep tested runbooks and backups so restores are routine, and add redundancy such as RAID 1 so a single drive failure does not cause an outage. After each incident, fix the root cause so the same failure does not return.

Read Why Regular Backups Matter and How to Set Them Up for the restore side.