Incident response
MTTA vs MTTR: definitions, formulas and a worked example
Incident response · Updated 2026-09-24 · 11 min read · By the CallHeim team
MTTA (mean time to acknowledge) is the average time between an alert triggering an incident and someone acknowledging it. MTTR is the average time between trigger and resolution — but “MTTR” is ambiguous in the industry, because the R is used for resolve, recover, repair or respond depending on who you ask. Define which one you mean before you compare a number against anyone else’s.
MTTA: mean time to acknowledge
MTTA measures responsiveness: how long, on average, between an alert opening an incident and a human acknowledging it. It is the cleanest of the standard metrics because “acknowledge” is usually a single, unambiguous timestamped action in most incident tools — there’s little room to argue about when it happened.
MTTA = (sum of acknowledge time − trigger time, across all incidents in the period) / number of incidents
MTTR: mean time to resolve — and its four meanings
MTTR is where most confusion happens, because the acronym is reused for four different measurements:
- Mean time to resolve. Trigger to the incident being marked resolved in the tool.
- Mean time to recover. Trigger to the affected service actually being back to normal for users — which can happen before or after the incident is formally marked resolved.
- Mean time to repair. Specifically for hardware or infrastructure failures: trigger to the faulty component being physically or logically repaired.
- Mean time to respond. Sometimes used as a synonym for MTTA, sometimes to mean trigger to the first substantive action taken, which is a different, fuzzier timestamp than acknowledgment.
Most incident-management tools, CallHeim included, compute mean time to resolve — trigger to the incident’s own resolved timestamp — when they report “MTTR.” That is the definition used for the rest of this guide, and it’s worth confirming explicitly whenever you compare a number from a different tool or team.
MTTR = (sum of resolved time − trigger time, across all incidents in the period) / number of incidents
Worked example
Example. Five incidents over one week, with their trigger, acknowledge and resolve timestamps:
| Incident | Trigger → Acknowledge | Trigger → Resolve |
|---|---|---|
| #101 | 3 min | 42 min |
| #102 | 1 min | 18 min |
| #103 | 9 min | 35 min |
| #104 | 2 min | 190 min |
| #105 | 5 min | 27 min |
MTTA = (3 + 1 + 9 + 2 + 5) / 5 = 4 minutes. MTTR = (42 + 18 + 35 + 190 + 27) / 5 = 62.4 minutes. Notice how far the mean MTTR sits from four of the five incidents — that single 190-minute outlier (#104) pulls the average up well past what a typical incident actually looked like, which is exactly the distortion percentiles are for.
Worth noticing separately: incident #104 also had one of the fastest acknowledge times (2 minutes) but by far the slowest resolution. A fast MTTA and a slow MTTR on the same incident usually means the responder was reachable and responsive, but the fix itself was genuinely hard — more diagnosis time, more people needed, or a change that had to be made carefully rather than quickly. Looking at MTTA and MTTR together, incident by incident, tells you more than either average does on its own.
Mean vs p50 vs p90
A mean (average) treats every incident equally and is dragged hard by outliers — one very bad incident can make the whole period look worse than it was for everyone else, or one incident resolved unusually fast can flatter a mean that hides a real problem elsewhere. Percentiles describe the distribution instead of collapsing it to one number:
- p50 (median). The middle value — half of incidents resolved faster, half slower. In the example above, sorting resolve times (18, 27, 35, 42, 190) gives a p50 of 35 minutes: a much more representative “typical incident” figure than the 62.4-minute mean.
- p90. The value below which 90% of incidents fall — a read on your worst-case handling, without letting a single extreme outlier define the whole metric the way a mean does with a small sample.
A team that only tracks the mean can look stable while its p90 quietly gets worse — the mean stays anchored by a majority of fast, easy incidents while a growing tail of slow ones goes unnoticed. Track mean and at least one percentile together, not one or the other.
Pitfalls that distort the numbers
- Reopened incidents. If an incident is resolved, then reopens, does the clock restart, or does the original trigger time still apply? Decide once and apply it consistently — measuring each reopen as its own response, from reopen to the next resolve, usually gives a truer picture of ongoing effort than stretching the original duration across a gap where nobody was actively working it.
- Merged duplicates. If two incidents are later merged as duplicates of the same underlying problem — the kind of repeat deduplication should have caught upstream — decide whether the merged-away incident’s duration counts at all; including it usually double-counts the same outage.
- Still-open incidents. An incident open at the moment you run the report has no resolve time yet. Excluding it from the average silently removes your worst-case incidents from the number; including it with “time so far” understates it differently. Flag open incidents separately rather than quietly folding them into either average.
- Business hours vs. wall-clock time. A P3 incident that sits overnight, unattended by design, will show a long MTTR in wall-clock terms even though nobody was expected to touch it until morning. Decide whether low-severity incidents should be measured against business hours or wall-clock time, and be consistent — mixing the two across severities makes cross-severity comparisons meaningless.
- Auto-resolved noise. An alert that resolves itself (a flap that clears, a monitor that recovers before anyone looks) can register as an incident with a near-zero MTTR, dragging the average down and making response times look better than they are for incidents a human actually worked.
How to improve each metric
To improve MTTA: make sure the right person is paged the first time — a bad escalation policy or the wrong on-call rotation is the most common cause of slow acknowledgment — reduce alert volume so a real page isn’t buried in noise, and keep the acknowledge action itself to one clear step.
To improve MTTR: invest in runbooks so the responder isn’t starting diagnosis from zero, make sure escalation reaches someone with the right context quickly rather than cycling through people who have to hand off again, and review your p90 tail specifically — the incidents dragging it up usually share a root cause worth fixing once rather than five separate incident write-ups.
Note that neither metric measures how long it took to detect the problem in the first place — that is usually called MTTD (mean time to detect), and it depends entirely on your monitoring coverage, not your incident response process. Detection time is outside the scope of MTTA and MTTR, which both start the clock at the moment an alert already triggered an incident. A slow or missing detection step distorts both metrics before they even start — see how severity and re-triage decisions can also skew a period’s averages if an incident’s level is never revisited.
How CallHeim handles this
MTTA and MTTR (mean, p50, p90) are computed from the incident’s own timestamps; reopened incidents are measured per response, merged duplicates are excluded and still-open incidents are flagged.
Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.
MTTA and MTTR sit at the center of CallHeim’s incident management software, alongside on-call software for scheduling and escalation.
Key takeaways
- MTTA measures time to acknowledge; MTTR is ambiguous — confirm whether it means resolve, recover, repair or respond before comparing numbers across tools.
- A single outlier incident can distort a mean far more than a percentile — track p50 and p90 alongside the average, not instead of it.
- Reopened incidents, merged duplicates, still-open incidents and auto-resolved noise all need an explicit, consistent rule, or your MTTA/MTTR trend is comparing different things week to week.
- Improving MTTA is mostly an escalation-policy and noise problem; improving MTTR is mostly a runbook and context problem.
- Detection time (MTTD) is a separate concern from MTTA/MTTR — it measures monitoring coverage, not incident response.
Questions
Common questions, answered.
What is a good MTTA?
There is no universal "good" number — a reasonable MTTA depends on your escalation timeouts, alert volume and team size. Rather than chasing an external benchmark, track your own trend over time and treat a rising MTTA as a signal to investigate, whether that is noisy alerts burying the real ones or an escalation policy that is not reaching the right person first.
What does MTTR stand for?
It depends who you ask: resolve, recover, repair or respond. Most incident-management tools report mean time to resolve — trigger to the incident’s own resolved timestamp — so confirm which definition a number uses before comparing it across tools or teams.
Why does one incident distort my average MTTR so much?
A mean treats every incident equally, so one very slow outlier can drag the whole period’s average up well past what a typical incident looked like. Track a percentile — p50 or p90 — alongside the mean rather than instead of it; a team can look stable on the mean while its p90 quietly gets worse.
Is MTTA or MTTR more important to improve first?
It depends on which is trending worse. A slow MTTA usually points to an escalation-policy or noise problem; a slow MTTR usually points to a runbook or context problem. The two are independent — one incident can have a fast acknowledge time and a slow resolution if the fix itself was genuinely hard.
Does MTTA or MTTR include the time it took to detect the problem?
No. Detection time is not part of either metric — it is a separate measurement, MTTD (mean time to detect), that depends on monitoring coverage rather than incident response. Both MTTA and MTTR start the clock only once an alert has already triggered an incident.
Does CallHeim compute MTTA and MTTR automatically?
MTTA and MTTR (mean, p50, p90) are computed from the incident’s own timestamps; reopened incidents are measured per response, merged duplicates are excluded and still-open incidents are flagged.
Try it
See this in your own on-call rotation.
14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.