Incident response
Incident response process: the lifecycle, roles and postmortems
Incident response · Updated 2026-09-24 · 11 min read · By the CallHeim team
An incident response process is the repeatable sequence a team follows from the moment something breaks to the moment the team has learned from it: detect, triage, acknowledge, mobilise, mitigate, resolve, close and learn. The value of writing it down isn’t the steps themselves — most teams could recite them — it’s that everyone follows the same sequence under pressure instead of improvising a slightly different one each time.
The lifecycle
- Detect / alert. Monitoring or a user report surfaces a problem, and an alert reaches the team through an alert source. Detection quality — how fast a real problem produces an alert — is a separate concern from everything that follows; a slow or missing alert here delays every later step regardless of how good the rest of the process is.
- Triage. Someone assesses what the alert actually means: is it real, how big is the impact, and what severity level does it warrant. Triage is also where duplicate or related alerts get connected to a single incident rather than tracked separately.
- Acknowledge. The on-call responder named by the escalation policy confirms they’re on it. This is the timestamp MTTA is measured against, and it stops escalation from continuing to page further tiers.
- Mobilise. For anything beyond a quick, solo fix, the responder pulls in whoever else is needed — a second engineer with relevant context, a communications owner, sometimes a formally named incident commander for the largest incidents. Mobilising early, even before the full scope is understood, is usually cheaper than mobilising late.
- Mitigate. The immediate goal is reducing user impact, which is not always the same as fixing the underlying cause — rolling back a bad deploy, failing over to a backup, or disabling a broken feature flag can mitigate an incident in minutes while the real fix takes hours.
- Resolve. The underlying problem is actually fixed, not just worked around, and the team confirms the affected system is genuinely back to normal rather than assuming it from the absence of new alerts.
- Close. The incident record is finalised — severity confirmed, timeline complete — and becomes the durable account of what happened, including for anyone who wasn’t in the room.
- Learn. The team reviews what happened and what should change, typically through a postmortem, so the same failure is less likely to repeat in the same form.
Not every incident needs every step to take long. A P4 that resolves itself in two minutes still passes through the same lifecycle — triage and acknowledge just happen quickly, and mobilise might be skipped entirely because one person is enough.
Example incident timeline
Example. A P2 incident, worked by one responder before a second engineer is pulled in — illustrative only, not a benchmark for how fast your own team should move.
| Stage | Elapsed (example) | What happens |
|---|---|---|
| Detect / alert | 0:00 | Monitoring fires; the alert reaches the on-call rotation. |
| Triage | 0:02 | Responder confirms a real, user-facing degradation and sets severity to P2. |
| Acknowledge | 0:03 | Responder acknowledges; escalation stops advancing to the next tier. |
| Mobilise | 0:08 | A second engineer with relevant context joins. |
| Mitigate | 0:20 | A feature flag is disabled, reducing user impact. |
| Resolve | 0:55 | Underlying cause fixed; the team confirms the service is back to normal. |
| Close | 1:10 | Incident record finalised, with the timeline and severity confirmed. |
“Learn” (the postmortem) happens after close and isn’t on this clock — see how MTTA and MTTR are measured from these same timestamps.
Roles
For anything beyond a small incident, naming roles explicitly — even informally — prevents the two most common coordination failures: nobody driving the response, and three different people telling stakeholders three different stories.
- Incident commander. Owns coordination, not necessarily the technical fix: keeps track of who’s doing what, makes the call on mitigation trade-offs, and decides when to pull in more help. On a small incident, the acknowledging responder plays this role by default.
- Communications owner. Keeps stakeholders — support, leadership, sometimes customers — updated on a predictable cadence, freeing the people actively working the problem from having to answer the same status question repeatedly.
- Subject-matter experts. Brought in for the specific system or component involved. Their job is diagnosis and mitigation, not coordination — mixing the two roles in one person works fine for a small incident and becomes a bottleneck on a large one.
- Scribe. On larger incidents, a dedicated notetaker keeps the timeline accurate in real time — who tried what, when, and what happened — so the incident commander and responders can focus on the problem instead of also trying to remember the sequence of events for the postmortem afterwards.
Roles are assigned based on the incident’s severity, not fixed job titles — the same engineer might be incident commander on one incident and a subject-matter expert on the next, and a P4 rarely needs more than one person wearing every hat at once.
Communication and status updates
Two audiences need different things, and conflating them produces updates that satisfy neither:
- Internal responders need frequent, technical, unfiltered updates — what’s been tried, what the current theory is, what’s next. A shared, running record of the incident that everyone actively working it can see is usually enough; the format matters less than having exactly one of it, so people aren’t reconstructing the timeline from memory afterwards.
- External or leadership audiences need less frequent, plain-language updates focused on impact and expected timing, not internal diagnostic detail. A public status page is a common way to serve this audience without pulling responders into repeated status calls.
A predictable cadence — an update every 30 minutes for a major incident, for example — matters more than the exact interval: stakeholders checking in less often when they know an update is coming on schedule is itself a form of reduced noise for the responding team. The right cadence also scales with severity: a P1 might warrant an update every 15–30 minutes, a P2 every hour, and a P3 a single update at resolution rather than a running commentary nobody is waiting on.
Postmortems
A postmortem (or retrospective) is where the “learn” step happens: a written account of what happened, why, and what the team is going to change as a result. Two practices make the difference between a postmortem that improves things and one that’s just paperwork:
- Blameless framing. The goal is understanding why the failure made sense given what people knew at the time — what the system, the alerts and the process allowed to happen — not identifying who to blame. A postmortem people are afraid to be honest in produces a worse account of what actually happened, which makes it a worse tool for preventing a repeat.
- Concrete, owned actions. A postmortem that ends in general intentions (“we should be more careful”) rarely changes anything. One that ends in specific, assigned, dated follow-up items — a monitor to add, a runbook to write, a dependency to remove — has something to check on later.
Postmortems are general incident-management practice, independent of any specific tool — the discipline is in how a team runs the conversation and tracks the resulting actions, not in the document format.
Runbook checklist
- Alert reaches the right on-call rotation and gets acknowledged within your target window, without contributing to alert fatigue.
- Severity is assigned (or confirmed) at triage and re-checked if the picture changes.
- A single running record of the incident exists that every active responder can see.
- For anything beyond a small incident, a coordinator role is explicitly named, even informally.
- Stakeholder updates go out on a predictable cadence, separate from the technical working notes.
- Mitigation is distinguished from resolution — reducing impact now is not the same as fixing the cause.
- The incident is formally closed, not just left to go quiet.
- A blameless postmortem happens for anything significant, with owned, dated follow-up actions.
How CallHeim handles this
Four incident states (Triggered, Acknowledged, Resolved, Closed) and five severities (P1 to P5). A resolved incident can be reopened; a closed incident is final.
Lifecycle changes (acknowledge, resolve, close), comments, escalation steps and suppressions are written to the incident timeline as they happen, and the timeline has no edit or delete in the product.
Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.
A workspace’s public status page lives at app.callheim.com/status/your-slug and your team updates it by hand.
The audit log records incident acknowledge and resolve, user invitations and deactivations, role changes, forced sign-out and workspace settings changes. Configuration changes to schedules, escalation policies, services, teams and integrations are recorded with before and after values.
This lifecycle is the core of CallHeim’s incident management software, built on the same on-call software foundation as scheduling and escalation.
Key takeaways
- The lifecycle — detect, triage, acknowledge, mobilise, mitigate, resolve, close, learn — applies to every incident; small incidents just move through it fast.
- Naming a coordination role separately from the technical fixer prevents the two most common coordination failures on larger incidents.
- Internal responders and external stakeholders need different kinds of updates, at different frequencies.
- Mitigating an incident (reducing impact) and resolving it (fixing the cause) are different steps — don’t confuse a mitigated incident with a closed one.
- A blameless postmortem with concrete, owned actions is what turns an incident into a process improvement rather than just a recovered outage.
Questions
Common questions, answered.
What are the stages of incident response?
Detect, triage, acknowledge, mobilise, mitigate, resolve, close and learn. Not every incident needs every stage to take long — a minor incident still passes through the same lifecycle, just quickly, and mobilise is often skipped when one person is enough.
What is the difference between mitigating and resolving an incident?
Mitigating reduces user impact — rolling back a deploy or disabling a feature flag — without necessarily fixing the underlying cause. Resolving means the underlying problem is actually fixed and the team has confirmed the system is genuinely back to normal, not just quiet.
Who should be the incident commander?
Whoever is best placed to coordinate, not necessarily the most senior person or the one fixing the issue — the role owns coordination, not the technical fix. On a small incident, the acknowledging responder plays this role by default.
How often should status updates go out during an incident?
A predictable cadence matters more than the exact interval, and it should scale with severity: a major incident might warrant an update every 15–30 minutes, a moderate one every hour, and a minor one a single update at resolution.
What makes a postmortem blameless?
Framing it around why the failure made sense given what people knew at the time, rather than who to blame — and ending it with concrete, owned, dated follow-up actions instead of general intentions like "we should be more careful".
Does CallHeim provide a public status page for incidents?
A workspace’s public status page lives at app.callheim.com/status/your-slug and your team updates it by hand.
Try it
See this in your own on-call rotation.
14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.