Incident response

Stage in the incident loop: RESPOND · RESOLVE · LEARN

Incident response workflow: run the incident, then measure what happened.

An incident management workflow with four states instead of a dozen, a timeline neither side of an argument can quietly edit, and response-time metrics computed from the incident's own clock rather than a dashboard's.

How it works

Four states, and no invented fifth.

Four incident states (Triggered, Acknowledged, Resolved, Closed) and five severities (P1 to P5). A resolved incident can be reopened; a closed incident is final.

  1. 01Triggered

    Created by the alert pipeline, or declared by a person. Escalation starts here.

  2. 02Acknowledged

    Somebody has it. Every remaining escalation tier and timer stops.

  3. 03Resolved

    The problem is over. It can still be reopened, back to Triggered.

  4. 04Closed

    Final. There is no fifth state, and no way back from here.

Severities are P1 through P5. An incident opened by an alert takes the source's own severity; CallHeim adds a suggested severity for a person to apply — it never decides one on its own.

Incidents opened by an alert get a sequential per-workspace number, shown as #2041 (an example, not a live number).

Roles

Role-gated, not a free-for-all.

Incident actions are role-gated: managers and stakeholders are read-only.

Responders act on the incident; managers and stakeholders can watch, comment and read the timeline without being able to change its state.

Timeline & audit

Two different records, both append-only.

Lifecycle changes (acknowledge, resolve, close), comments, escalation steps and suppressions are written to the incident timeline as they happen, and the timeline has no edit or delete in the product.

The audit log records incident acknowledge and resolve, user invitations and deactivations, role changes, forced sign-out and workspace settings changes. Configuration changes to schedules, escalation policies, services, teams and integrations are recorded with before and after values.

Keeping people outside the incident informed

A status page your team writes by hand.

A workspace’s public status page lives at app.callheim.com/status/your-slug and your team updates it by hand.

It is a hand-updated notice board, not a monitor: nothing on it changes because a system detected something. If you post it, it is because someone on your team decided to post it.

Learn

Then measure what happened, on every plan.

Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.

MTTA and MTTR (mean, p50, p90) are computed from the incident’s own timestamps; reopened incidents are measured per response, merged duplicates are excluded and still-open incidents are flagged.

  • MTTA and MTTR as mean, p50 and p90 — not just an average that one bad night ruins.
  • Alert volume by source, so you can see which integration is generating the noise.
  • Escalation rate: the share of incidents that went through escalation.
  • Period-over-period comparison, over windows of 7 to 365 days.

Analytics start when the alert reaches CallHeim. Detection time belongs to your monitoring, which noticed the problem first, so CallHeim does not report an MTTD figure.

Example incident #2041

Triggered → Acknowledged
2m 14s

MTTA

Triggered → Resolved
29m 32s

MTTR

  1. 02:14:06Incident created from 12 alerts · P1 · Triggered
  2. 02:14:08Paged the on-call responder by e-mail
  3. 02:16:20Acknowledged
  4. 02:20:48Comment · connection pool at its limit
  5. 02:43:38Resolved

Example scenario, not a customer incident.

CallHeim

Put a rule you can read between your alerts and your on-call.

CallHeim helps teams stay in control when critical systems are not. Explore the platform, connect one source, and send yourself a page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required