Incident response

Incident management software: what it is and how to evaluate it

Incident response · Updated 2026-09-24 · 12 min read · By the CallHeim team

Incident management software is the layer that sits between something going wrong and someone fixing it: it takes a signal from your monitoring stack, works out who should be told, gets that person paged in a way they will actually notice, and keeps a record of what happened while the team responds. This guide covers what the category actually does, the twelve capabilities worth evaluating before you buy, how the main vendor categories differ, and where CallHeim fits.

What incident management software does

Monitoring tools detect that something is wrong. Incident management software is the layer on top of that: it decides who needs to know, delivers the notification in a way that actually reaches them, tracks the response from first alert to resolution, and gives the team a record to learn from afterward. The two are often confused because some monitoring platforms bolt on basic alerting, but a dedicated incident management tool is built around the part monitoring tools treat as an afterthought — getting a human, reliably, even at 3 a.m., and coordinating what happens next.

The category splits into two overlapping jobs, and vendors bundle them differently. On-call software is the narrower piece — schedules, escalation and paging — covered in depth in our on-call software guide. Incident management is the wider job this guide covers: on-call plus the alert-handling pipeline in front of it (ingest, deduplication, routing) and the incident-lifecycle tooling behind it (states, timelines, status pages, analytics). A small team sometimes only needs the narrower piece; a team running several services with different owners usually needs the whole pipeline.

The through-line, whichever slice you need: the software’s only job is to turn a signal into a paged human and a record of what happened next, reliably enough that nobody has to wonder whether it worked.

12 capabilities to evaluate

Every vendor’s marketing page lists most of these. The list below is what to actually check in a trial or a demo — and the trap each one hides, based on how these evaluations typically go wrong.

Twelve capabilities to evaluate in incident management software
CapabilityWhat to checkCommon trap
Alert ingestDoes it accept your monitoring tools’ own payload over HTTPS, or does every source need a hand-built webhook mapping?A demo integration hides how much field mapping you’ll actually do for your own stack.
Deduplication & groupingDo repeats of the same alert collapse onto one incident automatically? Is grouping of related-but-different alerts rule-based, with a visible reason, or a model nobody can explain?Grouping you can’t explain is grouping you can’t trust the one time it groups the wrong two alerts together.
RoutingDoes an alert reach the right service or team without a growing if/else ruleset someone has to maintain by hand?Routing that only works because one person remembers every exception.
On-call schedulingLayers, overrides, time-off, shift swaps, follow-the-sun templates, and coverage-gap detection — in the time zones your team actually works in.A schedule builder that only demos well with three people in one time zone.
EscalationTiers, per-tier timeouts, and a mandatory fallback that is paged once every earlier tier has timed out.No fallback tier — the policy escalates to the end of its list and simply stops.
Notification channelsWhich channels are proven to reach a phone in production, not just listed on a features page.A channel that is “supported” but has never actually paged a real person.
Incident lifecycleClear states (triggered, acknowledged, resolved, closed), role-gated actions, and a timeline with no edit or delete after the fact.A “resolved” that quietly means three different things depending on who clicked it.
Status communicationA public status page your team controls, and clarity on who inside your org is allowed to post to it.No status page at all, so customer updates happen in a scramble of e-mails.
AnalyticsMTTA and MTTR computed from the incident’s own timestamps, and whether the analytics you need sit on every plan or only the expensive one.A dashboard that looks rich in the demo but is gated behind a tier you were not shown the price of.
IntegrationsA tested, maintained parser for your specific monitoring tools — not just a generic inbound webhook you map yourself.A large advertised integrations count that turns out to be mostly generic webhook templates rather than real, vendor-tested parsers.
SecurityTenant isolation on reads and writes, encryption at rest and in transit, role-based permissions, and an append-only audit log.Compliance language on the marketing site that the security page quietly contradicts.
Pricing modelA flat per-seat price, versus a base seat plus separately priced add-ons for the capabilities you actually need on day one.A cheap-looking seat price that becomes the expensive plan once grouping, analytics or status pages are added back in.

Deduplication and grouping, escalation, and on-call scheduling each have a full guide with worked examples: deduplication, escalation policy and on-call rotation.

How the vendor categories differ

Competitor information last verified September 2026. Sources are linked beside each fact.

Vendors in this space roughly split into four shapes, and knowing which shape you’re looking at explains most pricing and feature-gating surprises before you hit them.

Dedicated on-call and paging platforms sell the paging pipeline as the whole product, with per-seat pricing and, often, extra capability sold as add-ons on top. PagerDuty is the largest example: published tiers run from a free plan through a custom Enterprise tier, Per-seat pricing: Free $0; Professional $25/mo, or $21/mo on an annual plan; Business $49/mo, or $41/mo on an annual plan; Enterprise custom. (PagerDuty pricing page, checked 2026-09-24) and even its documented, mostly rule-based alert-grouping methods require a separately priced add-on regardless of tier. Four documented alert-grouping methods (Intelligent, Alert Content, Intelligent+Content, Time-only), three of them rule-based and configurable; grouping of any kind requires the AIOps add-on. (PagerDuty alert grouping documentation, checked 2026-09-24)

Broader incident-response platforms sell on-call as one module inside a larger product that also covers retrospectives, status communications and stakeholder updates — sometimes bundled into one seat, sometimes priced as a separate add-on on top of a base seat. Basic is free. Team is $19/mo, or $15/mo on an annual plan, per user, with on-call as a $10 add-on; Pro is $25/user with on-call at $20; on-call alone is $20/user; Enterprise is custom with a 99.99% uptime SLA. (incident.io pricing page, checked 2026-09-24) Rootly takes the add-on approach further still, selling Incident Response and On-Call as fully separate products with their own prices. Rootly prices Incident Response and On-Call as separate products, each Essentials $20/user/mo with an Enterprise "contact us" tier; a startup discount of up to 50% is offered to companies under 100 employees, under $50m raised and under 5 years old. (Rootly pricing page, checked 2026-09-24)

ITSM and helpdesk suites fold alerting into a much larger service-management product, which suits a team already committed to that ecosystem more than one evaluating on-call software on its own merits. Atlassian’s own transition away from its dedicated Opsgenie product and toward Jira Service Management is the clearest current example of this shape — Opsgenie stopped taking new customers on 4 June 2025 Opsgenie ended new sales on 4 June 2025 (announced 4 March 2025): no new purchases, sign-ups or plan changes. (Atlassian Opsgenie licensing page, checked 2026-09-24) with Atlassian naming JSM as the successor. Atlassian’s licensing and migration pages name Jira Service Management as the successor; the Opsgenie pricing page still (as checked) says "Jira Service Management or Compass". (Atlassian Opsgenie licensing page, checked 2026-09-24) (Full timeline in our Opsgenie end-of-support guide and Opsgenie alternatives guide.)

Smaller and newer entrants compete mainly on a simpler product and a lower or flatter price, trading breadth for a sharper focus. Grafana Cloud IRM, the paid successor to the now-archived Grafana OnCall open-source project, is a current example of this shape. Above the free tier, Grafana Cloud IRM Pro is $20 per active IRM user plus a $19/mo platform fee that covers the first 3 users; Enterprise is custom with a $25,000 minimum annual commitment. (Grafana pricing page, checked 2026-09-24) CallHeim, described below, is this fourth shape: newer, smaller, and explicit about what it trades away.

No category is inherently the right answer. Breadth costs money and configuration; a narrower tool costs you the features it hasn’t built yet. The PagerDuty alternatives guide and Opsgenie alternatives guide go deeper on named vendors in each shape, with sourced starting prices.

Build vs buy

The paging “happy path” — send a message when something breaks — is not the hard part to build. The hard part is everything that has to work the one time it matters most: a durable queue that survives a spike of alerts instead of dropping them, an escalation timer that reliably advances a missed tier on schedule, a delivery channel actually proven to reach a phone rather than merely implemented, and an audit trail a security review will accept. Building that well is a multi-quarter undertaking most teams don’t want to be in, because it isn’t the product they’re trying to ship.

Building makes sense when your routing or grouping logic is genuinely unusual to your business and a generic rules engine can’t express it, or when the team is small enough that a lightweight internal tool covers the one workflow you actually have. Buying makes sense for almost everyone else: maintaining paging reliability is a distraction from whatever your team is actually paid to build, and a vendor whose whole product is the paging pipeline has more reason to keep it working than a side project does.

It is rarely all-or-nothing. Plenty of teams build the parts specific to their own monitoring stack — a custom alert normaliser, a dashboard tailored to one service — and buy the parts that are hard to get right and identical across every company: durable ingest, escalation timing, and proven delivery. That durable-ingest piece is worth checking directly: an alert pipeline should queue accepted alerts rather than process them inline, so a burst doesn’t become a burst of drops. Inbound alerts are queued durably (SQS with a dead-letter queue and depth alarms) so a spike is buffered; the queue keeps messages for four days.

A selection checklist

Whichever vendor category you’re looking at, the same questions separate a real evaluation from a feature-list skim. Score your current tooling — or your leading candidate — against this list before you commit:

  • Whether you can read exactly why an alert paged someone, or why a repeat was grouped instead
  • How escalation policies are built: how many tiers, per-tier timeouts, and a fallback when nobody acknowledges
  • How an incident moves from triggered to acknowledged to resolved, and what its timeline records
  • The list price against your real seat count, and what is bundled versus sold as an add-on
  • How each of your alert sources connects: a documented payload mapping and setup guide, or a generic webhook you map yourself
  • How much of your current configuration a migration carries over, and whether you can run both tools in parallel during the move
  • Whether configuration changes are recorded with before-and-after values in an append-only audit log
  • Whether acknowledgement and resolution times, alert volume by source and schedule coverage gaps are visible at a glance
  • Whether the vendor states plainly, plan by plan, exactly what is included — not just what a feature-list checkmark implies.

Where CallHeim fits

CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

On escalation and scheduling: Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback. Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target. Teams with member rosters own schedules and link to escalation policies. Schedules use any IANA time zone. Schedules support layers, an override that sits above the layers, layer restrictions to weekdays, business hours or nights, follow-the-sun templates, time off, shift-swap requests with approval, and coverage-gap detection with suggested fills. See the on-call scheduling and alert ingest product pages for the full mechanics, or our on-call software guide for the narrower on-call slice of this category on its own.

On paging channels, stated plainly rather than as a features-page checkbox: Live in early access. Incident pages go out by e-mail with an acknowledge link, and sent and delivered status is recorded on the incident.

On the incident lifecycle and analytics: Four incident states (Triggered, Acknowledged, Resolved, Closed) and five severities (P1 to P5). A resolved incident can be reopened; a closed incident is final. Lifecycle changes (acknowledge, resolve, close), comments, escalation steps and suppressions are written to the incident timeline as they happen, and the timeline has no edit or delete in the product. MTTA and MTTR (mean, p50, p90) are computed from the incident’s own timestamps; reopened incidents are measured per response, merged duplicates are excluded and still-open incidents are flagged. Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.

On pricing: Starter $9, Pro $15 and Business $29 per user per month. Enterprise is negotiated. Paid plans are activated directly by the CallHeim team. 14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected. A flat per-seat rate, with noise handling and core analytics on every plan rather than sold back as an add-on. Full detail on the pricing page.

For a row-by-row comparison against named vendors, see CallHeim vs PagerDuty and CallHeim vs Opsgenie.

Key takeaways

  • Incident management software turns a monitoring signal into a paged human and a record of the response; on-call software (schedules, escalation, paging) is the narrower slice inside it.
  • Evaluate on all twelve capabilities together — ingest, dedup/grouping, routing, scheduling, escalation, channels, lifecycle, status, analytics, integrations, security and pricing model — not just the ones a demo happens to show first.
  • Vendors roughly split into four shapes: dedicated on-call platforms, broader incident-response suites, ITSM/helpdesk products with alerting attached, and smaller newer entrants — each trades breadth against cost and simplicity differently.
  • Building your own paging pipeline is a multi-quarter undertaking because of the failure paths, not the happy path; most teams are better served buying the delivery and escalation engine even if they build a custom alert source on top.
  • CallHeim is rule-based and deterministic with flat per-seat pricing, and e-mail paging is live in early access with a signed acknowledge link.

Questions

Common questions, answered.

What is the difference between incident management software and monitoring?

Monitoring detects that something is wrong. Incident management software is the layer on top: it decides who to notify, delivers a page reliably, tracks the response, and keeps a record — monitoring tools rarely do all of that well themselves.

Is on-call software the same as incident management software?

On-call software is the narrower piece — schedules, escalation and paging. Incident management is the wider job: on-call plus the alert pipeline in front of it (ingest, deduplication, routing) and the lifecycle tooling behind it (states, timelines, status pages, analytics).

Should we build our own incident management tool or buy one?

Buying usually makes more sense: the hard part is the failure paths (durable queueing, reliable escalation timing, proven delivery, an audit trail), not the happy path of sending a message. Building can make sense for a genuinely unusual routing need or a very small team with one simple workflow.

How much does incident management software cost?

It varies widely by vendor shape: some price a flat rate per seat, others price a base seat plus separate add-ons for grouping, status pages or advanced analytics. Always compare the total cost for the capabilities you actually need, not the entry-tier list price alone.

Do I need AI for good incident management?

No. CallHeim’s own alert handling is rule-based and deterministic — the same alert, settings and state produce the same decision — and no AI model reads incident data today. Explainable, rule-based grouping is a legitimate alternative to a model-based one.

What is alert deduplication versus alert grouping?

Deduplication collapses repeats of the same alert onto one incident. Grouping combines related but different alerts. Both matter, and it is worth checking whether grouping is rule-based and explainable or a black-box model — see our alert deduplication guide for the mechanics.

How long does it take to set up incident management software?

It depends mainly on how many alert sources you connect and how much of your escalation policy already exists on paper. A single service with one schedule can be configured in under an hour; a full rollout across many services and teams is usually a multi-week project.

Try it

See this in your own on-call rotation.

14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.