Learn
Learn: on-call and incident response guides
Plain-English guides to on-call scheduling, escalation, alert noise and incident response, written for the question you actually searched — with the mechanics named rather than glossed over.
Start here
The guides to read first
- ComparisonsPagerDuty alternatives: what to evaluate and who to considerPagerDuty alternatives compared on pricing, add-ons and explainable grouping, with an eight-point evaluation checklist and sourced vendor prices.11 min read
- ComparisonsPagerDuty pricing: plans, add-ons and what it costs at scalePagerDuty pricing by plan and add-on, with a worked 10-, 25- and 40-seat example, what to check before buying, and a short comparison with CallHeim.11 min read
- ComparisonsOpsgenie alternatives: PagerDuty, Jira Service Management and CallHeim comparedOpsgenie alternatives compared: dated shutdown facts, a buyer checklist, and how PagerDuty, JSM and CallHeim handle an Opsgenie replacement.11 min read
- Incident responseIncident management software: what it is and how to evaluate itWhat incident management software does, the 12 capabilities to evaluate, vendor categories, build vs buy, and a selection checklist.12 min read
- On-callOn-call software: schedules, escalation and paging, explainedWhat on-call software does: schedules, rotations, escalation, paging channels and fatigue, a buyer checklist and implementation steps.11 min read
2 guides
On-call
- Escalation policy: tiers, timeouts and fallbacks explainedHow an on-call escalation policy works: tiers, targets, timeouts and fallbacks, example policies for three team shapes, and common mistakes to avoid.11 min read
- On-call rotation: types, schedules and how to design a fair oneWeekly, daily and follow-the-sun on-call rotations compared, with layers, overrides, time-zone pitfalls and example schedules for teams of 4, 6 and 12.14 min read
3 guides
Incident response
- MTTA vs MTTR: definitions, formulas and a worked exampleMTTA and MTTR defined and calculated with a five-incident worked example, mean vs p50 vs p90, common measurement pitfalls, and how to improve each metric.11 min read
- Incident severity levels: P1–P5, SEV1–SEV5 and how to define themAn example P1–P5 (SEV1–SEV5) incident severity framework, severity vs priority vs urgency, who sets severity, and mapping monitoring-tool severities.10 min read
- Incident response process: the lifecycle, roles and postmortemsThe incident response lifecycle from detection to postmortem, incident commander and communications roles, status updates, and a runbook checklist.11 min read
2 guides
Alerting
- Alert fatigue: what causes it and how to measure itWhat causes alert fatigue, how to measure it with your own paging data, and practical fixes: dedup, routing, severity and rotation fairness.11 min read
- Alert deduplication vs grouping vs correlation, explainedHow alert deduplication, grouping and correlation differ, with worked examples, flapping detection and how to choose window and threshold settings.11 min read
Try it
See it on your own alerts.
14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.