On-call
On-call software: schedules, escalation and paging, explained
On-call · Updated 2026-09-24 · 11 min read · By the CallHeim team
On-call software is the part of an incident management stack that answers one specific question, reliably, at any hour: who is responsible for this alert right now, and how do we reach them if they don’t respond? This guide covers what it actually does, how schedules and escalation work together, which paging channels matter and why fatigue is worth measuring, a buyer checklist, a rollout sequence, and where CallHeim fits.
What on-call software does
At its core, on-call software does three things: it knows who is on call right now, it decides what happens if that person doesn’t respond, and it delivers the notification through a channel that actually reaches them. Those three jobs map to three pieces of the product — schedules (who is on call), escalation policies (what happens next), and paging channels (how the notification arrives) — and a weak implementation of any one of them undermines the other two. A perfect schedule paired with a paging channel nobody actually carries is not a working on-call system.
On-call software is usually one part of a wider incident management platform, sitting behind alert ingest and deduplication and in front of the incident lifecycle and status communication. If you’re evaluating the whole stack rather than just the paging layer, our incident management software guide covers the full picture — this guide stays deliberately narrow, on the scheduling-and-paging slice alone.
Before spreadsheets and shared calendars, this was often handled with a physical pager passed from desk to desk, or a phone list taped near a monitor — workable at a handful of engineers, and increasingly error-prone past that. The job hasn’t changed; the failure modes have just moved. A spreadsheet rotation doesn’t know when someone is on holiday. A shared calendar doesn’t escalate itself when an invite is declined. Dedicated on-call software exists to close exactly those gaps: the rotation, the escalation and the paging are one connected system instead of three things a human has to keep in sync by hand.
Schedules, rotations and escalation
A schedule names who is on call and when. The mechanics that separate a schedule that survives contact with a real team from one that only works in the demo: layered rotations (a base rotation plus an override layer that can substitute one person for a day without editing the whole pattern), holiday and time-off handling that doesn’t require manually re-building the rotation, restricting a layer to weekdays, business hours or nights, follow-the-sun templates for teams spread across time zones, shift-swap requests that need approval rather than a silent trade, and coverage-gap detection that flags a hole in the schedule before it becomes a 2 a.m. surprise. Our on-call rotation guide covers each of these with worked rotation patterns.
An escalation policy is the rule set layered on top of the schedule: it decides who gets paged first, and what happens if that person doesn’t acknowledge in time. It is built from tiers, each with a target (usually “whoever the schedule currently names,” not a fixed person) and a timeout, ending in a fallback that is paged once every earlier tier has timed out — a policy without one can escalate to the end of its list and simply stop. Acknowledging the incident should stop the escalation immediately, so a responder who is already working the problem doesn’t keep getting paged for it. The full mechanics, with example policies for different team shapes, are in our escalation policy guide.
The two work together deliberately: the schedule answers “who is primary this week,” and the escalation policy answers “what happens if primary doesn’t respond” — keeping them separate means a schedule change (a new hire, a reorg) never requires rebuilding the escalation logic on top of it.
Paging channels and fatigue
A paging channel only matters if it is proven to reach a person, not merely listed as supported. Push notifications, SMS, phone calls, e-mail and chat integrations each have different reliability characteristics — a push notification depends on a phone having a signal and the app not being force- closed by the OS; a phone call is harder to sleep through than a silent push; e-mail depends on notification settings actually being on. When you evaluate a vendor, ask specifically which channels have been exercised by real, incident-driven paging in production — not which channels exist in the settings screen.
Fatigue is the other half of the paging story, and it is easy to skip because it doesn’t show up until weeks in. A rotation that pages the same person too often, too late at night, or for alerts that turn out not to matter, produces slower responses over time even though every individual page looked fine in isolation. A published, explainable fatigue measure — one that weights interruptions during normal sleeping hours more heavily than a daytime page, for instance — gives a team something concrete to act on before burnout shows up as a missed real incident. See our alert fatigue guide for the fuller picture, including how alert volume itself (not just who gets paged) contributes to it.
Fatigue and channel choice interact, too. A responder who is paged by phone call for a genuine outage but by a quiet push notification for something routine learns, without anyone deciding it on purpose, which alerts to take seriously — so the channel a tier uses is itself a fatigue lever, not just a delivery detail. Matching channel weight to severity (a louder, harder-to-miss channel for the tiers that matter most) is a cheap adjustment most teams make only after the first missed page forces the conversation.
A buyer checklist
Before you commit to an on-call tool, check it against this list — ideally by actually doing each one in a trial, not by reading a features page:
- Send a real test page through every channel you plan to rely on, and confirm it reaches an actual device, not just that the settings screen shows it as enabled.
- Build one schedule with an override and a time-off entry, and confirm the on-call view updates correctly for the day the override applies.
- Deliberately let a test escalation time out through every tier and confirm the fallback actually fires.
- Check whether schedules support every time zone your team actually works in, not just the vendor’s home region.
- Ask what a fatigue or paging-load measure looks like, and whether it is visible to the people it measures, not just to managers.
- Confirm acknowledging an incident stops further escalation immediately.
- Check whether the number of tiers, the schedule complexity, or the channels you need are capped on a cheaper plan and only unlocked higher up.
- Ask directly which paging channels are proven in production today versus recently added — a vendor being specific about this is a good sign in itself.
Implementation steps
Working through these steps in order matters — each one depends on the step before it being right, and skipping ahead (connecting every alert source before a single schedule is tested end to end, for instance) is the most common way a rollout stalls:
- 01
Map services to alert sources
List every service that pages someone today, and which monitoring tool or system generates that page. This becomes your integration list.
- 02
Build the schedules
Start with who is actually on call this week, in their real time zones — not an idealised rotation nobody follows.
- 03
Write the escalation policy for each service
Primary tier, a secondary tier, timeouts that match how much downtime the service tolerates, and a fallback that is paged once every earlier tier has timed out.
- 04
Turn on and test every paging channel you plan to rely on
A channel “supported” in a settings screen is not the same as a channel proven to reach a phone — send a real test page through each one before trusting it.
- 05
Connect the alert sources
Wire in your monitoring tools one at a time, checking that the fields you need (severity, service, summary) map correctly before moving to the next.
- 06
Run a deliberate escalation drill
Let a test alert time out through every tier on purpose and confirm the fallback genuinely fires — this is the path real incidents exercise least often.
- 07
Review the first month of paging data
Volume by person, time-of-day distribution, and how often escalation actually advances past tier one — then adjust rotations and timeouts based on what really happened, not what you assumed going in.
Where CallHeim fits
On schedules: Teams with member rosters own schedules and link to escalation policies. Schedules use any IANA time zone. Schedules support layers, an override that sits above the layers, layer restrictions to weekdays, business hours or nights, follow-the-sun templates, time off, shift-swap requests with approval, and coverage-gap detection with suggested fills. Coverage gaps are shown in the console; they are not alerted in advance. See the full mechanics, with example rotation patterns, in our on-call rotation guide or the scheduling product page.
On escalation: Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback. Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target. Each alert source is bound to a service, and the service’s escalation policy (or its team’s) sets who is paged. Acknowledging an incident stops the escalation.
On paging channels, stated as plainly as the channel-reliability question above deserves: Live in early access. Incident pages go out by e-mail with an acknowledge link, and sent and delivered status is recorded on the incident.
On fatigue: A published on-call fatigue score: 0.40 sleep-window interruptions, 0.30 volume, 0.20 off-hours, 0.10 severity, over a rolling 14 days. It is a heuristic over your own paging history, not a wellbeing measure. Read more in our alert fatigue guide.
On pricing: Starter $9, Pro $15 and Business $29 per user per month. Enterprise is negotiated. Paid plans are activated directly by the CallHeim team. 14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected. Scheduling and escalation are baseline on every plan, not sold as an add-on. Full detail on the pricing page.
For the wider incident management picture beyond on-call alone, see the incident management software guide.
Key takeaways
- On-call software does three jobs together: knowing who is on call, deciding what happens if they don’t respond, and delivering a page through a channel that actually reaches them.
- Schedules and escalation policies are deliberately separate — the schedule says who is primary, the policy says what happens if primary misses it — so a roster change never means rebuilding escalation logic.
- A paging channel only counts if it’s proven to reach a device in production, not merely listed as supported; ask which channels are proven, specifically.
- Fatigue is measurable, not just a feeling — a published, explainable score that weights sleep-hour interruptions heavily gives a team something to act on before burnout causes a missed real incident.
- CallHeim’s schedules and escalation cover layers, overrides and fallbacks, flat-priced across every plan, with e-mail paging live in early access.
Questions
Common questions, answered.
What is the difference between on-call software and incident management software?
On-call software is the narrower piece: schedules, escalation and paging. Incident management software is the wider job — on-call plus the alert pipeline in front of it and the incident-lifecycle tooling behind it. See our incident management software guide for the full picture.
How many people should be in an on-call rotation?
There is no universal number — it depends on how much downtime the service tolerates and how sustainable the paging load is per person. What matters more than headcount is a mandatory fallback tier and a measured fatigue score, so an undersized rotation shows up as a number rather than a surprise.
What happens if nobody acknowledges a page?
A correctly built escalation policy advances to the next tier after its timeout, ending in a fallback tier that is paged once every earlier tier has timed out; delivery depends on the channel. A policy with no fallback can escalate to the end of its list and simply stop, which is the most common on-call design mistake.
Does on-call software prevent alert fatigue?
Not by itself — fatigue comes from alert volume and paging load, which the software can measure and help you reduce (through deduplication, grouping and fair rotations) but cannot eliminate on its own. See our alert fatigue guide for what to measure and how to act on it.
Which paging channels are most reliable?
Reliability varies by what "reliable" means to you: a phone call is harder to sleep through than a silent push notification, but depends on a working phone number; SMS and push depend on notification settings and connectivity. Always ask which channels are proven in production, not just listed as supported.
Can on-call schedules span multiple time zones?
Yes, in any properly built on-call tool — schedules should support any IANA time zone and follow-the-sun templates for teams spread across regions, so the next tier in an escalation naturally lands on someone already awake.
How much does on-call software cost?
Pricing models vary: some vendors charge a flat rate per seat, others charge a base seat plus separate add-ons for extra schedules, escalation tiers or analytics. CallHeim lists Starter, Pro and Business per user per month, with scheduling and escalation included on every plan.
Try it
See this in your own on-call rotation.
14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.