On-call

Escalation policy: tiers, timeouts and fallbacks explained

On-call · Updated 2026-09-24 · 11 min read · By the CallHeim team

An escalation policy is the rule set that decides who gets paged when an alert fires, and what happens if that person doesn’t respond in time. It is a chain of tiers, each with a target and a timeout, ending in a fallback that is paged once every earlier tier has timed out, so one missed page does not end the alert.

Why an escalation policy matters

Without one, every alert has exactly one chance to be seen: whoever it pages, however it reaches them, with no defined recovery if that fails. A phone on silent, a person in a meeting, a schedule with an accidental gap — any of those turns a single missed notification into an incident nobody responds to until a customer, or a second, unrelated alert, forces someone to notice. An escalation policy doesn’t prevent a missed page; it gives a missed page a defined next step, so it is not the end of the story.

It also does a second, quieter job: it separates “who is on-call” from “what happens if they don’t respond,” which lets a rotation change independently of the policy built on top of it. The schedule decides who is primary this week; the escalation policy decides what the system does once that person has been paged and the clock is running.

The anatomy of an escalation policy

  • Tiers. An ordered sequence of steps — primary responder, then secondary, then perhaps a manager or a wider team, each one tried in order.
  • Targets. What each tier actually pages: a specific person, a whole team, or whoever a schedule currently names as on-call. Targeting a schedule rather than a named person is what keeps a policy correct as the roster changes.
  • Timeouts. How long a tier waits for an acknowledgement before the policy advances to the next one. Too short, and people get escalated past before they can reasonably respond; too long, and a real outage sits unanswered.
  • Repeats. Whether a tier re-notifies its own target (a second push, a second call) before giving up and moving to the next tier, useful for a channel a person might have briefly missed.
  • Fallback. The last tier in the chain, paged once every earlier tier has timed out unanswered — a manager, a wider on-call alias, or another break-glass contact. Delivery still depends on the channel it uses. A policy without one can escalate to the end of its list and then simply stop.

Diagram: Tier 1 is paged. If Tier 1’s timeout expires unanswered, the policy escalates to Tier 2. If Tier 2’s timeout also expires unanswered, the policy pages the fallback tier next.

Design principles

  • Fewer, well-understood tiers beat many overlapping ones. A policy nobody can recite from memory is a policy nobody trusts during an actual incident.
  • Target roles and schedules, not named individuals. A policy that pages “whoever the on-call schedule currently names” keeps working after a reorg; one that hard-codes a person’s name pages a former employee six months later.
  • Set timeouts to the response you actually expect, not a round number picked without thinking about it — a service where five minutes of downtime matters needs a shorter timeout than one where thirty minutes is tolerable.
  • Escalate to whoever can act, not whoever is senior. A policy that mirrors the org chart instead of who can actually fix the problem adds a step that does nothing but delay.
  • Every tier must be reachable. A fallback that pages an inbox nobody watches is not a fallback — it is a policy that looks complete and isn’t.
  • Match the number of tiers to how much margin the service needs. A low-stakes internal tool can tolerate two tiers and a generous timeout; a service where minutes matter usually needs one more tier and a shorter one, so more chances exist to catch a missed page before it becomes a prolonged outage.
  • Watch for escalation as a symptom, not just a mechanism. A policy that escalates constantly is paging more people than it needs to, which is one of the more direct causes of alert fatigue — if Tier 2 gets paged on most alerts, the problem is usually upstream of the policy, not the policy itself.

Example policies

These are examples to adapt to your own team and services, not a template to copy exactly.

Small team (4–6 people)

Example escalation policy for a small team
TierTargetTimeout
1Primary on-call (schedule)5 minutes
2Secondary on-call (schedule)5 minutes
3 (fallback)Whole-team alert channelFinal — no further timeout

Follow-the-sun team

Example escalation policy for a follow-the-sun team
TierTargetTimeout
1Current region’s on-call (schedule)10 minutes
2Next region’s on-call, already in business hours10 minutes
3 (fallback)Engineering manager on dutyFinal

Routing to the next region’s schedule instead of a fixed backup means Tier 2 always lands on someone already awake and working — see our guide to on-call rotation patterns for how follow-the-sun rotations are usually built.

Critical service (e.g. payments)

Example escalation policy for a critical service
TierTargetTimeout
1Primary on-call (schedule)3 minutes
2Secondary on-call (schedule)3 minutes
3Service owner / team lead5 minutes
4 (fallback)Incident commander on dutyFinal

Shorter timeouts reflect a lower tolerance for downtime, and a fourth tier gives the policy somewhere to go if both on-call responders and the service owner all miss it — three tiers alone would leave no margin.

Common mistakes

  • No fallback tier. The policy escalates to the end of its list and simply stops, so an alert can go entirely unanswered if everyone named happens to miss it.
  • Timeouts set without thinking about the service. Too short, and people get escalated past before they can respond; too long, and a real outage sits unanswered for longer than it should.
  • Targeting a named person instead of a schedule. The policy quietly keeps paging someone after they change teams or leave, because nobody remembered to update it.
  • One policy reused for every service. A low-stakes internal tool and a customer-facing payment path rarely deserve the same timeouts or the same number of tiers.
  • Never tested end to end. A policy that looks correct in the configuration screen can still fail in practice — a disconnected phone number, a schedule with a gap, a fallback nobody actually monitors.
  • Confusing “repeat” with a longer timeout. Re-notifying the same target several times inside one long tier delays the moment a genuinely unavailable responder gets escalated past, which is worse than a shorter timeout with a real second tier behind it.

Testing your policy

A policy that has never been deliberately failed is untested, no matter how many real incidents it has handled — a string of successful pages only proves the primary tier works, and the primary tier is the one part of the chain least likely to fail in the first place.

  • Send a real test alert periodically and confirm it reaches the primary tier’s actual device, not just that the configuration looks right on screen.
  • Deliberately let a test alert time out through every tier and confirm the fallback genuinely fires — this is the path that gets exercised least often and fails silently most often.
  • Re-test after any schedule change, team membership change, or new hire/departure that could affect who a tier actually reaches.
  • Review timeouts against real incident response times periodically, rather than leaving them at whatever default was set when the policy was first created.

How CallHeim handles this

Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.

Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target.

Each alert source is bound to a service, and the service’s escalation policy (or its team’s) sets who is paged.

Acknowledging an incident stops the escalation.

A list of escalation policies showing their tiers, fallback target and repeat count
Escalation policy

What this looks like in CallHeim (example data).

See how escalation connects to alert ingest and on-call scheduling, and what happens after an incident is acknowledged in our incident response process guide. For the wider picture, see our complete guides to on-call software and incident management software.

Key takeaways

  • An escalation policy is tiers, targets, timeouts and a mandatory fallback — the fallback is the last tier, paged once every earlier tier has timed out.
  • Target schedules and roles, not named individuals, so the policy stays correct as the team changes.
  • Match tier count and timeouts to how much downtime the service can actually tolerate; a critical service usually needs shorter timeouts and one more tier than a low-stakes one.
  • The most common failure is an untested fallback — test the timeout path deliberately, not just the primary page.
  • Acknowledging an incident should stop the escalation immediately, so nobody keeps getting paged for something already being handled.

Questions

Common questions, answered.

How many tiers should an escalation policy have?

Most teams need two to four: a primary responder, a secondary backstop, and a fallback that is paged once every earlier tier has timed out. Add a tier only if it changes who gets paged — a service where minutes matter often adds one more tier with a shorter timeout than a low-stakes internal tool needs.

What is a good escalation timeout?

There is no universal number. Set each timeout to the response you actually expect from that tier, not a round default — a service where five minutes of downtime matters needs a shorter timeout than one where thirty minutes is tolerable. The only wrong answer is picking a number without thinking about the service behind it.

What happens if every tier in an escalation policy times out?

The policy pages its fallback tier, the last step in the chain, once every earlier tier has timed out; delivery depends on the channel. A policy without a genuine fallback simply stops once it reaches the end of its list, which is why a fallback is never optional.

Should an escalation policy page a person or a schedule?

A schedule or role, not a named individual. Targeting a schedule means the policy keeps paging whoever is actually on-call as the roster changes; targeting a named person means the policy quietly keeps paging them after they change teams or leave.

Does acknowledging an incident stop escalation?

Yes. Acknowledging an incident stops the escalation. That is what stops a tier from continuing to page further people once someone is already handling it.

How do you test an escalation policy without waiting for a real incident?

Send a real test alert to the primary tier and confirm it reaches the actual device, then deliberately let a separate test alert time out through every tier to confirm the fallback genuinely fires. The fallback path is exercised least often in real incidents, and it is the path most likely to fail silently.

Try it

See this in your own on-call rotation.

14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.