Alerting

Alert fatigue: what causes it and how to measure it

Alerting · Updated 2026-09-24 · 11 min read · By the CallHeim team

Alert fatigue is what happens when the volume, frequency or irrelevance of alerts wears down the people responding to them — until real problems get slower responses, or get missed, because they look just like the noise around them. It is a paging-system problem with a measurable shape, not a mood: the fix starts with counting what is actually arriving, not with a vague sense that things feel “too loud.”

What alert fatigue actually is

Alert fatigue is a form of habituation: after enough alerts that turned out not to matter, a person’s default response to the next one shifts from “investigate immediately” toward “probably nothing, check it later.” That shift is often the correct short-term adaptation to a genuinely noisy system — the problem is that it also delays the response the one time the alert is real. Fatigue shows up as slower acknowledgements, more alerts left unactioned until someone else notices the underlying problem, and, over time, people avoiding on-call duty altogether.

The cost is not evenly distributed either. A page at 3 p.m. and a page at 3 a.m. cost the same amount of attention in the moment, but the overnight page also costs sleep — and lost sleep compounds across a multi-day rotation in a way a single busy afternoon does not. That is why the practical fixes below separate “fewer alerts overall” from “fewer alerts at the worst possible time,” and why the measurement section that follows tracks both.

What causes it

  • Duplicate alerts. The same underlying issue re-firing as separate alerts — a disk-space warning that pages again every five minutes until someone resolves the disk, not the alert.
  • Flapping. A check that oscillates between healthy and unhealthy in quick succession, generating an alert (and a resolve, and another alert) for a condition that never stabilises long enough to act on.
  • Low-value alerts. Conditions that are technically true but rarely require a human response — informational-severity noise that was never tuned down after the monitor was created.
  • Poor routing. An alert that pages someone who cannot act on it, so it either gets ignored or forwarded manually, both of which cost time and erode trust in the paging system.
  • Paging for non-urgent conditions. Using an urgent, wake-someone-up channel for something that could reasonably wait for business hours, which trains people to treat every page as possibly non-urgent.
  • Alerts without enough context to act on. A page that names a symptom but not the service, the likely cause or a runbook forces the responder to spend the first several minutes just figuring out what they are looking at — which makes every alert feel heavier than it needs to, independent of how many there are.

These causes compound rather than operate independently. A flapping check that also lacks context is worse than either problem alone — it fires repeatedly and each firing costs extra time to interpret. Fixing the single loudest cause first, rather than trying to address all five at once, is usually the faster path to a measurable improvement, because it is also the easiest change to verify against your own numbers before moving on to the next one.

How to measure it with your own data

Skip the benchmark hunt — there is no universal “normal” number of alerts per shift that applies across teams with different systems, different traffic and different tolerances. What matters is your own trend over time, on a small set of numbers you can actually pull from your paging history:

Metrics to track for alert fatigue, and what each one tells you
MetricWhat it tells you
Alerts per on-call shiftWhether raw volume is the problem, and whether it is trending up or down over recent rotations.
Off-hours and overnight pagesHow much of that volume lands outside working hours, which is where the human cost is highest.
Sleep-window interruptionsA stronger signal than raw volume — five pages during the workday cost less than one at 3 a.m.
Acknowledge timeWhether responses are getting slower over successive shifts, often a leading indicator before someone burns out or stops responding.
Duplicate / repeat rateHow much of the volume is the same underlying issue re-alerting, versus genuinely distinct problems.

Example: a team that logs 40 pages across a week, 12 of them between 22:00 and 07:00, has enough to ask two concrete questions before touching any tooling — are those 12 night pages genuinely actionable right now, or could a routing change hold them until morning; and is the daytime volume of 28 alerts actually 28 separate problems, or five problems each re-triggering several times. Neither question needs a benchmark to be useful — it needs your own numbers, tracked across a few shifts so a single bad week doesn’t look like a trend. Acknowledge time is also one half of MTTA, so if you already track response metrics for incidents, the same data usually answers the fatigue question too.

Practical fixes

None of these require guessing at a target number in advance. Each one addresses a specific mechanism from the causes above, and the effect of applying it shows up directly in the metrics you are already tracking — fewer alerts per shift, fewer off-hours pages, or faster acknowledge times, depending on which mechanism it targets.

  • Deduplication windows. Collapse repeats of the same alert onto one incident while they keep arriving within a defined window, so one flapping check produces one incident instead of dozens. See how deduplication, grouping and correlation differ.
  • Grouping. Bundle related alerts from the same underlying event so a responder triages one incident instead of ten related ones separately.
  • Maintenance windows. Suppress expected alerts during planned work, so a deploy or a known-noisy migration doesn’t page anyone for conditions everyone already expects.
  • Severity and urgency, set deliberately. Reserve the urgent, wake-someone-up channel for conditions that actually need an immediate human response; route lower-severity conditions to a non-interrupting channel instead. See our guide to incident severity levels for how to define the split.
  • Routing to the right owner. An alert that reaches someone who can act on it gets resolved faster and erodes less trust than one that has to be manually forwarded.
  • Escalation hygiene. A clear escalation policy with sensible timeouts stops an unanswered alert from sitting silently, without turning every timeout into a fresh page to more people.
  • Rotation fairness. Uneven interruption load — the same person always drawing the noisiest week — concentrates fatigue on fewer people. See our guide to building a fair on-call rotation.

Alert fatigue checklist

Use this after you have pulled your own numbers, not instead of pulling them — a checklist ticked from memory tends to confirm whatever you already believed about your paging load.

  • You can pull alerts-per-shift, off-hours pages and acknowledge time from your own paging history, not just anecdotally.
  • Every alert source has a defined severity, and severities map to different channels rather than all paging the same way.
  • Repeats of the same condition collapse onto one incident instead of generating one alert each.
  • Planned work has a maintenance window instead of a flood of expected alerts.
  • Escalation timeouts are set deliberately, not left at whatever default nobody has revisited.
  • Interruption load, not just shift count, is reviewed periodically across the rotation.
  • Alerts with a chronically low action rate are reviewed for whether they should exist at all.

How CallHeim handles this

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. A repeat that arrives after a longer quiet gap opens a new incident.

A published on-call fatigue score: 0.40 sleep-window interruptions, 0.30 volume, 0.20 off-hours, 0.10 severity, over a rolling 14 days. It is a heuristic over your own paging history, not a wellbeing measure.

The dedup window is fixed at 300 seconds and the flap threshold at 4 state changes in 600 seconds — published defaults, not a percentage claim about how much quieter your particular alert stream will get, which depends entirely on how noisy it is today. See how the noise-reduction pipeline works end to end, and how it fits inside the wider on-call software workflow.

Key takeaways

  • Alert fatigue is measurable from your own paging history: volume per shift, off-hours pages, sleep-window interruptions and acknowledge time.
  • Duplicate alerts, flapping checks, low-value monitors, poor routing and paging for non-urgent conditions are the usual causes.
  • Deduplication and grouping address volume; severity, routing and escalation hygiene address whether the right alert reaches the right person at the right urgency.
  • Rotation fairness matters as much as raw noise reduction — the same person absorbing every noisy week burns out regardless of the team’s average.
  • Track your own numbers over several shifts before changing anything; a single bad week is not a trend.

Questions

Common questions, answered.

What is alert fatigue?

Habituation to alerts: after enough alerts that turned out not to matter, a person’s default response shifts from investigating immediately to assuming it is probably nothing. That shift is a reasonable short-term adaptation to a noisy system, but it also delays the response the one time the alert is real.

How many alerts per shift is too many?

There is no universal number that applies across teams with different systems and different traffic. What matters is your own trend: alerts per shift, off-hours pages and acknowledge time, tracked across several shifts so one bad week does not look like a trend.

What is the fastest fix for alert fatigue?

Usually deduplication or a maintenance window, because both attack raw volume directly and the effect shows up immediately in your own numbers. Fix the single loudest cause first rather than trying to address every cause at once — it is also the easiest change to verify before moving to the next one.

Does reducing alert volume also reduce burnout?

Volume is one factor, but timing matters as much: a page at 3 a.m. costs sleep in a way an afternoon page does not. Fairness — whether the same person absorbs every noisy week — matters independently of the team’s average volume.

How is alert fatigue different from a noisy on-call rotation?

Alert fatigue is about what arrives: too many alerts, too irrelevant, at the wrong time. A noisy rotation is about who absorbs it. The same alert stream can be fair if it is spread evenly across a team, or unfair if the same person always inherits it.

What deduplication and flap thresholds does CallHeim use?

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

Try it

See this in your own on-call rotation.

14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.