Incident response

Incident severity levels: P1–P5, SEV1–SEV5 and how to define them

Incident response · Updated 2026-09-24 · 10 min read · By the CallHeim team

Incident severity levels (commonly P1–P5 or SEV1–SEV5) rank incidents by impact so the right amount of urgency and the right people respond to each one, often driving which escalation policy a service uses. A P1 (or SEV1) is the highest-impact tier — typically a full outage of a core function — and lower numbers mean lower impact and a slower expected response. The exact criteria and response expectations at each level are a framework every team adapts to its own product, not a fixed standard.

Why severity levels exist

Without a shared severity scale, every incident gets triaged from scratch: who should be paged, how fast, and whether it’s worth waking someone up at 3 a.m. A severity level is a shortcut that answers all three questions at once, consistently, based on impact rather than on how alarming the alert text sounds or who happens to notice it first. It also gives everyone — engineering, support, leadership — a shared vocabulary for how bad something is without re-explaining it each time.

Severity levels do their job well only when they’re used consistently. A team that calls everything a P1 “to be safe” has, in practice, no severity scale at all — the whole point is that a P1 gets a different response than a P3, and that only works if most incidents genuinely aren’t P1s.

An example P1–P5 framework

Example framework — adapt the impact criteria and response expectations to your own product and organization; there is no universal standard.

Example P1 to P5 incident severity framework
LevelExample criteriaExample expected response
P1 / SEV1Core product is down or unusable for all or most users; data loss in progress.Immediate page, all hands, continuous work until mitigated.
P2 / SEV2A major feature is broken or badly degraded for a large share of users; no workaround.Immediate page, dedicated responder, updates on a fixed cadence.
P3 / SEV3A feature is degraded or broken for a subset of users, or a workaround exists.Paged during business hours or a short after-hours delay; single responder.
P4 / SEV4Minor, cosmetic or low-traffic issue with negligible user impact.Queued as normal work; no page.
P5 / SEV5No current user impact — a near-miss, a flaky alert, or an internal-only issue.Logged for review; no page, no dedicated response.

Some teams use three levels, some use five, some use named tiers instead of numbers. The number of levels matters less than having criteria specific enough that two different people would classify the same incident the same way.

Severity vs priority vs urgency

These three get conflated constantly, and keeping them separate avoids a lot of triage arguments:

  • Severity measures impact: how bad is this, objectively, right now — how many users, how core a function, is data at risk.
  • Priority measures order: given everything currently in flight, what gets worked on first. A P2 with nothing else happening might be worked before a P1 that’s already mitigated and just needs cleanup — priority is about the queue, severity is about the incident itself.
  • Urgency measures how much the cost of delay grows over time. A slow data-corruption bug might be low severity today but high urgency, because every hour it runs increases the eventual cleanup cost — the opposite of an incident that’s already at its worst and stable.

In practice, most teams collapse priority and urgency into the severity level for simplicity, and that’s a reasonable simplification as long as everyone knows that’s what they’ve done — the trouble starts when someone assumes severity alone tells them what to work on next without checking what else is in flight.

Who sets severity, and when it changes

Initial severity is usually set one of two ways: automatically, from a mapping configured on the alert source or monitoring tool, or manually, by whoever first triages the incident during the early steps of the incident response process. Either way, severity should be treated as a working assessment, not a permanent label:

  • Raise it when the blast radius turns out to be larger than first understood, or a workaround stops working.
  • Lower it once impact is confirmed contained, even if the incident is still technically open — this matters for accurate MTTA/MTTR reporting, since a P1 that spends most of its life at reduced impact skews severity-based averages if it’s never reclassified.
  • Record who changed it and when — a severity history is part of the incident’s own audit trail and useful in a postmortem.

Mapping monitoring-tool severities

Most monitoring and observability tools ship their own severity or priority scale — critical/warning/info, P1–P4, or a numeric score — and these rarely line up cleanly with your incident-management tool’s levels. A “critical” alert from one monitoring tool and a “critical” alert from another can mean very different things in practice. The fix is an explicit mapping per alert source: this source’s “critical” becomes your P2, that source’s “error” becomes your P3, and so on — reviewed periodically, since a source that changes its own alert thresholds can silently drift out of alignment with your mapping.

Example mapping. A worked example of how one team might translate three differently scaled monitoring tools onto a single five-level incident scale:

Example mapping from monitoring-tool severities to incident severity levels
Source severityExample tool labelMapped incident level
HighestCritical / Sev1 / P1P1 or P2, depending on scope
HighError / High urgencyP2 or P3
MediumWarning / Medium urgencyP3 or P4
Low or informationalInfo / Low urgencyP4 or P5, often no page

The mapping is rarely one-to-one — a source’s “highest” label might land on your P1 or your P2 depending on which service raised it, which is why the mapping is defined per source rather than once globally.

Severity and customer communication

Severity also drives who outside the response team hears about an incident, and how. A P1 usually warrants a public status update the moment impact is confirmed; a P4 usually doesn’t warrant any external communication at all. Keeping that mapping explicit — which levels trigger a status page update, which trigger a direct account-team notification, which stay entirely internal — avoids two opposite failure modes: customers finding out about an outage from social media before the status page updates, or a status page cluttered with routine, low-impact incidents nobody outside the team needed to know about.

Common mistakes

  • Severity inflation. Everything becomes a P1 because it’s the only way to guarantee a fast response, which defeats the purpose of having levels at all.
  • No re-triage. Severity is set once at trigger time and never revisited, even as the picture changes over the life of the incident.
  • Confusing severity with effort. A hard-to-fix bug isn’t automatically high severity if its user impact is small — difficulty belongs in the priority conversation, not the severity one.
  • One scale for every audience. Engineering’s severity scale and a customer-facing status page’s impact language don’t have to be the same words, but they do need to tell a consistent story about how bad things are.

How CallHeim handles this

An accepted alert is normalised to one record with five severity levels (P1 to P5) and a stable fingerprint.

CallHeim suggests a severity from a fixed keyword ruleset plus your own resolved incidents, shows the matched terms, and a person applies it. Suggestions never change severity or page anyone.

Severity is one input into the wider noise-reduction and incident management workflow — see how it interacts with deduplication and grouping upstream.

Key takeaways

  • Severity levels exist to make triage consistent and fast — they only work if most incidents aren’t the top tier.
  • Severity measures impact; priority is queue order; urgency is how fast the cost of delay grows. They often get collapsed into one number, but knowing the difference avoids triage disputes.
  • Treat severity as a working assessment — re-triage as the picture changes, and record the history.
  • Map each alert source’s own severity language explicitly onto your levels; don’t assume “critical” means the same thing everywhere.
  • Severity inflation (everything is a P1) is the most common way a severity scale stops being useful.

Questions

Common questions, answered.

How many severity levels should we use — three or five?

There is no fixed standard. Three levels are enough for many teams; five give finer granularity for larger organizations. What matters more than the count is that the criteria are specific enough that two different people would classify the same incident the same way.

What is the difference between severity, priority and urgency?

Severity measures impact: how bad is this, objectively, right now. Priority measures order: given everything currently in flight, what gets worked on first. Urgency measures how fast the cost of delay grows. Most teams collapse all three into one severity level for simplicity.

What is severity inflation, and why is it a problem?

Calling everything a P1 "to be safe." It defeats the purpose of having levels at all — the whole point of a severity scale is that a P1 gets a different response than a P3, which only works if most incidents genuinely are not P1s.

Who decides an incident’s severity?

Either automatically, from a mapping configured on the alert source, or manually, by whoever first triages the incident. Either way, treat it as a working assessment rather than a permanent label.

Should severity be re-evaluated during an incident?

Yes. Raise it if the blast radius turns out larger than first understood, lower it once impact is confirmed contained, and record who changed it and when as part of the incident’s own audit trail.

Does CallHeim suggest incident severity automatically?

CallHeim suggests a severity from a fixed keyword ruleset plus your own resolved incidents, shows the matched terms, and a person applies it. Suggestions never change severity or page anyone.

Try it

See this in your own on-call rotation.

14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.