Alerting
Alert deduplication vs grouping vs correlation, explained
Alerting · Updated 2026-09-24 · 11 min read · By the CallHeim team
Alert deduplication collapses repeats of the same alert into one incident using a fingerprint and a time window, so a flapping check doesn’t open ten incidents for one problem. It is a narrower, stricter mechanism than grouping (bundling related alerts from one source) or correlation (linking alerts across different sources or signals) — the three are often confused, and the difference determines what each setting can and can’t do for your noise.
Deduplication, grouping and correlation are three different things
These terms get used loosely, but they describe mechanisms with meaningfully different scope. Confusing them is how teams end up disappointed by a “deduplication” setting that was never going to solve a cross-tool alert fatigue problem, because that was never what deduplication does. All three sit inside a wider noise reduction pipeline, usually alongside escalation and paging.
Deduplication: same alert, repeated
Deduplication answers one question: is this the same alert firing again? It works by computing a fingerprint from stable fields of the incoming alert — typically the source, the check or rule name, and the affected resource, deliberately excluding fields that change on every firing like a timestamp or a free-text message. Two alerts with the same fingerprint, arriving within the deduplication window of each other, collapse onto the same incident instead of opening a new one.
The window itself can be built two ways:
- Sliding window. Every new repeat resets the clock, so the incident stays open as long as repeats keep arriving inside the window of the previous one. A monitor that pages once a minute for twenty minutes stays one incident throughout, even though twenty minutes is longer than any single window.
- Fixed window. The window is anchored to the first occurrence and doesn’t reset, so a burst that runs longer than the window opens a second incident partway through, even if alerts never stopped arriving. Fixed windows are simpler to reason about but can split one continuous problem into several incidents.
A sliding window is the more common default because it matches the intuitive definition of “still the same ongoing problem,” but it means an incident with no defined end can, in principle, stay open indefinitely as long as the alert keeps repeating.
Grouping: related alerts from one source, bundled
Grouping is looser than deduplication: instead of requiring an exact fingerprint match, it bundles alerts that are similar — usually measured by comparing alert titles or tags — coming from the same source within a shorter time window. Grouping catches the case deduplication misses: five different checks on the same service failing within seconds of each other, which aren’t the same alert repeating, but are clearly one underlying incident.
Grouping typically runs on a similarity threshold (how close two titles have to be to count as related) and its own, usually shorter, time window. Set the threshold too low and unrelated alerts get bundled together, hiding a second, genuinely new failure inside an existing incident. Set it too high and grouping never fires, leaving deduplication to do all the work alone.
Correlation: across sources or signals
Correlation is the broadest of the three: linking alerts that come from different sources, tools or signal types, on the theory that they describe the same underlying event — a database alert, an application error-rate alert and a customer-facing latency alert that all started within the same few minutes, for example. Correlation is harder to do well than deduplication or grouping because there’s no shared fingerprint or source to anchor on; it depends on timing, topology or learned relationships between services, and it’s the mechanism most likely to either miss a real link or invent one that isn’t there.
Worked examples
Example. A disk-usage check on db-primary-3 fires every two minutes while disk usage stays above the threshold.
| Time | Event | Outcome (5-minute sliding window) |
|---|---|---|
| 00:00 | Disk usage alert fires | New incident opened |
| 00:02 | Same fingerprint fires again | Collapses onto the open incident |
| 00:04 | Same fingerprint fires again | Collapses; window slides to 00:09 |
| 00:11 | Same fingerprint fires again (7 min after 00:04) | Outside the window — new incident opened |
The gap at 00:11 is longer than the window measured from the last repeat, so it reads as a new incident — this is the deduplication behaviour a sliding window produces, and it’s the correct outcome if the underlying problem genuinely recurred after a quiet spell rather than continuing uninterrupted.
Example. The same host also trips a CPU alert and a memory alert forty seconds after the disk alert, all from the same monitoring source. None of the three share a fingerprint, so deduplication leaves them as three separate incidents. Grouping, comparing alert titles from the same source within a short window, bundles the CPU and memory alerts with the disk incident as related — because they are close in time, from the same source, and plausibly one underlying host problem.
Flapping
Flapping is a specific failure mode dedup and grouping windows don’t fully solve on their own: a check that oscillates between OK and firing rapidly, generating a new alert (and, without a window, a new incident) on every transition. The usual fix is a separate flap detector: if a check changes state more than a set number of times within a set window, it’s flagged as flapping rather than treated as a stream of independent incidents, so the noise is named instead of hidden. A common definition — CallHeim publishes exactly this default — is four or more state transitions within a ten-minute window.
Trade-offs when choosing settings
Every dedup and grouping setting is a trade-off between two failure directions, and there is no setting that eliminates both:
- Windows too short, thresholds too strict: real recurrences of the same problem get split into separate incidents, each paging on-call again for something already known.
- Windows too long, thresholds too loose: a second, genuinely new failure arrives while an existing incident is still open and gets silently absorbed into it — the alert exists in the system, but nobody sees it as a distinct event needing its own attention.
Over-grouping is the more dangerous direction of the two, because it fails silently: an under-grouped setup produces annoying extra pages, which someone notices, while an over-grouped setup can hide a real, unrelated failure inside an incident everyone already believes is handled.
How to choose your settings
- Start from how long your noisiest, most repetitive alert typically takes to resolve on its own, and set the dedup window a little longer than that — not longer than it needs to be.
- Keep the grouping similarity threshold conservative until you’ve reviewed a week or two of grouped incidents and confirmed nothing distinct is being absorbed.
- Review flapping alerts separately from volume — a check that flaps four times in ten minutes needs a fix to the check or the underlying resource, not just a wider window.
- Re-check your settings after any change to alert volume or a new alert source — a threshold tuned for one noise profile can silently misbehave under a different one, and it changes the MTTA and MTTR numbers you’re tracking.
- Set severity independently of dedup and grouping settings — a well-deduplicated alert can still be the wrong severity level if the underlying mapping was never reviewed.
How CallHeim handles this
CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.
The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.
Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. A repeat that arrives after a longer quiet gap opens a new incident.
Each alert the pipeline processes is stored with what happened to it (opened an incident, grouped as a duplicate, or suppressed) and the reason.
Grouping in CallHeim compares alert titles from the same source using a similarity score, with a default threshold of 0.6 inside a 600-second window — that is same-source title similarity, not cross-source correlation. CallHeim does not link alerts across different sources or signal types today. Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro. This pipeline is part of the wider incident management workflow, alongside escalation and incident response.
Key takeaways
- Deduplication collapses exact repeats of the same alert by fingerprint; grouping bundles similar alerts from one source; correlation links across different sources — they solve different problems.
- A sliding window resets on every repeat; a fixed window doesn’t, and can split one ongoing problem into several incidents.
- Flapping (rapid state changes) needs its own detector, not just a longer dedup window.
- Over-grouping is the riskier failure direction: it hides new, distinct failures inside an incident that looks handled.
- Review your settings after any change in alert volume or monitoring sources, not just once at setup.
Questions
Common questions, answered.
What is the difference between alert deduplication and grouping?
Deduplication collapses exact repeats of the same alert, matched by a stable fingerprint, onto one incident. Grouping is looser: it bundles similar alerts from the same source within a shorter window — the case deduplication misses, such as five different checks on one service failing within seconds of each other.
What is alert flapping?
A check that oscillates rapidly between healthy and unhealthy, generating a new alert — and, without a window, a new incident — on every transition. The usual fix is a dedicated flap detector: if a check changes state more than a set number of times within a set window, it is flagged as flapping rather than treated as a stream of independent incidents.
Should I use a longer or shorter deduplication window?
Start from how long your noisiest, most repetitive alert typically takes to resolve on its own, and set the window a little longer than that. Too short splits one ongoing problem into repeated pages; too long risks silently absorbing a second, genuinely new failure into an incident that already looks handled.
Does alert correlation work across different monitoring tools?
Correlation — linking alerts across different sources — is the broadest and hardest of the three mechanisms to do well, because there is no shared fingerprint to anchor on. CallHeim does not link alerts across different sources or signal types today; grouping compares titles from the same source only.
What is a fingerprint in alert deduplication?
A value computed from an alert’s stable fields — typically the source, the check or rule name, and the affected resource — deliberately excluding fields that change on every firing, like a timestamp or a free-text message. Two alerts with the same fingerprint, arriving within the window, collapse onto one incident.
What are CallHeim’s default deduplication and grouping thresholds?
Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. Grouping compares alert titles from the same source using a similarity score, with a default threshold of 0.6 inside a 600-second window.
Try it
See this in your own on-call rotation.
14-day trial with up to 5 seats and no card required. Adding a user beyond a plan’s seat limit is refused (Trial 5, Starter 10, Pro 50, Business 200); existing users and paging are not affected.