Noise reduction

Stage in the incident loop: REDUCE

Reduce alert noise without a black box deciding for you.

Alert fatigue is not a mystery to be solved with a model. It is duplicates, flapping thresholds, cascades and unclear ownership — specific problems with specific, published rules.

Example

One noisy source, one incident.

Deduplication and grouping both key off the source. A repeating alert from one tool collapses onto one incident; alerts from different tools open separate incidents unless you route or merge them yourself.

Example, not a measured volume. A second source describing the same outage — a different monitoring tool, a status page, a customer report — does not join this incident automatically: grouping by fingerprint and by title similarity is same-source only (below). Cross-source correlation is something you route or merge by hand.

The pipeline

By default, seven ordered checks.

Orchestration rules run first, and tenant threshold rules can add further steps. Order matters as much as the rules do: a suppressed alert must not also open an incident, attach to one, or count toward a threshold — so suppression runs before anything else has a chance to act on it.

The decision pipeline. By default seven ordered checks run on every alert; orchestration rules run before them and tenant threshold rules can add more. Most checks divert the alert out of the sequence — a flapping series with nothing open to join is the one exception that still creates an incident.

  1. Maintenance window

    Suppress if a window covering this service, team or the whole workspace is open right now.

    active windowOutcome: suppress

  2. Suppress rules

    Suppress if any tenant Noise Rule matches the alert.

    rule matchOutcome: suppress

  3. Fingerprint dedup

    Attach to the open incident that already owns this fingerprint, while repeats keep arriving. The fingerprint is the tenant, the source, and the alert’s dedup key or identity — falling back to its title. The service is not part of it.

    300s sliding windowOutcome: attach

  4. Flap collapse

    A series that keeps changing state attaches to the open incident instead of paging again. With nothing open to join, it still opens an incident.

    4 changes / 600s (fixed)Outcome: group

  5. Title correlation

    Group with an open, still-Triggered alert from the same source and the same service when their titles overlap enough.

    Jaccard ≥ 0.6 default, 600sOutcome: group

  6. Group rules

    Apply the tenant’s own grouping rules over the window.

    600s windowOutcome: group

  7. Create incident

    Nothing above matched: open the incident with the source’s severity (or the severity an orchestration rule sets), pick the Escalation Policy for the service, and page.

    elseOutcome: create

Defaults shown above. Orchestration rules run first and tenant threshold rules can add further steps.

Maintenance windows and suppress rules

Maintenance windows suppress new alerts for a service, a team or the whole workspace, and recovery events are exempt.

Each window is scheduled with its own start and end time.

One deliberate exception: a recovery event is never suppressed. A resolve arriving during a maintenance window still closes the incident it belongs to, or the escalation keeps walking its tiers for a problem that has already ended.

Fingerprint deduplication

An accepted alert is normalised to one record with a stable fingerprint — the tenant, the source, and the alert’s dedup key or identity, falling back to its title. The service is not part of it.

Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. A repeat that arrives after a longer quiet gap opens a new incident. The window is 300s, and it slides: every new repeat resets it, rather than the window being fixed to the first alert’s arrival time.

Flap detection

A threshold sitting near its boundary crosses it repeatedly. Count the state changes inside a trailing window: at 4 in 600s the series is flapping, and it attaches to the open incident instead of paging again.

The incident stays open and visible while the series is grouped, so the first page stands and nothing is closed on your behalf. State changes, not events — an alert that simply keeps re-notifying in the same state is persistent, not flapping.

Title correlation

Two alerts from the same source and the same service, still open and unacknowledged, whose titles overlap by a Jaccard similarity of 0.6 or more, group together. That threshold — and whether the check runs at all — is a per-workspace setting, not a fixed constant: 0.6 is the default.

A second, separate measurement exists on the incident-intelligence surface for a person browsing related incidents. The two are deliberately distinct: this pipeline’s job is a cheap, explainable grouping decision on the hot path.

Each alert the pipeline processes is stored with what happened to it (opened an incident, grouped as a duplicate, or suppressed) and the reason.

Deterministic by design

Rule-based, not a model deciding for you

CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

CallHeim

Read the rules before you trust them with your pager.

CallHeim helps teams stay in control when critical systems are not. Explore the platform, connect one source, and send yourself a page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required