Incident intelligence

Stage in the incident loop: DECIDE · LEARN

Explainable incident intelligence.

CallHeim's alert handling and severity suggestions are rule-based and deterministic by default. You can see why an alert was grouped, what severity was suggested and why, and no incident data is sent to an external model to get there.

Rule-basedDeterministicAn AIOps alternative you can read the reasons for, not just the score.

How it works

Five deterministic steps, and what each one is for.

This is the same alert pipeline described on the alerting and noise-reduction pages, read from the angle of what it decides and reports rather than how it ingests.

  1. 01

    Fingerprint and collapse

    Every alert gets a fingerprint from the tenant, the source and its identity. Repeats collapse onto one incident while they keep arriving within 5 minutes of each other, on published, fixed thresholds.

    How grouping works →

  2. 02

    Severity suggestion

    A fixed keyword ruleset, weighted by severity, reads the alert and your own resolved-incident history and suggests a level with the matched terms shown. A person applies it.

  3. 03

    Related-incident insights

    On Pro and above, a related-incidents panel ranks other open incidents by how closely their alert text and service match — a starting point for someone triaging, not an automatic merge.

  4. 04

    Fatigue score

    A published formula scores on-call load from your own paging history, so an uneven rotation is visible before someone burns out.

  5. 05

    MTTA and MTTR

    Response-time metrics are computed from each incident’s own timestamps the moment it resolves — no model needed to know how long something took.

    How incident response measures it →

Private incident management, explained

A rule you can read beats a score you can’t.

CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.

That is what makes CallHeim an alternative to an AIOps model rather than a smaller version of one: a model returns a score, and a ruleset returns the terms it matched and their weights, so a wrong call has a reason attached to it instead of a confidence number.

Published defaults

Dedup window
300s
Flap threshold
4 in 600s
Title-similarity default
0.6

Dedup and flap are fixed. Title-similarity grouping is a per-workspace setting around this default.

Severity suggestion

Shown with its reasons, applied by a person.

CallHeim suggests a severity from a fixed keyword ruleset plus your own resolved incidents, shows the matched terms, and a person applies it. Suggestions never change severity or page anyone.

Example incident #12

An example, not a customer incident. Alert text: “checkout-api: payment authorization timeout.”

“p1”
P1 signal
“latency”
P2 signal
“timeout”
P2 signal

Suggested: P1. A person opens the incident and applies, changes or ignores it — the suggestion itself never changes a severity or pages anyone.

The suggestion shows the terms that matched and their weights, so you can see why it chose a severity. The keyword list itself is fixed by CallHeim — it is not a per-workspace setting — and you can override the result by hand or set severity with your own orchestration rules.

For the same alert, the same rules and the same incident history, the result is identical every time. Nothing in this pipeline trains on your incidents: the historical vote it reads from is your own resolved incidents, inside your own workspace.

The boundary

What runs by default, and what stays switched off.

By default, every suggestion is rule-based code running in CallHeim’s own AWS account. An optional model path exists in the code and is switched off.

Default, and what runs today

  • Severity suggestions, related-incident ranking, the fatigue score and MTTA/MTTR are all rule-based code running in our own AWS account.
  • No incident data is sent to any external LLM.
  • CallHeim ships no generative assistant and no autonomous agents. That is a deliberate choice.

Exists in the code, switched off

No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.

Page e-mail still carries incident content to your responders through Resend — that is delivery, not intelligence. See where data goes on the security page.

Alert correlation

Related incidents, ranked for a person to check.

Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.

The panel is a ranked list on the incident page for a person to read. CallHeim never combines incidents on its own.

Plan

Related-incident insights

Pro and above

Some features are tied to plan: the REST incident API and related-incident insights on Pro and above; email-to-alert (per-workspace inbound alert e-mail addresses) on Trial, Business and Enterprise, not Starter or Pro.

On-call fatigue

A heuristic you can see the weighting of.

A published on-call fatigue score: 0.40 sleep-window interruptions, 0.30 volume, 0.20 off-hours, 0.10 severity, over a rolling 14 days. It is a heuristic over your own paging history, not a wellbeing measure.

It is not a wellbeing measure and it does not diagnose anyone — it is a published formula over your own paging history, so an uneven rotation shows up on a page instead of staying invisible until someone complains.

CallHeim

Put a rule you can read between your alerts and your on-call.

CallHeim helps teams stay in control when critical systems are not. Explore the platform, connect one source, and send yourself a page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required