On-call & incident response

When production breaks, CallHeim gets the right person on it.

Alerts arrive over HTTPS from the tools you run — repeats of the same alert collapse onto one incident instead of paging again. The incident routes to its service’s escalation policy, then escalates tier by tier to a mandatory fallback until someone acknowledges — by rules you can read, not a black-box model.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required

Incident pages go out by e-mail, with a signed acknowledge link built in. See how paging works.

Signals

  • Datadog
  • Prometheus Alertmanager
  • Sentry
  • AWS CloudWatch Alarm (via SNS)
  • Grafana
  • Pingdom

+ 130 more sources

7 CHECKS · 300s
fingerprint
a91f…3c
collapsed
11 duplicates
jaccard
0.67 ≥ 0.6
severity suggested
P1

One incident

Severity P1Status Triggered#2041

[P1] checkout-api: payment authorization timeout

service checkout-api · 12 alerts · 1 page

JS

John Smith

tier 1 · email · 02:14:06

Example scenario · computed on synthetic data by CallHeim’s own rule code

The problem

A storm of alerts is not the same as a storm of problems.

Duplicates from the same monitor. A threshold flapping either side of its boundary. Nobody sure who owns it, because ownership never made it onto the alert. The pager stops meaning anything long before the failure is fixed — and the fix was never to send fewer alerts, it was to collapse the ones describing the same thing.

  • P1payment authorization timeout

    checkout-api · 12 alerts · one incident · one page

  • P2elevated 504 rate

    psp-gateway · 5 alerts · one incident · one page

  • P3retry queue depth rising

    checkout-web · 3 alerts · one incident · one page

An example shape, not a measured volume — the field is drawn dense on purpose and is not to scale with the twenty alerts in the incidents beside it. The rules that do the reducing are published in full, with their thresholds.

How CallHeim works

One loop, seven stages, and a rule you can read at every one.

Ingest, reduce, decide, page, respond, resolve, learn — each stage does one job, and each one's rule is published, not modelled.

  1. 01

    INGEST

    Alerts arrive over HTTPS from your monitoring, CI and cloud tools.

    Integrations · Catalog →

  2. 02

    REDUCE

    Repeats collapse inside a five-minute window; flapping and related alerts group by published rules.

    Noise Rules · Maintenance →

  3. 03

    DECIDE

    The source’s service picks the escalation policy. CallHeim suggests a severity; a person applies it.

    Orchestration Rules · Services →

  4. 04

    PAGE

    Tiers and timeouts advance to a mandatory fallback until someone acknowledges. E-mail paging today, in early access.

    Escalation Policies · Schedules →

  5. 05

    RESPOND

    Acknowledge, comment and work the incident on a timeline with no edit or delete.

    Incidents · Timeline →

  6. 06

    RESOLVE

    Triggered → Acknowledged → Resolved → Closed, with a public status page your team updates.

    Incidents · Status Page →

  7. 07

    LEARN

    MTTA and MTTR from each incident’s own timestamps, plus a published on-call fatigue score.

    Insights · Analytics →

Acknowledge an incident from the page e-mail through a signed, single-use link that opens a confirmation page. The link expires within 24 hours.

Why it's different

Four decisions we made on purpose.

Every one of these is a constraint we chose, not a feature we haven't built yet.

Explainable alert handling

CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

Private incident intelligence

No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.

How the suggestions work →

Escalation you can reason about

Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.

Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target.

Straightforward per-seat pricing

From $9 per user per month, list price. A 14-day trial, no card required.

Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.

Escalation, drawn from one routed decision

  1. Incident opens

    #2041 · P1 suggested · checkout-api

    The Service resolves to policy "Payments — Business Hours", and the escalation state machine starts.

  2. Tier 1 · John Smith

    email · 5m to answer

    The person this tier pages. A tier can also target a team or a schedule.

  3. ↓ wait 5 minutes · no answer

    Tier 2 · David Miller

    email · 5m to answer

    The person this tier pages. A tier can also target a team or a schedule.

  4. ↓ wait 5 minutes · no answer

    Tier 3 · Fallback: Head of Engineering

    email · no timeout, this is the last tier

    The mandatory fallback: one target, paged once the last tier times out.

  5. Acknowledged by John Smith

    at 02:16:20 · MTTA 2m 14s

    Acknowledgement stops the escalation. Tiers 2 and 3 are never paged, which is the entire point of the timeout.

The tiers, targets and timeouts of one Escalation Policy, from the same routed decision the engine produced. Tier 1 answered, so tiers 2 and 3 never fired — they are shown because the reason to configure them is the night tier 1 does not.

Five rule-based algorithms, shown in full

Inside your CallHeim environment

No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.

  • Severity classification

    Suggests — a person applies it

    weighted keywords + historical vote → P1…P4

    A fixed keyword ruleset scores the alert text into one of P1 to P4, then blends in a vote from your own resolved incidents with similar titles. Ties resolve to the more severe level. Writes a suggestion and shows the matched terms; a person applies it — the incident keeps the severity its source reported.

  • Related-incident insights

    Suggests — a person applies it

    0.7·cosine + 0.3·jaccard, 30-minute window

    On demand, ranks open incidents by 0.7 cosine plus 0.3 Jaccard similarity of title and body within a 30-minute window, with a same-service boost and a recency weight. Pro and above. This is a separate, on-demand insight — not the ingest pipeline’s title grouping, which uses a 600-second window.

  • On-call fatigue

    Suggests — a person applies it

    0.30·volume + 0.40·sleep + 0.20·off-hours + 0.10·severity

    A published heuristic over your own paging history across a rolling 14 days: 0.30 volume, 0.40 sleep-window interruptions, 0.20 off-hours, 0.10 severity. It is a heuristic, not a wellbeing measure, and it feeds a suggestion — nothing is sent on its own.

  • Rotation assist

    Suggests — a person applies it

    fatigue-aware fair rotation + greedy balancer

    Proposes a fairer rotation, weighted by the fatigue heuristic and a schedule balancer. A proposal, not a change — you approve it before anything moves.

  • Noise reduction

    Acts automatically

    the deterministic pipeline — no model of any kind

    Deduplication, flap collapse and title correlation run automatically, by the published default thresholds. This is the one surface here that acts without a person applying it.

Optional Amazon Bedrock enhancement

Disabled by default. The local result is always computed first and is what runs today.

Five rule-based algorithms, and the arithmetic behind each one. Only noise reduction acts on its own; the other four write a suggestion for a person to apply. All of it runs inside our own infrastructure; the optional Amazon Bedrock path is drawn because it exists in the code, and dashed because it is off.

Inside the console

What the loop actually looks like.

A list of escalation policies showing their tiers, fallback target and repeat count
Escalation policy

Escalation policy: tiers, per-tier timeouts and a mandatory fallback (example data).

An on-call schedule timeline showing rotation layers and the final schedule they resolve to
On-call schedule

On-call schedule: rotation layers and the final schedule they produce. Coverage gaps are shown in the console; they are not alerted in advance. (example data).

Incident response

Understand it, own it, coordinate it, close it.

An incident is a record with a timeline, an owner and a state, written as it happens — not a notification that disappears once it is dismissed.

  1. 02:14:06

    Incident created from 12 alerts · P1 · Triggered

    What happened, assembled from every alert that described it

  2. 02:14:07

    Escalation started · policy "Payments — Business Hours" tier 1

    Who owns it, set by the Service rather than by whoever is watching

  3. 02:14:08

    Paged John Smith · email

    The page itself, on the channel that reaches them

  4. 02:16:20

    Acknowledged by John Smith

    Acknowledged — the escalation stops here

  5. 02:20:48

    Comment · psp-gateway connection pool at its limit

    Coordination: comments written to the timeline as they happen

  6. 02:43:38

    Resolved · psp-gateway connection pool raised

    Resolved: MTTA and MTTR come from these timestamps

Example scenario, not a customer incident: one incident drawn on its own clock, each tick at its elapsed time, using the same example as the rest of this page. The MTTA and MTTR shown are illustrative. Acknowledgement is a sliver because the page reached someone; the rest is the work.

The rest of the platform

Integrations

Everything you monitor with, arriving at one place.

136 catalogued alert sources, all ending at the same deterministic pipeline.

Each alert source has a payload mapping for that tool’s webhook format and its own setup page, built and tested against sample payloads.

Browse the integration directory →

Switching from PagerDuty or Opsgenie

A way off, on your own schedule.

Opsgenie end of support is 5 April 2027: access is shut off then, and data not migrated by that date is deleted. Atlassian Opsgenie licensing page, checked 2026-09-24.

Run in parallel during a move: CallHeim receives PagerDuty Events API v1 and v2 bodies, Opsgenie outgoing webhooks, and alerts forwarded from Splunk On-Call, Squadcast, Zenduty and Grafana OnCall.

Import tooling (early access) reads your PagerDuty or Opsgenie configuration (US-region accounts); an optional dry run lets you review the result before anything is written to your workspace.

Security

Security built into every layer.

  • Reads and writes are scoped by the API authorization rule to the caller’s tenant, or to the caller’s own user for personal records, and custom operations enforce the tenant in server-side code. CallHeim is multi-tenant on a shared database.
  • The audit log is append-only: no client can modify or delete an entry.
  • Point-in-time recovery (35-day window) and deletion protection are enabled on CallHeim’s production data tables.
  • No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.

The full security and trust posture →

Pricing

Priced per seat, from $9 a month.

Starter, Pro and Business scale by seat count; Enterprise is negotiated. A 14-day trial, no card required.

Starter

$9per user / month

list price · up to 10 seats

For a small team that owns its own pager.

Pro

$15per user / month

list price · up to 50 seats

For teams that want API access and related-incident insights.

Business

$29per user / month

list price · up to 200 seats

For a growing organisation. The first paid plan with email-to-alert: an inbound alert address per workspace.

Paid plans are activated directly by the CallHeim team.

CallHeim

Put a rule you can read between your alerts and your on-call.

Explore the platform, connect one source, and watch it collapse into one page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required