On-call & incident response
When production breaks, CallHeim gets the right person on it.
Alerts arrive over HTTPS from the tools you run — repeats of the same alert collapse onto one incident instead of paging again. The incident routes to its service’s escalation policy, then escalates tier by tier to a mandatory fallback until someone acknowledges — by rules you can read, not a black-box model.
Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required
Incident pages go out by e-mail, with a signed acknowledge link built in. See how paging works.
Signals
- Datadog
- Prometheus Alertmanager
- Sentry
- AWS CloudWatch Alarm (via SNS)
- Grafana
- Pingdom
+ 130 more sources
- fingerprint
- a91f…3c
- collapsed
- 11 duplicates
- jaccard
- 0.67 ≥ 0.6
- severity suggested
- P1
One incident
[P1] checkout-api: payment authorization timeout
service checkout-api · 12 alerts · 1 page
John Smith
tier 1 · email · 02:14:06
Example scenario · computed on synthetic data by CallHeim’s own rule code
Start where you are
Built for teams responsible for critical systems
It already speaks to what you run.
136 catalogued sources, each with a payload mapping tested on sample payloads, and its own page. Search for yours.
SRE and platform teams
You own the pager for services other teams depend on, and your monitoring already knows before you do.
Read the fit →IT operations
Your alerts arrive from an ITSM tool and an inbox as often as from a monitoring agent, and on-call still has to work.
Read the fit →Security operations
Detections need routing and escalation without their contents being sent anywhere to be scored.
Read the fit →
The problem
A storm of alerts is not the same as a storm of problems.
Duplicates from the same monitor. A threshold flapping either side of its boundary. Nobody sure who owns it, because ownership never made it onto the alert. The pager stops meaning anything long before the failure is fixed — and the fix was never to send fewer alerts, it was to collapse the ones describing the same thing.
- P1payment authorization timeout
checkout-api · 12 alerts · one incident · one page
- P2elevated 504 rate
psp-gateway · 5 alerts · one incident · one page
- P3retry queue depth rising
checkout-web · 3 alerts · one incident · one page
How CallHeim works
One loop, seven stages, and a rule you can read at every one.
Ingest, reduce, decide, page, respond, resolve, learn — each stage does one job, and each one's rule is published, not modelled.
- 01
INGEST
Alerts arrive over HTTPS from your monitoring, CI and cloud tools.
- 02
REDUCE
Repeats collapse inside a five-minute window; flapping and related alerts group by published rules.
- 03
DECIDE
The source’s service picks the escalation policy. CallHeim suggests a severity; a person applies it.
- 04
PAGE
Tiers and timeouts advance to a mandatory fallback until someone acknowledges. E-mail paging today, in early access.
- 05
RESPOND
Acknowledge, comment and work the incident on a timeline with no edit or delete.
- 06
RESOLVE
Triggered → Acknowledged → Resolved → Closed, with a public status page your team updates.
- 07
LEARN
MTTA and MTTR from each incident’s own timestamps, plus a published on-call fatigue score.
Acknowledge an incident from the page e-mail through a signed, single-use link that opens a confirmation page. The link expires within 24 hours.
Why it's different
Four decisions we made on purpose.
Every one of these is a constraint we chose, not a feature we haven't built yet.
Explainable alert handling
CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today.
The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.
Private incident intelligence
No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.
Escalation you can reason about
Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.
Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target.
Straightforward per-seat pricing
From $9 per user per month, list price. A 14-day trial, no card required.
Core analytics (MTTA, MTTR, alert volume by source, escalation rate and period-over-period comparison, over windows of 7 to 365 days) are not gated by plan; related-incident and responder insights start at Pro.
Escalation, drawn from one routed decision
Incident opens
#2041 · P1 suggested · checkout-api
The Service resolves to policy "Payments — Business Hours", and the escalation state machine starts.
Tier 1 · John Smith
email · 5m to answer
The person this tier pages. A tier can also target a team or a schedule.
↓ wait 5 minutes · no answer
Tier 2 · David Miller
email · 5m to answer
The person this tier pages. A tier can also target a team or a schedule.
↓ wait 5 minutes · no answer
Tier 3 · Fallback: Head of Engineering
email · no timeout, this is the last tier
The mandatory fallback: one target, paged once the last tier times out.
Acknowledged by John Smith
at 02:16:20 · MTTA 2m 14s
Acknowledgement stops the escalation. Tiers 2 and 3 are never paged, which is the entire point of the timeout.
Five rule-based algorithms, shown in full
Inside your CallHeim environment
No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.
Severity classification
Suggests — a person applies it
weighted keywords + historical vote → P1…P4
A fixed keyword ruleset scores the alert text into one of P1 to P4, then blends in a vote from your own resolved incidents with similar titles. Ties resolve to the more severe level. Writes a suggestion and shows the matched terms; a person applies it — the incident keeps the severity its source reported.
Related-incident insights
Suggests — a person applies it
0.7·cosine + 0.3·jaccard, 30-minute window
On demand, ranks open incidents by 0.7 cosine plus 0.3 Jaccard similarity of title and body within a 30-minute window, with a same-service boost and a recency weight. Pro and above. This is a separate, on-demand insight — not the ingest pipeline’s title grouping, which uses a 600-second window.
On-call fatigue
Suggests — a person applies it
0.30·volume + 0.40·sleep + 0.20·off-hours + 0.10·severity
A published heuristic over your own paging history across a rolling 14 days: 0.30 volume, 0.40 sleep-window interruptions, 0.20 off-hours, 0.10 severity. It is a heuristic, not a wellbeing measure, and it feeds a suggestion — nothing is sent on its own.
Rotation assist
Suggests — a person applies it
fatigue-aware fair rotation + greedy balancer
Proposes a fairer rotation, weighted by the fatigue heuristic and a schedule balancer. A proposal, not a change — you approve it before anything moves.
Noise reduction
Acts automatically
the deterministic pipeline — no model of any kind
Deduplication, flap collapse and title correlation run automatically, by the published default thresholds. This is the one surface here that acts without a person applying it.
Optional Amazon Bedrock enhancement
Disabled by default. The local result is always computed first and is what runs today.
Inside the console
What the loop actually looks like.

Escalation policy: tiers, per-tier timeouts and a mandatory fallback (example data).

On-call schedule: rotation layers and the final schedule they produce. Coverage gaps are shown in the console; they are not alerted in advance. (example data).
Incident response
Understand it, own it, coordinate it, close it.
An incident is a record with a timeline, an owner and a state, written as it happens — not a notification that disappears once it is dismissed.
Triggered → Acknowledged
2m 14sMTTA
Paged, and somebody took it.
Acknowledged → Resolved
29m 32sMTTR
Comments, timeline, and the fix.
- 02:14:06
Incident created from 12 alerts · P1 · Triggered
What happened, assembled from every alert that described it
- 02:14:07
Escalation started · policy "Payments — Business Hours" tier 1
Who owns it, set by the Service rather than by whoever is watching
- 02:14:08
Paged John Smith · email
The page itself, on the channel that reaches them
- 02:16:20
Acknowledged by John Smith
Acknowledged — the escalation stops here
- 02:20:48
Comment · psp-gateway connection pool at its limit
Coordination: comments written to the timeline as they happen
- 02:43:38
Resolved · psp-gateway connection pool raised
Resolved: MTTA and MTTR come from these timestamps
The rest of the platform
Alerting & escalation
HTTPS webhooks with opt-in HMAC signing, routed to a service and its escalation policy — up to 8 tiers, per-tier timeouts and a mandatory fallback.
Read more →Noise reduction
Duplicate collapse, flap grouping and title correlation on published default thresholds — deterministic, no model in the path.
Read more →Incident intelligence
A severity suggestion with the matched terms shown, and related-incident insights on Pro and above. A person applies every suggestion.
Read more →On-call scheduling
Layers, overrides, time off, approved shift swaps and coverage-gap detection, in any IANA time zone.
Read more →Incident response
A timeline with no edit or delete, role-gated actions, a public status page your team updates, and MTTA/MTTR analytics.
Read more →
Integrations
Everything you monitor with, arriving at one place.
136 catalogued alert sources, all ending at the same deterministic pipeline.
Each alert source has a payload mapping for that tool’s webhook format and its own setup page, built and tested against sample payloads.
Switching from PagerDuty or Opsgenie
A way off, on your own schedule.
Opsgenie end of support is 5 April 2027: access is shut off then, and data not migrated by that date is deleted. Atlassian Opsgenie licensing page, checked 2026-09-24.
Run in parallel during a move: CallHeim receives PagerDuty Events API v1 and v2 bodies, Opsgenie outgoing webhooks, and alerts forwarded from Splunk On-Call, Squadcast, Zenduty and Grafana OnCall.
Import tooling (early access) reads your PagerDuty or Opsgenie configuration (US-region accounts); an optional dry run lets you review the result before anything is written to your workspace.
Security
Security built into every layer.
- Reads and writes are scoped by the API authorization rule to the caller’s tenant, or to the caller’s own user for personal records, and custom operations enforce the tenant in server-side code. CallHeim is multi-tenant on a shared database.
- The audit log is append-only: no client can modify or delete an entry.
- Point-in-time recovery (35-day window) and deletion protection are enabled on CallHeim’s production data tables.
- No incident data is sent to any external LLM. CallHeim’s AI features are rule-based code running in our own AWS account, and CallHeim does not train any model on your incidents. An optional Amazon Bedrock path exists in the code and is switched off.
Pricing
Priced per seat, from $9 a month.
Starter, Pro and Business scale by seat count; Enterprise is negotiated. A 14-day trial, no card required.
Starter
$9per user / month
list price · up to 10 seats
For a small team that owns its own pager.
Pro
$15per user / month
list price · up to 50 seats
For teams that want API access and related-incident insights.
Business
$29per user / month
list price · up to 200 seats
For a growing organisation. The first paid plan with email-to-alert: an inbound alert address per workspace.
Paid plans are activated directly by the CallHeim team.
Put a rule you can read between your alerts and your on-call.
Explore the platform, connect one source, and watch it collapse into one page.
Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required