For SRE, platform and DevOps

On-call for SRE teams: your monitoring already knows. CallHeim makes sure a person does.

You did not buy Datadog so it could wake four people about one dependency. The gap between “something fired” and “the right person is looking at it, with a record of why” is the part most tools ask you to take on faith.

Your stack

The tools already on your dashboards.

Each alert source has a payload mapping for that tool’s webhook format and its own setup page, built and tested against sample payloads.

Browse every integration →

Alert fatigue, not just alert volume

Published noise rules, not a black box.

You are the one who has to explain, in the postmortem, why a page did or did not arrive. These are the rules you would be explaining.

  • Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other.
  • A repeat that arrives after a longer quiet gap opens a new incident.
  • Alerts from the same source and service whose titles share enough words (title similarity 0.6 or higher) group instead of paging twice, for as long as the first one is still open.
  • Maintenance windows suppress new alerts for a service, a team or the whole workspace, and recovery events are exempt.

Published, not proprietary

The thresholds behind the noise handling are published: a 300-second dedup window, a flap threshold of 4 state changes in 600 seconds, and title-similarity grouping at a default of 0.6.

CallHeim’s alert handling is rule-based and deterministic: given the same alert, settings and state it makes the same decision, and no AI model reads your alert data today. A series that changes state four times in ten minutes groups onto its open incident instead of re-paging, on fixed thresholds.

How noise reduction works →

On-call that respects sleep

The rotation is the retention problem

Everyone measures pages per person. It is the least useful of the numbers that matter, because being woken at 3 a.m. is not equivalent to being pinged at 3 p.m. — and treating them as equivalent is how a rotation quietly becomes unfair.

A published on-call fatigue score: 0.40 sleep-window interruptions, 0.30 volume, 0.20 off-hours, 0.10 severity, over a rolling 14 days. It is a heuristic over your own paging history, not a wellbeing measure.

The score can feed a fair-rotation suggestion that a person reviews and applies; it does not reorder a live rotation automatically.

On-call scheduling →

Fatigue weighting, published

  • 0.40

    Sleep-window interruptions (22:00–07:00 default)

  • 0.30

    Page volume

  • 0.20

    Off-hours pages

  • 0.10

    Severity pressure

Escalation and the record

A state machine you can reason about, and a timeline that does not get edited.

Escalation

  • Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.
  • Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target.
  • Each alert source is bound to a service, and the service’s escalation policy (or its team’s) sets who is paged.
  • Acknowledging an incident stops the escalation.

Alerting & escalation →

Timeline and MTTA/MTTR

  • Lifecycle changes (acknowledge, resolve, close), comments, escalation steps and suppressions are written to the incident timeline as they happen, and the timeline has no edit or delete in the product.
  • MTTA and MTTR (mean, p50, p90) are computed from the incident’s own timestamps; reopened incidents are measured per response, merged duplicates are excluded and still-open incidents are flagged.

Incident response →

How you get paged

E-mail paging, live in early access.

Page

Live in early access. Incident pages go out by e-mail with an acknowledge link, and sent and delivered status is recorded on the incident.

Acknowledge

A signed, single-use link in the page e-mail opens a confirmation page where you acknowledge the incident. The link expires within 24 hours. Resolving is done in the console.

CallHeim

Put a rule you can read between your alerts and your on-call.

CallHeim helps teams stay in control when critical systems are not. Explore the platform, connect one source, and send yourself a page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required