Alerting & escalation

Stage in the incident loop: INGEST · DECIDE · PAGE

On-call alerting and escalation policies that reach the right person.

Three jobs that get bundled together and shouldn't be. Ingest is a queue problem. Ownership is a routing problem. Reaching someone at 3 a.m. is a persistence problem. CallHeim treats them as three.

Ingest

However your tools send it

Send alerts to your CallHeim integration URL over HTTPS, in CallHeim’s alert format or through a field mapping you define. CallHeim queues each accepted alert durably and processes it asynchronously; acceptance is an HTTP 202.

An accepted alert is normalised to one record with five severity levels (P1 to P5) and a stable fingerprint. Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. A repeat that arrives after a longer quiet gap opens a new incident.

PagerDuty Events API v1 and v2 request bodies are also accepted, for a parallel run during a move off PagerDuty. Every sender needs a CallHeim integration URL and key; existing PagerDuty routing keys are not reused.

Browse all 136 integrations →

The ingest contract

endpoint
POST /v1/alerts
accept
An HTTP 202: the alert is queued durably and processed asynchronously, not handled inline.
body
CallHeim’s own alert format, or a field mapping you define for whatever your tool sends.
signing
CallHeim verifies an HMAC-SHA256 signature on each source’s requests once you enable signing on that integration (the sending tool must be able to sign).

Decide

Who owns this alert?

A rule an evaluator can trace end to end, not a model that has to be trusted.

Each alert source is bound to a service, and the service’s escalation policy (or its team’s) sets who is paged.

A two-tier ruleset — workspace-wide rules first, then per-service rules — can route an alert to a different service, change its severity or priority, suppress it, or pause paging, before ownership is decided this way.

CallHeim also suggests a severity from a fixed keyword ruleset and your own resolved incidents, and shows the matched terms so a person can apply it. The incident itself keeps the severity its source reported. How CallHeim handles severity →

Page

Escalation that reaches the right person.

Escalation runs on an AWS Step Functions state machine: an unanswered tier times out and advances to the next tier or the fallback.

  • Escalation policies have up to 8 tiers, per-tier timeouts from 1 to 240 minutes, and a mandatory fallback target.
  • Acknowledging an incident stops the escalation.
  • Teams with member rosters own schedules and link to escalation policies, in any IANA time zone
Example · Escalation Policy · Payments — Business Hours
T+0:00

Tier 1

The responder on call right now, paged on their own notification order.

  1. T+0:00

    Tier 1

  2. T+5:00

    No acknowledgement → tier 2

  3. T+10:00

    Still nothing → the tier chain repeats

  4. Repeats exhausted

    One fallback target, once

  5. Acknowledged

    The walk stops

A list of escalation policies showing their tiers, fallback target and repeat count
Escalation policy

Escalation policy: tiers, per-tier timeouts and a mandatory fallback (example data).

Paging

E-mail paging, live in early access.

E-mail paging is live in early access: incident pages go out through Resend with an acknowledge link, and sent and delivered status is recorded on the incident.

Acknowledge an incident from the page e-mail through a signed, single-use link that opens a confirmation page. The link expires within 24 hours.

CallHeim

Put a rule you can read between your alerts and your on-call.

CallHeim helps teams stay in control when critical systems are not. Explore the platform, connect one source, and send yourself a page.

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required