Metrics & monitoring

Send Prometheus Alertmanager (direct) alerts to on-call with CallHeim.

Alertmanager pointed straight at CallHeim (no Prometheus relay). CallHeim maps the payload, collapses repeats within five minutes, and pages whoever is on call for the service it belongs to.

alertmanageralso accepts: am

Alertmanager groups

A group is split into one alert per member. If several alerts from one group are on the same incident, it stays open until every one of them has recovered.

Alerts pushed straight into Alertmanager — from Cortex, Thanos Ruler, Mimir's ruler, VictoriaMetrics' vmalert, or a script posting to Alertmanager's own HTTP API — become CallHeim alerts: repeats of the same alert collapse onto one incident, and the service the integration is bound to pages whoever is on call for it. CallHeim includes a payload mapping for Alertmanager’s webhook format, built and tested against sample payloads. This is the direct-receiver catalogue entry: if you run a Prometheus server evaluating rules in front of Alertmanager, the Prometheus Alertmanager guide is the more specific one to read — the wire format and this integration's setup are otherwise identical.

How the alert reaches on-call

  1. Step 1A rule engine other than Prometheus itself — Cortex, Thanos Ruler, Mimir's ruler, VictoriaMetrics' vmalert — or a script calling Alertmanager's own API evaluates a condition and posts the alert straight to Alertmanager, or someone posts one by hand through amtool.
  2. Step 2Alertmanager groups it with other alerts sharing the same labels, exactly as it would a Prometheus-originated alert.
  3. Step 3Alertmanager POSTs the group as one JSON body to your CallHeim ingest URL, on the interval its route config sets.
  4. Step 4CallHeim splits the group into one alert per member, maps each alert's severity label, and stores the identity namespaced under am: so it cannot collide with a Prometheus-routed alert.
  5. Step 5The integration's bound service selects an escalation policy, which pages the on-call tier.

Setting it up

  1. 01

    Create the Alertmanager (direct) integration in CallHeim.

    This catalogue entry has its own ingest URL, already tied to the alertmanager parser — no query-param workaround needed.

  2. 02

    Add a webhook_configs receiver to Alertmanager's configuration, pointed at that URL.

    The same alertmanager.yml receivers block as any other webhook_config; nothing about the alert's origin changes the syntax.

  3. 03

    Set send_resolved: true on the receiver.

    Makes Alertmanager POST again when alerts resolve, which is what resolves the CallHeim incident — see Recovery below.

  4. 04

    Route the alerts you push directly into Alertmanager to this receiver.

    A route (or route tree) matching the labels your ruler or script sets — the same routing mechanism Alertmanager always uses, regardless of who created the alert.

  5. 05

    Reload or restart Alertmanager to apply the config.

Example payloadJSON
{
  "version": "4",
  "groupKey": "{}:{alertname=\"ReplicationLagHigh\"}",
  "status": "firing",
  "receiver": "callheim-direct",
  "commonLabels": { "alertname": "ReplicationLagHigh", "severity": "critical" },
  "alerts": [
    {
      "status": "firing",
      "labels": { "alertname": "ReplicationLagHigh", "severity": "critical", "instance": "pg-replica-2" },
      "annotations": { "summary": "Replica lag on pg-replica-2 exceeds 60s", "description": "Pushed by a custom checker script, not a Prometheus rule" },
      "startsAt": "2026-06-27T10:00:00Z",
      "fingerprint": "b7c8d9e0f1a2"
    }
  ]
}
Example shape written by us, matching Alertmanager's documented webhook_config format — the same shape whether the alert came from a Prometheus server, another ruler, or a direct API push; check the Alertmanager documentation for the current schema.

What CallHeim reads from the payload

Payload fields and what each one maps to on the incident
Payload fieldMaps toNote
annotations.summary / labels.alertname / annotations.descriptionincident title
labels.severityseveritycritical→P1, warning→P3, info→P4; a severity label literally set to "page" also maps to P1 — identical mapping to the Prometheus-routed integration.
annotations.description / summary / messageincident body
alerts[].fingerprintidentity signalNamespaced under am:, so a direct Alertmanager alert and a Prometheus-routed one sharing the same underlying fingerprint never collapse onto each other if a tenant runs both integrations.

How CallHeim processes it

Routing

Each alert source is bound to a service, and the service’s escalation policy (or its team’s) sets who is paged.

Deduplication

Repeats of the same alert collapse onto one incident while they keep arriving within five minutes of each other. A repeat that arrives after a longer quiet gap opens a new incident. The fingerprint is Alertmanager's own per-alert fingerprint when present, else the group key, else the alert name and instance — namespaced under am:, distinct from the Prometheus integration's own prom: namespace even when both would otherwise see the identical alert.

Recovery

When Alertmanager sends a resolved notification, CallHeim resolves the incident, exactly as for the Prometheus-routed integration. If several alerts from one group are on the same incident, it stays open until every one of them has recovered. Checked with a real Alertmanager (v0.28.1) on 26 September 2026, including groups where one alert recovers while another is still firing. An incident that also holds an alert which never sends a recovery stays open until a person resolves it.

Troubleshooting

Common problems and how to fix them
ProblemFix
An incident stays open after an alert resolved.Expected while another alert on the same incident is still firing — it resolves when the last one recovers (see Recovery above). If every alert has recovered, check that send_resolved: true is set on the receiver, or resolve it by hand.
HTTP 400 on send.Check the Alertmanager version — the webhook payload schema (for example the values field) changed between major versions, independent of who is pushing alerts into it.
Alerts from a script or ruler never reach CallHeim.Confirm the pushed alert actually matches a route pointing at this receiver — an alert with no matching route falls through to Alertmanager's default route, same as a Prometheus-originated one would.
A directly-pushed alert has no summary text.Whatever created the alert (script, ruler) did not set a summary or description annotation on it — CallHeim falls back to the bare alertname.
HTTP 401 once you enabled signing.Alertmanager's http_headers sends a fixed value, not a per-request signature — it can't satisfy CallHeim's HMAC check. Leave signing off for this integration unless you add a proxy in front that can compute and add the header per request.

Security

CallHeim verifies an HMAC-SHA256 signature on each source’s requests once you enable signing on that integration (the sending tool must be able to sign). A current Alertmanager accepts a custom header through the webhook receiver's http_config (http_headers), but that value is static, and a static header can never be a valid per-request HMAC-SHA256 signature of a changing body. Leave CallHeim's signing off for a direct Alertmanager webhook and rely on the ingest URL itself as the credential, unless you add a signing proxy in front that can compute the header per request.

Vendor documentation checked

The thresholds it passes through

Dedup window
300s
Flap threshold
4 transitions / 600s
Title correlation
similarity ≥ 0.6, same source and service
Group window default
600s

All defaults are published. You can turn title correlation off or change its threshold, and set the window on your own noise rules; the dedup window and the flap settings are fixed. How alerts are processed →

CallHeim

Point Prometheus Alertmanager (direct) at CallHeim and see what it does with your alerts.

Explore the platform, connect one source, and send yourself a test page by e-mail (early access).

Early access · every workspace starts with a 14-day trial for up to 5 seats, no card required