checkout-api p99 latency > 2s
Example · the pipeline's own mechanics
Watch an alert become an incident.
One alert, walked through the checks above, with the numbers those checks actually use.
01 · A webhook arrives
{
"alert_type": "error",
"title": "checkout-api p99 latency > 2s",
"host": "checkout-api-canary",
}02 · The alert gets a fingerprint
- fingerprint
- sha256(tenant · source · dedup key)
Tenant, source and the alert’s dedup key or identity, falling back to its title. The service is not an input.
- window
- 300s, sliding
Resets on every new repeat.
03 · Duplicates collapse, grouping does not cross sources
- repeats
- same fingerprint
Attach to the open incident; no new page.
- title grouping
- Jaccard ≥ 0.6 default
Same source, same service, still Triggered, within 600s. Never across sources.
04 · Severity is the source's, plus a suggestion
- incident severity
- P1
Mapped from what Datadog sent (or an orchestration override).
- suggestion
- P1, matched “latency”, “p99”
A separate suggestion from a keyword ruleset and your resolved incidents. A person applies it; it never changes severity on its own.
05 · Routed and paged
How incident severity levels work →What an escalation policy is →