SaviourOps

Incidents

A promotion, not a forwarded alert.

An incident is created only when alerts are sustained and correlated enough to earn one. It arrives with a likely cause, every contributing alert attached as evidence, and one owner.

firingpayments-api · 7 alertsconfidence: high
Likely causeRollout: payments-api v2.13 → v2.14
04:09:14Image changed on payments-api04:10:02p99 crossed SLO — 1.6s > 800ms04:11:47Error rate 44.9% ≥ threshold 10%04:12:38Promoted from 7 alerts · checkout-web affected downstream
InvestigateAcknowledge

Lifecycle

Four states, and the fourth one matters: marking something a false positive is first-class, because a promotion engine you cannot correct is a promotion engine you stop trusting.

Open
Promoted from one or more alerts and currently firing. Severity is scored low, medium, high or critical from the composing alerts rather than set by hand.
Acknowledged
Someone has picked it up. Acknowledgement stops the escalation ladder from advancing to the next level.
Resolved
Closed out, with the full timeline and evidence retained for the postmortem and for similar-incident matching later.
False positive
Marked as noise. This is the correction signal — it tells you where promotion was too eager.
Notes
Free-text running commentary on the incident, kept alongside the machine-generated timeline.

Leading with the conclusion

The verdict is derived deterministically from what actually changed, so it still works when the model is unavailable. Enrichment is additive, never load-bearing.

Likely cause
Image rollouts, deploys and Kubernetes patterns are checked in priority order. When nothing changed in the window it says so and leads with the symptom instead of inventing a cause.
Confidence
High for an image change or deploy, medium for a recognised Kubernetes pattern, low when the verdict falls back to the symptom.
Root cause analysis
A causal chain built from the correlated signals, with the reasoning shown rather than asserted.
Correlated timeline
Deploys, image changes, Kubernetes events, metric breaches and the composing alerts on one axis.
Blast radius
Which services downstream were affected, and how far the failure actually travelled.
Kubernetes events
Cluster events from the incident window attached directly to the incident, not a tab away.

Grouping and correlation

The part that decides whether your phone rings. Grouping happens before notification, never after.

Cross-source grouping
Uptime checks, Kubernetes events, traces, host metrics and custom rules all produce the same alert object, so they can be correlated with each other rather than siloed.
Storm collapse
Many workloads failing together on one node collapse into a single root-cause callout instead of one incident per symptom.
Similar incidents
Past incidents matched by embedding, so you can see whether this has happened before and what closed it last time.
Anomalies
Detected anomalies surfaced against the incident and resolvable independently of it.

After the fact

Runbooks
Generated from the incident, and matched against existing runbooks so an established procedure wins over a new suggestion.
Postmortems
Drafted from the timeline and evidence, then edited and moved through review states by a human.
Retained evidence
Every contributing alert, its reasons and its thresholds stay attached after resolution.

Generated runbooks, postmortems and similar-incident search are model-backed and gated behind a feature flag. Everything else on this page — promotion, grouping, the deterministic verdict, the timeline and blast radius — runs without a model.

Find out how much of your paging was noise.

Connect one cluster and watch a week of alerts sort themselves into the handful that actually mattered.