SaviourOps

Incident intelligence

Most of what fires doesn’t deserve a human.

Every firing rule becomes a cheap, disposable alert. Only the ones that are sustained, correlated and worth waking someone for get promoted to an incident.

Private betaLimited spots — no credit card required.

Incidents › #714C5637 › payments-apiCriticalFiring

payments-api is critical

Container OOMKilled 3 times · firing for 5m

payments-api · ip-10-0-1-23 · Sep 4, 16:09:19

SilenceAdd NoteOpen LogsView ClusterAcknowledgeMark Resolved
Likely causeHigh confidence

Rollout v2.13 → v2.14 pushed memory past the 512Mi limit

deployed 16:03 · memory ceiling 16:09 · error rate 19.82% ≥ 10%

Investigate →
Correlated timeline16:00 – 16:20
NORMALDEGRADEDOOMKilledERROR RATEMEMORYDEPLOYKUBERNETESTRACES19.82%512Mi limitv2.13 → v2.14restart ×1restart ×316:0016:0516:1016:1516:20
Alerts (2)
Criticalpayments-apiKubernetesOOMKilled · restart ×3
Criticalpayments-apiAPMerror rate 19.82% ≥ 10%

One health model

Every signal in. One conclusion out.

Logs, traces, Kubernetes events, metrics, uptime checks and network telemetry all produce the same kind of alert, scored by the same health model. No per-source special-casing, so no two screens can disagree about whether something is wrong.

Logs
ERR payments-api "pool exhausted 20/20"WRN checkout-web "retry 3/3 failed"INF auth-svc "token refresh OK"ERR payments-api "503 upstream timeout"WRN cart-worker "high latency 2.4s"ERR payments-api "pool exhausted 20/20"WRN checkout-web "retry 3/3 failed"INF auth-svc "token refresh OK"ERR payments-api "503 upstream timeout"WRN cart-worker "high latency 2.4s"ERR payments-api "pool exhausted 20/20"WRN checkout-web "retry 3/3 failed"INF auth-svc "token refresh OK"ERR payments-api "503 upstream timeout"WRN cart-worker "high latency 2.4s"
Traces
trace:7c3f8e → checkout → payments → stripetrace:1a5d2b → order → inventory → dbtrace:4e8f1c → api-gw → search → elastictrace:9e1b4d → auth → session-storetrace:7c3f8e → checkout → payments → stripetrace:1a5d2b → order → inventory → dbtrace:4e8f1c → api-gw → search → elastictrace:9e1b4d → auth → session-storetrace:7c3f8e → checkout → payments → stripetrace:1a5d2b → order → inventory → dbtrace:4e8f1c → api-gw → search → elastictrace:9e1b4d → auth → session-store
Kubernetes
pod/payments-api-7d4b8 OOMKilled restart:3node/ip-10-0-1-23 MemoryPressuredeploy/payments-api image v2.13 → v2.14hpa/api-gw cpu:78% target:60%pod/payments-api-7d4b8 OOMKilled restart:3node/ip-10-0-1-23 MemoryPressuredeploy/payments-api image v2.13 → v2.14hpa/api-gw cpu:78% target:60%pod/payments-api-7d4b8 OOMKilled restart:3node/ip-10-0-1-23 MemoryPressuredeploy/payments-api image v2.13 → v2.14hpa/api-gw cpu:78% target:60%
Metrics
p99 payments-api 1.6s ▲ 340%error_rate checkout-web 12.4% ▲cpu db-primary 92%connections db-pool 20/20 saturatedp99 payments-api 1.6s ▲ 340%error_rate checkout-web 12.4% ▲cpu db-primary 92%connections db-pool 20/20 saturatedp99 payments-api 1.6s ▲ 340%error_rate checkout-web 12.4% ▲cpu db-primary 92%connections db-pool 20/20 saturated
Uptime
api.demo.com · UP 99.97%payments-api · DEGRADED 94.2%webhook-ingress · UP 100%heartbeat nightly-batch · missedapi.demo.com · UP 99.97%payments-api · DEGRADED 94.2%webhook-ingress · UP 100%heartbeat nightly-batch · missedapi.demo.com · UP 99.97%payments-api · DEGRADED 94.2%webhook-ingress · UP 100%heartbeat nightly-batch · missed
Network
dns api.prod.internal 340mstcp rst count 847 last 5m ▲ingress 10.2.0.1 → 10.0.3.42:8080packet loss us-east-1 0.02%dns api.prod.internal 340mstcp rst count 847 last 5m ▲ingress 10.2.0.1 → 10.0.3.42:8080packet loss us-east-1 0.02%dns api.prod.internal 340mstcp rst count 847 last 5m ▲ingress 10.2.0.1 → 10.0.3.42:8080packet loss us-east-1 0.02%
P1 · payments-api degradedopened 04:12:38Cause: rollout v2.13 → v2.14confidence highPaged one engineer, not seven7 alerts → 1 incident

The mechanism

Alerts are cheap. Incidents are earned.

Every breach becomes an alert, allowed to be noisy, transient and self-resolving. An incident is a deliberate promotion of those alerts — and that promotion is the only place an incident is ever born.

01
uptimek8s eventtrace
↓
one alert object

Everything that breaches becomes an alert

Uptime checks, Kubernetes events, OTLP traces, host metrics and your own rules all produce the same object, scored by one health model. No per-source special-casing, so no two screens can disagree.

02
cart-worker · pod restarted
↓ 31s later
resolved on its own · nobody paged

Most of them resolve without you

A pod that restarts and recovers in thirty-one seconds fires an alert and clears it. So does a one-minute upstream hiccup. They stay visible in the inbox and they never reach a phone.

03
payments-api · error rate 44.9%payments-api · p99 over SLOcheckout-web · error rate 12.4%
↓
1 incident · 7 alerts attached

What is sustained and correlated gets promoted

Alerts that persist, cluster around one entity, or describe one failure collapse into a single incident — with every contributing alert attached as evidence, an owner, and a page.

01Detect

A wall of green, and the one slash that isn't.

Monitors across HTTP, TCP, ping, DNS, gRPC and SSL certificates, plus heartbeat checks for the cron jobs that should have reported in.

Demo APIupHTTP
Up · 100% (24h) · checked 12s ago · every 60shttps://api.demo.com
Uptime record99.62%1 incident24h7d30d
30 days agoNow
Uptime (SLA)24h100%7d99.94%30d99.62%
Response time105 msp95 111ms
SSL certificateValid · renews in 38d
Checksevery 60s · timeout 30sOpens an incident after 1 confirmed failure
One confirmed failure opens an incident

Not one slow response. A degraded check is amber in the record and nothing else — the product’s own check policy, stated in the panel.

SLA windows people actually report on

24h, 7d and 30d, colour-coded by threshold. Not “one-hour uptime.”

Heartbeats for the things that go quiet

Cron jobs and workers that should have checked in and didn’t fire the same kind of alert as a failed request.

Services and infrastructure too

APM service health, the dependency map and cluster verdicts all feed the same alert spine.

02Explain

It opens with a conclusion, not a chart.

Every incident leads with a likely cause and the confidence behind it, derived from what actually changed in the window.

The alerts inbox: seven firing alerts across Kubernetes, APM and uptime sources, each with the threshold that tripped it, and a Promote action on the degraded ones
Deterministic first, model second

Image rollouts, deploys and Kubernetes patterns produce the verdict. The model only enriches it, so it still works when the model is unavailable.

Honest when nothing changed

If there was no change in the window it says so and leads with the symptom instead of inventing a cause.

Evidence attached, not linked

Every contributing alert, the correlated timeline, the blast radius and similar past incidents sit on the incident itself.

Explorer for everything else

Logs, metrics, HTTP, DNS and database queries in one search surface, with facets, live tail and pattern grouping.

03Respond

One page, to the person actually on call.

Schedules and rotations, multi-level escalation, overrides and shift swaps — and a public status page that updates from the same incident.

Escalation policy · payments
Level 0On-call engineer · from the payments rotationemail · push
+5 minLevel 1 · platform-oncall groupemail · push · phone
+15 minLevel 2 · engineering leadsphone · Slack channel
Publicstatus.yourdomain.com · degraded, updatingauto
Overrides and shift swaps

Cover a shift, hand it back, and see who is actually on call right now rather than who the rota says should be.

Every delivery, logged

Which channel, to whom, when, and whether it was acknowledged. No more “did the page go out?”

Slack incident channels

A channel per incident with the verdict and timeline posted in, plus PagerDuty, SMS, phone, Discord, Telegram and webhooks.

Included, not billed separately

On-call is in the free tier. If you would rather keep PagerDuty, it is a supported channel.

04Kubernetes

One cause, not seventeen symptoms.

A cluster is where alert storms are born. One node runs out of memory and every workload on it dies at once — this is the case promotion was built for.

A Kubernetes cluster overview: a verdict line naming what needs attention, node and pod vitals, and the workload mosaic
Stricter grouping than it looks

Five failing pods of one Deployment count as one symptom — workloads are keyed by root owner, so a thrashing ReplicaSet cannot inflate the count. A storm needs three distinct workloads within thirty minutes on one node.

Systemic storms too

ImagePullBackOff across four namespaces reads as a registry problem, not four unrelated deployment failures.

A verdict line on every cluster

All quiet — 30 workloads healthy · 1/1 nodes · last change 42m ago, or what needs attention and for how long.

Drift answers “what changed”

Image rollouts, manifest edits and scale events — the change history that makes a likely-cause verdict possible at all.

Questions

Before you install anything.

How is this different from an alerting rule in Datadog or Alertmanager?+

Those tools fire on a rule and forward the result. There is no object representing “currently firing” that can be deduplicated or correlated before it becomes a page. SaviourOps makes that object first-class, so grouping and suppression happen before anyone is notified rather than after.

Do I have to instrument my applications?+

No. The host agent collects HTTP, DNS, database-query and connection events from the kernel with eBPF, so you get request-level detail without touching application code. If you already emit OpenTelemetry traces, point them at the ingestion endpoint and they are correlated alongside.

What happens to the verdict if the AI service is down?+

It still works. Likely cause is derived deterministically from image changes, deploys and Kubernetes patterns; the model only enriches that result. If nothing changed in the window, the verdict says so and leads with the symptom instead of inventing a cause.

Can I keep the data in my own infrastructure?+

On Enterprise, yes — the platform runs as a set of services against your own PostgreSQL and ClickHouse. On the cloud version, an in-cluster OpenTelemetry gateway keeps the ingestion key out of individual app deployments.

Does it replace PagerDuty?+

It covers the same ground — schedules, rotations, multi-level escalation, overrides and shift swaps — and it is included rather than billed separately. If you would rather keep PagerDuty, it is a supported alert channel.

Find out how much of your paging was noise.

Connect one cluster and watch a week of alerts sort themselves into the handful that actually mattered.