Anomaly detection vs root-cause analysis

Anomaly detection tells you a metric changed, and root-cause analysis tells you where the change sits and what happened around the time it started.

Updated · Zello Labs

Anomaly detection tells you a metric changed more than its normal pattern allows. Root-cause analysis comes after it. It tells you where the change sits, whether a rate moved or the mix shifted, and what happened close to the start. Detection gives you an alert. Root-cause analysis gives you a place to start looking.

What does anomaly detection tell you?

Anomaly detection compares a metric’s latest value with the value it expected, and flags a value outside the normal range. A good detector knows the metric’s habits. It removes the weekly pattern and the trend first, so a quiet Sunday or a slow climb doesn’t count as a change. Why weekly patterns cause false alarms covers that part in detail.

The output is short. You get a metric, a time, the expected value, the actual value and the gap between them. For example, activation rate is 9.4% below expected since Wednesday at 10:00 UTC. That is accurate and worth knowing. It is also where detection stops.

Why do alerts alone create triage work?

An alert tells you something moved and leaves every next question to a person. Someone slices the metric by region, then by plan, then by platform, looking for the part that moved. Then they ask around to learn what shipped or changed that week. That takes hours, and it only happens for the metrics someone had time to look at.

Alerts that fire on normal variation make it worse. When most alerts turn out to be a Monday dip or a holiday, people stop reading them and then switch them off. A real problem then goes unnoticed, because the alert that would have caught it is off.

What comes after anomaly detection?

Root-cause analysis does the slicing and the asking around, and it does them the same way every time. It adds four things to the alert.

Where the change sits

A segment can carry most of a change just by being most of the business. Metron tests each segment of each column you list against that segment’s own history, which separates the segment that moved from the one that is simply large. You learn which segment holds most of the change. If one segment stands out, Metron looks inside it and breaks it down by the remaining columns.

Whether the rate moved or the mix shifted

For a ratio metric like conversion, a drop can come from two places. The rate inside a segment fell, or volume moved toward a segment that always converts lower. You fix those in different ways, so Metron reports which one happened. Rate effect vs mix effect walks through an example.

What happened near the start

Metron checks which events fall close to the start of the change. Deploys and releases arrive from GitHub by webhook. You log the rest yourself, such as a price change or a campaign. Each event is described by elapsed time, such as 30 minutes before, with no score and no “likely cause” label.

A write-up

Metron puts the findings into a few plain sentences with the numbers in them. A language model you supply words the result. It sees the engine’s findings and nothing else, never your raw tables. If you ask a follow-up about a segment, Metron answers from the same evidence.

Anomaly detection vs root-cause analysis, side by side

Anomaly detection Root-cause analysis
Question it answers Did this metric move outside its normal range? Where does the change sit, and what happened around it?
Input One metric over time The flagged metric, its segments and an event log
Output An alert with expected and actual values A write-up with the segment, rate or mix, nearby events and a confidence for each claim
What you do next Start investigating Check the segment and the events it points to
Where it stops It can’t say which part of the business moved It can’t prove what caused the change

What can’t root-cause analysis tell you?

Root-cause analysis finds association. It can show that most of a drop sits in one segment and that a release went out 52 minutes before the drop began. It can’t prove the release caused the drop. Warehouse data records what happened and when. It doesn’t record why. Two things can line up in time by chance, and some causes never get logged.

A tool that blurs this line will sooner or later state a wrong cause with confidence, and one wrong cause costs more trust than many right answers earn. Metron scores where and when as separate claims, each with its own confidence. Cause is marked as not established, every time. See how Metron scores each claim.

How Metron runs both steps

Metron runs inside your environment and queries your Postgres or BigQuery warehouse with a read-only role. Detection judges each change against the metric’s own normal, after removing the weekly pattern and the trend. A new metric gets a simpler model until there is enough history. When a run checks many metrics at once, Metron raises the bar for each one, so a long list of metrics doesn’t turn into a long list of false alarms.

An investigation can also start without an alert. Someone asks why a metric moved last Tuesday, and Metron runs the same steps and returns the same kind of evidence.

Metron is in private beta, and we are setting up a small number of teams by hand. If your alerts already tell you something moved and you want the investigation that comes next, request beta access.

Common questions

What is the difference between anomaly detection and root cause analysis?

Anomaly detection tells you a metric moved outside its normal range. Root-cause analysis starts from that alert and works out where the change sits, whether a rate moved or the mix shifted, and which events landed near the start. Detection gives you the what and the when. Root-cause analysis adds the where and the context.

What comes after anomaly detection?

An investigation comes next. Someone has to find the segment that moved, check whether the rate changed or the mix shifted, and learn what shipped or changed near the start. Root-cause analysis runs those steps on the data and writes up what it found, so your team starts from evidence and a short list of places to look.

Can root cause analysis prove what caused a metric to change?

No. It can show where a change sits and which events happened near its start, and that is association. Warehouse data can't prove cause, because events can line up by chance and some causes never get logged. Metron scores where and when as separate claims and leaves the question of cause to your team.

Why do teams turn off anomaly alerts?

Usually because the alerts fire on normal variation, such as a weekly dip or a holiday, and each one still needs someone to investigate by hand. After enough false alarms, people stop reading them. Removing the weekly pattern and the trend before judging a change cuts false alarms, and attaching a write-up cuts the triage.