Data observability vs business metric monitoring
Data observability checks that your data arrived intact, and metric root-cause analysis explains why a business metric moved once you trust the numbers.
Updated · Zello Labs
Data observability watches the health of your data, such as freshness, volume, schema changes and failed pipeline runs. Anomaly detection on business metrics watches the numbers themselves. Metric root-cause analysis takes the next step. It assumes the data is correct and explains why a metric like revenue or conversion moved, and where the change sits.
What does data observability watch?
Data observability tools watch your data on its way into the warehouse. The usual checks answer a handful of questions about each table.
- Did it update on time?
- Did roughly the expected number of rows arrive?
- Did a column change type, get renamed or disappear?
- Did a job fail, or a test on the data stop passing?
- Are null values showing up in a field that is usually filled?
When a check fails, the alert points at a table, a column or a job. The person who acts on it is usually a data engineer or whoever owns that pipeline. The fix is a rerun, a backfill or a schema update.
Data observability often uses anomaly detection under the hood. It learns the normal row count or load time for a table and flags a departure from it. The method looks a lot like anomaly detection on a business metric. The signal is different. A data check asks whether the data is complete and on time. It has no view on whether the business had a good week.
What does metric root-cause analysis do?
Metric root-cause analysis starts from the business number, such as revenue, activation rate or refund rate. It takes the data behind that number as correct and asks three questions in order.
- Did the metric really move, or is this normal variation for it?
- Where does the change sit?
- What happened close to the time it started?
The first question is anomaly detection, applied to the business metric. The second and third are the root-cause part. The output is an explanation that someone outside the data team can read, with the numbers in it. For the steps in more detail, read our guide to working out why a metric dropped.
Data observability vs metric root-cause analysis
| Data observability | Metric root-cause analysis | |
|---|---|---|
| Question it answers | Is my data fresh, complete and in the shape I expect? | Why did this business metric move, and where is the change? |
| What it watches | Tables, columns, jobs and pipelines, including freshness, row counts, schema and null rates | Business metrics such as revenue, conversion or refund rate, split by the columns you slice them by |
| Who uses it | Data engineers and pipeline owners | Analysts, data teams and the people who ask them why a number moved |
| Typical output | An alert on a table or job, such as a stale table, a schema change or a failed run | A short write-up with the size of the change, the segment that holds most of it and events near the start |
| What it assumes | Nothing about the business | The data behind the metric is correct |
One drop, two readings
Take a hypothetical SaaS activation rate that falls 9.4% below expected on a Wednesday.
A data observability tool checks the tables behind it. The signups table loaded on time. Row counts look normal. No column changed. Every check passes, and that is the right result, because the data is fine. Activation really fell.
Metric root-cause analysis picks up from there. It tests each segment of each column you slice by and finds that 68% of the drop sits in signups through Google SSO, where activation fell from 41% to 29%. It also finds a release of the auth service that went out 52 minutes before the drop began. The release and the drop line up in time. Whether one caused the other is your team’s call. See how Metron scores each of those claims.
In the reverse case, a nightly load fails halfway and yesterday’s orders table holds half its usual rows. Revenue looks like it fell by half. Data observability catches this at the source. A metric tool that took the number at face value would send you after a business problem that doesn’t exist.
Can you use both?
Yes. Each one covers a gap the other leaves. Data observability tells you when to distrust a number. Metric root-cause analysis tells you what changed when the number is right. Keep pipeline alerts with the people who own the pipelines, and send metric write-ups to the people who own the metric.
The two meet when a key number moves. Start by asking whether the data is intact. If a data check failed for the same day, look there first. If every check passed, treat the change as real and start from the root-cause write-up. For what happens once you know a change is real, see what root-cause analysis adds after an anomaly alert.
How Metron handles bad data
Metron is a metric root-cause tool, and it does not watch your pipelines. It runs four narrow checks on the data behind each metric, covering freshness, a collapse in row count, a spike in null values and a segment that disappears from the breakdown. When one fires, Metron shows the change as a data issue and keeps it off the list of business anomalies.
These checks have limits. For a metric that is just a count of rows, a real drop and a missing load look the same, so Metron can’t tell them apart. The checks exist to stop a broken load from reading as a business change. For full coverage of your pipelines, use a tool built for that job.
Metron is in private beta, and we are setting up a small number of teams by hand. If you run Postgres or BigQuery and have a metric your team can’t explain, request beta access.
Common questions
Is data observability the same as anomaly detection?
No. Anomaly detection is a method, and data observability is one place it gets used. Data observability applies it to signals about the data itself, such as row counts and load times. Business metric monitoring applies it to the number you report, such as revenue or conversion, and asks whether that value moved outside its normal range.
Can data observability tell me why revenue dropped?
Only when the cause is the data itself, such as a late load or a broken join. If the pipeline is healthy and revenue really fell, a data check has nothing to flag. You need a tool that works on the revenue number, finds the segment that holds the drop, and checks what happened near the start.
Do I need data observability before metric root-cause analysis?
You don't need one before the other. If your pipelines break often, fix that first, because root-cause analysis assumes the numbers are right. If your data is mostly reliable and the open question is why revenue or conversion moved, root-cause analysis answers that directly. Teams with both problems can run both.
Does Metron check data quality?
Metron runs four narrow checks on the data behind each metric. It looks at freshness, a collapse in row count, a spike in null values and a segment that disappears. A change that trips one is shown as a data issue and never reported as a business anomaly. The checks don't replace a data observability tool.