Self-hosted anomaly detection and warehouse-native root cause analysis
Metron runs in your environment, reads your warehouse in place with a read-only role, and hands the language model you supply its findings only.
Updated · Zello Labs
Self-hosted anomaly detection means the engine that reads your metrics runs inside your own environment. Metron applies this to root cause analysis. It queries your warehouse in place with a read-only role, keeps no second copy, runs the statistics next to your data, and gives the language model you supply its findings only. The write-up is what leaves.
Why does it matter where root cause analysis runs?
Root cause analysis needs more than the top-line number. To say where a change sits, the engine slices the metric by every column you list and compares each segment against its own history. Those breakdowns are often the most sensitive view of the business: revenue by customer tier, conversion by country, refunds by fulfillment center.
A tool that runs in someone else’s cloud needs those rows sent to it, which means a pipeline to build, a copy to secure and a vendor review to pass. Running the engine next to the warehouse removes all three. The engine does the analysis where the data already lives, and the outputs go only where you send them.
Metron queries in place with a read-only role
You connect Metron with a role that can SELECT the tables behind your metrics. Metron never writes to your warehouse.
- You add a connection to PostgreSQL or Google BigQuery. Metron tests it before it saves anything.
- Metron stores the credentials in a HashiCorp Vault secrets backend, each secret at its own path, outside its own application database. API responses never return them.
- You define a metric by pointing at real columns. Metron validates each expression when you save it, against the table’s real columns and a fixed grammar. A metric definition can only ever ask your warehouse for an aggregate over real columns.
- When a run starts, Metron reads the history it needs straight from your tables, as far back as the data goes. A metric built on a table with months of rows starts with months of history.
There is no replication pipeline and no second copy of your data to keep in step. Each run reads the warehouse as it is at that moment.
Metron keeps its own records in an application database that runs alongside it. That covers connections, metric definitions, investigations with their evidence, logged events and your feedback.
Where does the language model fit?
The engine is the analyst and the language model is the interface. Detection, the segment tests and the event matching all happen in the engine. The model receives a structured set of findings: the metric, the expected and actual values, the segment that holds the change and its share, and the events near the start with their time gaps. It never sees a raw table, so it has nothing to invent a reason from.
You supply the model endpoint. Metron talks to an OpenAI-compatible chat completions API, so you decide where that text goes. It can be a provider your company already uses or a model running inside your own network.
Metron checks the model’s text before anyone reads it. Every number and named entity has to appear in the findings, and causal language is blocked. If the text fails, Metron asks once more. If it fails again, Metron falls back to a plain template built straight from the findings, and the same template covers a deployment with no model configured. Follow-up questions are answered from the same evidence.
What stays in your environment, and what leaves
| What stays in your environment | What leaves |
|---|---|
| Raw rows, queried in place in your warehouse | The engine’s findings, sent to the model endpoint you supply. If that endpoint runs inside your network, they stay too. |
| Warehouse credentials, held in a HashiCorp Vault secrets backend | The write-up, sent by webhook or email to the destination you set |
| Detection, segment tests and event matching | |
| Metric definitions, investigations, logged events and feedback |
Events come in the other direction. GitHub sends deploy and release events to Metron by webhook, and Metron checks each payload’s signature before it records anything. If no webhook secret is configured, Metron refuses every request. You log other events, such as a price change or a campaign, by hand.
What warehouse-native root cause analysis returns
Running in place still gives you the full investigation. Each one reports three claims and scores them separately.
- Something changed. Metron removes the weekly pattern and the trend and judges the change against the metric’s own normal. A new metric gets a simpler model until it has enough history.
- Where it sits. Metron tests each segment of each column against that segment’s own history and says which one holds most of the change, and whether the rate changed or the mix shifted.
- What happened near it. Deploys and logged events close to the start of the change, with the time gap. Metron labels this as timing.
Metron never claims cause. Warehouse data can’t prove one, so the “why” stays with your team. The homepage section on how Metron scores each claim shows all three side by side.
Your feedback on an investigation is stored with the evidence behind it. It does not retrain detection.
What you take on when you self-host
Self-hosting moves some work to your side. You provide a place for Metron to run that can reach your warehouse, a read-only role and a language model endpoint. Today a run starts when you trigger it or when someone asks about a change.
Metron supports PostgreSQL and Google BigQuery today. Snowflake is coming soon, and Amazon Redshift and Azure Synapse are planned. For the difference between watching pipeline health and explaining a business metric, read data observability vs metric root cause analysis.
Metron is a Zello Labs product in private beta, and we set up each team by hand. If you need root cause analysis that stays inside your environment, tell us your warehouse and one metric.
Common questions
What is self-hosted anomaly detection?
It is anomaly detection that runs inside your own environment instead of a vendor's cloud. Metron is self-hosted root-cause analysis for business metrics. It queries your warehouse in place with a read-only role, runs the statistics next to your data, and sends only its findings to the language model endpoint you choose.
Does Metron copy my warehouse data?
No. There is no replication pipeline and no second copy to keep in step. Metron queries PostgreSQL or BigQuery in place each time it runs. It keeps its own records, such as metric definitions, investigations, events and feedback, in an application database that runs in your environment with it.
What does the language model see?
Only the engine's findings: the metric, the size of the change, the segment that holds it and the events near it. The model never reads your tables. Metron checks that every number and name in the model's text appears in those findings, and falls back to a plain template when the check fails.
Where are my warehouse credentials stored?
In a HashiCorp Vault secrets backend, each secret at its own path. They are not kept in Metron's application database or in environment variables, and API responses never include them. Metron tests a connection before it saves anything, and deleting a connection deletes its secret.