What it takes to trust an operations agent
The moment that taught me most this month was a failure. Our operations agent was working a broken cluster. Somewhere upstream a feature flag had been flipped, and the noisy symptom was showing up in a different service entirely. We’d told the agent, in plain words, to check flag state before naming a culprit. It did. It fetched the config, had the answer sitting right there in its context, and then submitted its first guess anyway. Three times out of three.
It preferred its own hypothesis to its own evidence. Anyone who has run an incident call has met a person like that. The difference is that we could change the agent.
The benchmark, and the result
We’re building an operations agent as part of Voxeltron, the half of Deep Thinking that runs the software AI builds. Its job is to notice when something is wrong, work out where and why, and propose a fix. Before we trust it with anything, we want to know how good it is, and that means measuring it against something we didn’t write.
AIOpsLab, Microsoft’s benchmark for autonomous incident response, is the closest public thing to that work. It deploys real microservice applications onto a real Kubernetes cluster, injects a real fault, and hands the agent a shell with the same tools an on-call engineer would use. It’s graded on detection, localisation, root-cause analysis and mitigation, and a mitigation only counts if the cluster actually recovers. There’s no answer to paste in. The pods come back healthy or they don’t.
Our starting scaffold scored 76.1%. Five days later, detection, localisation and mitigation together came in at 91.89%, across 222 scenario runs, with a 95% confidence interval of 87.5% to 94.8%. The best published score we know of is Cerebral Systems’ 91.39%, from their benchmark page. Their number sits inside our interval, so the careful claim is that we’re at or above it, not that we beat it.
A few honest caveats before anything else. This is an internal result. We ran it ourselves and nobody has reproduced it. Cerebral’s figure includes root-cause analysis and ours above does not. Our root-cause score was 53.8%, and with it included the four-task composite is 86.2%, below theirs. I’ll come back to why. And their localisation score is higher than ours, which frankly interests me more than our own headline.
So read this as a description of a week’s work and what we learned, not as a leaderboard claim.
Measure before you change anything
When you’re fifteen points behind, the tempting move is to guess. Rewrite the prompt, run the whole benchmark, see if the number moves. A full run takes most of a day, and a two-point swing on one run is indistinguishable from noise. Guess and rerun burns a day per guess and teaches you very little.
So we stopped doing it. Every idea had to survive cheap checks first: does the evidence it relies on actually exist in transcripts we already had? Does it help on a few scenarios picked at random rather than by hand? Most ideas died at the first check, in minutes, for almost nothing. One whole theory about timeouts fell over as soon as we looked properly and found the timeouts had happened after grading.
Over five days, we ran the full benchmark twice. Everything else was small and fast. The expensive measurement was something an idea had to earn.
Put the rule in code, not in the prompt
The flag story wasn’t a one-off. We re-learned the same lesson several times that week, because apparently we needed to. If a step is only in the prompt, a confident model will skip it at exactly the moment it matters.
What fixed it was taking the step out of the model’s hands. Checks the agent had to perform became checks the system performed around the agent, in code, with the result put in front of it in terms it couldn’t misread. The model kept its judgement. It lost the option of skipping the homework.
I think this is the most transferable thing in the whole exercise. When you evaluate any agent that acts on real systems, ask which of its safety steps are instructions and which are enforced. Instructions are hopes.
Spend where the model is good
By the end we had measurements for two models on every task, and they didn’t say what the price list implies. The bigger model wasn’t better everywhere. The cheaper one was better at noticing that something was wrong, partly because the bigger one kept investigating a yes-or-no question long after it had the answer. The bigger one was much better at fixing things live, which is the genuinely hard reasoning. So each task goes to the model that measured best at it.
If your agent platform can’t do that, you’re paying top prices for work a smaller model does better.
Don’t game the answer key
This is the part I care most about, because it’s where benchmark results quietly stop meaning anything.
Some network faults in the benchmark leave no trace in logs, metrics or traces. The only visible evidence is the fault injection itself, which sits in the cluster as ordinary configuration. An on-call engineer looking at that namespace would see it immediately and would be negligent to ignore it. It’s also the benchmark’s own fingerprint. We went back and forth on it, and settled on two rules. The agent can read live cluster state, as a real engineer could, but it never reads the benchmark’s answer files or scenario mappings, and a test enforces that. And every session that relied on that kind of evidence is marked, so the split is visible. On the final run, localisation scored 40 of 46 on marked sessions and 37 of 38 on unmarked ones.
Root-cause analysis is where the line cost us. It asks the agent to pick a layer and fault type from fixed menus, and on a handful of scenarios the benchmark’s expected label disagrees with what, as far as we can tell, every strong model reads from the same evidence. We tried several approaches and none moved those cases. The known shortcut is to teach the agent the expected label for each scenario. That’s memorising the answer key, and a score bought that way would say nothing about how the agent handles your incidents. So we left root cause at 53.8% and published it next to the rest.
We’d make the same trade again. I’d also suggest that a vendor who won’t show you their equivalent split is asking for a lot of faith.
What this means if you’re evaluating one
An operations agent is something you’ll eventually let touch production, so the questions are less about the score and more about how it was earned.
- Was it measured against live systems, where a fix only counts if the system recovers?
- Which of its checks are enforced by the system around it, and which are just instructions?
- Can it show you which evidence each decision relied on?
- What does it do when a fix is outside the limits you’ve set? It should stop and ask.
- Has anyone outside the vendor reproduced the result?
That last one applies to us too. Voxeltron and its operations agent are in development, we don’t have customers yet, and this result is ours alone until someone else runs it. If you’d like to be one of the people who does, or you’re working on the same problem, we’d like to hear from you.
Related reading: DeepSeek V4 just reset the price floor for agentic AI, human approval gates that don’t bottleneck, and cost governance before the invoice.