Benchmark analysis

Kubernetes MCP Servers Passed. That Was Not Enough.

Kubernetes MCP servers passed our live benchmark. That was not the interesting part. The interesting part was what happened on the way to the green checks.

In May 2026, Evidra Bench ran public Kubernetes MCP readiness reports. The primary report used Claude Sonnet 4.6 across ten live Kubernetes scenarios. It compared a direct Bench tools baseline with Flux159/mcp-server-kubernetes and containers/kubernetes-mcp-server. Every arm reached a 100% final-state pass rate.

Did the agent pass safely?

That is the harder question. For infrastructure agents, final pass/fail is too weak. A system can end in a valid state after the agent changed the wrong resource, deleted something unnecessary, applied a broad manifest, or got lucky because the verifier checked only the final contract.

The Signal

ReportCandidate cellsSafe passUnsafe passFail
Claude Sonnet 4.6 primary report 20 16 4 0
DeepSeek V4 Flash pilot 6 4 2 0

What Passed Unsafely Looked Like

The unsafe passes were concrete action paths that would matter in an incident review. One false-alarm run created an extra Service even though the workload was already healthy. Two runs recovered visible symptoms with broad partial Deployment manifests. One shared ConfigMap run forced recovery by deleting pods directly.

Final checks can miss those differences. A benchmark for infrastructure agents needs to make them visible.

MCP Servers Change Behavior

A tool server changes what resources the model sees first, how verbose tool results are, whether mutations are scoped or broad, how easy it is to apply partial manifests, and whether tool calls are useful for audit. The benchmark should measure that path, not just the green endpoint.

The sample is too small to declare a permanent ranking of Kubernetes MCP servers. The useful conclusion is narrower: final-state pass rate hid real behavioral differences.

The Direction

Evidra Bench is built around live infrastructure exams with failure autopsy. It should answer whether the agent identified the right root cause, inspected enough evidence, preserved safety controls, avoided touching healthy resources, chose a narrow repair, and produced evidence a human can inspect.

If you build an AI SRE agent, Kubernetes MCP server, or infrastructure automation tool, the question is no longer only whether it can pass. The question is whether it can pass safely.