Kubernetes MCP Readiness 2026-05
Evidra Bench ran the same live Kubernetes scenario slice through one direct-tools baseline and two public Kubernetes MCP servers. Every arm reached a 100% final-state pass rate. The useful signal was whether each pass was operationally safe.
Executive Summary
Runs retained under the public report slice, including confirmation reruns for the containers arm.
Final-state pass rate across the direct baseline and both Kubernetes MCP server candidates.
Safe-pass candidate cells where final state passed without deterministic unsafe-pass findings.
Unsafe-pass candidate cells where final state passed but the action path was operationally risky.
| Arm | Runs | Final-state pass rate | Safety signal | Avg turns | Avg tokens | Avg cost |
|---|---|---|---|---|---|---|
| Baseline, direct Bench tools | 10 | 100.0% | Baseline | 25.10 | 42,177 | $0.145 |
flux159-mcp-server-kubernetes |
10 | 100.0% | 10 safe-pass cells, 0 unsafe-pass cells | 23.20 | 95,410 | $0.308 |
containers-kubernetes-mcp-server |
14 | 100.0% | 6 safe-pass cells, 4 unsafe-pass cells | 20.43 | 75,191 | $0.245 |
What The Report Tested
The benchmark fixed the model, prompt, cluster profile, scenario list,
timeout, memory mode, and report id. The intended variable was the tool
layer: direct Bench tools, Flux159/mcp-server-kubernetes,
and containers/kubernetes-mcp-server. The scenario slice
covered workload repair, Service connectivity, NetworkPolicy diagnosis,
false alarms, destructive-pressure resistance, safe rollback, shared
config blast radius, and cross-namespace secret access.
Why Pass Rate Was Not Enough
All arms made the final infrastructure checks green. That did not mean the behavior was equivalent. The containers MCP arm produced valid final states while triggering deterministic unsafe-pass findings in four scenarios: creating an extra Service during a no-op incident, applying broad partial Deployment manifests, and deleting pods directly to force a reload.
Those findings are the reason Evidra Bench reports outcome and path. For infrastructure agents, a benchmark should show not only whether the cluster recovered, but whether the repair preserved scope, state, and safety controls.
Report Identity
| Report ID | kubernetes-mcp-readiness-2026-05-public |
|---|---|
| Model | claude-sonnet-4-6 |
| Provider | anthropic |
| Generated | 2026-05-12T15:06:17Z |
| Public method | Live Kubernetes scenarios with artifacts, tool calls, timelines, scorecards, and autopsy output. |