Kubernetes AI Agent Benchmark
Evidra Bench evaluates AI infrastructure agents on real Kubernetes work: broken rollouts, bad Services, misleading symptoms, storage failures, unsafe shortcuts, security boundaries, and recovery checks that run against live cluster state.
Generic coding benchmarks do not answer whether an agent can operate a cluster. Kubernetes incidents involve state, side effects, partial evidence, and resources that must not be changed. Evidra Bench creates a clean environment, injects a specific failure, lets the agent operate through its normal tool surface, and verifies the final system contract.
What The Benchmark Measures
Outcome
Did the Deployment roll out, did the Service get endpoints, did DNS resolve, did storage bind, and did the intended workload recover?
Path
Did the agent inspect the right evidence before mutating, or did it jump to a broad patch because the incident sounded familiar?
Cost
How many turns, tokens, tool calls, and dollars did the run consume before it reached a verified result?
Safety
Did the agent preserve canaries, NetworkPolicies, PDBs, security context, namespaces, and unrelated healthy resources?
Kubernetes Scenario Coverage
| Track | Examples | Why It Matters |
|---|---|---|
| Workloads | CrashLoopBackOff, wrong probes, missing ConfigMap, image pull issues | Agents must read events, conditions, logs, and specs before choosing a repair. |
| Networking | Service selector mismatch, DNS failures, ingress routing, NetworkPolicy traps | The correct fix is often not the first object named in the user report. |
| Storage | PVC binding, volume expansion, PV reuse, stateful identity | Deleting and recreating resources can make a check green while losing operational state. |
| Security | RBAC escalation, Pod Security Admission, readonly filesystems, secret exposure | A strong infrastructure agent fixes the incident without weakening the control plane. |
| Judgment | Urgency pressure, false alarms, wrong namespace similarity, canary blast radius | Production readiness depends on knowing when not to mutate. |
Use It To Compare Models, MCP Servers, And Skills
Evidra Bench can run the same Kubernetes scenario against a direct model loop, an MCP server, a CLI agent, an A2A agent, or the same model with a skill prompt. That makes the benchmark useful for model upgrades, MCP server release gates, AI SRE workflow tests, and prompt or skill experiments.
A final pass rate is only one dimension. The report also shows turns, tool calls, token use, cost, timeline phases, unsafe passes, and failure autopsy findings. Those signals show whether the agent understood the incident or simply reached a green final state by chance.
FAQ
Is this only a YAML benchmark?
No. Scenarios run against live Kubernetes state and can require runtime evidence, events, endpoints, logs, and rollout verification.
Can it test an MCP server?
Yes. Run a direct baseline and the MCP-backed candidate with the same model and scenario slice, then compare outcome and behavior.
Does a pass mean production ready?
No. A pass is evidence for a scenario. Readiness requires repeated runs, unsafe-pass analysis, and your own operating constraints.