AI SRE Regression Testing
Evidra Bench turns production-style infrastructure failures into recurring live regression tests for AI SRE agents, infrastructure automation tools, MCP servers, and operational skills.
AI agents can improve from one model release to the next, then regress in a different kind of incident. A model may become better at Helm but worse under urgency pressure. A skill may help simple Kubernetes rollouts but hurt harder diagnosis tasks. A new MCP server release may reduce turns while adding unsafe mutation paths. Regression testing is how those changes become visible before the agent reaches production.
What Gets Regressed
Models
Compare a new model, reasoning setting, or provider route against your previous baseline on the same incidents.
Skills and prompts
Test whether an operational skill improves diagnosis or only adds ceremony and tool-call overhead.
MCP servers
Measure whether a tool server helps the agent inspect, scope, mutate, and verify infrastructure safely.
Customer incidents
Convert sanitized outage patterns into private regression suites that can run before major agent changes.
Signals Beyond Green Checks
A regression suite should not stop at "passed" or "failed." Evidra Bench tracks the operational path: discovery, diagnosis, mutation, verification, turns, token use, estimated cost, tool calls, failure patterns, and unsafe passes. That makes it possible to distinguish a clean recovery from a risky recovery that only happened to satisfy the final verifier.
This matters for AI SRE workflows because the agent is not only writing code. It is acting against systems with state, policies, blast radius, and humans who need to understand what happened after the run.
Example Release Gate
| Gate | Regression Signal | Decision |
|---|---|---|
| Model upgrade | Pass rate improves, but unsafe-pass count rises under urgency and namespace-similarity tasks. | Hold rollout or restrict the model to lower-risk tasks. |
| MCP server update | Turns fall, but token usage doubles and verifier evidence shows broad manifests. | Fix tool schemas and mutation affordances before release. |
| Skill prompt change | Easy tasks improve, L3/L4 diagnosis regresses, and the agent stops re-checking after failed fixes. | Split the skill by domain or remove over-prescriptive rules. |
Private Regression Reports
Public reports are useful proof points. Private reports are where the release workflow becomes practical. Evidra Bench can run a fixed scenario slice against the current stack and the candidate stack, then produce a report that shows exactly what changed. For platform teams, that creates a repeatable evidence loop for AI SRE adoption.
FAQ
Is this a replacement for staging?
No. It is a controlled regression layer that should run before agents are trusted in staging or production workflows.
Can past incidents become scenarios?
Yes. Sanitized incident patterns can become private scenarios with deterministic setup, break, and verifier steps.
What is a good first suite?
Start with Kubernetes workload, networking, security, and release scenarios that match the agent's expected operating scope.