Agent evaluation must test the complete system: model, instructions, context, tools, permissions and failure handling. This template organizes evidence before autonomy expands.
Identification
Record name, version, owner, outcome, users, environments, model, instructions, tools, data sources and date. A relevant change should trigger a new evaluation run.
Test cases
Include case ID, type, input, expected outcome, expected tool, prohibited actions, approval requirement and failure severity. Maintain normal, ambiguous, incomplete, adversarial and unavailable-tool cases.
Metrics
Measure correct completion, tool and parameter accuracy, valid evidence, permission compliance, steps, cost, latency, escalation and error recovery.
Decision
Set thresholds before running the evaluation. Classify failures by impact and record correction, owner and retest. Outcomes may be block, experiment, suggestion mode, approval-gated launch or bounded execution.
The download uses stable columns and an explicitly marked example row. Adapt thresholds to the use case and connect the decision to the AI governance guide.
Use these resources
AI Agents
Understand, build, evaluate and operate AI agents with clear goals and boundaries.