Agent evaluation must test the complete system: model, instructions, context, tools, permissions and failure handling. This template organizes evidence before autonomy expands.

Identification

Record name, version, owner, outcome, users, environments, model, instructions, tools, data sources and date. A relevant change should trigger a new evaluation run.

Test cases

Include case ID, type, input, expected outcome, expected tool, prohibited actions, approval requirement and failure severity. Maintain normal, ambiguous, incomplete, adversarial and unavailable-tool cases.

Metrics

Measure correct completion, tool and parameter accuracy, valid evidence, permission compliance, steps, cost, latency, escalation and error recovery.

Decision

Set thresholds before running the evaluation. Classify failures by impact and record correction, owner and retest. Outcomes may be block, experiment, suggestion mode, approval-gated launch or bounded execution.

The download uses stable columns and an explicitly marked example row. Adapt thresholds to the use case and connect the decision to the AI governance guide.

[ PRACTICAL RESOURCES ]

Use these resources

[ CONTINUE EXPLORING ]

AI Agents

Understand, build, evaluate and operate AI agents with clear goals and boundaries.

Primary sources