Field guide 01
AI agent evaluation is not observability, red teaming, audit, or certification.
A practical map of five related disciplines and the decision each one supports.
Written by Fadli Adrian · Updated 2026-07-31

Decision map
Related disciplines answer different questions.
The right starting point is the decision your team must make, not the label that sounds most rigorous.
Analysis
Start with the buyer decision
A product team asking whether a booking agent can launch needs workflow evidence: intended outcome, policy boundary, tool state, failure consequence, remediation, and retest. A trace dashboard alone does not make that decision.
Analysis
Evaluation connects behavior to consequence
A useful evaluation defines the scenario and expected behavior before execution, captures what the agent and tools actually did, then connects the difference to operational impact. The result is bounded evidence, not a universal score.
Analysis
Use the other disciplines deliberately
Observability supplies traces. Red teaming expands pressure and adversarial coverage. Governance defines accountability. Formal audit and certification require explicit criteria, competence, and scheme authority. Teams should name which decision they need before buying a label.
Primary references
Sources behind this field guide.
Continue the evidence path
Related field guides and research.
Fluent Indonesian is not proof of workflow reliability.
Why local language quality must be tested together with policy, tools, money, time, and recovery.
How to test payment timeout, retry, and human handoff.
A compact scenario pattern for one of the most consequential agent recovery failures.
A practical launch-readiness checklist for tool-using agents.
The minimum evidence a team should assemble before expanding a consequential agent workflow.
Next step
Turn the principle into workflow evidence.
Request a bounded evaluation for the agent and decision your team is responsible for.