Field guide 01

AI agent evaluation is not observability, red teaming, audit, or certification.

A practical map of five related disciplines and the decision each one supports.

Written by Fadli Adrian · Updated 2026-07-31

Crystalline connections representing adjacent assurance disciplines

Decision map

Related disciplines answer different questions.

The right starting point is the decision your team must make, not the label that sounds most rigorous.

Discipline
Primary question
Useful output
Observability
What happened?
Traces and operational signals
Red teaming
How can it be broken?
Adversarial failure paths
Evaluation
Did it meet the expected outcome?
Behavior tied to consequence
Audit
Does evidence satisfy criteria?
Review against an explicit basis
Certification
Can conformity be formally stated?
Scheme-bound assurance
01

Analysis

Start with the buyer decision

A product team asking whether a booking agent can launch needs workflow evidence: intended outcome, policy boundary, tool state, failure consequence, remediation, and retest. A trace dashboard alone does not make that decision.

02

Analysis

Evaluation connects behavior to consequence

A useful evaluation defines the scenario and expected behavior before execution, captures what the agent and tools actually did, then connects the difference to operational impact. The result is bounded evidence, not a universal score.

03

Analysis

Use the other disciplines deliberately

Observability supplies traces. Red teaming expands pressure and adversarial coverage. Governance defines accountability. Formal audit and certification require explicit criteria, competence, and scheme authority. Teams should name which decision they need before buying a label.

Primary references

Sources behind this field guide.

Continue the evidence path

Related field guides and research.

Read the flagship research report

Next step

Turn the principle into workflow evidence.

Request a bounded evaluation for the agent and decision your team is responsible for.