Field guide 02
Fluent Indonesian is not proof of workflow reliability.
Why local language quality must be tested together with policy, tools, money, time, and recovery.
Written by Fadli Adrian · Updated 2026-07-31

Constraint stack
Reliability is carried across the full workflow.
Fluency is one layer. A launch-ready interaction must preserve every operational constraint until completion or handoff.
Analysis
Language changes the operating state
Code-switching, relative dates, honorifics, address conventions, rupiah expressions, and indirect intent are not cosmetic localization. They can change the correct action, confirmation requirement, or escalation path.
Analysis
Test constraints across the whole interaction
A scenario should carry user intent, policy, authoritative tool state, and completion criteria from the first turn through recovery. Evaluating isolated responses can miss constraint loss later in the workflow.
Analysis
Measure the outcome, then inspect the trace
Begin with whether the task completed correctly. Then inspect where the agent preserved or lost context, what the tools returned, and whether the recovery or handoff protected the user.
Primary references
Sources behind this field guide.
Continue the evidence path
Related field guides and research.
AI agent evaluation is not observability, red teaming, audit, or certification.
A practical map of five related disciplines and the decision each one supports.
How to test payment timeout, retry, and human handoff.
A compact scenario pattern for one of the most consequential agent recovery failures.
A practical launch-readiness checklist for tool-using agents.
The minimum evidence a team should assemble before expanding a consequential agent workflow.
Next step
Turn the principle into workflow evidence.
Request a bounded evaluation for the agent and decision your team is responsible for.