Language in context
Code-switching, local phrasing, and ambiguous intent must survive the full workflow.
Verune Labs independently stress-tests voice and chat agents in real workflows, turns observed failures into engineering priorities, and retests the changes before launch. Starting with consequential workflows in Indonesia.
The readiness gap
Fluency can hide failures in policy enforcement, tool state, payment recovery, scheduling, and human escalation.
Code-switching, local phrasing, and ambiguous intent must survive the full workflow.
Rupiah, relative dates, identity formats, and policy edges change the correct outcome.
The agent must recover from tool-state conflicts and know when a human should take over.
We evaluate whether the agent completes the intended outcome while preserving policy, tool state, user context, and recovery behavior.
We run normal, edge, and adversarial Indonesian interactions across language, policy, tool use, escalation, and recovery—not just isolated prompts.
Evaluation surfaces
Six non-redundant evaluation lenses connect fluent interaction to the operational evidence behind it.
Bisa reschedule ke Jumat, but keep the same driver, ya?
A fluent answer is only the surface. The trace below connects the interaction to an operational launch decision.
Preserve the Friday reschedule and same-driver constraint across both languages, then confirm the combined intent.
The agent replies fluently but reschedules without carrying the same-driver constraint into the tool call.
Natural language quality hides a material change to the requested service.
Extract constraints before response generation and bind them to the structured action payload.
Bilingual constraint preserved
Each finding connects observed behavior to business impact, reproducible evidence, and a concrete next step.
Clear setup, inputs, expected behavior, and observed outcome.
Transcripts and traces connected to each finding.
Severity, business impact, and concrete engineering direction.
Initial behavior versus the result after your changes.

A real voice or chat flow we can exercise
qualification signal
A bounded task with a real launch or remediation decision
qualification signal
Enough context to reproduce and fix failures
qualification signal

Founder-led from Jakarta
Fadli Adrian works directly with teams to scope launch risk, review evidence, and translate observed failures into an engineering decision.
Meet Fadli on LinkedInFAQ
Sprint request
Best fit: a team with a working voice or chat agent, one consequential workflow, a decision to make, and enough access to reproduce meaningful interactions safely.
01 Founder-led fit and scope review
02 Scope and engagement terms confirmed before access
03 No production access required by default