Language in context
Code-switching, local phrasing, and ambiguous intent must survive the full workflow.
Verune stress-tests voice and chat agents against Indonesian language, workflows, policies, tool use, and escalation—before launch.
The readiness gap
A convincing demo can still fail on the messy details that decide whether a customer gets the right outcome.
Code-switching, local phrasing, and ambiguous intent must survive the full workflow.
Rupiah, relative dates, identity formats, and policy edges change the correct outcome.
The agent must recover from tool-state conflicts and know when a human should take over.
A convincing Indonesian response is not enough. We test whether the agent completes the real workflow under operational pressure.
We run normal, edge, and adversarial Indonesian interactions across language, policy, tool use, escalation, and recovery—not just isolated prompts.
Evaluation surfaces
Six non-redundant evaluation lenses connect fluent interaction to the operational evidence behind it.
Bisa reschedule ke Jumat, but keep the same driver, ya?
A fluent answer is only the surface. The trace below connects the interaction to an operational launch decision.
Preserve the Friday reschedule and same-driver constraint across both languages, then confirm the combined intent.
The agent replies fluently but reschedules without carrying the same-driver constraint into the tool call.
Natural language quality hides a material change to the requested service.
Extract constraints before response generation and bind them to the structured action payload.
Bilingual constraint preserved
Each finding connects observed behavior to business impact, reproducible evidence, and a concrete next step.
Clear setup, inputs, expected behavior, and observed outcome.
Transcripts and traces connected to each finding.
Severity, business impact, and concrete engineering direction.
Initial behavior versus the result after your changes.

A real voice or chat flow we can exercise
qualification signal
A concrete workflow, customer, or launch risk
qualification signal
Enough context to reproduce and fix failures
qualification signal

Founder-led from Jakarta
Fadli Adrian works directly with design partners to scope the launch risk, review the evidence, and translate findings into an engineering decision.
Meet Fadli on LinkedInFAQ
Design-partner application
Best fit: a team with a working voice or chat agent, a real Indonesian launch or workflow, and enough access to reproduce meaningful interactions.
01 Founder-led qualification
02 Scope and pilot terms confirmed after fit review
03 No production access required by default