Skip to content
Independent AI agent evaluation

Agents thatwork in the real world.

Verune Labs independently stress-tests voice and chat agents in real workflows, turns observed failures into engineering priorities, and retests the changes before launch. Starting with consequential workflows in Indonesia.

15–30core and edge scenarios
7–10business days, normally
1retest after your fixes

The readiness gap

Speaking Indonesian is not the same as completing Indonesian workflows.

Fluency can hide failures in policy enforcement, tool state, payment recovery, scheduling, and human escalation.

01

Language in context

Code-switching, local phrasing, and ambiguous intent must survive the full workflow.

02

Local operating detail

Rupiah, relative dates, identity formats, and policy edges change the correct outcome.

03

Recovery & escalation

The agent must recover from tool-state conflicts and know when a human should take over.

What we test

Workflow reliability
under pressure.

We evaluate whether the agent completes the intended outcome while preserving policy, tool state, user context, and recovery behavior.

01

Localized workflow stress-testing

We run normal, edge, and adversarial Indonesian interactions across language, policy, tool use, escalation, and recovery—not just isolated prompts.

7readiness dimensions
Methodology

Scope.Test.Decide.

ILLUSTRATIVENOT A CUSTOMER RESULTFRAME 00

Evidence,
not confidence.

Evaluation surfaces

Choose the signals your agent must hold together.

Six non-redundant evaluation lenses connect fluent interaction to the operational evidence behind it.

Input trace · Language contextIllustrative

Code-switching without losing the constraint

Bisa reschedule ke Jumat, but keep the same driver, ya?

A fluent answer is only the surface. The trace below connects the interaction to an operational launch decision.

Expected behavior01

Preserve the Friday reschedule and same-driver constraint across both languages, then confirm the combined intent.

Observed failure02

The agent replies fluently but reschedules without carrying the same-driver constraint into the tool call.

Business risk03

Natural language quality hides a material change to the requested service.

Fix direction04

Extract constraints before response generation and bind them to the structured action payload.

Illustrative retest

Bilingual constraint preserved

Evidence linked
Task completionPolicy complianceTool-use accuracyEscalation judgmentConversation recovery
Engineering-ready deliverables

Findings your team
can act on.

Each finding connects observed behavior to business impact, reproducible evidence, and a concrete next step.

Reproducible scenarios

Clear setup, inputs, expected behavior, and observed outcome.

Annotated evidence

Transcripts and traces connected to each finding.

Prioritized fixes

Severity, business impact, and concrete engineering direction.

Retest comparison

Initial behavior versus the result after your changes.

Data handlingBoundaries agreed before testing
Staging-first by default
Synthetic or sanitized data preferred
Minimum access for the agreed window
Artifacts separated by engagement
Launch Evidence Sprint

Bring a real
launch decision.

Crystalline bird in flight
01

Working agent

A real voice or chat flow we can exercise

Required

qualification signal

  • Usable demo or staging path
  • Defined user outcome
  • Representative tool behavior
  • Known policy boundaries
Discuss your agent
Strongest signal
02

Consequential workflow

A bounded task with a real launch or remediation decision

Strong fit

qualification signal

  • Target Indonesian users
  • Real operating constraints
  • Owner for policy decisions
  • Launch or rollout timing
  • Named business or domain owner
Request a scope review
03

Engineering access

Enough context to reproduce and fix failures

Required

qualification signal

  • Logs, traces, or transcripts
  • Tool and state visibility
  • A technical point of contact
  • Ability to ship fixes
  • One regression retest
Discuss your workflow
Fadli Adrian

Founder-led from Jakarta

Fadli Adrian works directly with teams to scope launch risk, review evidence, and translate observed failures into an engineering decision.

Meet Fadli on LinkedIn
Request the sprint

FAQ

Before we start.

Sprint request

Bring us the agent you need to trust.

Best fit: a team with a working voice or chat agent, one consequential workflow, a decision to make, and enough access to reproduce meaningful interactions safely.

01 Founder-led fit and scope review

02 Scope and engagement terms confirmed before access

03 No production access required by default

Submitting an application does not create a service agreement or guarantee acceptance.