Evaluation methodology

Evidence begins with the decision—not the demo.

Verune works backward from a real operational outcome, then tests the language, policy, tool, and recovery behavior required to reach it reliably.

Crystalline network representing connected evaluation scenarios

Four-stage loop

A bounded process with an explicit finish line.

01

Define the decision

Document the workflow outcome, user intent, policy constraints, system boundaries, known risks, and evidence needed to support a launch decision.

Scope
02

Build the scenario set

Translate the workflow into normal, boundary, ambiguous, and adversarial interactions grounded in Indonesian language and operating context.

Design
03

Run and adjudicate

Capture agent behavior and tool state, compare it with the expected outcome, classify failures, and record reproducible evidence.

Test
04

Remediate and retest

Prioritize fixes by impact, re-run critical cases, and update the recommendation based on observed behavior rather than intent.

Verify

Evaluation lenses

Fluency is only one part of a reliable outcome.

Each scenario can exercise multiple lenses at once because real failures rarely remain inside a single category.

01

Language in context

Code-switching, local phrasing, ambiguity, and intent preservation across the complete workflow.

02

Local operating detail

Rupiah, relative dates, time context, address formats, identity conventions, and policy edges.

03

Tool-state integrity

Whether agent claims, confirmations, and next actions remain consistent with authoritative system responses.

04

Policy boundaries

Approval limits, restricted actions, required confirmations, and behavior under pressure or manipulation.

05

Human handoff

Whether escalation happens at the right time with enough context for a person to continue safely.

06

Failure recovery

How the agent handles uncertainty, partial completion, conflicting state, and repeated attempts.

Evidence standard

A finding must be reproducible, consequential, and actionable.

01

Input

The exact scenario, state, constraints, and interaction context that produced the behavior.

02

Observation

What the agent said or did, including relevant tool response and workflow state.

03

Impact

Why the behavior matters to the user, operation, policy, or launch decision.

04

Next action

A concrete fix, control, owner, or retest condition tied back to the evidence.

Next step

Turn a risky workflow into a testable one.

Bring the use case and the launch decision. We will work backward to the evidence your team needs.