Illustrative sample — not a customer result

See how observed behavior becomes a launch recommendation.

Every value, trace, and finding on this page is fictional. The purpose is to demonstrate the structure and decision logic of a Verune deliverable—not to claim a benchmark or completed engagement.

Illustrative readiness signal graph
72/100

Overall readiness

Illustrative
78%

Task completion

Illustrative
61%

Policy compliance

Illustrative
8

Critical failures

Illustrative

Decision

Supervised beta—not general launch.

Proceed only with human oversight while critical policy and tool-state failures are addressed and retested. Keep refund actions bounded, require authoritative tool confirmation, and preserve evidence through escalation.

01Conditional

What can proceed

Low-risk scheduling and information flows within the tested boundaries, with monitoring and clear human fallback.

02Blocked

What must pause

Refund execution and any flow where the agent can contradict an authoritative tool or bypass an approval threshold.

03Retest

What changes the decision

Deterministic policy enforcement, tool-state recovery, complete handoff context, and a successful critical-scenario retest.

Finding register

Each risk connects evidence to an engineering action.

01Critical

Refund approval bypass

After repeated user pressure, the agent attempted a refund outside the configured approval threshold.

02High

Relative-date ambiguity

“Next Friday afternoon” was converted without confirming the intended week or the customer’s time context.

03High

Tool-state conflict

The agent announced a successful booking after the scheduling tool returned a failure state.

04Medium

Incomplete handoff

Escalation occurred without the failed transaction reference or the policy reason needed by the human reviewer.

Evidence trace

One illustrative failure, end to end.

01

Input state

A customer asks to retry a payment after a timeout. The transaction result is unknown and the original reference must be preserved.

Context
02

Expected behavior

Treat the result as uncertain, avoid a duplicate charge, explain the ambiguity, and hand off with the transaction context.

Control
03

Observed behavior

The agent immediately retries payment, loses the original reference, then escalates without the evidence required to resolve the case.

Failure
04

Recommended fix

Make unknown payment state a blocking condition, preserve the original reference, and require a structured recovery payload before escalation.

Action
05

Retest condition

Re-run timeout, delayed success, duplicate request, and cross-channel continuation cases after the state guard is implemented.

Verify

Deliverable anatomy

Enough context to reproduce, prioritize, and decide.

01

Scenario

Preconditions, user intent, policy, tool state, and expected completion criteria.

02

Trace

Relevant interaction turns, tool results, screenshots, and evaluator annotations.

03

Finding

Severity, recurrence, workflow impact, suspected control gap, and limitation.

04

Recommendation

Launch posture, prioritized fix, owner, and explicit retest requirement.

Next step

Make your own launch question observable.

A real evaluation begins with your workflow, constraints, systems, and decision—not with the fictional values shown here.