Skip to content

Every conversation. Under examination.

Human simulation for voice agents. Match experienced test agents to the conversation, then examine what actually happens.

Caller and agent traces compared with the expected outcome. Thursday afternoon was requested, but Tuesday at ten was confirmed.

Behavioural testing for voice agents.

Inside the evaluation

Choose an example
1 / 4 · Set up0%
nForge/A change of planIllustrative walkthrough · 00:28
01 / Define the test

A specific caller. A measurable condition.

Persona

A caller with a changing request

Connection

Voice or text conversation

Test condition

Confirm the latest preference before creating a booking.

Real runs can generate a suite from your agent description, then execute the selected personas and scenarios.

About this example

Define the caller, the task and what success means.

Authored example · 28 seconds. No audio measurements or evaluation scores are claimed here.

01

Define the caller, the task and what success means.

00:00 / 00:28

Authored example, not a live run. No audio measurements or evaluation scores are claimed here.

AgentNet

A profile.
A history.
A better fit.

A professional network for simulated callers. Each agent brings a distinct persona, voice and behaviour. As it takes part in more tests, its experience builds into its profile.

Match the test to an agent’s skills and relevant experience. Let nForge assign the caller, or choose one yourself.

Meet the test agents
AgentNet / Profile anatomySimulated caller

Lisa

The caller who changes her mind.

“Actually, wait — could we make that Thursday?”

Behaviour
Thinks aloud. Revises answers. Second-guesses a choice.
Useful for
Context retention · Request changes · Confirmation
Experience profile
Tests participated in, scenarios encountered and the results they produced.
Selection
Match relevant experience to the next test.

Illustrative profile structure. No test history or scores are shown.

Who calls.
Where it happens.
What must hold.

Three reusable parts. Combined into a test you can run again.

Agent / Who

Agent / Who

Select a caller by persona, behaviour and relevant testing experience.

Scenario / Where

Scenario / Where

Set the conversation stage, context and required state.

Test / What

Test / What

Define the goal, objections and conditions that must pass.

From intent to a repeatable test

Describe it.
Define it.
Prove it.

Start with a bot description. Generate a suite, then make the expected behaviour explicit in a structured test definition.

Test DSLIllustrative definition
Agent
Caller who changes their mind
Scenario
Appointment chosen, not yet confirmed
Goal
Move the request to Thursday afternoon
Must hold
Confirm the latest preference.
Do not act on the superseded request.

A domain-specific language (DSL) makes the test explicit and reusable. This shows its intent, not executable syntax.

Evaluate the conversationOutcomePerformanceQuality

Keep the gains.
Catch the regressions.

01

Assertions

Define what must hold.

02

Baselines

Compare each run.

03

CI gates

Block failures. Surface warnings.

Voice or text / WebSocket / Twilio / OpenAI Realtime

View sample report