All benchmarks

Whissle conversation-flow suite

Default flows across six agent types

Whether the flow a customer gets out of the box holds up on a real voice call against a caller who improvises — across support, intake, reception, scheduling, rentals and collections, with no per-domain tuning.

61.2%30 of 49 tasks
95% CI 47.2–73.6%
VoicePreliminary16 of 65 excluded (24.6%)Our own judge — Whissle deterministic flow analyzer

Ran on

claude-haiku-4-5

the cheapest and fastest model in its family

provider Anthropic

The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.

Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.

Spoken through

recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record which

synthesis the production synthesis configuration; the run does not record which engine served it

Whissle metadata head — not in the path

Whissle's metadata head — the emotion, intent and entity tags our cascade emits beside the transcript — is not deployed in production, so it was not in the path for any spoken run on this page. These numbers were set without it.

Established by an audit of the production speech configuration, not recorded by the run. The runs log no engine per session, so we can state the configuration they ran against but not which engine served any individual turn.

No published comparator

Run 7 August 2026.

Results

Every metric this run produced

Headline first, then the breakdown. Each row carries the sample it was computed over and its interval — or an explicit statement that there isn't one, which on a small run is the more useful fact.
Headline and per-dimension metrics for this run, with sample size and confidence interval for each.
MetricValueSampleInterval
Task success across six domainsHeadline95% Wilson score interval, computed here from the pass count. Sessions that never produced a turn are excluded; counting every one of them as a failure instead gives 46.2%.61.2%30 of 49 tasks95% CI 47.2–73.6%
Customer supportNo sessions excluded — every one produced a conversation. The denominator is small enough that the interval is wide; read the interval, not the point.63.6%7 of 11 tasks95% CI 35.4–84.8%
Headache intake1 of 10 sessions never produced a turn and are excluded — audio was flowing but no transcript events arrived. Every one of these denominators is small enough that the interval is wide; read the interval, not the point.77.8%7 of 9 tasks95% CI 45.3–93.7%
Dental reception2 of 11 sessions never produced a turn and are excluded — audio was flowing but no transcript events arrived. Every one of these denominators is small enough that the interval is wide; read the interval, not the point.55.6%5 of 9 tasks95% CI 26.7–81.1%
Appointment schedulingNo sessions excluded — every one produced a conversation. The denominator is small enough that the interval is wide; read the interval, not the point.45.5%5 of 11 tasks95% CI 21.3–72.0%
Car rental6 of 11 sessions never produced a turn and are excluded — our own credit gate cut off the simulated caller. Every one of these denominators is small enough that the interval is wide; read the interval, not the point.100.0%5 of 5 tasks95% CI 56.6–100.0%
Debt collection7 of 11 sessions never produced a turn and are excluded — our own credit gate cut off the simulated caller. Every one of these denominators is small enough that the interval is wide; read the interval, not the point.25.0%1 of 4 tasks95% CI 4.6–69.9%
Reproduce

Run this yourself

The harness, the task suites and the raw trajectories are public, and the API these ran against is the same one your agents run on. You do not have to take the number.

You will need: Node 18+ and uv, and a Whissle workspace key (Settings → API keys on whissle.ai). The simulated caller and the flow analyzer both route through Whissle's own model API, so there is no third-party key to obtain — but the workspace does need credit on it, which is the thing that bit this run.

git clone https://github.com/WhissleAI/tau2-bench-w && cd tau2-bench-w
uv sync
export WHISSLE_BASE=https://aws-gateway-backend.whissle.ai/bot WHISSLE_API_KEY=...

./run_flow_sim.sh --agent-type customer_support
./run_flow_sim.sh --agent-type headache_enrollment
# ...and the remaining four types

Budget for the credit the simulated caller spends. 16 of the 65 sessions in the published run died because the workspace ran dry mid-suite.

Running a benchmark straight from the Whissle CLI is on the roadmap and does not exist yet — the CLI manages agents, calls and records today, and we are not going to print a command on this page that would fail when you paste it. Until it lands, the harness above is the path.

Comparison

What we measured against

A comparison is only a comparison when the two numbers come from the same experiment, on stated arms. We split the figures accordingly — produced on this harness with only the agent swapped, or published elsewhere — and both sides name their model.

Our arm in every row below

Whissle ran on

claude-haiku-4-5

the cheapest and fastest model in its family

provider Anthropic

The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.

Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.

Scoring 61.2% over 49 scored tasks. Where a row below names a different model, that difference is listed with the others.

No setup-matched baseline yet

We have not yet run another model through this exact harness — same domain, same task set, same user simulator, same pass^1, same concurrency — so there is no head-to-head to show. Published figures for other systems appear below for context, greyed, with the reasons they are not directly comparable. When a matched run lands it will appear here as a real comparison.

Nobody publishes a comparable figure for this benchmark at this modality, so we show none. An invented comparator would be worse than an empty column.

Exclusions

What was thrown away

The easiest way to improve a benchmark score is to drop the cases that went badly. Here is the arithmetic, so you can check we didn't.

Attempted

65

Scored

49

Excluded

16

16 of 65 excluded (24.6%)

Reasons cases were excluded from scoring, with the count for each.
ReasonCasesWhose fault
The simulated caller's own model calls were refused for insufficient credit, so the session never started. Our billing gate, our workspace, nothing to do with the agent under test.13Ours
Audio was flowing from the agent but no transcript events ever arrived — a dead data channel on the session.3Ours

16 of 65 sessions (24.6%) never produced a single conversational turn, and all of them are our own infrastructure rather than the agent's behaviour. That is a high exclusion rate and it is concentrated: two of the six domains carry 13 of the 16. Because dropping them moves the number in our favour, the floor is published beside it — if every excluded session were counted as an outright failure the figure would be 46.2%. The truth is bounded by 46.2% and 61.2%.

Comparability

What this number can and cannot be put beside

Derived from the run's own configuration rather than written by hand, so it cannot drift away from the result it describes.

Nobody else publishes this, so there is nothing to compare it against, and no external figure is quoted. The judge is ours: a deterministic rule analyzer, but one we wrote, running against a specification we also wrote. Read it as our own instrumentation. The per-domain denominators are 4 to 11 sessions — every interval on this card is wide, and the spread between domains is not yet distinguishable from noise.

Comparable with

  • Other runs of this same benchmark on this page, at the same modality (Voice).
  • Runs of this benchmark on the same arm — claude-haiku-4-5 (the cheapest and fastest model in its family). A run on a different model is a different experiment.
  • Other Pass^1 figures — one attempt per task, no best-of-n.

Not comparable with

  • Any external leaderboard — no published figure exists for this suite at this modality, so we show none.
  • Text-modality leaderboards. The tasks are the same but the input is spoken, which is a strictly harder problem.
  • Any figure produced with Whissle's metadata head in the path. It was not in this one — the run used the production speech configuration, and that head is not deployed there.
  • Independently-judged results. A rule-based auditor that replays the session against the flow's declared spec. It is deterministic rather than an opinion — but we built it and we run it, so it is not an independent verdict. Treat these numbers as our own instrumentation, not as third-party validation.
  • Runs scored on the full suite — 16 of 65 excluded (24.6%) here.
Sample cases

What passing and failing look like

A pass rate tells you how often. These tell you what happened. Excerpts are recorded artifacts from the run — where we have no publishable transcript we say so rather than reconstructing one.

No published cases for this run

We only publish excerpts we actually recorded. This run's trajectories haven't been prepared for publication, and inventing an illustrative transcript would defeat the purpose of the section. Cases land here with the next run.
History

The same benchmark over time

Whether we are actually getting better is a question one number cannot answer. Every run we publish stays on the record, including the ones that went backwards.

One run on record

One run on record — there is no trend to read yet. The next run lands here alongside it. A trend needs at least two runs on the same scenario set — anything less is a point, not a direction.
Every recorded run of this benchmark, oldest first, with its score, sample size, exclusions and the model configuration it ran on.
Run dateScoreScoredExcludedRan onStatus
7 Aug 202661.2%(30 of 49)4916claude-haiku-4-5Preliminary
Methodology

How this run was produced

Enough detail to argue with, and enough to reproduce.
Agent under test
Six shipped Whissle agent types, each booted with the default conversation flow its type ships with — no hand-authored flow, no per-domain tuning — driven over a real voice session.
How it was run
Each session puts a persona with a goal on a real voice call against an agent running its type's default flow. The caller improvises — it disputes, gives wrong information first, changes its mind, refuses to verify — so the flow takes branches a scripted test never would. A session counts as a success only if the caller's goal was actually met. Alongside the score, a deterministic analyzer replays the session against the flow's declared specification and reports illegal transitions, tool leakage, variable desync and dead ends; those findings are the reason the suite exists, and the score is the summary.
Model config
Model claude-haiku-4-5Provider Anthropicthe cheapest and fastest model in its family, in its own vendor's lineup.The harness pinned nothing — no model, no provider, no reasoning-depth or speed setting. Those request controls did not exist on the benchmark endpoint when these runs were made, so every arm inherited the production deployment default.Read from the deployment default in force on each run date, not from the run's own record: the endpoint did not report which model served a turn until after these runs, and no run artifact names one. Automatic failover to the secondary provider (Gemini 2.5 Flash) is logged but not recorded per turn, so we cannot rule out that a small fraction of turns was served by it.
Speech path
Recognition the production speech configuration — a third-party engine chosen by language routing; the run does not record whichSynthesis the production synthesis configuration; the run does not record which engine served itWhissle metadata head — not in the path.Whissle's metadata head — the emotion, intent and entity tags our cascade emits beside the transcript — is not deployed in production, so it was not in the path for any spoken run on this page. These numbers were set without it.Established by an audit of the production speech configuration, not recorded by the run. The runs log no engine per session, so we can state the configuration they ran against but not which engine served any individual turn.
Judge
Our own judge — Whissle deterministic flow analyzer. A rule-based auditor that replays the session against the flow's declared spec. It is deterministic rather than an opinion — but we built it and we run it, so it is not an independent verdict. Treat these numbers as our own instrumentation, not as third-party validation.
Sampling
Full scenario set per domain — every authored caller persona, no sub-selection Population: All authored caller personas across the six agent types the suite covers. No seed — the selection was not randomised. 1 attempt per task.
Run date
7 August 2026
Harness commit
Not recorded for this run. Newer runs pin the commit.
Cost
Not recorded for this run.
Default flows across six agent types — Whissle conversation-flow suite (Voice) · Whissle